Notation
- $n$ – Sample size
- $N$ – Population size
- $x_i, y_i$ – $i$-th observation of variables $X$ and $Y$
- $\overline{x}, \overline{y}$ – Sample means of $X$ and $Y$
- $\mu_x, \mu_y$ – Population means of $X$ and $Y$
- $s_x, s_y$ – Sample standard deviations of $X$ and $Y$
- $\sigma_x, \sigma_y$ – Population standard deviations of $X$ and $Y$
- $f_j$ – Frequency of $j$-th category
Numeric / Count Data
Central Tendency
Mean (Arithmetic Mean)
The arithmetic mean is the sum of all observed values divided by the total number of observations. It represents the centre of gravity of a distribution.
Sample Mean:
$$\overline{x} = \frac{1}{n} \sum_{i=1}^{n} x_i$$
Population Mean:
$$\mu = \frac{1}{N} \sum_{i=1}^{N} x_i$$
Sensitive to outliers; applicable only to interval and ratio scales.
Median
The median is the value of the point which has half the data smaller than that point and half the data larger than that point.
$$\text{Median} = 50^{th} \text{ percentile or }Q2$$
Robust to outliers; applicable to ordinal, interval, and ratio scales.
Mode
The mode is the value (or set of values) occurring with the greatest frequency in a dataset. It is not necessarily unique.
$$\text{Mode} = \underset{x}{\arg\max}\ f(x)$$
The only measure of central tendency applicable to nominal (categorical) data. It is the category occurring with the highest frequency
Spread / Dispersion
Range
The range is the difference between the maximum and minimum observed values – the simplest measure of statistical dispersion.
$$R = x_{\max} – x_{\min}$$
Highly sensitive to outliers as it depends solely on the two extreme values.
Variance
Variance measures the average squared deviation of each observation from the mean, quantifying spread in squared units of the original variable.
Sample Variance
$$s^2 = \frac{1}{n-1} \sum_{i=1}^{n} (x_i – \overline{x})^2$$
Population Variance:
$$\sigma^2 = \frac{1}{N} \sum_{i=1}^{N} (x_i – \mu)^2$$
The denominator $n-1$ ensures the sample variance is an unbiased estimator of the population variance.
Standard Deviation
The standard deviation is the positive square root of variance. It expresses dispersion in the same units as the original observations, making it directly interpretable alongside the mean.
Sample:
$$s = \sqrt{\frac{1}{n-1} \sum_{i=1}^{n} (x_i – \overline{x})^2}$$
Population:
$$\sigma = \sqrt{\frac{1}{N} \sum_{i=1}^{N} (x_i – \mu)^2}$$
Interquartile Range (IQR)
The IQR is the difference between the third quartile ($Q_3$, 75th percentile) and the first quartile ($Q_1$, 25th percentile), representing the spread of the central 50% of the data.
$$\text{IQR} = Q_3 – Q_1$$
Values below $Q_1 – 1.5 \cdot \text{IQR}$ or above $Q_3 + 1.5 \cdot \text{IQR}$ are flagged as outliers.
Coefficient of Variation (CV)
The CV is a dimensionless measure of relative dispersion expressed as the ratio of the standard deviation to the mean. It enables comparison of variability across datasets with different units or scales.
$$CV = \frac{s}{\overline{x}} \times 100\%$$
Meaningful only when the mean is positive and data are on a ratio scale.
Position & Shape
Quartiles & Percentiles
The $p$-th percentile is the value below which $p\%$ of observations fall. Quartiles divide the distribution into four equal parts at $Q_1$ (25th), $Q_2$ (50th), and $Q_3$ (75th) percentiles.
Positional Index:
$$L_p = \frac{p}{100} \times (n + 1)$$
Interpolation between adjacent order statistics is applied when $L_p$ is non-integer.
Skewness
Skewness is the third standardised central moment, measuring the degree and direction of asymmetry in a distribution. Zero indicates symmetry; positive values indicate a right (positive) tail; negative values indicate a left (negative) tail.
$$g_1 = \frac{\dfrac{1}{n} \displaystyle\sum_{i=1}^{n} (x_i – \overline{x})^3}{\left[\dfrac{1}{n} \displaystyle\sum_{i=1}^{n} (x_i – \overline{x})^2 \right]^{3/2}}$$
Kurtosis
Kurtosis is the fourth standardised central moment, measuring tail heaviness relative to a normal distribution. Excess kurtosis (kurtosis $-$ 3) is commonly reported: positive = heavy tails (leptokurtic), negative = light tails (platykurtic).
Kurtosis:
$$\text{Kurt} = \frac{\dfrac{1}{n} \displaystyle\sum_{i=1}^{n} (x_i – \overline{x})^4}{\left[\dfrac{1}{n} \displaystyle\sum_{i=1}^{n} (x_i – \overline{x})^2 \right]^{2}}$$
Excess Kurtosis:
$$\gamma_2 = \text{Kurt} – 3$$
The normal distribution has kurtosis $= 3$ and excess kurtosis $= 0$.
Bivariate
Covariance
Covariance is a measure of the joint variability of two random variables. It quantifies the degree to which two variables change together — that is, whether an increase in one variable tends to be associated with an increase or decrease in the other. A positive covariance indicates that the variables tend to move in the same direction; a negative covariance indicates they tend to move in opposite directions; a value of zero suggests no linear relationship.
Sample Covariance:
$$\text{Cov}(X, Y) = s_{xy} = \frac{1}{n-1} \sum_{i=1}^{n} (x_i – \overline{x})(y_i – \overline{y})$$
Population Covariance:
$$\text{Cov}(X, Y) = \sigma_{xy} = \frac{1}{N} \sum_{i=1}^{N} (x_i – \mu_x)(y_i – \mu_y)$$
The denominator $n-1$ makes the sample covariance an unbiased estimator of the population covariance.
Key Properties:
- $\text{Cov}(X, X) = \text{Var}(X)$
- $\text{Cov}(X, Y) = \text{Cov}(Y, X)$ — symmetric
- Units are the product of the units of $X$ and $Y$, making it hard to interpret across different scales.
Pearson Correlation Coefficient
The Pearson correlation coefficient is a normalised measure of the linear association between two continuous variables. It is defined as the covariance of the two variables divided by the product of their standard deviations, scaling the result to a dimensionless value in the range $[-1, +1]$.
A value of $+1$ indicates a perfect positive linear relationship, $-1$ a perfect negative linear relationship, and $0$ indicates no linear relationship.
Sample Pearson Correlation:
$$r_{xy} = \frac{\displaystyle\sum_{i=1}^{n}(x_i – \overline{x})(y_i – \overline{y})}{\sqrt{\displaystyle\sum_{i=1}^{n}(x_i – \overline{x})^2} \cdot \sqrt{\displaystyle\sum_{i=1}^{n}(y_i – \overline{y})^2}}$$
Equivalently expressed using covariance:
$$r_{xy} = \frac{s_{xy}}{s_x \cdot s_y}$$
Population Pearson Correlation:
$$\rho_{xy} = \frac{\sigma_{xy}}{\sigma_x \cdot \sigma_y}$$
$r$ is the sample estimator; $\rho$ (rho) denotes the population parameter.
Interpretation Guide:
- $0.00 – 0.19$ $\rightarrow$ Negligible
- $0.20 – 0.39$ $\rightarrow$ Weak
- $0.40 – 0.59$ $\rightarrow$ Moderate
- $0.60 – 0.79$ $\rightarrow$ Strong
- $0.80 – 1.00$ $\rightarrow$ Very Strong
Measures only linear association. A value near 0 does not rule out a non-linear relationship.
Spearman Rank Correlation
The Spearman rank correlation coefficient is a non-parametric measure of monotonic association between two variables. It is computed by applying the Pearson formula to the ranked values of the observations rather than the raw values, making it robust to outliers and applicable to ordinal data.
Let $R(x_i)$ and $R(y_i)$ denote the ranks of $x_i$ and $y_i$ respectively, and $d_i = R(x_i) – R(y_i)$ be the rank difference.
Spearman Correlation (simplified – no tied ranks):
$$r_s = 1 – \frac{6 \displaystyle\sum_{i=1}^{n} d_i^2}{n(n^2 – 1)}$$
General Form (via Pearson on ranks):
$$r_s = \frac{\displaystyle\sum_{i=1}^{n}(R(x_i) – \overline{R}_x)(R(y_i) – \overline{R}_y)}{\sqrt{\displaystyle\sum_{i=1}^{n}(R(x_i) – \overline{R}_x)^2} \cdot \sqrt{\displaystyle\sum_{i=1}^{n}(R(y_i) – \overline{R}_y)^2}}$$
The simplified formula with $d_i^2$ is exact only when there are no tied ranks. Use the general form when ties are present.
Categorical Data
Frequency Analysis
Frequency Count
The absolute frequency of a category $c_j$ is the number of times it appears in the dataset. It is the fundamental building block of all categorical summary statistics.
$$f_j = \sum_{i=1}^{n} \mathbf{1}[x_i = c_j], \quad j = 1, 2, \ldots, k$$
$\mathbf{1}[\cdot]$ is the indicator function, equal to 1 when the condition is true and 0 otherwise.
Relative Frequency (Proportion)
The relative frequency of a category is the ratio of its absolute frequency to the total number of observations. It estimates the probability of that category in the underlying population.
$$p_j = \frac{f_j}{n}, \qquad \text{subject to} \quad \sum_{j=1}^{k} p_j = 1$$
Percentage Frequency
Percentage frequency scales the relative frequency to a base of 100, expressing each category’s share of the whole.
$$\%_j = \frac{f_j}{n} \times 100$$
Cumulative Frequency
Cumulative frequency is the running total of frequencies up to and including a given category, meaningful when categories have a natural order (e.g., ordinal data).
Absolute Cumulative Frequency:
$$F_j = \sum_{l=1}^{j} f_l$$
Cumulative Relative Frequency:
$$P_j = \frac{F_j}{n} = \sum_{l=1}^{j} p_l$$