Descriptive Statistics : a glimpse

  • $n$ – Sample size
  • $N$ – Population size
  • $x_i, y_i$ – $i$-th observation of variables $X$ and $Y$
  • $\overline{x}, \overline{y}$ – Sample means of $X$ and $Y$
  • $\mu_x, \mu_y$ – Population means of $X$ and $Y$
  • $s_x, s_y$ – Sample standard deviations of $X$ and $Y$
  • $\sigma_x, \sigma_y$ – Population standard deviations of $X$ and $Y$
  • $f_j$ – Frequency of $j$-th category

The arithmetic mean is the sum of all observed values divided by the total number of observations. It represents the centre of gravity of a distribution.

Sample Mean:

$$\overline{x} = \frac{1}{n} \sum_{i=1}^{n} x_i$$

Population Mean:

$$\mu = \frac{1}{N} \sum_{i=1}^{N} x_i$$

 Sensitive to outliers; applicable only to interval and ratio scales.


The median is the value of the point  which has half the data smaller than that point and half the data larger than that point.

$$\text{Median} = 50^{th} \text{ percentile or }Q2$$

Robust to outliers; applicable to ordinal, interval, and ratio scales.


The mode is the value (or set of values) occurring with the greatest frequency in a dataset. It is not necessarily unique.

$$\text{Mode} = \underset{x}{\arg\max}\ f(x)$$

The only measure of central tendency applicable to nominal (categorical) data. It is the category occurring with the highest frequency


The range is the difference between the maximum and minimum observed values – the simplest measure of statistical dispersion.

$$R = x_{\max} – x_{\min}$$

Highly sensitive to outliers as it depends solely on the two extreme values.


Variance measures the average squared deviation of each observation from the mean, quantifying spread in squared units of the original variable.

Sample Variance

$$s^2 = \frac{1}{n-1} \sum_{i=1}^{n} (x_i – \overline{x})^2$$

Population Variance:

$$\sigma^2 = \frac{1}{N} \sum_{i=1}^{N} (x_i – \mu)^2$$

The denominator $n-1$ ensures the sample variance is an unbiased estimator of the population variance.


The standard deviation is the positive square root of variance. It expresses dispersion in the same units as the original observations, making it directly interpretable alongside the mean.

Sample:

$$s = \sqrt{\frac{1}{n-1} \sum_{i=1}^{n} (x_i – \overline{x})^2}$$

Population:

$$\sigma = \sqrt{\frac{1}{N} \sum_{i=1}^{N} (x_i – \mu)^2}$$


The IQR is the difference between the third quartile ($Q_3$, 75th percentile) and the first quartile ($Q_1$, 25th percentile), representing the spread of the central 50% of the data.

$$\text{IQR} = Q_3 – Q_1$$

Values below $Q_1 – 1.5 \cdot \text{IQR}$ or above $Q_3 + 1.5 \cdot \text{IQR}$ are flagged as outliers.


The CV is a dimensionless measure of relative dispersion expressed as the ratio of the standard deviation to the mean. It enables comparison of variability across datasets with different units or scales.

$$CV = \frac{s}{\overline{x}} \times 100\%$$

Meaningful only when the mean is positive and data are on a ratio scale.


The $p$-th percentile is the value below which $p\%$ of observations fall. Quartiles divide the distribution into four equal parts at $Q_1$ (25th), $Q_2$ (50th), and $Q_3$ (75th) percentiles.

Positional Index:

$$L_p = \frac{p}{100} \times (n + 1)$$

Interpolation between adjacent order statistics is applied when $L_p$ is non-integer.


Skewness is the third standardised central moment, measuring the degree and direction of asymmetry in a distribution. Zero indicates symmetry; positive values indicate a right (positive) tail; negative values indicate a left (negative) tail.

$$g_1 = \frac{\dfrac{1}{n} \displaystyle\sum_{i=1}^{n} (x_i – \overline{x})^3}{\left[\dfrac{1}{n} \displaystyle\sum_{i=1}^{n} (x_i – \overline{x})^2 \right]^{3/2}}$$


Kurtosis is the fourth standardised central moment, measuring tail heaviness relative to a normal distribution. Excess kurtosis (kurtosis $-$ 3) is commonly reported: positive = heavy tails (leptokurtic), negative = light tails (platykurtic).

Kurtosis:

$$\text{Kurt} = \frac{\dfrac{1}{n} \displaystyle\sum_{i=1}^{n} (x_i – \overline{x})^4}{\left[\dfrac{1}{n} \displaystyle\sum_{i=1}^{n} (x_i – \overline{x})^2 \right]^{2}}$$

Excess Kurtosis:

$$\gamma_2 = \text{Kurt} – 3$$

The normal distribution has kurtosis $= 3$ and excess kurtosis $= 0$.


Covariance is a measure of the joint variability of two random variables. It quantifies the degree to which two variables change together — that is, whether an increase in one variable tends to be associated with an increase or decrease in the other. A positive covariance indicates that the variables tend to move in the same direction; a negative covariance indicates they tend to move in opposite directions; a value of zero suggests no linear relationship.

Sample Covariance:

$$\text{Cov}(X, Y) = s_{xy} = \frac{1}{n-1} \sum_{i=1}^{n} (x_i – \overline{x})(y_i – \overline{y})$$

Population Covariance:

$$\text{Cov}(X, Y) = \sigma_{xy} = \frac{1}{N} \sum_{i=1}^{N} (x_i – \mu_x)(y_i – \mu_y)$$

The denominator $n-1$ makes the sample covariance an unbiased estimator of the population covariance.

Key Properties:

  • $\text{Cov}(X, X) = \text{Var}(X)$
  • $\text{Cov}(X, Y) = \text{Cov}(Y, X)$ — symmetric
  • Units are the product of the units of $X$ and $Y$, making it hard to interpret across different scales.

The Pearson correlation coefficient is a normalised measure of the linear association between two continuous variables. It is defined as the covariance of the two variables divided by the product of their standard deviations, scaling the result to a dimensionless value in the range $[-1, +1]$.

A value of $+1$ indicates a perfect positive linear relationship, $-1$ a perfect negative linear relationship, and $0$ indicates no linear relationship.

Sample Pearson Correlation:

$$r_{xy} = \frac{\displaystyle\sum_{i=1}^{n}(x_i – \overline{x})(y_i – \overline{y})}{\sqrt{\displaystyle\sum_{i=1}^{n}(x_i – \overline{x})^2} \cdot \sqrt{\displaystyle\sum_{i=1}^{n}(y_i – \overline{y})^2}}$$

Equivalently expressed using covariance:

$$r_{xy} = \frac{s_{xy}}{s_x \cdot s_y}$$

Population Pearson Correlation:

$$\rho_{xy} = \frac{\sigma_{xy}}{\sigma_x \cdot \sigma_y}$$

 $r$ is the sample estimator; $\rho$ (rho) denotes the population parameter.

Interpretation Guide:

  • $0.00 – 0.19$ $\rightarrow$ Negligible
  • $0.20 – 0.39$ $\rightarrow$ Weak
  • $0.40 – 0.59$ $\rightarrow$ Moderate
  • $0.60 – 0.79$ $\rightarrow$ Strong
  • $0.80 – 1.00$ $\rightarrow$ Very Strong

Measures only linear association. A value near 0 does not rule out a non-linear relationship.


The Spearman rank correlation coefficient is a non-parametric measure of monotonic association between two variables. It is computed by applying the Pearson formula to the ranked values of the observations rather than the raw values, making it robust to outliers and applicable to ordinal data.

Let $R(x_i)$ and $R(y_i)$ denote the ranks of $x_i$ and $y_i$ respectively, and $d_i = R(x_i) – R(y_i)$ be the rank difference.

Spearman Correlation (simplified – no tied ranks):

$$r_s = 1 – \frac{6 \displaystyle\sum_{i=1}^{n} d_i^2}{n(n^2 – 1)}$$

General Form (via Pearson on ranks):

$$r_s = \frac{\displaystyle\sum_{i=1}^{n}(R(x_i) – \overline{R}_x)(R(y_i) – \overline{R}_y)}{\sqrt{\displaystyle\sum_{i=1}^{n}(R(x_i) – \overline{R}_x)^2} \cdot \sqrt{\displaystyle\sum_{i=1}^{n}(R(y_i) – \overline{R}_y)^2}}$$

The simplified formula with $d_i^2$ is exact only when there are no tied ranks. Use the general form when ties are present.


The absolute frequency of a category $c_j$ is the number of times it appears in the dataset. It is the fundamental building block of all categorical summary statistics.

$$f_j = \sum_{i=1}^{n} \mathbf{1}[x_i = c_j], \quad j = 1, 2, \ldots, k$$

$\mathbf{1}[\cdot]$ is the indicator function, equal to 1 when the condition is true and 0 otherwise.


The relative frequency of a category is the ratio of its absolute frequency to the total number of observations. It estimates the probability of that category in the underlying population.

$$p_j = \frac{f_j}{n}, \qquad \text{subject to} \quad \sum_{j=1}^{k} p_j = 1$$


Percentage frequency scales the relative frequency to a base of 100, expressing each category’s share of the whole.

$$\%_j = \frac{f_j}{n} \times 100$$


Cumulative frequency is the running total of frequencies up to and including a given category, meaningful when categories have a natural order (e.g., ordinal data).

Absolute Cumulative Frequency:

$$F_j = \sum_{l=1}^{j} f_l$$

Cumulative Relative Frequency:

$$P_j = \frac{F_j}{n} = \sum_{l=1}^{j} p_l$$


Scroll to Top