Statistical Inference Preliminaries

This notes provides few ideas of sampling distribution of sample mean and variance

  • Theoretical Distributions
  • Target Population
  • Sampled Population
  • Random Sample
  • Statistics
  • Order statistics
  • Sampling Distribution

Let $X_1,\cdots,X_n$ be a random sample from a density $f(X|\theta$), then some statistics are

  $\overline {X} = \frac{1}{n}\sum X_i$

  $M’_r = \frac{1}{n}\sum X^r_i$  

  $M_r  = \frac{1}{n}\sum(X_i-\overline {X})^r$

  ${S^2}  = \frac{1}{(n-1)}\sum(X_i-\overline {X})^2$

For a random variable X, consider population $r^{th}$ moments

1. Non central moments: $\mu’_r = E[X^r]$

2. Central moments:$\mu_r = E[(X-\mu)^r]$

In particular, the Mean of X is $\mu = \mu_1′ = E(X)$

Now, Consider $M’_r$

  $E[M’_r] = E\Big[\frac{1}{n}\sum_{i=1}^n {X_i}^r\Big]$

  $=\frac{1}{n}\sum E({X_i}^r)$

  $=\frac{1}{n}\sum {\mu’_r}$

  $\Rightarrow E[M’_r] = {\mu’_r}$

Also $V[M’_r] = V \Big[\frac{1}{n}\sum {X_i}^r\Big]$ where V denotes the variance

   $= \frac{1}{n^2}\sum V({X_i}^r)$

   $= \frac{1}{n}\Big[E({X_i}^{2r}) – E({X_i^r})^2\Big]$

   $V\Big[{M’_r}\Big] = \frac{1}{n}\Big[{\mu’_{2r}}-{\mu’_r}^2\Big]$

  In Particular if r = 1,

1. $E[M’_1] = E[\overline {X}] = \mu’_1 = \mu$

2. $V[M’_1] = V[\overline {X}]$

   $= \frac{1}{n}[\mu’_2 – ({\mu’_1})^2]$

  $= \frac{1}{n}{\sigma}^2$

This holds for any $f(X~|~\theta)$ with $\mu = E(X)$ and $\sigma^2 = V(X)$

Regarding Sample Variance:

${S^2}  = \frac{1}{n-1}\sum(X_i-\overline {X})^2$

$E[S^2] = \frac{1}{n}\sum_{i=1}^n E{({X_i}-\overline {X})}^2$

Now, $\sum{({X_i}-\mu)}^2 = \sum{({X_i}-\bar{X}+\overline {X}-\mu)}^2$

 $= \sum[{({X_i}-\overline {X})}^2 + {(\overline {X}-\mu)}^2 + 2({X_i}-\overline {X})(\overline {X}-\mu)]$

 $= \sum{({X_i}-\overline {X})}^2 + n{(\overline {X}-\mu)}^2 + 2(\overline {X}-\mu)\sum({X_i} -\overline {X})$

Since,$2(\overline {X}-\mu)\sum({X_i}-\overline {X}) = 0$

$\sum {({X_i}-\mu)}^2$

$= \sum{({X_i}-\overline {X})}^2 + n{(\overline {X}-\mu)}^2$

So

$E[S^2] = \frac{1}{n-1} E\Big[\sum {({X_i}-\mu)}^2  – n{(\overline {X}-\mu)}^2\Big]$

$= \frac{1}{n-1} [\sum E{({X_i}-\mu)}^2  – nE{(\overline {X}-\mu)}^2 ]$

$= \frac{1}{n-1} \Big[\sum {\sigma}^2  – nV(\overline {X})\Big]$

$= \frac{1}{n-1} [n{\sigma}^2  – n \frac{{\sigma}^2}{n}]$

$= {\sigma}^2$

Hence for $X_1, X_2, \cdots, X_n \sim f(X~|~\theta)$  then

  $E[S^2] = {\sigma}^2$

Also, $V[S^2] =  \frac{1}{n}\Big[\mu_4 – \frac{n-3}{n-1}\mu_2^2\Big]$

Let $X_1, X_2,\cdots, X_n \sim \text{Bern}~(\theta)$

  Using the properties of sample mean,

  $E[\overline {X}]=\mu = \theta$

  $V[\overline {X}]=\frac{\sigma^2}{n} = \frac{\theta(1-\theta)}{n}$

Let $X_1, X_2, \cdots, X_n \sim \text{Poisson}~(\theta)$

  Using the properties of sample mean,

  $E[\overline {X}]=\mu = \theta$

  $V[\overline {X}]=\frac{\sigma^2}{n} = \frac{\theta}{n}$

Let $X_1, X_2, \cdots, X_n \sim \text{Expo} (\theta)$

Here, $\theta$ is the rate parameter(=$\frac{1}{scale}$)

This example illustrates the distribution of sample mean.

  $\sum X_i \sim \text{Gamma}~(n,\theta)$

PDF is $\frac{{\theta}^n}{\sqrt{n}}z^{n-1} e^{-\theta z}$

where  $Z = \sum {X_i}>0$

$\Rightarrow p\Big[\sum X_i\leq y\Big]=\int_0^y ~\frac{\theta^n}{\sqrt{n}} z^{n-1} e^{-\theta z}~dz$

$p[\overline {X} \leq \frac{y}{n}] = \int_0^y\frac{{\theta}^n}{\sqrt{n}}z^{n-1} e^{-\theta z} ~dz$

Now, $x =\frac{y}{n} \Rightarrow  y = nx$

$p[\overline {X} \leq {x}] = \int_0^{nx}\frac{{\theta}^n}{\sqrt{n}}z^{n-1} e^{-\theta z}~dz$

Let $u =\frac{\sum{x_i}}{n}=\frac{z}{n}$

When z = 0, then u = 0; when z = nx then u = x.  Also z = un implies $dz = n du$

$p[\overline {X} \leq {x}] = \int_0^{x}\frac{{\theta}^n}{\sqrt{n}} {(un)}^{n-1}e^{-n~\theta~u}~n~du$

$= \int_0^{x}\frac{{(n\theta)}^n}{\sqrt{n}} {u}^{n-1}e^{-n\theta u}~du$

Therefore $\overline {X}$, sample mean follows Gamma(n,n$\theta$). Hence, we have the following results from the summaries of a Gamma distribution with rate parameter $\theta$

  1. Mean of $\overline {X}$ is $E\Big[\overline {X}\Big]=\frac{n}{n\theta}=\frac{1}{\theta}$
  2. Variance of $\overline {X}$ is $V\Big[\overline {X}\Big]=\frac{n}{n^2\theta^2}=\frac{1}{n\theta^2}$

Further one can easily understand that these results are the same when it is obtained from the properties of sample mean, when $X_1, X_2, \cdots, X_n \sim \text{Expo} (\theta)$.

Results regarding Normal distribution/Random sample from Normal population

Let $X_1, X_2, \cdots, X_n \sim N({\mu},~{\sigma}^2)$ then

  1. $\frac{\overline X-\mu}{\sigma} \sim N(0,\frac{1}{n})$
  2. $\overline X$ and $\sum(X_i-\overline X)^2$ are mutually independent
  3. $\frac{(n-1)s^2}{\sigma^2} \sim {\chi^2_{(n-1)}}$

These results exemplify the way sampling distributions are defined for a statistic

Scroll to Top