Normal-Normal Conjugate Model – 1

These notes develop the Bayesian conjugate analysis for Normal-Normal — following a unified five-step approach:

  1. specifying the likelihood
  2. defining the prior
  3. deriving Bayes’ theorem from first principles
  4. computing the marginal likelihood in full
  5. substituting to obtain the exact posterior.

The Normal-Normal model is treated via two approaches — through the sufficient statistic $\bar{y}$ and through the full joint likelihood — both yielding identical posterior inference. Throughout, the role of prior hyperparameters, conjugacy, and the influence of sample size on posterior updating are made explicit.

Assumptions

  • We observe $n$ independent data points $y_1, y_2, \dots, y_n$
  • Each $y_i \sim N(\mu, \sigma^2)$ where $\sigma^2$ is known and fixed
  • The unknown parameter of interest is $\mu$ alone
  • The prior on $\mu$ is Normal: $\mu \sim N(\mu_0, \tau^2)$ where $\mu_0$ and $\tau^2$ are known hyperparameters

We present two approaches. Approach A exploits the sufficiency of $\bar{y}$ under the Normal model. Approach B works through the full joint likelihood of all $n$ observations directly. Both lead to identical posterior inference.

Approach A: Via the Sample Mean $\bar{y}$ as Sufficient Statistic

Step 1. The Likelihood

Since $y_1, \dots, y_n$ are independent and each $y_i \sim N(\mu, \sigma^2)$, the sample mean is:

$$\bar{y} = \frac{1}{n}\sum_{i=1}^n y_i$$

By the property that a sum of independent Normal variates is again Normal:

$$\sum_{i=1}^n y_i \sim N(n\mu,\ n\sigma^2)$$

and therefore:

$$\bar{y} \sim N\left(\mu,\ \frac{\sigma^2}{n}\right)$$

The sample mean $\bar{y}$ is a sufficient statistic for $\mu$ — it captures all information about $\mu$ contained in the data. We therefore work with $\bar{y}$ as a single effective observation. The likelihood is:

$$p(\bar{y} \mid \mu) = \frac{1}{\sqrt{2\pi\sigma^2/n}} \exp\left(-\frac{(\bar{y}-\mu)^2}{2\sigma^2/n}\right)$$

Step 2. The Prior

$$\mu \sim N(\mu_0,\ \tau^2)$$

$$p(\mu) = \frac{1}{\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_0)^2}{2\tau^2}\right)$$

Step 3. Bayes’ Theorem from First Principles

From the definition of conditional probability, the joint distribution factors two ways:

$$p(\mu,\ \bar{y}) = p(\bar{y} \mid \mu)\cdot p(\mu)$$

$$p(\mu,\ \bar{y}) = p(\mu \mid \bar{y})\cdot p(\bar{y})$$

Setting these equal and dividing both sides by $p(\bar{y})$:

$$\boxed{p(\mu \mid \bar{y}) = \frac{p(\bar{y} \mid \mu)\cdot p(\mu)}{p(\bar{y})}}$$

where:

$$p(\bar{y}) = \int_{-\infty}^{\infty} p(\bar{y} \mid \mu)\, p(\mu)\, d\mu$$

Step 4. The Marginal Likelihood — Full Derivation

Substituting:

$$p(\bar{y}) = \int_{-\infty}^{\infty} \frac{1}{\sqrt{2\pi\sigma^2/n}} \exp\left(-\frac{(\bar{y}-\mu)^2}{2\sigma^2/n}\right) \cdot \frac{1}{\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_0)^2}{2\tau^2}\right) d\mu$$

Pulling out constants:

$$p(\bar{y}) = \frac{1}{2\pi\sqrt{\sigma^2/n}\sqrt{\tau^2}} \int_{-\infty}^{\infty} \exp\left(-\frac{(\bar{y}-\mu)^2}{2\sigma^2/n} – \frac{(\mu-\mu_0)^2}{2\tau^2}\right) d\mu$$

Combining the exponent. Let us work on the expression in the exponent:

$$Q = \frac{(\bar{y}-\mu)^2}{2\sigma^2/n} + \frac{(\mu-\mu_0)^2}{2\tau^2}$$

Expanding each term:

$$Q = \frac{\bar{y}^2 – 2\bar{y}\mu + \mu^2}{2\sigma^2/n} + \frac{\mu^2 – 2\mu\mu_0 + \mu_0^2}{2\tau^2}$$

Collecting terms in $\mu^2$ and $\mu$:

$$Q = \mu^2\left(\frac{1}{2\sigma^2/n} + \frac{1}{2\tau^2}\right) – \mu\left(\frac{\bar{y}}{\sigma^2/n} + \frac{\mu_0}{\tau^2}\right) + \left(\frac{\bar{y}^2}{2\sigma^2/n} + \frac{\mu_0^2}{2\tau^2}\right)$$

Define:

$$\frac{1}{\tau_n^2} = \frac{1}{\sigma^2/n} + \frac{1}{\tau^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2}$$

and:

$$\mu_n = \tau_n^2\left(\frac{\bar{y}}{\sigma^2/n} + \frac{\mu_0}{\tau^2}\right) = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$

Then completing the square in $\mu$:

$$Q = \frac{1}{2\tau_n^2}(\mu – \mu_n)^2 + C$$

where $C$ is a term that depends only on $\bar{y}$, $\mu_0$, $\sigma^2$, $\tau^2$, $n$ — not on $\mu$. Explicitly:

$$C = \frac{\bar{y}^2}{2\sigma^2/n} + \frac{\mu_0^2}{2\tau^2} – \frac{\mu_n^2}{2\tau_n^2}$$

Substituting back into the integral:

$$p(\bar{y}) = \frac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\, e^{-C} \int_{-\infty}^{\infty} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right) d\mu$$

The remaining integral is a Gaussian integral:

$$\int_{-\infty}^{\infty} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right) d\mu = \sqrt{2\pi\tau_n^2}$$

Therefore:

$$\boxed{p(\bar{y}) = \frac{\sqrt{2\pi\tau_n^2}}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\, e^{-C} = \frac{1}{\sqrt{2\pi\left(\sigma^2/n + \tau^2\right)}}\exp\left(-\frac{(\bar{y}-\mu_0)^2}{2(\sigma^2/n+\tau^2)}\right)}$$

This is recognisable as $\bar{y} \sim N(\mu_0,\ \sigma^2/n + \tau^2)$ — the marginal distribution of $\bar{y}$ averaging over the prior on $\mu$.

Step 5. Substituting into Bayes’ Theorem

Numerator:

$$p(\bar{y} \mid \mu)\cdot p(\mu) = \frac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\exp\left(-\frac{(\bar{y}-\mu)^2}{2\sigma^2/n} – \frac{(\mu-\mu_0)^2}{2\tau^2}\right)$$

$$= \frac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2} – C\right)$$

Denominator:

$$p(\bar{y}) = \frac{1}{\sqrt{2\pi(\sigma^2/n+\tau^2)}}\exp\left(-\frac{(\bar{y}-\mu_0)^2}{2(\sigma^2/n+\tau^2)}\right)$$

Dividing:

$$p(\mu \mid \bar{y}) = \frac{\dfrac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\exp\left(-\dfrac{(\mu-\mu_n)^2}{2\tau_n^2} – C\right)}{\dfrac{1}{\sqrt{2\pi(\sigma^2/n+\tau^2)}}\exp\left(-\dfrac{(\bar{y}-\mu_0)^2}{2(\sigma^2/n+\tau^2)}\right)}$$

The term $e^{-C}$ in the numerator and the denominator exponential cancel exactly — both encode the same function of $\bar{y}$ alone. The constants also simplify:

$$\frac{\dfrac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}}{\dfrac{1}{\sqrt{2\pi(\sigma^2/n+\tau^2)}}} = \frac{\sqrt{2\pi(\sigma^2/n+\tau^2)}}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}} = \frac{1}{\sqrt{2\pi\tau_n^2}}$$

Therefore:

$$\boxed{p(\mu \mid \bar{y}) = \frac{1}{\sqrt{2\pi\tau_n^2}}\exp\left(-\frac{(\mu – \mu_n)^2}{2\tau_n^2}\right)}$$

$$\therefore \quad \mu \mid \mathbf{y} \sim N(\mu_n,\ \tau_n^2)$$

where:

$$\frac{1}{\tau_n^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2} \qquad \text{and} \qquad \mu_n = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$

Note: Approach A exploits the fact that $\bar{y}$ is a sufficient statistic for $\mu$ under the Normal model. The entire dataset of $n$ observations is compressed into a single quantity $\bar{y} \sim N(\mu, \sigma^2/n)$ without any loss of information about $\mu$. This is what makes Approach A compact.

Next Approach B

Scroll to Top