These notes develop the Bayesian conjugate analysis for Normal-Normal — following a unified five-step approach:
- specifying the likelihood
- defining the prior
- deriving Bayes’ theorem from first principles
- computing the marginal likelihood in full
- substituting to obtain the exact posterior.
The Normal-Normal model is treated via two approaches — through the sufficient statistic $\bar{y}$ and through the full joint likelihood — both yielding identical posterior inference. Throughout, the role of prior hyperparameters, conjugacy, and the influence of sample size on posterior updating are made explicit.
Assumptions
- We observe $n$ independent data points $y_1, y_2, \dots, y_n$
- Each $y_i \sim N(\mu, \sigma^2)$ where $\sigma^2$ is known and fixed
- The unknown parameter of interest is $\mu$ alone
- The prior on $\mu$ is Normal: $\mu \sim N(\mu_0, \tau^2)$ where $\mu_0$ and $\tau^2$ are known hyperparameters
We present two approaches. Approach A exploits the sufficiency of $\bar{y}$ under the Normal model. Approach B works through the full joint likelihood of all $n$ observations directly. Both lead to identical posterior inference.
Approach A: Via the Sample Mean $\bar{y}$ as Sufficient Statistic
Step 1. The Likelihood
Since $y_1, \dots, y_n$ are independent and each $y_i \sim N(\mu, \sigma^2)$, the sample mean is:
$$\bar{y} = \frac{1}{n}\sum_{i=1}^n y_i$$
By the property that a sum of independent Normal variates is again Normal:
$$\sum_{i=1}^n y_i \sim N(n\mu,\ n\sigma^2)$$
and therefore:
$$\bar{y} \sim N\left(\mu,\ \frac{\sigma^2}{n}\right)$$
The sample mean $\bar{y}$ is a sufficient statistic for $\mu$ — it captures all information about $\mu$ contained in the data. We therefore work with $\bar{y}$ as a single effective observation. The likelihood is:
$$p(\bar{y} \mid \mu) = \frac{1}{\sqrt{2\pi\sigma^2/n}} \exp\left(-\frac{(\bar{y}-\mu)^2}{2\sigma^2/n}\right)$$
Step 2. The Prior
$$\mu \sim N(\mu_0,\ \tau^2)$$
$$p(\mu) = \frac{1}{\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_0)^2}{2\tau^2}\right)$$
Step 3. Bayes’ Theorem from First Principles
From the definition of conditional probability, the joint distribution factors two ways:
$$p(\mu,\ \bar{y}) = p(\bar{y} \mid \mu)\cdot p(\mu)$$
$$p(\mu,\ \bar{y}) = p(\mu \mid \bar{y})\cdot p(\bar{y})$$
Setting these equal and dividing both sides by $p(\bar{y})$:
$$\boxed{p(\mu \mid \bar{y}) = \frac{p(\bar{y} \mid \mu)\cdot p(\mu)}{p(\bar{y})}}$$
where:
$$p(\bar{y}) = \int_{-\infty}^{\infty} p(\bar{y} \mid \mu)\, p(\mu)\, d\mu$$
Step 4. The Marginal Likelihood — Full Derivation
Substituting:
$$p(\bar{y}) = \int_{-\infty}^{\infty} \frac{1}{\sqrt{2\pi\sigma^2/n}} \exp\left(-\frac{(\bar{y}-\mu)^2}{2\sigma^2/n}\right) \cdot \frac{1}{\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_0)^2}{2\tau^2}\right) d\mu$$
Pulling out constants:
$$p(\bar{y}) = \frac{1}{2\pi\sqrt{\sigma^2/n}\sqrt{\tau^2}} \int_{-\infty}^{\infty} \exp\left(-\frac{(\bar{y}-\mu)^2}{2\sigma^2/n} – \frac{(\mu-\mu_0)^2}{2\tau^2}\right) d\mu$$
Combining the exponent. Let us work on the expression in the exponent:
$$Q = \frac{(\bar{y}-\mu)^2}{2\sigma^2/n} + \frac{(\mu-\mu_0)^2}{2\tau^2}$$
Expanding each term:
$$Q = \frac{\bar{y}^2 – 2\bar{y}\mu + \mu^2}{2\sigma^2/n} + \frac{\mu^2 – 2\mu\mu_0 + \mu_0^2}{2\tau^2}$$
Collecting terms in $\mu^2$ and $\mu$:
$$Q = \mu^2\left(\frac{1}{2\sigma^2/n} + \frac{1}{2\tau^2}\right) – \mu\left(\frac{\bar{y}}{\sigma^2/n} + \frac{\mu_0}{\tau^2}\right) + \left(\frac{\bar{y}^2}{2\sigma^2/n} + \frac{\mu_0^2}{2\tau^2}\right)$$
Define:
$$\frac{1}{\tau_n^2} = \frac{1}{\sigma^2/n} + \frac{1}{\tau^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2}$$
and:
$$\mu_n = \tau_n^2\left(\frac{\bar{y}}{\sigma^2/n} + \frac{\mu_0}{\tau^2}\right) = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$
Then completing the square in $\mu$:
$$Q = \frac{1}{2\tau_n^2}(\mu – \mu_n)^2 + C$$
where $C$ is a term that depends only on $\bar{y}$, $\mu_0$, $\sigma^2$, $\tau^2$, $n$ — not on $\mu$. Explicitly:
$$C = \frac{\bar{y}^2}{2\sigma^2/n} + \frac{\mu_0^2}{2\tau^2} – \frac{\mu_n^2}{2\tau_n^2}$$
Substituting back into the integral:
$$p(\bar{y}) = \frac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\, e^{-C} \int_{-\infty}^{\infty} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right) d\mu$$
The remaining integral is a Gaussian integral:
$$\int_{-\infty}^{\infty} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right) d\mu = \sqrt{2\pi\tau_n^2}$$
Therefore:
$$\boxed{p(\bar{y}) = \frac{\sqrt{2\pi\tau_n^2}}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\, e^{-C} = \frac{1}{\sqrt{2\pi\left(\sigma^2/n + \tau^2\right)}}\exp\left(-\frac{(\bar{y}-\mu_0)^2}{2(\sigma^2/n+\tau^2)}\right)}$$
This is recognisable as $\bar{y} \sim N(\mu_0,\ \sigma^2/n + \tau^2)$ — the marginal distribution of $\bar{y}$ averaging over the prior on $\mu$.
Step 5. Substituting into Bayes’ Theorem
Numerator:
$$p(\bar{y} \mid \mu)\cdot p(\mu) = \frac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\exp\left(-\frac{(\bar{y}-\mu)^2}{2\sigma^2/n} – \frac{(\mu-\mu_0)^2}{2\tau^2}\right)$$
$$= \frac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2} – C\right)$$
Denominator:
$$p(\bar{y}) = \frac{1}{\sqrt{2\pi(\sigma^2/n+\tau^2)}}\exp\left(-\frac{(\bar{y}-\mu_0)^2}{2(\sigma^2/n+\tau^2)}\right)$$
Dividing:
$$p(\mu \mid \bar{y}) = \frac{\dfrac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}\exp\left(-\dfrac{(\mu-\mu_n)^2}{2\tau_n^2} – C\right)}{\dfrac{1}{\sqrt{2\pi(\sigma^2/n+\tau^2)}}\exp\left(-\dfrac{(\bar{y}-\mu_0)^2}{2(\sigma^2/n+\tau^2)}\right)}$$
The term $e^{-C}$ in the numerator and the denominator exponential cancel exactly — both encode the same function of $\bar{y}$ alone. The constants also simplify:
$$\frac{\dfrac{1}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}}}{\dfrac{1}{\sqrt{2\pi(\sigma^2/n+\tau^2)}}} = \frac{\sqrt{2\pi(\sigma^2/n+\tau^2)}}{2\pi\sqrt{\sigma^2/n}\,\sqrt{\tau^2}} = \frac{1}{\sqrt{2\pi\tau_n^2}}$$
Therefore:
$$\boxed{p(\mu \mid \bar{y}) = \frac{1}{\sqrt{2\pi\tau_n^2}}\exp\left(-\frac{(\mu – \mu_n)^2}{2\tau_n^2}\right)}$$
$$\therefore \quad \mu \mid \mathbf{y} \sim N(\mu_n,\ \tau_n^2)$$
where:
$$\frac{1}{\tau_n^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2} \qquad \text{and} \qquad \mu_n = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$
Note: Approach A exploits the fact that $\bar{y}$ is a sufficient statistic for $\mu$ under the Normal model. The entire dataset of $n$ observations is compressed into a single quantity $\bar{y} \sim N(\mu, \sigma^2/n)$ without any loss of information about $\mu$. This is what makes Approach A compact.
Next Approach B