Normal-Normal Conjugate Model – 2

Approach B: Via the Full Joint Likelihood

Recall from Approach A
The Normal-Normal model is treated via two approaches — through the sufficient statistic $\bar{y}$ and through the full joint likelihood — both yielding identical posterior inference. Throughout, the role of prior hyperparameters, conjugacy, and the influence of sample size on posterior updating are made explicit.

Assumptions

  • We observe $n$ independent data points $y_1, y_2, \dots, y_n$
  • Each $y_i \sim N(\mu, \sigma^2)$ where $\sigma^2$ is known and fixed
  • The unknown parameter of interest is $\mu$ alone
  • The prior on $\mu$ is Normal: $\mu \sim N(\mu_0, \tau^2)$ where $\mu_0$ and $\tau^2$ are known hyperparameters

This approach works directly with the full joint likelihood of all $n$ observations and must complete the square across all $n$ terms — yet arrives at exactly the same posterior.

Step 1. The Likelihood

$$y_i \mid \mu \sim N(\mu, \sigma^2), \quad i = 1, \dots, n, \quad \text{independently}$$

The joint likelihood is:

$$\mathscr{L}(\mu \mid \mathbf{y}) = \prod_{i=1}^n \frac{1}{\sqrt{2\pi\sigma^2}} \exp\left(-\frac{(y_i – \mu)^2}{2\sigma^2}\right)$$

$$= \left(\frac{1}{\sqrt{2\pi\sigma^2}}\right)^n \exp\left(-\frac{1}{2\sigma^2}\sum_{i=1}^n(y_i-\mu)^2\right)$$

$$= \frac{1}{(2\pi\sigma^2)^{n/2}} \exp\left(-\frac{1}{2\sigma^2}\sum_{i=1}^n(y_i-\mu)^2\right)$$

Step 2. The Prior

$$p(\mu) = \frac{1}{\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_0)^2}{2\tau^2}\right)$$

Step 3. Bayes’ Theorem from First Principles

Identical to Approach B:

$$p(\mu \mid \mathbf{y}) = \frac{p(\mathbf{y} \mid \mu)\cdot p(\mu)}{p(\mathbf{y})}$$

where:

$$p(\mathbf{y}) = \int_{-\infty}^{\infty} p(\mathbf{y} \mid \mu)\, p(\mu)\, d\mu$$

Step 4. The Marginal Likelihood — Full Derivation

$$p(\mathbf{y}) = \int_{-\infty}^{\infty} \frac{1}{(2\pi\sigma^2)^{n/2}} \exp\left(-\frac{1}{2\sigma^2}\sum_{i=1}^n(y_i-\mu)^2\right) \cdot \frac{1}{\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_0)^2}{2\tau^2}\right) d\mu$$

Pulling constants out:

$$p(\mathbf{y}) = \frac{1}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}} \int_{-\infty}^{\infty} \exp\left(-\frac{\sum_{i=1}^n(y_i-\mu)^2}{2\sigma^2} – \frac{(\mu-\mu_0)^2}{2\tau^2}\right) d\mu$$

Expanding the sum in the exponent. First expand $\sum_{i=1}^n(y_i – \mu)^2$:

$$\sum_{i=1}^n(y_i-\mu)^2 = \sum_{i=1}^n y_i^2 – 2\mu\sum_{i=1}^n y_i + n\mu^2 = \sum_{i=1}^n y_i^2 – 2n\bar{y}\mu + n\mu^2$$

So the full exponent is:

$$Q = \frac{\sum_{i=1}^n y_i^2 – 2n\bar{y}\mu + n\mu^2}{2\sigma^2} + \frac{\mu^2 – 2\mu\mu_0 + \mu_0^2}{2\tau^2}$$

Collecting terms in $\mu^2$ and $\mu$:

$$Q = \mu^2\left(\frac{n}{2\sigma^2} + \frac{1}{2\tau^2}\right) – \mu\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right) + \left(\frac{\sum_{i=1}^n y_i^2}{2\sigma^2} + \frac{\mu_0^2}{2\tau^2}\right)$$

Define exactly as before:

$$\frac{1}{\tau_n^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2}$$

$$\mu_n = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$

Completing the square in $\mu$:

$$Q = \frac{(\mu – \mu_n)^2}{2\tau_n^2} + C^*$$

where $C^*$ depends only on the data and hyperparameters, not on $\mu$:

$$C^* = \frac{\sum_{i=1}^n y_i^2}{2\sigma^2} + \frac{\mu_0^2}{2\tau^2} – \frac{\mu_n^2}{2\tau_n^2}$$

Substituting back:

$$p(\mathbf{y}) = \frac{e^{-C^*}}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}} \int_{-\infty}^{\infty} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right) d\mu$$

Evaluating the Gaussian integral:

$$\int_{-\infty}^{\infty} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right) d\mu = \sqrt{2\pi\tau_n^2}$$

Therefore:

$$\boxed{p(\mathbf{y}) = \frac{\sqrt{2\pi\tau_n^2}}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}\, e^{-C^*}}$$

Step 5. Substituting into Bayes’ Theorem

Numerator:

$$p(\mathbf{y} \mid \mu)\cdot p(\mu) = \frac{1}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2} – C^*\right)$$

Denominator:

$$p(\mathbf{y}) = \frac{\sqrt{2\pi\tau_n^2}}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}\, e^{-C^*}$$

Dividing:

$$p(\mu \mid \mathbf{y}) = \frac{\dfrac{1}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}\exp\left(-\dfrac{(\mu-\mu_n)^2}{2\tau_n^2} – C^*\right)}{\dfrac{\sqrt{2\pi\tau_n^2}}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}e^{-C^*}}$$

The term $\dfrac{1}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}$ cancels, and $e^{-C^*}$ cancels:

$$p(\mu \mid \mathbf{y}) = \frac{1}{\sqrt{2\pi\tau_n^2}}\exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right)$$

$$\therefore \quad \mu \mid \mathbf{y} \sim N(\mu_n,\ \tau_n^2)$$

with:

$$\frac{1}{\tau_n^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2} \qquad \text{and} \qquad \mu_n = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$

Conclusion

Both approaches yield exactly the same posterior:

$$\boxed{\mu \mid \mathbf{y} \sim N(\mu_n,\ \tau_n^2)}$$

$$\frac{1}{\tau_n^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2} \qquad \mu_n = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$

Both approaches lead to the fact that both derivations produce identical $\mu_n$ and $\tau_n^2$ confirms that no information about $\mu$ is lost when summarising the data through $\bar{y}$ — which is precisely what sufficiency means.

Next Normal-Gamma Model

Scroll to Top