Approach B: Via the Full Joint Likelihood
Recall from Approach A
The Normal-Normal model is treated via two approaches — through the sufficient statistic $\bar{y}$ and through the full joint likelihood — both yielding identical posterior inference. Throughout, the role of prior hyperparameters, conjugacy, and the influence of sample size on posterior updating are made explicit.
Assumptions
- We observe $n$ independent data points $y_1, y_2, \dots, y_n$
- Each $y_i \sim N(\mu, \sigma^2)$ where $\sigma^2$ is known and fixed
- The unknown parameter of interest is $\mu$ alone
- The prior on $\mu$ is Normal: $\mu \sim N(\mu_0, \tau^2)$ where $\mu_0$ and $\tau^2$ are known hyperparameters
This approach works directly with the full joint likelihood of all $n$ observations and must complete the square across all $n$ terms — yet arrives at exactly the same posterior.
Step 1. The Likelihood
$$y_i \mid \mu \sim N(\mu, \sigma^2), \quad i = 1, \dots, n, \quad \text{independently}$$
The joint likelihood is:
$$\mathscr{L}(\mu \mid \mathbf{y}) = \prod_{i=1}^n \frac{1}{\sqrt{2\pi\sigma^2}} \exp\left(-\frac{(y_i – \mu)^2}{2\sigma^2}\right)$$
$$= \left(\frac{1}{\sqrt{2\pi\sigma^2}}\right)^n \exp\left(-\frac{1}{2\sigma^2}\sum_{i=1}^n(y_i-\mu)^2\right)$$
$$= \frac{1}{(2\pi\sigma^2)^{n/2}} \exp\left(-\frac{1}{2\sigma^2}\sum_{i=1}^n(y_i-\mu)^2\right)$$
Step 2. The Prior
$$p(\mu) = \frac{1}{\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_0)^2}{2\tau^2}\right)$$
Step 3. Bayes’ Theorem from First Principles
Identical to Approach B:
$$p(\mu \mid \mathbf{y}) = \frac{p(\mathbf{y} \mid \mu)\cdot p(\mu)}{p(\mathbf{y})}$$
where:
$$p(\mathbf{y}) = \int_{-\infty}^{\infty} p(\mathbf{y} \mid \mu)\, p(\mu)\, d\mu$$
Step 4. The Marginal Likelihood — Full Derivation
$$p(\mathbf{y}) = \int_{-\infty}^{\infty} \frac{1}{(2\pi\sigma^2)^{n/2}} \exp\left(-\frac{1}{2\sigma^2}\sum_{i=1}^n(y_i-\mu)^2\right) \cdot \frac{1}{\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_0)^2}{2\tau^2}\right) d\mu$$
Pulling constants out:
$$p(\mathbf{y}) = \frac{1}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}} \int_{-\infty}^{\infty} \exp\left(-\frac{\sum_{i=1}^n(y_i-\mu)^2}{2\sigma^2} – \frac{(\mu-\mu_0)^2}{2\tau^2}\right) d\mu$$
Expanding the sum in the exponent. First expand $\sum_{i=1}^n(y_i – \mu)^2$:
$$\sum_{i=1}^n(y_i-\mu)^2 = \sum_{i=1}^n y_i^2 – 2\mu\sum_{i=1}^n y_i + n\mu^2 = \sum_{i=1}^n y_i^2 – 2n\bar{y}\mu + n\mu^2$$
So the full exponent is:
$$Q = \frac{\sum_{i=1}^n y_i^2 – 2n\bar{y}\mu + n\mu^2}{2\sigma^2} + \frac{\mu^2 – 2\mu\mu_0 + \mu_0^2}{2\tau^2}$$
Collecting terms in $\mu^2$ and $\mu$:
$$Q = \mu^2\left(\frac{n}{2\sigma^2} + \frac{1}{2\tau^2}\right) – \mu\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right) + \left(\frac{\sum_{i=1}^n y_i^2}{2\sigma^2} + \frac{\mu_0^2}{2\tau^2}\right)$$
Define exactly as before:
$$\frac{1}{\tau_n^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2}$$
$$\mu_n = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$
Completing the square in $\mu$:
$$Q = \frac{(\mu – \mu_n)^2}{2\tau_n^2} + C^*$$
where $C^*$ depends only on the data and hyperparameters, not on $\mu$:
$$C^* = \frac{\sum_{i=1}^n y_i^2}{2\sigma^2} + \frac{\mu_0^2}{2\tau^2} – \frac{\mu_n^2}{2\tau_n^2}$$
Substituting back:
$$p(\mathbf{y}) = \frac{e^{-C^*}}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}} \int_{-\infty}^{\infty} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right) d\mu$$
Evaluating the Gaussian integral:
$$\int_{-\infty}^{\infty} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right) d\mu = \sqrt{2\pi\tau_n^2}$$
Therefore:
$$\boxed{p(\mathbf{y}) = \frac{\sqrt{2\pi\tau_n^2}}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}\, e^{-C^*}}$$
Step 5. Substituting into Bayes’ Theorem
Numerator:
$$p(\mathbf{y} \mid \mu)\cdot p(\mu) = \frac{1}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}} \exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2} – C^*\right)$$
Denominator:
$$p(\mathbf{y}) = \frac{\sqrt{2\pi\tau_n^2}}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}\, e^{-C^*}$$
Dividing:
$$p(\mu \mid \mathbf{y}) = \frac{\dfrac{1}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}\exp\left(-\dfrac{(\mu-\mu_n)^2}{2\tau_n^2} – C^*\right)}{\dfrac{\sqrt{2\pi\tau_n^2}}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}e^{-C^*}}$$
The term $\dfrac{1}{(2\pi\sigma^2)^{n/2}\sqrt{2\pi\tau^2}}$ cancels, and $e^{-C^*}$ cancels:
$$p(\mu \mid \mathbf{y}) = \frac{1}{\sqrt{2\pi\tau_n^2}}\exp\left(-\frac{(\mu-\mu_n)^2}{2\tau_n^2}\right)$$
$$\therefore \quad \mu \mid \mathbf{y} \sim N(\mu_n,\ \tau_n^2)$$
with:
$$\frac{1}{\tau_n^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2} \qquad \text{and} \qquad \mu_n = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$
Conclusion
Both approaches yield exactly the same posterior:
$$\boxed{\mu \mid \mathbf{y} \sim N(\mu_n,\ \tau_n^2)}$$
$$\frac{1}{\tau_n^2} = \frac{n}{\sigma^2} + \frac{1}{\tau^2} \qquad \mu_n = \tau_n^2\left(\frac{n\bar{y}}{\sigma^2} + \frac{\mu_0}{\tau^2}\right)$$
Both approaches lead to the fact that both derivations produce identical $\mu_n$ and $\tau_n^2$ confirms that no information about $\mu$ is lost when summarising the data through $\bar{y}$ — which is precisely what sufficiency means.
Next Normal-Gamma Model