Normal Model: Both $\mu$ and $\sigma^2$ are unknown
In the one-parameter models, only $\mu$ or only $\sigma^2$ is unknown, so the prior is placed on that single parameter alone and Bayes’ rule gives a direct closed-form update.
Recall One parameter model
Once both $\mu$ and $\sigma^2$ are unknown together, we need a genuine joint prior $p(\mu,\sigma^2)$ on the pair, not two separate one-parameter priors — this is the NIG model, where the joint prior is built as $p(\mu\mid\sigma^2),p(\sigma^2)$ so that Bayes’ rule can be applied jointly and still yield a closed-form joint posterior.
Likelihood
Observe $y_1, y_2, \dots, y_n$ iid, each $y_i \mid \mu, \sigma^2 \sim N(\mu, \sigma^2)$.
Joint likelihood:
$$\mathscr{L}(\mu, \sigma^2 \mid \mathbf{y}) = (2\pi\sigma^2)^{-n/2} \exp\left(-\frac{1}{2\sigma^2} \sum_{i=1}^n (y_i – \mu)^2\right)$$
Using the identity $\sum_{i=1}^n (y_i – \mu)^2 = nS^2 + n(\bar{y} – \mu)^2$
where $S^2 = \frac{1}{n}\sum_{i=1}^n(y_i – \bar{y})^2$:
$$\mathscr{L}(\mu, \sigma^2 \mid \mathbf{y}) = (2\pi\sigma^2)^{-n/2} \exp\left(-\frac{nS^2 + n(\bar{y}-\mu)^2}{2\sigma^2}\right)$$
The sufficient statistics are $(\bar{y}, S^2)$.
Joint Prior: Normal-Inverse Gamma
The conjugate joint prior on $(\mu, \sigma^2)$ is specified hierarchically:
$$\mu \mid \sigma^2 \sim N\left(\mu_0,\ \frac{\sigma^2}{\kappa_0}\right), \qquad \sigma^2 \sim \text{IG}(\alpha_0, \beta_0)$$
Written explicitly as $p(\mu, \sigma^2) = p(\mu \mid \sigma^2)\, p(\sigma^2)$:
$$p(\mu \mid \sigma^2) = \left(\frac{\kappa_0}{2\pi\sigma^2}\right)^{1/2} \exp\left(-\frac{\kappa_0(\mu – \mu_0)^2}{2\sigma^2}\right)$$
$$p(\sigma^2) = \frac{\beta_0^{\alpha_0}}{\Gamma(\alpha_0)}\,(\sigma^2)^{-(\alpha_0+1)}\,e^{-\beta_0/\sigma^2}$$
Joint prior kernel:
$$p(\mu, \sigma^2) \propto (\sigma^2)^{-(\alpha_0 + 3/2)} \exp\left(-\frac{1}{\sigma^2}\left[\beta_0 + \frac{\kappa_0(\mu-\mu_0)^2}{2}\right]\right)$$
Hyperparameters: $\mu_0$ (prior mean of $\mu$), $\kappa_0$ (prior sample size),
$\alpha_0$ (prior shape), $\beta_0$ (prior scale).
Bayes’ Rule
$$p(\mu, \sigma^2 \mid \mathbf{y}) \propto \mathscr{L}(\mu, \sigma^2 \mid \mathbf{y})\cdot p(\mu, \sigma^2)$$
$$p(\mathbf{y}) = \int_0^\infty \int_{-\infty}^\infty \mathscr{L}(\mu, \sigma^2 \mid \mathbf{y})\, p(\mu, \sigma^2)\, d\mu\, d\sigma^2$$
We state $p(\mathbf{y})$ exists in closed form but do not derive it;
the posterior follows by identifying the kernel.
Combining Likelihood and Prior
The posterior kernel is:
$$p(\mu, \sigma^2 \mid \mathbf{y}) \propto (\sigma^2)^{-n/2}\exp\left(-\frac{nS^2 + n(\bar{y}-\mu)^2}{2\sigma^2}\right)\cdot(\sigma^2)^{-(\alpha_0+3/2)}\exp\left(-\frac{\kappa_0(\mu-\mu_0)^2}{2\sigma^2} – \frac{\beta_0}{\sigma^2}\right)$$
Step 1: Collect powers of $\sigma^2$
$$(\sigma^2)^{-n/2} \cdot (\sigma^2)^{-(\alpha_0+3/2)} = (\sigma^2)^{-(n/2 + \alpha_0 + 3/2)}$$
Step 2: Collect all exponent terms
$$\text{Exponent} = -\frac{1}{2\sigma^2}\left[nS^2 + n(\bar{y}-\mu)^2 + \kappa_0(\mu-\mu_0)^2 + 2\beta_0\right]$$
Step 3: Expand the quadratic in $\mu$
$$n(\bar{y}-\mu)^2 + \kappa_0(\mu-\mu_0)^2$$
$$= n\bar{y}^2 – 2n\bar{y}\mu + n\mu^2 + \kappa_0\mu_0^2 – 2\kappa_0\mu_0\mu + \kappa_0\mu^2$$
$$= (\kappa_0 + n)\mu^2 – 2(\kappa_0\mu_0 + n\bar{y})\mu + (n\bar{y}^2 + \kappa_0\mu_0^2)$$
Step 4: Complete the square in $\mu$
Let $\kappa_n = \kappa_0 + n$ and $\mu_n = \dfrac{\kappa_0\mu_0 + n\bar{y}}{\kappa_n}$.
$$(\kappa_0+n)\mu^2 – 2(\kappa_0\mu_0+n\bar{y})\mu + (n\bar{y}^2+\kappa_0\mu_0^2)$$
$$= \kappa_n\left(\mu – \mu_n\right)^2 – \kappa_n\mu_n^2 + (n\bar{y}^2 + \kappa_0\mu_0^2)$$
The residual constant is:
$$n\bar{y}^2 + \kappa_0\mu_0^2 – \kappa_n\mu_n^2= n\bar{y}^2 + \kappa_0\mu_0^2 – \frac{(\kappa_0\mu_0 + n\bar{y})^2}{\kappa_n}$$
Expanding the numerator of the last term:
$$(\kappa_0\mu_0 + n\bar{y})^2 = \kappa_0^2\mu_0^2 + 2\kappa_0 n\mu_0\bar{y} + n^2\bar{y}^2$$
So the residual becomes:
$$\frac{\kappa_n(n\bar{y}^2 + \kappa_0\mu_0^2) – \kappa_0^2\mu_0^2 – 2\kappa_0 n\mu_0\bar{y} – n^2\bar{y}^2}{\kappa_n}$$
Numerator:
$$(\kappa_0+n)(n\bar{y}^2+\kappa_0\mu_0^2) – \kappa_0^2\mu_0^2 – 2\kappa_0 n\mu_0\bar{y} – n^2\bar{y}^2$$
$$= \kappa_0 n\bar{y}^2 + n^2\bar{y}^2 + \kappa_0^2\mu_0^2 + \kappa_0 n\mu_0^2- \kappa_0^2\mu_0^2 – 2\kappa_0 n\mu_0\bar{y} – n^2\bar{y}^2$$
$$= \kappa_0 n\bar{y}^2 – 2\kappa_0 n\mu_0\bar{y} + \kappa_0 n\mu_0^2$$
$$= \kappa_0 n(\bar{y} – \mu_0)^2$$
Therefore the residual constant is $\dfrac{\kappa_0 n}{\kappa_n}(\bar{y}-\mu_0)^2$.
Step 5: Reassemble the full exponent
$$-\frac{1}{2\sigma^2}\left[\kappa_n(\mu-\mu_n)^2 + nS^2 + \frac{\kappa_0 n}{\kappa_n}(\bar{y}-\mu_0)^2 + 2\beta_0\right]$$
Step 6: Define A and B
$$A = \beta_0 + \frac{nS^2}{2} + \frac{\kappa_0 n}{2\kappa_n}(\bar{y}-\mu_0)^2, \qquad B = \frac{\kappa_n}{2}$$
so the exponent is $-\dfrac{A + B(\mu-\mu_n)^2}{\sigma^2}$ and the joint posterior kernel is:
$$\boxed{p(\mu, \sigma^2 \mid \mathbf{y}) \propto (\sigma^2)^{-(\alpha_n + 3/2)}\exp\left(-\frac{A + B(\mu – \mu_n)^2}{\sigma^2}\right)}$$ where $\alpha_n = \alpha_0 + n/2$.
$A$ collects all information that updates the scale of $\sigma^2$:
within-sample variance plus disagreement between prior mean and data mean.
$B$ is half the posterior sample size $\kappa_n$, governing the precision of $\mu$ given $\sigma^2$.
Posterior Distributions
Conditional Posterior of $\mu \mid \sigma^2$
The $\mu$-dependent part of the joint kernel is:
$$\exp\left(-\frac{B(\mu-\mu_n)^2}{\sigma^2}\right) = \exp\left(-\frac{\kappa_n(\mu-\mu_n)^2}{2\sigma^2}\right)$$
This is a Gaussian kernel in $\mu$, so:
$$\mu \mid \sigma^2, \mathbf{y} \sim N\left(\mu_n,\ \frac{\sigma^2}{\kappa_n}\right)$$
Marginal Posterior of $\sigma^2$
Integrate out $\mu$ from the joint kernel:
$$\int_{-\infty}^{\infty} \exp\left(-\frac{B(\mu-\mu_n)^2}{\sigma^2}\right) d\mu = \left(\frac{\pi\sigma^2}{B}\right)^{1/2} \propto (\sigma^2)^{1/2}$$
Absorbing this into the power of $\sigma^2$:
$$p(\sigma^2 \mid \mathbf{y}) \propto (\sigma^2)^{-(\alpha_n+3/2)} \cdot (\sigma^2)^{1/2} \cdot \exp\left(-\frac{A}{\sigma^2}\right) = (\sigma^2)^{-(\alpha_n+1)} \exp\left(-\frac{A}{\sigma^2}\right)$$
This is the kernel of $\text{IG}(\alpha_n, A)$. Since $A = \beta_n$:
$$\sigma^2 \mid \mathbf{y} \sim \text{IG}(\alpha_n, \beta_n)$$
Marginal Posterior of $\mu$
Start from the joint posterior kernel (derived earlier):
$$p(\mu, \sigma^2 \mid \mathbf{y}) \propto (\sigma^2)^{-(\alpha_n+3/2)} \exp\left(-\frac{A + B(\mu-\mu_n)^2}{\sigma^2}\right)$$ where $A = \beta_n$ and $B = \kappa_n/2$.
Set up the integral over sigma^2, holding mu fixed
$$p(\mu \mid \mathbf{y}) \propto \int_0^\infty (\sigma^2)^{-(\alpha_n+3/2)} \exp\left(-\frac{C(\mu)}{\sigma^2}\right) d\sigma^2$$ where $C(\mu) = A + B(\mu-\mu_n)^2$ is treated as a single constant during this integration, since it does not depend on $\sigma^2$.
The standard Inverse Gamma-type integral
$$\int_0^\infty (\sigma^2)^{-(a+1)} e^{-C/\sigma^2}\, d\sigma^2 = \Gamma(a)\, C^{-a}$$
Matching the power: $\alpha_n + 3/2 = a+1 \Rightarrow a = \alpha_n + 1/2$.
Apply the formula
$$p(\mu \mid \mathbf{y}) \propto \Gamma\!\left(\alpha_n+\tfrac{1}{2}\right) \cdot C(\mu)^{-(\alpha_n+1/2)}$$
The Gamma function is a constant in $\mu$, so:
$$p(\mu \mid \mathbf{y}) \propto \big[A + B(\mu-\mu_n)^2\big]^{-(\alpha_n+1/2)}$$
Factor out A to expose the t-kernel shape
$$A + B(\mu-\mu_n)^2 = A\left[1 + \frac{B}{A}(\mu-\mu_n)^2\right]$$
$$p(\mu \mid \mathbf{y}) \propto \left[1 + \frac{B}{A}(\mu-\mu_n)^2\right]^{-(\alpha_n+1/2)}$$
(the constant $A^{-(\alpha_n+1/2)}$ is absorbed into the proportionality.)
Match against the standard Student-t kernel
Standard non-central Student-t density, $\nu$ degrees of freedom, location $m$, scale$^2 = s^2$:
$$p(x) \propto \left[1 + \frac{1}{\nu}\frac{(x-m)^2}{s^2}\right]^{-(\nu+1)/2}$$
Matching powers: $\dfrac{\nu+1}{2} = \alpha_n + \dfrac{1}{2} \;\Rightarrow\; \nu = 2\alpha_n$
Matching the bracket coefficient: $\dfrac{1}{\nu s^2} = \dfrac{B}{A} = \dfrac{\kappa_n/2}{\beta_n} \;\Rightarrow\; s^2 = \dfrac{\beta_n}{\alpha_n \kappa_n}$
This $\implies \boxed{\mu \mid \mathbf{y} \sim t_{2\alpha_n}\!\left(\mu_n,\ \frac{\beta_n}{\alpha_n \kappa_n}\right)}$,
$2\alpha_n$ degrees of freedom, location $\mu_n$, scale$^2 = \beta_n/(\alpha_n\kappa_n)$.
Summary of Updated Hyperparameters
$$\mu_n = \frac{\kappa_0\mu_0 + n\bar{y}}{\kappa_0 + n}$$
$$\kappa_n = \kappa_0 + n$$
$$\alpha_n = \alpha_0 + \frac{n}{2}$$
$$\beta_n = A = \beta_0 + \frac{nS^2}{2} + \frac{\kappa_0 n}{2\kappa_n}(\bar{y}-\mu_0)^2$$
Final Note: Conjugacy Holds at the Conditional Level
The Normal-Inverse Gamma prior is conjugate for the joint and the conditional: $p(\mu\mid\sigma^2)$ is
Normal both as prior and posterior, and $p(\sigma^2)$ is Inverse Gamma both as
prior and posterior. This is the notion of conditional conjugacy.
But once we marginalise — integrate out the other parameter — conjugacy breaks.
$p(\sigma^2\mid\mathbf{y})$ marginal stays Inverse Gamma (shown earlier, since
integrating out $\mu$ only multiplies in a power of $\sigma^2$, not a new
functional form).
But $p(\mu\mid\mathbf{y})$ marginal is not Normal — it
is Student-t. The prior on $\mu$ was Normal; the marginal posterior on $\mu$ is
a different family entirely.
So “conjugacy” here is a property of the
conditional structure of the model, not something that survives marginalisation
of every parameter.
A note – kernel
A kernel is a probability density function up to its normalising constant — it has the correct shape but does not integrate to 1 by itself.
Writing $p(\theta\mid\mathbf{y}) \propto f(\theta)$ means $f(\theta)$ is the kernel; the true density is $\frac{f(\theta)}{\int f(\theta)\,d\theta}$.