Normal Models – Two Parameters Case

In the one-parameter models, only $\mu$ or only $\sigma^2$ is unknown, so the prior is placed on that single parameter alone and Bayes’ rule gives a direct closed-form update.

Recall One parameter model

Once both $\mu$ and $\sigma^2$ are unknown together, we need a genuine joint prior $p(\mu,\sigma^2)$ on the pair, not two separate one-parameter priors — this is the NIG model, where the joint prior is built as $p(\mu\mid\sigma^2),p(\sigma^2)$ so that Bayes’ rule can be applied jointly and still yield a closed-form joint posterior.

Observe $y_1, y_2, \dots, y_n$ iid, each $y_i \mid \mu, \sigma^2 \sim N(\mu, \sigma^2)$.

$$\mathscr{L}(\mu, \sigma^2 \mid \mathbf{y}) = (2\pi\sigma^2)^{-n/2} \exp\left(-\frac{1}{2\sigma^2} \sum_{i=1}^n (y_i – \mu)^2\right)$$

Using the identity $\sum_{i=1}^n (y_i – \mu)^2 = nS^2 + n(\bar{y} – \mu)^2$

where $S^2 = \frac{1}{n}\sum_{i=1}^n(y_i – \bar{y})^2$:

$$\mathscr{L}(\mu, \sigma^2 \mid \mathbf{y}) = (2\pi\sigma^2)^{-n/2} \exp\left(-\frac{nS^2 + n(\bar{y}-\mu)^2}{2\sigma^2}\right)$$

The sufficient statistics are $(\bar{y}, S^2)$.


The conjugate joint prior on $(\mu, \sigma^2)$ is specified hierarchically:

$$\mu \mid \sigma^2 \sim N\left(\mu_0,\ \frac{\sigma^2}{\kappa_0}\right), \qquad \sigma^2 \sim \text{IG}(\alpha_0, \beta_0)$$

Written explicitly as $p(\mu, \sigma^2) = p(\mu \mid \sigma^2)\, p(\sigma^2)$:

$$p(\mu \mid \sigma^2) = \left(\frac{\kappa_0}{2\pi\sigma^2}\right)^{1/2} \exp\left(-\frac{\kappa_0(\mu – \mu_0)^2}{2\sigma^2}\right)$$

$$p(\sigma^2) = \frac{\beta_0^{\alpha_0}}{\Gamma(\alpha_0)}\,(\sigma^2)^{-(\alpha_0+1)}\,e^{-\beta_0/\sigma^2}$$

$$p(\mu, \sigma^2) \propto (\sigma^2)^{-(\alpha_0 + 3/2)} \exp\left(-\frac{1}{\sigma^2}\left[\beta_0 + \frac{\kappa_0(\mu-\mu_0)^2}{2}\right]\right)$$

Hyperparameters: $\mu_0$ (prior mean of $\mu$), $\kappa_0$ (prior sample size),

$\alpha_0$ (prior shape), $\beta_0$ (prior scale).

Bayes’ Rule

$$p(\mu, \sigma^2 \mid \mathbf{y}) \propto \mathscr{L}(\mu, \sigma^2 \mid \mathbf{y})\cdot p(\mu, \sigma^2)$$

$$p(\mathbf{y}) = \int_0^\infty \int_{-\infty}^\infty \mathscr{L}(\mu, \sigma^2 \mid \mathbf{y})\, p(\mu, \sigma^2)\, d\mu\, d\sigma^2$$

We state $p(\mathbf{y})$ exists in closed form but do not derive it;

the posterior follows by identifying the kernel.

Combining Likelihood and Prior


$$p(\mu, \sigma^2 \mid \mathbf{y}) \propto (\sigma^2)^{-n/2}\exp\left(-\frac{nS^2 + n(\bar{y}-\mu)^2}{2\sigma^2}\right)\cdot(\sigma^2)^{-(\alpha_0+3/2)}\exp\left(-\frac{\kappa_0(\mu-\mu_0)^2}{2\sigma^2} – \frac{\beta_0}{\sigma^2}\right)$$

Step 1: Collect powers of $\sigma^2$

$$(\sigma^2)^{-n/2} \cdot (\sigma^2)^{-(\alpha_0+3/2)} = (\sigma^2)^{-(n/2 + \alpha_0 + 3/2)}$$

Step 2: Collect all exponent terms

$$\text{Exponent} = -\frac{1}{2\sigma^2}\left[nS^2 + n(\bar{y}-\mu)^2 + \kappa_0(\mu-\mu_0)^2 + 2\beta_0\right]$$

Step 3: Expand the quadratic in $\mu$

$$n(\bar{y}-\mu)^2 + \kappa_0(\mu-\mu_0)^2$$

$$= n\bar{y}^2 – 2n\bar{y}\mu + n\mu^2 + \kappa_0\mu_0^2 – 2\kappa_0\mu_0\mu + \kappa_0\mu^2$$

$$= (\kappa_0 + n)\mu^2 – 2(\kappa_0\mu_0 + n\bar{y})\mu + (n\bar{y}^2 + \kappa_0\mu_0^2)$$

Step 4: Complete the square in $\mu$

Let $\kappa_n = \kappa_0 + n$ and $\mu_n = \dfrac{\kappa_0\mu_0 + n\bar{y}}{\kappa_n}$.

$$(\kappa_0+n)\mu^2 – 2(\kappa_0\mu_0+n\bar{y})\mu + (n\bar{y}^2+\kappa_0\mu_0^2)$$

$$= \kappa_n\left(\mu – \mu_n\right)^2 – \kappa_n\mu_n^2 + (n\bar{y}^2 + \kappa_0\mu_0^2)$$

The residual constant is:

$$n\bar{y}^2 + \kappa_0\mu_0^2 – \kappa_n\mu_n^2= n\bar{y}^2 + \kappa_0\mu_0^2 – \frac{(\kappa_0\mu_0 + n\bar{y})^2}{\kappa_n}$$

Expanding the numerator of the last term:

$$(\kappa_0\mu_0 + n\bar{y})^2 = \kappa_0^2\mu_0^2 + 2\kappa_0 n\mu_0\bar{y} + n^2\bar{y}^2$$

So the residual becomes:

$$\frac{\kappa_n(n\bar{y}^2 + \kappa_0\mu_0^2) – \kappa_0^2\mu_0^2 – 2\kappa_0 n\mu_0\bar{y} – n^2\bar{y}^2}{\kappa_n}$$

Numerator:

$$(\kappa_0+n)(n\bar{y}^2+\kappa_0\mu_0^2) – \kappa_0^2\mu_0^2 – 2\kappa_0 n\mu_0\bar{y} – n^2\bar{y}^2$$

$$= \kappa_0 n\bar{y}^2 + n^2\bar{y}^2 + \kappa_0^2\mu_0^2 + \kappa_0 n\mu_0^2- \kappa_0^2\mu_0^2 – 2\kappa_0 n\mu_0\bar{y} – n^2\bar{y}^2$$

$$= \kappa_0 n\bar{y}^2 – 2\kappa_0 n\mu_0\bar{y} + \kappa_0 n\mu_0^2$$

$$= \kappa_0 n(\bar{y} – \mu_0)^2$$

Therefore the residual constant is $\dfrac{\kappa_0 n}{\kappa_n}(\bar{y}-\mu_0)^2$.

Step 5: Reassemble the full exponent

$$-\frac{1}{2\sigma^2}\left[\kappa_n(\mu-\mu_n)^2 + nS^2 + \frac{\kappa_0 n}{\kappa_n}(\bar{y}-\mu_0)^2 + 2\beta_0\right]$$

Step 6: Define A and B

$$A = \beta_0 + \frac{nS^2}{2} + \frac{\kappa_0 n}{2\kappa_n}(\bar{y}-\mu_0)^2, \qquad B = \frac{\kappa_n}{2}$$

so the exponent is $-\dfrac{A + B(\mu-\mu_n)^2}{\sigma^2}$ and the joint posterior kernel is:

$$\boxed{p(\mu, \sigma^2 \mid \mathbf{y}) \propto (\sigma^2)^{-(\alpha_n + 3/2)}\exp\left(-\frac{A + B(\mu – \mu_n)^2}{\sigma^2}\right)}$$ where $\alpha_n = \alpha_0 + n/2$.

$A$ collects all information that updates the scale of $\sigma^2$:

within-sample variance plus disagreement between prior mean and data mean.

$B$ is half the posterior sample size $\kappa_n$, governing the precision of $\mu$ given $\sigma^2$.


Conditional Posterior of $\mu \mid \sigma^2$

The $\mu$-dependent part of the joint kernel is:

$$\exp\left(-\frac{B(\mu-\mu_n)^2}{\sigma^2}\right) = \exp\left(-\frac{\kappa_n(\mu-\mu_n)^2}{2\sigma^2}\right)$$

This is a Gaussian kernel in $\mu$, so:

$$\mu \mid \sigma^2, \mathbf{y} \sim N\left(\mu_n,\ \frac{\sigma^2}{\kappa_n}\right)$$

Integrate out $\mu$ from the joint kernel:

$$\int_{-\infty}^{\infty} \exp\left(-\frac{B(\mu-\mu_n)^2}{\sigma^2}\right) d\mu = \left(\frac{\pi\sigma^2}{B}\right)^{1/2} \propto (\sigma^2)^{1/2}$$

Absorbing this into the power of $\sigma^2$:

$$p(\sigma^2 \mid \mathbf{y}) \propto (\sigma^2)^{-(\alpha_n+3/2)} \cdot (\sigma^2)^{1/2} \cdot \exp\left(-\frac{A}{\sigma^2}\right) = (\sigma^2)^{-(\alpha_n+1)} \exp\left(-\frac{A}{\sigma^2}\right)$$

This is the kernel of $\text{IG}(\alpha_n, A)$. Since $A = \beta_n$:

$$\sigma^2 \mid \mathbf{y} \sim \text{IG}(\alpha_n, \beta_n)$$


Start from the joint posterior kernel (derived earlier):

$$p(\mu, \sigma^2 \mid \mathbf{y}) \propto (\sigma^2)^{-(\alpha_n+3/2)} \exp\left(-\frac{A + B(\mu-\mu_n)^2}{\sigma^2}\right)$$ where $A = \beta_n$ and $B = \kappa_n/2$.

Set up the integral over sigma^2, holding mu fixed

$$p(\mu \mid \mathbf{y}) \propto \int_0^\infty (\sigma^2)^{-(\alpha_n+3/2)} \exp\left(-\frac{C(\mu)}{\sigma^2}\right) d\sigma^2$$ where $C(\mu) = A + B(\mu-\mu_n)^2$ is treated as a single constant during this integration, since it does not depend on $\sigma^2$.

The standard Inverse Gamma-type integral

$$\int_0^\infty (\sigma^2)^{-(a+1)} e^{-C/\sigma^2}\, d\sigma^2 = \Gamma(a)\, C^{-a}$$

Matching the power: $\alpha_n + 3/2 = a+1 \Rightarrow a = \alpha_n + 1/2$.

Apply the formula

$$p(\mu \mid \mathbf{y}) \propto \Gamma\!\left(\alpha_n+\tfrac{1}{2}\right) \cdot C(\mu)^{-(\alpha_n+1/2)}$$

The Gamma function is a constant in $\mu$, so:

$$p(\mu \mid \mathbf{y}) \propto \big[A + B(\mu-\mu_n)^2\big]^{-(\alpha_n+1/2)}$$

Factor out A to expose the t-kernel shape

$$A + B(\mu-\mu_n)^2 = A\left[1 + \frac{B}{A}(\mu-\mu_n)^2\right]$$

$$p(\mu \mid \mathbf{y}) \propto \left[1 + \frac{B}{A}(\mu-\mu_n)^2\right]^{-(\alpha_n+1/2)}$$

(the constant $A^{-(\alpha_n+1/2)}$ is absorbed into the proportionality.)

Match against the standard Student-t kernel

Standard non-central Student-t density, $\nu$ degrees of freedom, location $m$, scale$^2 = s^2$:

$$p(x) \propto \left[1 + \frac{1}{\nu}\frac{(x-m)^2}{s^2}\right]^{-(\nu+1)/2}$$

Matching powers: $\dfrac{\nu+1}{2} = \alpha_n + \dfrac{1}{2} \;\Rightarrow\; \nu = 2\alpha_n$

Matching the bracket coefficient: $\dfrac{1}{\nu s^2} = \dfrac{B}{A} = \dfrac{\kappa_n/2}{\beta_n} \;\Rightarrow\; s^2 = \dfrac{\beta_n}{\alpha_n \kappa_n}$

This $\implies \boxed{\mu \mid \mathbf{y} \sim t_{2\alpha_n}\!\left(\mu_n,\ \frac{\beta_n}{\alpha_n \kappa_n}\right)}$,

$2\alpha_n$ degrees of freedom, location $\mu_n$, scale$^2 = \beta_n/(\alpha_n\kappa_n)$.


$$\mu_n = \frac{\kappa_0\mu_0 + n\bar{y}}{\kappa_0 + n}$$

$$\kappa_n = \kappa_0 + n$$

$$\alpha_n = \alpha_0 + \frac{n}{2}$$

$$\beta_n = A = \beta_0 + \frac{nS^2}{2} + \frac{\kappa_0 n}{2\kappa_n}(\bar{y}-\mu_0)^2$$

The Normal-Inverse Gamma prior is conjugate for the joint and the conditional: $p(\mu\mid\sigma^2)$ is
Normal both as prior and posterior, and $p(\sigma^2)$ is Inverse Gamma both as
prior and posterior. This is the notion of conditional conjugacy.

But once we marginalise — integrate out the other parameter — conjugacy breaks.
$p(\sigma^2\mid\mathbf{y})$ marginal stays Inverse Gamma (shown earlier, since
integrating out $\mu$ only multiplies in a power of $\sigma^2$, not a new
functional form).

But $p(\mu\mid\mathbf{y})$ marginal is not Normal — it
is Student-t. The prior on $\mu$ was Normal; the marginal posterior on $\mu$ is
a different family entirely.

So “conjugacy” here is a property of the
conditional structure of the model, not something that survives marginalisation
of every parameter.


A note – kernel

A kernel is a probability density function up to its normalising constant — it has the correct shape but does not integrate to 1 by itself.

Writing $p(\theta\mid\mathbf{y}) \propto f(\theta)$ means $f(\theta)$ is the kernel; the true density is $\frac{f(\theta)}{\int f(\theta)\,d\theta}$.


Scroll to Top