A Note about Prior Distribution

In Bayesian analysis, the prior distribution plays a crucial role, as it treats parameters as random variables. The selection of an appropriate prior depends on the nature of the parameter, enabling the incorporation of relevant information into the analysis.

When prior distributions lack a clear basis in population data, they can be challenging to define. This has led to persistent interest in identifying prior distributions that exert minimal influence on the posterior distribution. Such priors are commonly referred to as “reference priors”, and their corresponding prior densities are described as vague, flat, diffuse, or noninformative. The rationale for using noninformative priors is often framed as allowing the data to “speak for itself,” meaning that inferences should be driven solely by the data, without external influence.

A related concept is the weakly informative prior, which incorporates some information — enough to “regularize” the posterior distribution, ensuring it remains within reasonable bounds — without fully encoding the underlying scientific knowledge of the parameter.

A prior density $p(\theta)$ is considered proper if it does not depend on the data and integrates to 1. If the integral of the prior does not equal 1 but instead converges to a positive finite value, it is referred to as an unnormalized density. Such a density can be renormalized by multiplying it by a constant to ensure it integrates to 1. If the prior density does not integrate to 1 or any finite value, it is considered improper.

For continuous parameters, a common approach is to select a uniform distribution as a noninformative prior for the parameter $\theta$. One well-known method for defining noninformative prior distributions was introduced by Jeffreys, which involves considering one-to-one transformations of the parameter.

Jeffreys’ principle leads to defining the noninformative prior density as $p(\theta) \propto \sqrt{I(\theta)}$, where $I(\theta)$ is the Fisher information of the model at $\theta$.

Fisher information is obtained from the Hessian matrix as:
$$I(\theta) = -E\left[\frac{d^2 l(\theta)}{d\theta^2}\right]$$

Here $l(\theta)$ refers to the natural logarithm of the likelihood function $\mathscr{L}(\theta)$; that is, $l(\theta) = \ln \mathscr{L}(\theta)$.


Informative Prior Distribution

An informative prior incorporates existing knowledge or beliefs about a parameter before observing the data. It is typically chosen based on previous research, expert opinion, or historical data, and is expected to significantly influence the posterior distribution. Informative priors are used when strong prior knowledge is available, and they guide the analysis in a way that reflects this knowledge.

One common approach to defining informative priors is the moment-matching prior. This method involves selecting a prior distribution that matches certain moments (such as the mean and variance) of the target distribution. Moment matching ensures that the prior reflects the characteristics of the parameter as known or estimated from previous data or theoretical considerations. This approach formalizes the incorporation of prior knowledge while ensuring that the prior is statistically consistent with the expected behavior of the parameter.


Conjugate Priors

A conjugate prior is a prior distribution that, when combined with a likelihood function, results in a posterior distribution of the same family as the prior. This property simplifies Bayesian analysis by making the posterior distribution analytically tractable, which is especially useful for computational efficiency.

Conjugate priors are particularly advantageous in situations where computational simplicity and analytical solutions are preferred. However, their main limitation is that they may not always accurately reflect the true nature of the parameter, as they impose a structure that might not fully align with prior knowledge.

In summary, conjugate priors facilitate Bayesian inference by ensuring that the posterior belongs to the same distributional family as the prior, thereby streamlining the process of updating beliefs.

Hyperparameters

In Bayesian analysis, hyperparameters are parameters that control the prior distributions of other parameters in a hierarchical model. They play a crucial role in shaping the prior information before any data is observed, enabling a more flexible representation of uncertainty.


An Example

Please refer this example on Binomial Likelihood with Beta Prior model

The Binomial-Beta model is a classic example of a hierarchical model, where a Binomial likelihood is combined with a Beta prior. The Binomial distribution is commonly used for modeling binary outcomes (success/failure), while the Beta distribution is a conjugate prior for the Binomial, meaning the posterior distribution also follows a Beta distribution.

In this context, the parameter of interest is the probability of success $\theta$, which is assumed to follow a Beta distribution:

$$\theta \sim \text{Beta}(\alpha, \beta)$$

where $\alpha$ and $\beta$ are the hyperparameters of the Beta distribution, determining the shape of the prior. These hyperparameters encapsulate prior beliefs about the probability of success. For instance:

  • If $\alpha = \beta = 1$, the prior is uniform, indicating no preference for any value of $\theta$.
  • Larger values of $\alpha$ and $\beta$ correspond to stronger prior beliefs about $\theta$, with $\alpha$ controlling the weight of successes and $\beta$ controlling the weight of failures.

The hyperparameters $\alpha$ and $\beta$ influence how observed data update the belief about $\theta$, and thus play a critical role in determining posterior inference.

The Binomial-Beta model demonstrates how hyperparameters control the prior belief about a parameter of interest and how they shape the posterior distribution in light of observed data.

Binomial-Beta Model: Complete Derivation

Now let us examine the mathematical derivation of the conjugacy property of the Binomial–Beta model

The Likelihood Function

We observe $y$ successes in $n$ independent Bernoulli trials, where $\theta$ is the probability of success. The number of successes $y$ follows a Binomial distribution:

$$y \mid \theta \sim \text{Binomial}(n, \theta)$$

The probability mass function is:

$$p(y \mid \theta) = \binom{n}{y} \theta^y (1-\theta)^{n-y}, \quad y = 0, 1, \dots, n$$

where $\dbinom{n}{y} = \dfrac{n!}{y!(n-y)!}$ is the binomial coefficient, counting the number of ways $y$ successes can occur in $n$ trials. Viewed as a function of $\theta$ for fixed observed $y$, this is the likelihood function:

$$\mathscr{L}(\theta \mid y) = \binom{n}{y} \theta^y (1-\theta)^{n-y}$$


The Prior Distribution

Before observing any data, we encode our prior belief about $\theta$ through a prior distribution. Since $\theta \in (0,1)$ is a probability, a natural and flexible choice is the Beta distribution:

$$\theta \sim \text{Beta}(\alpha, \beta)$$

The Beta distribution has the probability density function:

$$p(\theta) = \frac{1}{B(\alpha,\beta)}, \theta^{\alpha-1}(1-\theta)^{\beta-1}, \quad 0 < \theta < 1$$

where $\alpha > 0$ and $\beta > 0$ are the hyperparameters, and $B(\alpha, \beta)$ is the Beta function defined as:

$$B(\alpha,\beta) = \int_0^1 \theta^{\alpha-1}(1-\theta)^{\beta-1}, d\theta$$

The Beta function has a closed form in terms of the Gamma function $\Gamma(\cdot)$:

$$B(\alpha,\beta) = \frac{\Gamma(\alpha),\Gamma(\beta)}{\Gamma(\alpha+\beta)}$$

where for any positive integer $k$, $\Gamma(k) = (k-1)!$, and more generally $\Gamma(k) = \displaystyle\int_0^\infty t^{k-1}e^{-t},dt$ for real $k > 0$.


Bayes’ Theorem

To derive Bayes’ theorem from first principles, we begin with the definition of conditional probability. For any two events $A$ and $B$:

$$P(A \mid B) = \frac{P(A \cap B)}{P(B)}$$

Applied to $\theta$ and $y$, the joint distribution can be written in two equivalent ways:

$$p(\theta, y) = p(y \mid \theta)\cdot p(\theta)$$

$$p(\theta, y) = p(\theta \mid y)\cdot p(y)$$

Setting these equal:

$$p(\theta \mid y)\cdot p(y) = p(y \mid \theta)\cdot p(\theta)$$

Dividing both sides by $p(y)$, provided $p(y) > 0$:

$$p(\theta \mid y) = \frac{p(y \mid \theta)\cdot p(\theta)}{p(y)}$$

Recognizing $p(y \mid \theta) = \mathscr{L}(\theta \mid y)$, we write:

$$\boxed{p(\theta \mid y) = \frac{\mathscr{L}(\theta \mid y)\cdot p(\theta)}{p(y)}}$$

This is Bayes’ theorem, where:

  • $p(\theta \mid y)$ is the posterior distribution — our updated belief about $\theta$ after observing $y$
  • $\mathscr{L}(\theta \mid y)$ is the likelihood — how probable the observed data are for a given $\theta$
  • $p(\theta)$ is the prior distribution — our belief about $\theta$ before observing data
  • $p(y)$ is the marginal likelihood — the probability of the observed data averaged over all possible values of $\theta$

The Marginal Likelihood

The marginal likelihood $p(y)$ is obtained by integrating the joint distribution $p(y, \theta)$ over all possible values of $\theta$:

$$p(y) = \int_0^1 p(y \mid \theta), p(\theta), d\theta$$

Substituting the Binomial likelihood and Beta prior:

$$p(y) = \int_0^1 \binom{n}{y} \theta^y (1-\theta)^{n-y} \cdot \frac{1}{B(\alpha,\beta)}, \theta^{\alpha-1}(1-\theta)^{\beta-1}, d\theta$$

Since $\dbinom{n}{y}$ and $\dfrac{1}{B(\alpha,\beta)}$ do not depend on $\theta$, they are pulled outside the integral:

$$p(y) = \frac{\dbinom{n}{y}}{B(\alpha,\beta)} \int_0^1 \theta^y(1-\theta)^{n-y}\cdot \theta^{\alpha-1}(1-\theta)^{\beta-1}, d\theta$$

Combining the powers of $\theta$ and $(1-\theta)$ inside the integral:

$$p(y) = \frac{\dbinom{n}{y}}{B(\alpha,\beta)} \int_0^1 \theta^{(\alpha+y)-1}(1-\theta)^{(\beta+n-y)-1}, d\theta$$

Now we recognize this integral as a Beta function with parameters $a = \alpha + y$ and $b = \beta + n – y$:

$$\int_0^1 \theta^{(\alpha+y)-1}(1-\theta)^{(\beta+n-y)-1}, d\theta = B(\alpha+y,, \beta+n-y)$$

Substituting back:

$$p(y) = \frac{\dbinom{n}{y}}{B(\alpha,\beta)} \cdot B(\alpha+y,, \beta+n-y)$$

Now expanding both Beta functions using the Gamma function relationship $B(a,b) = \dfrac{\Gamma(a)\Gamma(b)}{\Gamma(a+b)}$:

$$B(\alpha, \beta) = \frac{\Gamma(\alpha),\Gamma(\beta)}{\Gamma(\alpha+\beta)}, \qquad B(\alpha+y,,\beta+n-y) = \frac{\Gamma(\alpha+y),\Gamma(\beta+n-y)}{\Gamma(\alpha+\beta+n)}$$

Substituting:

$$p(y) = \frac{\dbinom{n}{y}}{\dfrac{\Gamma(\alpha)\Gamma(\beta)}{\Gamma(\alpha+\beta)}} \cdot \frac{\Gamma(\alpha+y),\Gamma(\beta+n-y)}{\Gamma(\alpha+\beta+n)}$$

$$\boxed{p(y) = \binom{n}{y}, \frac{\Gamma(\alpha+\beta)}{\Gamma(\alpha),\Gamma(\beta)} \cdot \frac{\Gamma(\alpha+y),\Gamma(\beta+n-y)}{\Gamma(\alpha+\beta+n)}}$$


Substituting into Bayes’ Theorem

We now substitute the likelihood, prior, and marginal likelihood into Bayes’ theorem:

$$p(\theta \mid y) = \frac{\mathscr{L}(\theta \mid y)\cdot p(\theta)}{p(y)}$$

Numerator:

$$\mathscr{L}(\theta \mid y)\cdot p(\theta) = \binom{n}{y}\theta^y(1-\theta)^{n-y} \cdot \frac{1}{B(\alpha,\beta)},\theta^{\alpha-1}(1-\theta)^{\beta-1}$$

$$= \binom{n}{y}\cdot\frac{1}{B(\alpha,\beta)}\cdot \theta^{\alpha+y-1}(1-\theta)^{\beta+n-y-1}$$

Denominator:

$$p(y) = \binom{n}{y},\frac{\Gamma(\alpha+\beta)}{\Gamma(\alpha),\Gamma(\beta)}\cdot\frac{\Gamma(\alpha+y),\Gamma(\beta+n-y)}{\Gamma(\alpha+\beta+n)}$$

Dividing numerator by denominator:

$$p(\theta \mid y) = \frac{\dbinom{n}{y}\cdot\dfrac{1}{B(\alpha,\beta)}\cdot \theta^{\alpha+y-1}(1-\theta)^{\beta+n-y-1}}{\dbinom{n}{y}\cdot\dfrac{\Gamma(\alpha+\beta)}{\Gamma(\alpha)\Gamma(\beta)}\cdot\dfrac{\Gamma(\alpha+y)\Gamma(\beta+n-y)}{\Gamma(\alpha+\beta+n)}}$$

The $\dbinom{n}{y}$ cancels from numerator and denominator:

$$p(\theta \mid y) = \frac{\dfrac{1}{B(\alpha,\beta)}\cdot \theta^{\alpha+y-1}(1-\theta)^{\beta+n-y-1}}{\dfrac{\Gamma(\alpha+\beta)}{\Gamma(\alpha)\Gamma(\beta)}\cdot\dfrac{\Gamma(\alpha+y)\Gamma(\beta+n-y)}{\Gamma(\alpha+\beta+n)}}$$

Since $\dfrac{1}{B(\alpha,\beta)} = \dfrac{\Gamma(\alpha+\beta)}{\Gamma(\alpha)\Gamma(\beta)}$, this term cancels between numerator and denominator:

$$p(\theta \mid y) = \frac{\theta^{\alpha+y-1}(1-\theta)^{\beta+n-y-1}}{\dfrac{\Gamma(\alpha+y),\Gamma(\beta+n-y)}{\Gamma(\alpha+\beta+n)}}$$

Bringing the denominator up:

$$p(\theta \mid y) = \frac{\Gamma(\alpha+\beta+n)}{\Gamma(\alpha+y),\Gamma(\beta+n-y)}\cdot\theta^{\alpha+y-1}(1-\theta)^{\beta+n-y-1}$$

Recognizing that $\dfrac{\Gamma(\alpha+\beta+n)}{\Gamma(\alpha+y),\Gamma(\beta+n-y)} = \dfrac{1}{B(\alpha+y,,\beta+n-y)}$:

$$\boxed{p(\theta \mid y) = \frac{1}{B(\alpha+y,,\beta+n-y)},\theta^{\alpha+y-1}(1-\theta)^{\beta+n-y-1}}$$


This is exactly the probability density function of a $\text{Beta}(\alpha+y,,\beta+n-y)$ distribution. Therefore:

$$\boxed{\theta \mid y \sim \text{Beta}(\alpha + y,; \beta + n – y)}$$

This confirms that the Beta prior is conjugate to the Binomial likelihood — the posterior belongs to the same Beta family as the prior, with hyperparameters updated from $(\alpha, \beta)$ to $(\alpha + y,, \beta + n – y)$.

It is also worth noting that the posterior distribution of $\theta$ — given a Beta prior and Binomial likelihood — remains a Beta distribution with updated hyperparameters: $\theta \mid y \sim \text{Beta}(\alpha + y,\ \beta + n – y)$. This is an example of a conjugate prior.

The advantage of conjugate priors is that they allow for analytical derivation of the posterior distribution, making it easier and more efficient to obtain posterior summaries.

Interpretation of Updated Hyperparameters

HyperparameterPriorPosteriorInterpretation
Successes weight$\alpha$$\alpha + y$Prior successes + observed successes
Failures weight$\beta$$\beta + n-y$Prior failures + observed failures
Total weight$\alpha + \beta$$\alpha + \beta + n$Prior strength + sample size

This is the elegance of conjugacy — the observed data simply add onto the prior hyperparameters, updating beliefs in a clean and interpretable way.

Next More from Posterior

For other Conjugate Models,click Model 1 Model 2 Model 3

Scroll to Top