Bayesian Estimation_1

Before you proceed, please refer to Bayes theorem and inferential preliminaries

Introduction to Bayesian Approaches for Statistical Inference

Bayesian statistics offers a probabilistic framework for statistical inference, which allows for the incorporation of prior knowledge or beliefs into the analysis. Unlike traditional frequentist methods, Bayesian inference treats unknown parameters as random variables and updates the belief about these parameters based on observed data.

Key Concepts:

  • Prior Distribution: Represents the initial belief or knowledge about a parameter $\theta$ before any data $\mathbf{X} = (X_1, \ldots, X_n)$ is observed. It expresses what is known about the parameter in the absence of data. This prior can be based on previous studies, expert knowledge, or assumptions.
  • Likelihood Function: Let $f(\mathbf{x} \mid \theta)$ denote the joint pdf or pmf of the sample $\mathbf{X}$. Then, given that $\mathbf{X} = \mathbf{x}$ is observed, the function of $\theta$ defined by $L(\theta \mid \mathbf{x}) = f(\mathbf{x} \mid \theta)$ is called the likelihood function. The likelihood function measures how well a particular value of the parameter $\theta$ explains the observed data $\mathbf{X} $. Likelihood is not a probability distribution over $\theta$. It need not integrate to 1 over the range of $\theta$, and no probabilistic interpretation is assigned to $\theta$ itself.
  • Posterior Distribution: The posterior distribution represents the updated belief about the parameter $\theta$ after observing the data. It is computed by combining the prior and the likelihood of the observed data using Bayes’ theorem.

Posterior Summaries

In Bayesian analysis, posterior summaries are key statistics that provide insight into the distribution of a parameter of interest after observing the data. These summaries help interpret the posterior distribution, which is often complex and multidimensional. Common posterior summaries include the Maximum A Posteriori (MAP) estimate, posterior mean, standard deviation (SD), percentiles, and credible intervals.

It should be noted that these summaries can be obtained from the posterior distribution. So, obviously, knowledge of the posterior distribution “completely” seems to be compulsory or helpful. To a greater extent this is important, but in general it is not true. However, here we consider the “complete” state of knowledge.

Let us revisit the continuous case discussed at this link. The posterior is a Beta distribution, which is a known theoretical distribution. Its mean and variance are completely known. Its relation with the F-distribution will help in finding the percentiles/quantiles and the bounds of credible intervals for a specific level of probability. Such “closed form” expressions play a vital role in Bayesian Inference.


Example: Beta Distribution as Posterior

$$\theta \mid \mathbf{x} \sim \text{Beta}(a, b)$$

Mean and Variance

$$E(\theta) = \frac{a}{a+b}$$

$$\text{Var}(\theta) = \frac{ab}{(a+b)^2(a+b+1)}$$


Relation Between Beta and F Distribution

If $\theta \sim \text{Beta}(a, b)$, then the transformation

$$F = \frac{b \cdot \theta}{a(1 – \theta)}$$

follows an $F$-distribution with degrees of freedom $d_1 = 2a$ and $d_2 = 2b$:

$$F \sim F(2a,, 2b)$$

Equivalently, if $F \sim F(2a, 2b)$, then

$$\theta = \frac{a \cdot F}{b + a \cdot F} \sim \text{Beta}(a, b)$$


Using the F-Table for Posterior Inference

Percentiles of the Posterior

To find the $p$-th percentile $\theta_p$ of $\text{Beta}(a, b)$:

  1. Look up $F_p(2a, 2b)$ from the F-table.
  2. Convert back:

$$\theta_p = \frac{a \cdot F_p(2a,, 2b)}{b + a \cdot F_p(2a,, 2b)}$$

Credible Interval at Level $(1-\alpha)$

$$\theta_L = \frac{a \cdot F_{\alpha/2}(2a,, 2b)}{b + a \cdot F_{\alpha/2}(2a,, 2b)}$$

$$\theta_U = \frac{a \cdot F_{1-\alpha/2}(2a,, 2b)}{b + a \cdot F_{1-\alpha/2}(2a,, 2b)}$$

where $F_{\alpha/2}$ and $F_{1-\alpha/2}$ are read from the F-table with degrees of freedom $(2a, 2b)$.

Revisiting Hero Vs Honda Example – Continuous case

The continuous case uses Case 1: Uniform Prior, giving posterior $\theta \mid X \sim \text{Beta}(36, 66)$ with $n=100$, $x=35$. Here are all the calculations using the F-table relationship.


Numerical Example: Posterior $\theta \mid X \sim \text{Beta}(36,\ 66)$

Mean and Variance

$$E(\theta) = \frac{36}{36+66} = \frac{36}{102} \approx 0.3529$$

$$\text{Var}(\theta) = \frac{36 \times 66}{102^2 \times 103} = \frac{2376}{1071612} \approx 0.002217$$

$$\text{SD}(\theta) \approx 0.0471$$


F-Table Setup

With $a = 36$, $b = 66$, the transformation is:

$$F = \frac{66 \cdot \theta}{36(1-\theta)} \sim F(72,\ 132)$$

and the inverse:

$$\theta = \frac{36 \cdot F}{66 + 36 \cdot F}$$


90% Credible Interval $(\alpha = 0.10)$

From F-table: $F_{0.05}(72, 132) \approx 0.686$ and $F_{0.95}(72, 132) \approx 1.436$

$$\theta_L = \frac{36 \times 0.686}{66 + 36 \times 0.686} = \frac{24.70}{90.70} \approx 0.272$$

$$\theta_U = \frac{36 \times 1.436}{66 + 36 \times 1.436} = \frac{51.70}{117.70} \approx 0.439$$

90% Credible Interval: (0.272, 0.439)


95% Credible Interval $(\alpha = 0.05)$

From F-table: $F_{0.025}(72, 132) \approx 0.634$ and $F_{0.975}(72, 132) \approx 1.526$

$$\theta_L = \frac{36 \times 0.634}{66 + 36 \times 0.634} = \frac{22.82}{88.82} \approx 0.257$$

$$\theta_U = \frac{36 \times 1.526}{66 + 36 \times 1.526} = \frac{54.94}{120.94} \approx 0.454$$

95% Credible Interval is (0.257, 0.454)


Posterior Probability: $P(\theta > 0.62 \mid X)$

Convert $\theta = 0.62$ to the F scale:

$$F^* = \frac{66 \times 0.62}{36 \times (1 – 0.62)} = \frac{40.92}{13.68} \approx 2.993$$

From F-table: $P(F(72, 132) > 2.993)$ — since $F_{0.995}(72, 132) \approx 2.00$ and $F_{0.999}(72,132) \approx 2.50$, a value of 2.993 lies well in the extreme upper tail.

$$P(\theta > 0.62 \mid X) \approx 0.001 \quad \text{(negligible)}$$

This is expected — the posterior is centred around 0.35 with SD $\approx 0.047$, so $\theta = 0.62$ is more than 5 SDs above the mean.


Posterior Probability: $P(\theta > 0.40 \mid X)$

Convert $\theta = 0.40$:

$$F^* = \frac{66 \times 0.40}{36 \times 0.60} = \frac{26.4}{21.6} \approx 1.222$$

From F-table: $P(F(72, 132) > 1.222) \approx 0.17$

$$P(\theta > 0.40 \mid X) \approx 0.17$$

Posterior Summaries Across Different Priors

Data: $n = 100$, $x = 35$ (Binomial)


Posteriors

PriorBeta ParametersPosterior
Uniform Beta(1,1)$a=36,\ b=66$Beta(36, 66)
U-shaped Beta(0.5,0.5)$a=35.5,\ b=65.5$Beta(35.5, 65.5)
Bell Beta(5,5)$a=40,\ b=70$Beta(40, 70)
Skewed Beta(2,8)$a=37,\ b=73$Beta(37, 73)

F-Distribution Degrees of Freedom: $(2a,\ 2b)$

Prior$d_1 = 2a$$d_2 = 2b$
Uniform72132
U-shaped71131
Bell80140
Skewed74146

Formulas Used

$$E(\theta) = \frac{a}{a+b}, \qquad \text{Var}(\theta) = \frac{ab}{(a+b)^2(a+b+1)}, \qquad \theta = \frac{a \cdot F}{b + a \cdot F}$$


Comparative Posterior Summary Table

SummaryUniform Beta(36,66)U-shaped Beta(35.5,65.5)Bell Beta(40,70)Skewed Beta(37,73)
Mean$\frac{36}{102} = 0.3529$$\frac{35.5}{101} = 0.3515$$\frac{40}{110} = 0.3636$$\frac{37}{110} = 0.3364$
Variance$0.002217$$0.002211$$0.002207$$0.002031$
SD$0.0471$$0.0470$$0.0470$$0.0451$
5th Percentile $(\theta_{0.05})$$0.278$$0.277$$0.288$$0.261$
95th Percentile $(\theta_{0.95})$$0.434$$0.433$$0.443$$0.415$
2.5th Percentile $(\theta_{0.025})$$0.257$$0.256$$0.267$$0.241$
97.5th Percentile $(\theta_{0.975})$$0.454$$0.453$$0.464$$0.435$
90% Credible Interval$(0.278,\ 0.434)$$(0.277,\ 0.433)$$(0.288,\ 0.443)$$(0.261,\ 0.415)$
95% Credible Interval$(0.257,\ 0.454)$$(0.256,\ 0.453)$$(0.267,\ 0.464)$$(0.241,\ 0.435)$
$P(\theta > 0.40)$$\approx 0.170$$\approx 0.172$$\approx 0.224$$\approx 0.121$
$P(\theta > 0.62)$$\approx 0.001$$\approx 0.001$$\approx 0.001$$\approx 0.001$

Key Observations

  • The U-shaped prior (Case 2) has negligible effect — with $n=100$ observations, even an extreme prior is overridden. Summaries are nearly identical to the Uniform case.
  • The Bell prior (Case 3) pulls the posterior mean slightly upward ($0.3636$) toward its prior belief of $0.50$, and widens the upper credible bound marginally.
  • The Skewed prior (Case 4) pulls the posterior mean slightly downward ($0.3364$), reflecting its pessimistic prior belief, and shifts the credible interval leftward.
  • All four posteriors agree: $P(\theta > 0.62) \approx 0$, confirming that a pass rate above 62% is essentially ruled out by the data.
  • With 100 observations, the prior choice matters only modestly — the data dominates in all cases.

Key Takeaways

  • The posterior is the complete basis for all Bayesian inference — everything flows from it.
  • Likelihood is not a probability over $\theta$; it ranks parameter values by how well they explain the data.
  • Posterior summaries — mean, SD, credible intervals, tail probabilities — are only as tractable as the form of the posterior.
  • When the posterior is a known theoretical distribution (Closed form), all summaries follow analytically.
  • The Beta–F relationship is the key that unlocks percentiles and credible intervals directly from standard statistical tables.
  • Four different priors, same data — the summaries barely moved. Data dominates when $n$ is large.
  • The prior is your starting point. The data updates it. The posterior is the updated state of knowledge about $\theta$.

Scroll to Top