This Quick Bites is a beginner-friendly primer on Bayesian Inference, carefully scoped to build intuition before formalism. It opens with the core ideas — what Bayesian inference is, how Bayes’ Theorem works, and what prior and posterior distributions mean — before walking through the Beta-Binomial model as a concrete, closed-form example.
Key distinctions such as likelihood vs.probability, discrete vs. continuous updating, and informative vs. non-informative priors are addressed as standalone questions to avoid conceptual gaps.
The set closes with sequential updating, capturing one of the most elegant ideas in Bayesian thinking: that learning is continuous, and every posterior is simply the next prior.
1. What is Bayesian Inference?
Bayesian Inference is a statistical method for making inferences about unknown parameters using Bayes’ Theorem. It updates prior beliefs about parameters — expressed as a prior distribution — with observed data via the likelihood, yielding a posterior distribution that represents updated beliefs.
2. Bayes’ Theorem
Bayes’ Theorem is the foundation of Bayesian inference. It describes how to update the probability of a hypothesis given new data:
$$
P(\theta \mid X) = \frac{P(X \mid \theta)\, P(\theta)}{P(X)}
$$
$P(\theta \mid X)$ – Posterior – Updated belief after seeing data
$P(X \mid \theta)$ – Likelihood – Probability of the data given the parameter |
$P(\theta)$ – Prior – Initial belief about the parameter
$P(X)$ – Normalizing constant – Ensures the posterior is a valid probability distribution |
3. Prior and Posterior Distributions
A prior distribution encodes initial beliefs about a parameter before observing any data, drawing on expert knowledge, prior studies, or assumptions. Common choices include uniform priors, normal priors, and the Beta distribution.
A posterior distribution is the updated belief after combining the prior with observed data through Bayes’ Theorem. It reflects remaining uncertainty about the parameter and serves as the central output of Bayesian analysis.
4. The Beta Distribution as a Conjugate Prior
The Beta distribution is the standard prior choice for binomial proportions because it is a conjugate prior for the binomial likelihood — meaning the posterior is also a Beta distribution, keeping computation tractable. It is also naturally bounded on $[0, 1]$, making it well-suited for modeling probabilities and proportions.
5. Bayesian vs. Frequentist Inference
| Bayesian | Frequentist | |
|---|---|---|
| Parameters | Treated as random variables with distributions | Treated as fixed but unknown constants |
| Prior knowledge | Explicitly incorporated | Not used |
| Output | Posterior distribution | Point estimates and confidence intervals |
| Uncertainty | Quantified via the posterior | Expressed through repeated-sampling properties |
6. Interpreting the Posterior
The posterior distribution gives a complete picture of plausible parameter values after observing the data. Key summaries include:
- Point estimates — the posterior mean or mode
- Credible intervals — e.g., a 95% credible interval directly states that the parameter has a 95% probability of lying within that range, given the data and prior (unlike frequentist confidence intervals, which are statements about long-run sampling behavior)
7. Benefits of Bayesian Methods
- Incorporates prior knowledge — useful when data are scarce or prior research is available
- Quantifies uncertainty — the posterior distribution captures full parameter uncertainty, not just a point estimate
- Flexible — applicable to a wide range of models, including complex hierarchical structures
8. Prediction: Posterior Predictive Distribution
Bayesian methods support principled prediction via the posterior predictive distribution, which integrates over all possible parameter values weighted by their posterior probabilities. This yields not just a predicted value, but a full distribution over future observations, naturally encoding predictive uncertainty.
9. What is the Likelihood Function?
The likelihood function $L(\theta \mid X)$ measures how probable the observed data $X$ are for a given value of the parameter $\theta$. Crucially, it is viewed as a function of $\theta$, not of $X$ — the data are fixed and the parameter varies.
For a binomial experiment with $n$ trials and $k$ successes, the likelihood is:
$$
L(\theta \mid k) = \binom{n}{k} \theta^k (1-\theta)^{n-k}
$$
The likelihood is not a probability distribution over $\theta$ — it does not integrate to 1. Its role in Bayes’ Theorem is to transfer information from the data into the posterior.
10. How Does Bayes’ Theorem Work in the Discrete Case?
When $\theta$ takes a finite set of values, the posterior probability of each value is computed directly. For each candidate $\theta_i$:
$$
P(\theta_i \mid X) = \frac{P(X \mid \theta_i)\, P(\theta_i)}{\sum_j P(X \mid \theta_j)\, P(\theta_j)}
$$
The denominator is simply a normalizing sum — it ensures all posterior probabilities add to 1. This discrete form is the cleanest way to see Bayes’ Theorem in action: each hypothesis is re-weighted by how well it explains the data, relative to all other hypotheses.
11. How Does the Denominator Change in the Continuous Case?
When $\theta$ is continuous, the normalizing sum becomes an integral:
$$
P(\theta \mid X) = \frac{P(X \mid \theta)\, P(\theta)}{\int P(X \mid \theta)\, P(\theta)\, d\theta}
$$
This integral is often analytically intractable. Since it is just a constant with respect to $\theta$, it is common practice to work with the proportionality shortcut:
$$
P(\theta \mid X) \;\propto\; P(X \mid \theta)\; P(\theta)
$$
This reads: the posterior is proportional to the likelihood times the prior. The normalizing constant is recovered separately — or avoided entirely when a conjugate prior is used.
12. What are Informative, Weakly Informative, and Non-Informative Priors?
Priors are classified by how strongly they shape the posterior before data are seen.
- Informative prior — encodes strong, specific prior knowledge (e.g., $\text{Beta}(10, 2)$ asserts the proportion is likely high). Dominates the posterior when data are scarce.
- Weakly informative prior — provides gentle regularisation without imposing strong beliefs (e.g., $\text{Beta}(2, 2)$). A practical default in many applied settings.
- Non-informative prior — attempts to let the data speak freely, with minimal prior influence (e.g., $\text{Beta}(1, 1)$, the uniform distribution on $[0,1]$).
A key principle: as $n$ grows, the likelihood dominates and the posterior becomes increasingly insensitive to the choice of prior. With small samples, the prior matters considerably — making its specification a deliberate and transparent modelling decision.
13. What is Sequential Updating in Bayesian Inference?
One of the most powerful features of Bayesian inference is that today’s posterior becomes tomorrow’s prior. As new data arrive, beliefs are updated incrementally without reprocessing past observations.
For the Beta-Binomial model, if the current posterior is $\text{Beta}(\alpha, \beta)$ and a new batch of data yields $k$ successes in $n$ trials, the updated posterior is simply:
$$
\text{Beta}(\alpha + k,\quad \beta + n – k)
$$
This sequential coherence is a natural consequence of Bayes’ Theorem and reflects how rational belief revision works — each new observation refines the picture without discarding what was learned before.