Prior × likelihood → posterior
Bayes' rule changes your mind in proportion to the evidence. Start with a distribution over the unknown — the prior — multiply it by how well each value explains the data you saw — the likelihood — and renormalise. What comes out is the posterior: everything you should believe now, given what you believed before and what you have just seen. For a coin with unknown bias $p$ the arithmetic never leaves a single family: a Beta prior fed Binomial data returns a Beta posterior. That is conjugacy, and it lets this page draw the prior, the scaled likelihood and the posterior together, sliding and narrowing as each flip arrives.
The question
What should you believe about a coin you have barely flipped?
Take a coin with an unknown bias: the probability $p$ that it lands heads. You cannot see $p$, but each flip is a noisy clue about it. The Bayesian answer to "what do I believe about $p$ now?" is a probability distribution over the possible values of $p$, and Bayes' rule turns the distribution you hold into the one you hold after the next observation. Writing the unknown as $\theta$ and the data so far as $y$:
The prior says which values were plausible before the data. The likelihood says how well each value explains the data that actually arrived. Multiplying them gives the posterior, which is the prior rewritten in the light of the evidence; the normalising constant does not depend on $\theta$, so it cannot change the answer's shape. That is the whole of Bayesian inference, and this part is that one line made visible.
This part leans directly on the probability course. It assumes you are comfortable with a random variable that lives on $[0,1]$ with a density rather than a list of probabilities, and with the Beta family in particular. If "a distribution over a parameter" is new, the probability guide is the prerequisite; here we put that machinery to work.
One framing note. In the frequentist picture the parameter is fixed and only the estimate is random. Bayesian inference makes no claim about the world here: the coin may be fixed, but your belief about it is uncertain, and representing that belief as a random variable is honest. The randomness is in the state of knowledge, not in the coin.
Start with a prior
A distribution over the unknown, chosen out in the open
The natural prior for a bias is $\mathrm{Beta}(a, b)$, whose density on $[0, 1]$ is proportional to $\theta^{\,a-1}(1-\theta)^{\,b-1}$. Its mean is $a/(a+b)$, so the parameters set both where the belief sits and how sharply: larger $a$ pulls toward one, larger $b$ pulls toward zero, and raising both keeps the mean roughly fixed while narrowing the curve. That is why $a$ and $b$ read as pseudo-counts: a $\mathrm{Beta}(a, b)$ prior behaves like having already seen about $a-1$ heads and $b-1$ tails. The flat $\mathrm{Beta}(1, 1)$ is the uniform distribution — no opinion — while $\mathrm{Beta}(30, 10)$ is a confident claim of a heads-heavy coin.
Choosing a prior is choosing, in public, what you expected. The demo lets you change it by hand. The two handles sit on the base line at the prior's 10% and 90% points: drag them apart for a vague, wide smear, together for a tall spike, or both to the right to bet on heads. The sliders set $a$ and $b$ directly. A different analyst would start elsewhere, but with enough data both converge to the same posterior, as step five shows.
Feed it flips
The likelihood rewrites the prior, one coin at a time
After $n$ flips with $s$ heads and $f = n - s$ tails, independence makes the likelihood a product, one factor per flip: $\theta$ for a head, $1-\theta$ for a tail, so $p(y \mid \theta) \propto \theta^{\,s}(1-\theta)^{\,f}$. Multiply by the Beta prior and the exponents add, giving $\theta^{\,a+s-1}(1-\theta)^{\,b+f-1}$ — exactly a $\mathrm{Beta}(a+s,\; b+f)$ density. This is conjugacy: prior and posterior share a family, and the update is just adding counts. A head adds one to $a$; a tail adds one to $b$.
Press Flip heads or Flip tails for a single observation, or Flip 10 to draw ten flips from a hidden coin whose true bias is $0.65$ — a coin you cannot see, as in reality. The solid curve is the posterior, the light curve is the prior you dragged, and the dashed curve is the likelihood scaled to its own peak, drawn only to show what the data alone point at. Watch the solid curve start near the prior and migrate onto the dashed hill.
Prior (light solid), likelihood scaled to its peak (dashed), posterior (solid). Drag the two handles along the base to reshape the prior; the vertical lines mark the prior and posterior means.
Two things to notice. The posterior mean is always a compromise between the prior mean and the data proportion $s/n$, weighted by how much each carries: with $a+b=4$ the prior is worth about four pseudo-flips, so one real flip barely moves it, while forty flips leave the prior a small voice. And the posterior is not only moving, it is narrowing — more data means a sharper curve, which is the Bayesian standard error appearing as the width of a distribution.
The posterior is the whole answer
A distribution, not a number, with every summary you need
Bayesian inference hands back a distribution, which carries more than any single number. When you need one, you choose the summary that suits the question. The mean $a/(a+b)$ minimises squared error and is the natural choice for prediction; the mode $(a-1)/(a+b-2)$ is the most probable single value; the median splits the belief in half. The readout shows mean and mode side by side so you can see them differ whenever the posterior is skewed.
For an interval, the posterior gives a credible interval: the central 95%, read off as the 2.5% and 97.5% quantiles. Its meaning is literal — given the model and prior, there is a 95% probability the unknown lies in this range. A confidence interval instead describes the long-run behaviour of a random interval. The two often look alike numerically but say different things, and Part 18 sets them side by side.
The posterior also answers the question you care about next. The probability that the next flip is heads is the average of $\theta$ under the posterior — the posterior mean. This posterior predictive probability, shown in the readout, is $(a+s)/(a+b+n)$ and already blends prior and data in the right proportions. Reporting only the mean would discard the width, and the width is what tells you how much you actually know.
The prior fades, the data takes over
Bernstein–von Mises, and the prior as regularisation
Keep pressing Flip 10. As counts grow, the posterior concentrates around the observed proportion $s/n$ and the prior's contribution becomes negligible. This is general, not just a Beta trick: the Bernstein–von Mises theorem says that a posterior from $n$ independent observations approaches a normal centred at the maximum-likelihood estimate with variance equal to the inverse Fisher information — exactly the quantity Part 8 derived for the frequentist estimate. Bayesian and frequentist answers can differ wildly at small $n$ and are forced into agreement as $n$ grows. The prior contributes a fixed amount of information while the data contribute an amount growing with $n$, so the prior's relative weight falls like $1/n$.
That is why a prior is a head start rather than a rigged outcome. The hidden coin's bias is $0.65$; start from a prior insisting on $0.2$, hold Flip 10 for a while, and watch the posterior walk to the truth. Reset and flip once or twice, and see how much of the answer is still the prior talking.
From the other side, the prior is a regulariser. Adding pseudo-counts is the same kind of move as a penalty that stops an estimate running off when data are scarce; the ridge penalty in the least-squares part is that idea in another costume. The Bayesian version is a probability statement, so it comes with an interval; a penalty is a prescription without a distribution. With plenty of data the two coincide. The frequentist contrast in one sentence: maximum likelihood reports the single number $s/n$, while Bayesian inference replaces it with a curve, using a prior to say what it expected and taking on the obligation to defend that prior.
Where this shows up
The update rule, running under the hood
Bayesian state estimation
A robot estimating its pose runs this update: the belief over the pose is the prior, a sensor reading contributes a likelihood, and the posterior becomes the new belief. The Kalman filter in the SLAM part is the Gaussian version of the canvas above — a Gaussian prior fused with a Gaussian likelihood, narrowed by each measurement.
Preference modelling
Preference models in the RLHF chapter are Bayesian at heart: a prior over which responses people prefer, a likelihood from the comparisons observed, and a posterior used to guide the policy. The tension you feel dragging the handles is the same one.
The update is universal. A held-out evaluation score is a likelihood on a benchmark, and reasoning about a model's true error means reasoning about a posterior over it — the evaluation chapter turns that into a confidence statement, and Part 18 makes it a credible one. Wherever there is a quantity you cannot observe and a clue you can, there is a prior, a likelihood and a posterior, and the arithmetic is the same addition of evidence you have been watching.
Further reading
These references treat Bayesian inference as a procedure for updating belief rather than a pile of formulas. If you take away one thing, take away the picture of a prior multiplied by a likelihood and renormalised into a posterior, with the counts adding up.
Wasserman gives conjugacy in a page; McElreath builds the coin example with priors defended as modelling choices; Gelman and colleagues is the standard graduate treatment.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapter 11 — priors, posteriors and conjugacy.
- Richard McElreath, Statistical Rethinking, chapters 1–3 — the coin-flip update and the case for priors.
- Andrew Gelman et al., Bayesian Data Analysis, chapters 1–2 — the posterior as the object of inference, and the Beta–Binomial example.
- Grant Sanderson, "Bayes' theorem, the geometry of changing beliefs", 3Blue1Brown — the proportional update, drawn.
Cheat sheet
| Term | Meaning here |
|---|---|
| Prior $p(\theta)$ | Belief before the data; a distribution, not a point |
| Likelihood $p(y\mid\theta)$ | How well each value of the unknown explains the data |
| Posterior $p(\theta\mid y)$ | The renormalised prior × likelihood; the whole answer |
| Conjugacy | Beta prior + Binomial data → Beta posterior |
| $\mathrm{Beta}(a, b)$ | Density on $[0,1]$ with mean $a/(a+b)$; $a, b$ are pseudo-counts |
| Posterior mean | $(a+s)/(a+b+n)$; minimises squared error and gives the predictive probability |
| Credible interval | Central quantiles of the posterior; a probability statement about the parameter |
| Bernstein–von Mises | As $n$ grows the posterior becomes normal around the MLE; the prior's pull fades like $1/n$ |