Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

What should you believe about a coin you have barely flipped?

Take a coin with an unknown bias: the probability $p$ that it lands heads. You cannot see $p$, but each flip is a noisy clue about it. The Bayesian answer to "what do I believe about $p$ now?" is a probability distribution over the possible values of $p$, and Bayes' rule turns the distribution you hold into the one you hold after the next observation. Writing the unknown as $\theta$ and the data so far as $y$:

$$\underbrace{p(\theta \mid y)}_{\text{posterior}} \;\propto\; \underbrace{p(\theta)}_{\text{prior}} \;\times\; \underbrace{p(y \mid \theta)}_{\text{likelihood}}$$

The prior says which values were plausible before the data. The likelihood says how well each value explains the data that actually arrived. Multiplying them gives the posterior, which is the prior rewritten in the light of the evidence; the normalising constant does not depend on $\theta$, so it cannot change the answer's shape. That is the whole of Bayesian inference, and this part is that one line made visible.

This part leans directly on the probability course. It assumes you are comfortable with a random variable that lives on $[0,1]$ with a density rather than a list of probabilities, and with the Beta family in particular. If "a distribution over a parameter" is new, the probability guide is the prerequisite; here we put that machinery to work.

One framing note. In the frequentist picture the parameter is fixed and only the estimate is random. Bayesian inference makes no claim about the world here: the coin may be fixed, but your belief about it is uncertain, and representing that belief as a random variable is honest. The randomness is in the state of knowledge, not in the coin.

💡 By the end of this part you'll see why Bayes' rule is multiplication followed by renormalisation, why a Beta prior with Binomial data stays Beta, why the posterior is the whole answer rather than a single number, and why the prior's pull fades as flips accumulate.
2

Start with a prior

A distribution over the unknown, chosen out in the open

The natural prior for a bias is $\mathrm{Beta}(a, b)$, whose density on $[0, 1]$ is proportional to $\theta^{\,a-1}(1-\theta)^{\,b-1}$. Its mean is $a/(a+b)$, so the parameters set both where the belief sits and how sharply: larger $a$ pulls toward one, larger $b$ pulls toward zero, and raising both keeps the mean roughly fixed while narrowing the curve. That is why $a$ and $b$ read as pseudo-counts: a $\mathrm{Beta}(a, b)$ prior behaves like having already seen about $a-1$ heads and $b-1$ tails. The flat $\mathrm{Beta}(1, 1)$ is the uniform distribution — no opinion — while $\mathrm{Beta}(30, 10)$ is a confident claim of a heads-heavy coin.

Choosing a prior is choosing, in public, what you expected. The demo lets you change it by hand. The two handles sit on the base line at the prior's 10% and 90% points: drag them apart for a vague, wide smear, together for a tall spike, or both to the right to bet on heads. The sliders set $a$ and $b$ directly. A different analyst would start elsewhere, but with enough data both converge to the same posterior, as step five shows.

3

Feed it flips

The likelihood rewrites the prior, one coin at a time

After $n$ flips with $s$ heads and $f = n - s$ tails, independence makes the likelihood a product, one factor per flip: $\theta$ for a head, $1-\theta$ for a tail, so $p(y \mid \theta) \propto \theta^{\,s}(1-\theta)^{\,f}$. Multiply by the Beta prior and the exponents add, giving $\theta^{\,a+s-1}(1-\theta)^{\,b+f-1}$ — exactly a $\mathrm{Beta}(a+s,\; b+f)$ density. This is conjugacy: prior and posterior share a family, and the update is just adding counts. A head adds one to $a$; a tail adds one to $b$.

Press Flip heads or Flip tails for a single observation, or Flip 10 to draw ten flips from a hidden coin whose true bias is $0.65$ — a coin you cannot see, as in reality. The solid curve is the posterior, the light curve is the prior you dragged, and the dashed curve is the likelihood scaled to its own peak, drawn only to show what the data alone point at. Watch the solid curve start near the prior and migrate onto the dashed hill.

Prior (light solid), likelihood scaled to its peak (dashed), posterior (solid). Drag the two handles along the base to reshape the prior; the vertical lines mark the prior and posterior means.

Two things to notice. The posterior mean is always a compromise between the prior mean and the data proportion $s/n$, weighted by how much each carries: with $a+b=4$ the prior is worth about four pseudo-flips, so one real flip barely moves it, while forty flips leave the prior a small voice. And the posterior is not only moving, it is narrowing — more data means a sharper curve, which is the Bayesian standard error appearing as the width of a distribution.

⚠ The prior matters most when data are few. With a confident prior and no flips, the posterior is the prior. That is not a bug: with almost no data, almost all belief comes from the prior, and an honest posterior says so. A sensible prior is also a restraint that keeps a tiny sample from producing an extreme conclusion.
4

The posterior is the whole answer

A distribution, not a number, with every summary you need

Bayesian inference hands back a distribution, which carries more than any single number. When you need one, you choose the summary that suits the question. The mean $a/(a+b)$ minimises squared error and is the natural choice for prediction; the mode $(a-1)/(a+b-2)$ is the most probable single value; the median splits the belief in half. The readout shows mean and mode side by side so you can see them differ whenever the posterior is skewed.

For an interval, the posterior gives a credible interval: the central 95%, read off as the 2.5% and 97.5% quantiles. Its meaning is literal — given the model and prior, there is a 95% probability the unknown lies in this range. A confidence interval instead describes the long-run behaviour of a random interval. The two often look alike numerically but say different things, and Part 18 sets them side by side.

The posterior also answers the question you care about next. The probability that the next flip is heads is the average of $\theta$ under the posterior — the posterior mean. This posterior predictive probability, shown in the readout, is $(a+s)/(a+b+n)$ and already blends prior and data in the right proportions. Reporting only the mean would discard the width, and the width is what tells you how much you actually know.

5

The prior fades, the data takes over

Bernstein–von Mises, and the prior as regularisation

Keep pressing Flip 10. As counts grow, the posterior concentrates around the observed proportion $s/n$ and the prior's contribution becomes negligible. This is general, not just a Beta trick: the Bernstein–von Mises theorem says that a posterior from $n$ independent observations approaches a normal centred at the maximum-likelihood estimate with variance equal to the inverse Fisher information — exactly the quantity Part 8 derived for the frequentist estimate. Bayesian and frequentist answers can differ wildly at small $n$ and are forced into agreement as $n$ grows. The prior contributes a fixed amount of information while the data contribute an amount growing with $n$, so the prior's relative weight falls like $1/n$.

That is why a prior is a head start rather than a rigged outcome. The hidden coin's bias is $0.65$; start from a prior insisting on $0.2$, hold Flip 10 for a while, and watch the posterior walk to the truth. Reset and flip once or twice, and see how much of the answer is still the prior talking.

From the other side, the prior is a regulariser. Adding pseudo-counts is the same kind of move as a penalty that stops an estimate running off when data are scarce; the ridge penalty in the least-squares part is that idea in another costume. The Bayesian version is a probability statement, so it comes with an interval; a penalty is a prescription without a distribution. With plenty of data the two coincide. The frequentist contrast in one sentence: maximum likelihood reports the single number $s/n$, while Bayesian inference replaces it with a curve, using a prior to say what it expected and taking on the obligation to defend that prior.

6

Where this shows up

The update rule, running under the hood

Robotics & Vision

Bayesian state estimation

A robot estimating its pose runs this update: the belief over the pose is the prior, a sensor reading contributes a likelihood, and the posterior becomes the new belief. The Kalman filter in the SLAM part is the Gaussian version of the canvas above — a Gaussian prior fused with a Gaussian likelihood, narrowed by each measurement.

ML / AI

Preference modelling

Preference models in the RLHF chapter are Bayesian at heart: a prior over which responses people prefer, a likelihood from the comparisons observed, and a posterior used to guide the policy. The tension you feel dragging the handles is the same one.

The update is universal. A held-out evaluation score is a likelihood on a benchmark, and reasoning about a model's true error means reasoning about a posterior over it — the evaluation chapter turns that into a confidence statement, and Part 18 makes it a credible one. Wherever there is a quantity you cannot observe and a clue you can, there is a prior, a likelihood and a posterior, and the arithmetic is the same addition of evidence you have been watching.

Further reading

These references treat Bayesian inference as a procedure for updating belief rather than a pile of formulas. If you take away one thing, take away the picture of a prior multiplied by a likelihood and renormalised into a posterior, with the counts adding up.

Wasserman gives conjugacy in a page; McElreath builds the coin example with priors defended as modelling choices; Gelman and colleagues is the standard graduate treatment.

Cheat sheet

TermMeaning here
Prior $p(\theta)$Belief before the data; a distribution, not a point
Likelihood $p(y\mid\theta)$How well each value of the unknown explains the data
Posterior $p(\theta\mid y)$The renormalised prior × likelihood; the whole answer
ConjugacyBeta prior + Binomial data → Beta posterior
$\mathrm{Beta}(a, b)$Density on $[0,1]$ with mean $a/(a+b)$; $a, b$ are pseudo-counts
Posterior mean$(a+s)/(a+b+n)$; minimises squared error and gives the predictive probability
Credible intervalCentral quantiles of the posterior; a probability statement about the parameter
Bernstein–von MisesAs $n$ grows the posterior becomes normal around the MLE; the prior's pull fades like $1/n$
7

Check your understanding

0/4 answered