Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

How does evidence change a belief in a parameter?

Suppose you hand me a coin and I want to know its bias. I could flip it a hundred times and report the fraction of heads. That is an estimate, and it is a good one, but it is a single number with no statement attached about how wrong it might be. Worse, before the first flip I have nothing at all — the likelihood of $n=0$ observations is flat, so it ranks every value of $\theta$ equally and refuses to choose.

What I usually have is not nothing. I have seen coins before, I can feel their weight, I know that a coin manufactured to be fair is the default. That knowledge is a prior distribution over $\theta$: not a guess at one value but a spread of plausibility across all of them, high near one half and tapering toward the absurd extremes. Then I flip, and each flip is evidence. The likelihood tells me how well each candidate $\theta$ explains the flips I saw; the prior tells me how plausible each candidate was before. Multiplying them and renormalising gives the posterior: what I should believe now.

The output is a distribution, not a number. That is the whole point. From a posterior over $\theta$ I can read a best guess (its mean or its peak), an uncertainty (its spread, or an interval that holds 95% of its mass), and the probability of any statement I care about — that the coin is biased toward heads, that it is within a cent of fair, that two coins differ. None of those questions even parse if all I have is a point estimate.

There is also a practical delight hiding here. For the coin, the posterior has a shape from the same family as the prior, so updating is arithmetic rather than integration: heads add to one exponent, tails add to the other, and the distribution stays a Beta. That property is called conjugacy, and it is why this pair is the standard first example of Bayesian inference. It also makes streaming trivial — today's posterior is tomorrow's prior, one flip at a time, and the answer does not depend on the order the flips arrived.

💡 By the end of this part you'll see why the posterior is a distribution over $\theta$ rather than a likelihood in disguise, how the conjugate update $a\to a+k$ and $b\to b+n-k$ turns each flip into an increment, and why enough data makes the choice of prior stop mattering.
2

Bayes' rule for a parameter

Likelihood times prior, then renormalise

Write $\theta$ for the unknown bias and $x$ for the data — here a sequence of heads and tails. Two ingredients go in. The likelihood $p(x\mid\theta)$ says how probable the observed data would be if the bias really were $\theta$. The prior $p(\theta)$ says how plausible each $\theta$ was before any data. Bayes' rule multiplies them and divides by a constant that makes the result integrate to one.

$$p(\theta\mid x)=\frac{p(x\mid\theta)\,p(\theta)}{p(x)},\qquad p(x)=\int p(x\mid\theta)\,p(\theta)\,d\theta,\qquad\text{so}\qquad p(\theta\mid x)\;\propto\;p(x\mid\theta)\,p(\theta).$$

The denominator $p(x)$ is the evidence or marginal likelihood: the probability of the data averaged over the prior. It does not depend on $\theta$ at all, so once the data are in hand it is just a number, and the shape of the posterior is entirely the product of the other two factors. Computing it is the hard part in general, which is exactly why the conjugate cases matter so much.

Compare this with the likelihood from the previous part. The likelihood $p(x\mid\theta)$ is a function of $\theta$, but for a continuous parameter it is not a density over $\theta$: its integral over $\theta$ has no reason to be one, and the maximum-likelihood estimate is just its peak. The prior is a density over $\theta$, and dividing by the evidence transfers that property to the posterior. That is the formal difference the last part warned about: $p(x\mid\theta)$ is evidence about $\theta$, while $p(\theta\mid x)$ is belief about $\theta$. Only the second one can be integrated to give probabilities of statements about the parameter.

Read the product picture literally and the intuition is immediate. The prior puts its mass where $\theta$ is plausible in advance; the likelihood puts its mass where $\theta$ explains the flips. Where both are large, the product is large. The posterior is their overlap, renormalised — a compromise that leans on the prior when data are scarce and on the likelihood when data are plentiful. The whole rest of this part is that sentence, made quantitative for the coin.

3

Conjugate priors: the Beta–binomial model

A prior family that survives the update unchanged

A prior family is conjugate to a likelihood when the posterior lands back in the same family. For the coin it does. Put a Beta prior on $\theta$, multiply by the Bernoulli/binomial likelihood of heads and tails, and the exponents simply add. The Beta has two shape parameters, $a$ and $b$, and its density on $[0,1]$ is

$$\text{Beta}(\theta;a,b)=\frac{\theta^{\,a-1}(1-\theta)^{\,b-1}}{B(a,b)},\qquad B(a,b)=\frac{\Gamma(a)\,\Gamma(b)}{\Gamma(a+b)},\qquad \theta\in[0,1].$$

Think of $a$ and $b$ as imaginary heads and tails. Setting $a=b=1$ gives the flat uniform density: you have seen no flips and every bias is equally plausible. Raising $a$ pulls mass toward one; raising $b$ pulls it toward zero; raising both makes the distribution taller and narrower, as if you had seen many flips already and were confident. The prior mean is $a/(a+b)$, and the quantity $a+b$ is the prior strength: the number of real observations the prior is worth.

Now observe $k$ heads in $n$ flips, so the likelihood is proportional to $\theta^{\,k}(1-\theta)^{\,n-k}$. Multiplying is just adding exponents, which gives the heart of the whole part:

$$p(\theta)\sim\text{Beta}(a,b),\qquad k\text{ heads in }n\text{ flips}\quad\Longrightarrow\quad p(\theta\mid x)\sim\text{Beta}(a+k,\;b+n-k).$$

The update is literally counting. Every head increments $a$; every tail increments $b$. The posterior mean inherits a beautifully interpretable form, a weighted average of the prior mean and the sample proportion with weights equal to their strengths:

$$\mathbb{E}[\theta\mid x]=\frac{a+k}{a+b+n}=\underbrace{\frac{a+b}{a+b+n}}_{\text{prior weight}}\cdot\frac{a}{a+b}\;+\;\underbrace{\frac{n}{a+b+n}}_{\text{data weight}}\cdot\frac{k}{n}.$$

With $a+b$ small relative to $n$, the data weight dominates and the posterior mean is essentially the observed frequency. With few flips, the prior holds its ground. The demo below makes both regimes visible at once. Drag the two handles on the prior to move its mean and to widen or narrow it, set how biased the coin really is, then stream in flips and watch the posterior sharpen.

Drag the prior mean handle left and right, and the spread handle to widen or narrow the prior. The dashed blue curve is the prior, the grey curve is the likelihood of the flips so far (drawn to the same peak height so the shapes can be compared), and the bold pink curve is the posterior. The shaded band is the middle 95% of the posterior.

Watch the three curves. At zero flips the likelihood is flat and the posterior is the prior exactly — with no evidence, belief is unchanged. After a handful of flips the grey likelihood has a clear peak, the posterior has shifted toward it and narrowed, and its mean has been dragged partway from the prior mean. After a few dozen flips the prior's pull is small and the posterior sits almost on the observed frequency. The credible interval readout shrinks with each flip, and that shrinkage is the honest version of what a standard error was trying to say.

4

Streaming updates and prior sensitivity

Today's posterior is tomorrow's prior

Bayes' rule does not care whether the data arrive all at once or one at a time. Apply it repeatedly, treating each posterior as the prior for the next observation, and the product telescopes:

$$p(\theta\mid x_1,\dots,x_n)\;\propto\;p(x_n\mid\theta)\,p(x_{n-1}\mid\theta)\cdots p(x_1\mid\theta)\,p(\theta).$$

Because multiplication commutes, the order of the flips is irrelevant; the running posterior after $n$ flips depends only on $k$, the number of heads. For the Beta this is especially transparent: each head does $a\to a+1$ and each tail does $b\to b+1$, so streaming the data is exactly the conjugate update read one step at a time. There is no need to store history beyond the two counters, which is why conjugate models are the workhorse of online settings; the posterior is a sufficient statistic in two numbers.

How much does the starting point matter? The posterior variance shows the tug of war. For $\text{Beta}(a+k,\,b+n-k)$ it is

$$\operatorname{Var}[\theta\mid x]=\frac{(a+k)(b+n-k)}{(a+b+n)^2\,(a+b+n+1)},$$

which shrinks like $1/n$ as flips accumulate. The prior enters only through the fixed offsets $a$ and $b$, while the data enter through growing $k$ and $n-k$. So the prior's influence is a constant that gets divided by a growing sample: two people who start with wildly different beliefs, but who agree on the model, end up with posteriors that agree more and more closely as they see the same data. That is the phenomenon the second demo exposes. It is also the first appearance of a theme that Part 3 makes precise when it asks whether an estimator is any good: with enough data the likelihood dominates, and the answer stops depending on where you started.

The same stream of flips is replayed under four different priors. Move the observation slider from zero upward: at zero the four curves are the priors, and by the right-hand end they have collapsed onto one another and onto the dashed true θ.

Early on, the priors disagree loudly. A prior biased toward tails puts its mass low and needs real evidence to move; a prior already convinced the coin is very biased toward heads barely budges until it is contradicted. But contradictions accumulate, and the shared data win. Watch the four means in the readout close in on the observed frequency, and note that it is the likelihood — identical for all four — that does the winning. Prior choice is a statement about how much you knew before, and its cost is bounded by exactly that: a fixed number of pseudo-counts that real data dilute away.

5

Where this shows up

Beta–binomial in the wild

AI / ML

Calibrated model outputs

A model's stated probability can be read as a belief about a rate, and smoothing that belief with a Beta prior is a standard defence against overconfidence on small counts. That perspective is what calibration and evaluation measure, and what proper scoring rules such as log loss grade.

AI / ML

Preferences as evidence

Preference learning in RLHF is Bayesian in spirit: a prior over what a good response looks like, updated by human comparisons, with the posterior standing in for the reward signal. Conjugate updates are the tractable special case; general posteriors need the sampling machinery of Part 11.

Math

The general theorem

The coin is one instance of the rule that flips the conditioning in Bayes' theorem: $p(\theta\mid x)\propto p(x\mid\theta)p(\theta)$ applies to any unknown quantity, not just a bias. The Beta is simply the case where the algebra stays closed-form.

Math

What the likelihood could not do

The previous part built the likelihood surface and showed why it is not a distribution over the parameter. Everything here is the repair: multiply by a prior, renormalise, and the same surface becomes a belief you can integrate.

6

Cheat sheet

Every formula in one place

IdeaFormulaReading
Bayes for a parameter$p(\theta\mid x)\propto p(x\mid\theta)\,p(\theta)$Posterior is proportional to likelihood times prior.
Evidence$p(x)=\int p(x\mid\theta)\,p(\theta)\,d\theta$The normalising constant; the average likelihood under the prior.
Beta density$\theta^{a-1}(1-\theta)^{b-1}/B(a,b)$Two exponents; $a=b=1$ is uniform.
Conjugate update$\text{Beta}(a,b)\to\text{Beta}(a+k,\,b+n-k)$Heads add to $a$, tails add to $b$.
Posterior mean$(a+k)/(a+b+n)$Prior mean and sample proportion blended by strength.
Prior strength$a+b$How many observations the prior is worth.
Posterior variance$\frac{(a+k)(b+n-k)}{(a+b+n)^2(a+b+n+1)}$Shrinks like $1/n$; the prior is a fixed offset.
Credible interval$[q_{0.025},\,q_{0.975}]$Holds 95% of the posterior mass for $\theta$.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered