Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Two intervals, one dataset, two meanings

Part 17 left you with a posterior distribution: everything the prior and the likelihood jointly say about an unknown parameter. That is a complete answer, and almost always an unusable one. A decision-maker wants an interval they can quote — "the rate is somewhere between 0.55 and 0.82" — and very often a single number to type into a report or a controller. Attaching an interval to an estimate is the first job of this part; picking the number is the second.

The interval comes in two families that look nearly identical on screen and are built from completely different logic. A credible interval is a Bayesian object. It is a subset of the parameter space that carries $1-\alpha$ of the posterior probability, so it is legitimate to say "there is a 95% probability the parameter lies in here, given what I saw." A confidence interval is a frequentist object. It is produced by a rule that, averaged over repeated experiments, captures the parameter $1-\alpha$ of the time. The 95% refers to the performance of the rule across experiments, not to the particular interval in front of you. Both are defensible, both are taught, and they answer different questions.

The single-number question is subtler than it looks. A skewed posterior — the normal case, not the exception — has three different natural centres: its mean, its median and its mode. They disagree, and which one you write down changes the number you report. Decision theory turns that choice from a matter of taste into a matter of cost: name what it costs to be wrong, and the best summary follows. An interval and a point estimate are both summaries, and both are choices.

💡 By the end of this part you'll see why a credible interval is a probability over the parameter while a confidence interval is a coverage promise over experiments, why the Bayesian can say "95% probability the parameter is here" and the frequentist cannot, and why the mean, the median and the mode are the Bayes-optimal reports under squared, absolute and all-or-nothing loss respectively.
2

A probability over the parameter

The credible interval is the posterior, readable

Bayesian inference treats the parameter $\theta$ as a random variable, so it can carry probability. Start from a prior $p(\theta)$, observe data $x$, and update to the posterior

$$p(\theta \mid x) \;=\; \frac{p(x \mid \theta)\,p(\theta)}{p(x)}.$$

A credible interval at level $1-\alpha$ is any interval $[l, u]$ that holds $1-\alpha$ of that posterior mass:

$$P\!\left(\theta \in [l, u] \mid x\right) \;=\; 1-\alpha.$$

Every term in that equation is a probability given the data you actually have, which is exactly why the Bayesian sentence — "there is a 95% posterior probability that the rate lies in this interval" — is not a slip of the tongue but the literal content of the formula. The interval is not unique, because many different intervals can carry the same mass. The two standard choices, equal-tailed and highest-density, are the subject of step 4; for now take the equal-tailed version, which chops $\alpha/2$ of the probability off each tail, so $l$ is the $\alpha/2$ posterior quantile and $u$ is the $1-\alpha/2$ quantile.

The demo below fixes a Beta prior and a batch of Bernoulli trials and computes both kinds of interval from that one dataset. The upper canvas is the Bayesian picture: the posterior density with its credible region shaded, and the highest-density band drawn as a second alternative. The lower canvas is the frequentist picture built by repeating the whole experiment many times. Drag the slider to change how many successes were seen, and press a level button to move from 90% to 95% to 99%; both intervals and the coverage count follow along.

Top: the Beta posterior over the success rate, with the equal-tailed credible interval shaded and the highest-density interval marked. Bottom: 100 repeated experiments, each producing one confidence interval; the dashed line is the true rate, red bars are the intervals that missed it.

Credible interval: a posterior probability. Confidence interval: a coverage rate of the procedure. Same data, different meanings.

The two intervals usually land close together, which is why the distinction is easy to miss. That closeness is a coincidence of this model and this prior: change the prior, or move to a model with a badly skewed posterior, and the credible interval tracks the shape of the posterior while the confidence interval does whatever its formula was built to do.

3

Coverage, not probability

The interval is random; the parameter is not

A frequentist confidence interval is built from a rule that takes the data and returns two numbers, $l(X)$ and $u(X)$. Both endpoints are random, because the data are random. The parameter $\theta$ is a fixed, unknown constant that does not move. The defining property of a $1-\alpha$ confidence procedure is

$$P_\theta\!\left(l(X) \le \theta \le u(X)\right) \;=\; 1-\alpha \quad \text{for every } \theta.$$

Read the subscript carefully: the probability is over the data $X$, computed at a fixed $\theta$. It says that if you ran the experiment over and over and applied the rule each time, then in $1-\alpha$ of the repetitions the resulting interval would contain the fixed truth. The bottom canvas shows exactly this. Each horizontal bar is one simulated experiment with its own interval; the dashed line is the true rate; the bars that miss it are drawn in the warning colour. At 95% roughly five bars in a hundred should miss, and the readout keeps the running count.

This is why the frequentist cannot say "there is a 95% probability that $\theta$ is in this interval." Once the data are in hand, the endpoints are just two numbers, and $\theta$ is a constant: either it lies between them or it does not, and there is no probability left to speak of. The 95% was a property of the rule before the data arrived, a long-run frequency, not a degree of belief about the realised interval. Saying otherwise is the single most common misreading in applied statistics, and it is not a quibble — it is the entire difference between the two columns.

The Bayesian is allowed the sentence the frequentist is not, and the reason is structural. A Bayesian puts a prior on $\theta$, which makes $\theta$ random; the posterior is that distribution conditioned on the data, so a credible interval is a genuine $1-\alpha$ probability event in it. The price is the prior: the posterior statement is only as good as the model that produced it. The frequentist avoids the prior and pays by being unable to condition on the observed interval.

⚠ A confidence interval can be 95% and useless. Coverage is an average over experiments, and it says nothing about the particular interval you got. A rule that produces a tiny interval 95% of the time and an enormous one 5% of the time still has 95% coverage, but a user who is handed the enormous one learns almost nothing. Coverage is a necessary property of a procedure, not a certificate that this estimate is precise.
4

Equal-tailed versus highest-density

Two ways to cut the same $1-\alpha$ of probability

An interval that carries $1-\alpha$ of a posterior is not unique, and the two sensible choices differ precisely when the posterior is skewed. The equal-tailed interval leaves $\alpha/2$ of the probability below it and $\alpha/2$ above, so its endpoints are the $\alpha/2$ and $1-\alpha/2$ quantiles. The highest-density interval, or HPD, is the shortest interval that carries $1-\alpha$ of the mass. Shortest is a strong condition: inside an HPD interval every point has higher posterior density than every point outside it, which is why an HPD region is also called a highest-posterior-density region.

For a symmetric posterior the two coincide exactly — there is no side to be lopsided on. For a skewed posterior they separate. The top canvas makes the difference visible: the shaded band is the equal-tailed interval, whose left end cuts into the thick part of the curve so that both tails are drained equally, and the dashed pair of lines is the HPD, which lets the thin right tail carry more of the excluded probability and keeps the interval tight around the bulk. When the posterior piles up near a boundary, as a Beta posterior does for rare events, the equal-tailed interval can extend far into a region of negligible density, while the HPD declines to.

Shortest sounds strictly better, but HPD intervals carry a quiet defect: they are not invariant under reparameterisation. Summarise a rate on the probability scale, summarise the same uncertainty on the log-odds scale, and the equal-tailed intervals will transform into one another — quantiles commute with monotone maps — while the HPD intervals will not. There is no free lunch in choosing a summary, which is the theme that runs into the next step.

For reporting, the practical rule is to say which interval you computed. "The 95% equal-tailed credible interval is $[0.55, 0.82]$" is complete; leaving out "equal-tailed" leaves the reader guessing.

5

Choosing a number is a decision

Loss, risk, and the Bayes action

Now shrink the interval to a point. You must write down one value $a$ to stand for the unknown $\theta$, and the honest way to choose it is to say what it costs to be wrong. That cost is a loss function $L(\theta, a)$: the penalty incurred when the truth is $\theta$ and you report $a$. Since $\theta$ is uncertain, you cannot compute the loss of a report directly; you average it over the posterior to get the risk,

$$R(a) \;=\; \mathbb{E}\!\left[L(\theta, a) \mid x\right] \;=\; \int L(\theta, a)\,p(\theta \mid x)\,d\theta,$$

and you report the value that minimises it. That minimiser is the Bayes action. Three standard losses each produce one of the three classic summaries, and the derivation is short enough to do in your head.

Under squared-error loss, $L(\theta,a) = (\theta-a)^2$, expanding the square splits the risk into a part that depends on $a$ and a part that does not:

$$R(a) \;=\; \operatorname{Var}(\theta \mid x) + \bigl(a - \mathbb{E}[\theta \mid x]\bigr)^2.$$

The variance is fixed, so the only way to lower the risk is to kill the squared term, which happens at $a = \mathbb{E}[\theta \mid x]$. Squared error is minimised by the posterior mean. Under absolute-error loss, $L(\theta,a) = |\theta-a|$, the derivative of the risk is $\tfrac{d}{da}R(a) = P(\theta < a \mid x) - P(\theta > a \mid x)$, which is zero exactly when the two sides carry equal probability — that is, at the posterior median. Under all-or-nothing loss, which gives zero when you are exactly right and a constant otherwise, the risk is minimised by the single most probable value, the posterior mode.

Pick a loss below and watch the posterior answer change without the data changing at all. The curve is a deliberately skewed Beta posterior; the three candidate reports are the mean, the median and the mode. The dashed lines mark all three, the thick line is the Bayes action for the chosen loss, and the readout lists the risk of every candidate so you can see that the selected one really is the cheapest.

A skewed posterior with its mean, median and mode marked. The thick line is the report that minimises the selected loss; the shaded window is the tolerance used by the all-or-nothing loss.

Squared error → mean, absolute error → median, 0/1 loss → mode. The data are identical; only the cost of error changed.

The all-or-nothing case deserves its caveat. For a continuous posterior the probability of hitting the parameter exactly is zero, so a literal $L=0$ when correct makes the risk one for every report and the problem vacuous. The demo therefore uses the standard fix: count a report as correct if it lands within a window of width $2\epsilon$, and the mode becomes the best place to centre that window. "The mode is optimal under 0/1 loss" presupposes a notion of being close enough, and you should say what it is.

Once loss is on the table, the summaries stop looking canonical. If over-reporting a rate is three times as expensive as under-reporting it — a chemotherapy dose, a stress test threshold — the optimal report is no longer the mean, median or mode but a quantile of the posterior, pulled to the safe side by the asymmetry. The mean is the right answer only to the question "what is the fair number under squared error." A point estimate is a decision rather than a mathematical given: it encodes the loss, whether you chose it deliberately or by habit.

6

Where this shows up

Uncertainty you have to act on

Robotics

Reporting a pose with uncertainty

A pose estimate is never a point; it is a distribution, usually a Gaussian with a covariance. Dead reckoning carries that distribution forward, and the SLAM and pose-graph formulations refine it. But a planner needs one pose to steer toward. Whether you feed it the mean, the median or a cautious quantile is a loss decision: if a collision costs far more than a detour, you report a pessimistic quantile of the position posterior rather than its mean, and the robot hugs the safe side of the corridor. The credible ellipse is the interval; the reported pose is the Bayes action.

ML / AI

Reporting a metric with a range

A benchmark score or a latency figure is an estimate with uncertainty. The evaluation chapter turns a finite test set into an accuracy number, and the serving metrics chapter reports latency under load. A credible range turns each into an interval you can quote, and the ship-or-not decision is a loss problem: deploying a model that regresses quality usually costs more than waiting for a small improvement, so the threshold should be set on the posterior, not on the point estimate. Where the cost is asymmetric, the serving policy is a quantile of the posterior, not its mean.

The pattern recurs across the site. A kinematics solver reporting a joint angle must decide which summary of a noisy estimate to command; an optimizer that stops on a noisy gradient is implicitly trading a loss against uncertainty; a probability calculation is only useful once someone converts a distribution to an interval or a point; and a calibration routine reporting a focal length is choosing, usually silently, a loss function for the sensor. The numerics part adds the reminder that even the endpoints are computed with finite precision.

The unifying move is to ask, whenever a single number is reported, what its loss was. If the answer is "we always report the mean", the loss was squared error, chosen by habit. That is fine, but it is a choice, and this part has put the choice back in your hands.

Further reading

The references below treat intervals and point estimates as two ends of one problem: how to summarise a distribution for someone who has to act on it. If you take away one thing, take away that every reported number is the solution to an optimisation, whether or not anyone named the objective.

Gelman and colleagues is the standard modern reference for credible intervals and their decision-theoretic framing; Berger is the compact account of loss, risk and the Bayes action; Robert shows why the choice of summary is never neutral; and Wasserman is the sharpest short treatment of what a confidence interval does and does not claim.

Cheat sheet

TermMeaning here
Credible intervalA set carrying $1-\alpha$ of the posterior probability; a probability statement about $\theta$ given the data
Equal-tailed intervalCuts $\alpha/2$ of the posterior probability from each tail; endpoints are quantiles
HPD intervalThe shortest interval with $1-\alpha$ posterior mass; inside density exceeds outside density, but it is not reparameterisation-invariant
Confidence intervalA rule whose intervals cover the fixed $\theta$ with probability $1-\alpha$ over repeated experiments
CoverageThe long-run fraction of experiments in which the rule's interval contains the truth
Loss $L(\theta,a)$The penalty for reporting $a$ when the truth is $\theta$
Risk $R(a)$The posterior-expected loss of the report $a$
Bayes actionThe report that minimises the risk
Squared error → mean$R(a) = \operatorname{Var}(\theta) + (a-\mathbb{E}\theta)^2$ is minimised at the posterior mean
Absolute error → medianThe risk's derivative balances the two tail probabilities, giving the posterior median
0/1 loss → modeThe most probable value minimises all-or-nothing loss (with a tolerance for a continuous posterior)
7

Check your understanding

0/4 answered