Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

How do you report uncertainty?

Suppose you poll n=60 people and k=42 say yes. The obvious estimate is the sample proportion $\hat p = 42/60 = 0.70$. That number is the maximum-likelihood estimate of the true yes-rate $\theta$, and Part 1 derived exactly why: it is the value that makes the observed data most probable. What it is not is a claim that $\theta = 0.70$. Draw a different sixty people and you get 0.68, or 0.73, or 0.65. The estimate has a distribution, and a single draw from it carries no error bar of its own.

Part 3 built the machinery for describing that distribution: the sampling distribution of an estimator, its bias, its variance, and the Fisher information that sets a floor on how tight it can be. The variance of $\hat p$ is $\theta(1-\theta)/n$, and its square root is the standard error. That single quantity is the raw material of every interval in this part. An interval is nothing more than an estimate plus a multiple of its standard error, and the only interesting question is what the multiplier and the interval mean.

Here is where the two schools split. The frequentist says: treat the parameter as fixed and unknown, treat the data as random, and build an interval whose procedure covers the fixed truth with a stated long-run frequency. The Bayesian says: treat the parameter as a random variable with a prior, condition on the data, and read the interval straight off the posterior distribution. Both produce two numbers, a and b, and both will call [a,b] a 95% interval. The sentences that follow are not the same sentence.

This matters far beyond philosophy. A medical trial that reports a 95% confidence interval and a reader who interprets it as "there is a 95% chance the effect lies in here" are using two incompatible logics, and the second one is usually wrong. A robot localisation system that keeps a Gaussian belief over its pose is using the Bayesian object whether or not its author would say so. Getting the distinction straight is what lets you choose the right tool, and read the tool you were handed.

💡 By the end of this part you'll see why a confidence interval is a statement about a procedure while a credible interval is a statement about the parameter, how the two are computed from the same estimate and standard error, and why they usually agree numerically while meaning entirely different things.
2

Frequentist confidence intervals

Estimate, standard error, multiplier

The recipe is always the same three ingredients. Take an estimator $\hat\theta$, find its standard error $\operatorname{se}(\hat\theta)$ — the standard deviation of its sampling distribution — and multiply the standard error by a constant that depends only on the confidence level you want. For an estimator that is approximately normal, that constant is the standard normal quantile.

$$\text{CI}_{1-\alpha}=\hat\theta \pm z_{1-\alpha/2}\,\operatorname{se}(\hat\theta),\qquad z_{0.975}=1.96,\qquad z_{0.95}=1.645,\qquad z_{0.995}=2.576.$$

For a proportion the standard error is the plug-in quantity $\operatorname{se}(\hat p)=\sqrt{\hat p(1-\hat p)/n}$. Plugging in $\hat p$ for the unknown $\theta$ is itself an approximation, and it is the one that gives the Wald interval. For n=60 and $\hat p=0.70$ the standard error is $\sqrt{0.70\cdot 0.30/60}\approx 0.0592$, so the 95% Wald interval is $0.70 \pm 1.96\cdot 0.0592$, roughly [0.584, 0.816]. The interval is wide because sixty people is not very many, and the width scales like $1/\sqrt n$: to halve it you need four times the data.

For a mean the same shape holds, with the sample standard deviation in place of the Bernoulli standard error. When n is small and the population standard deviation is itself estimated, the normal quantile is replaced by a t quantile with n-1 degrees of freedom, which fattens the interval just enough to pay for estimating the scale. As n grows the t distribution collapses onto the normal one, and the two intervals become indistinguishable. This is the interval you meet in every introductory course: $\bar x \pm t_{n-1,1-\alpha/2}\,s/\sqrt n$.

The crucial structural fact is which quantities are random. In the frequentist picture $\theta$ is a fixed constant. It has no distribution, so you cannot attach a probability to it being in a particular place. The endpoints $\hat\theta \pm z\,\operatorname{se}(\hat\theta)$, by contrast, are functions of the data and therefore random: every new sample slides the interval to a new location and stretches it to a new width. The randomness lives in the interval, not in the parameter. That is the whole content of the next section, and it is why the correct statement is a statement about a procedure rather than about a single [a,b].

$$\text{Frequentist:}\quad P_\theta\big(\theta\in[\hat\theta-z_{1-\alpha/2}\operatorname{se}(\hat\theta),\;\hat\theta+z_{1-\alpha/2}\operatorname{se}(\hat\theta)]\big)=1-\alpha.$$

The subscript on $P_\theta$ is doing real work. The probability is computed over the sampling distribution of the data for a fixed $\theta$; it is a statement about how often the random interval traps that fixed value. Part 3's Cramér–Rao bound told you the smallest standard error any unbiased estimator can achieve, so a confidence interval built on an efficient estimator is as short as the frequentist logic allows. The estimators part is where that floor was derived.

3

What "95% confidence" really means

A property of the procedure, not of one interval

Say it precisely. Before you collect data, the interval you are about to build is a random object. The claim is that the procedure that builds it produces an interval containing $\theta$ in 95% of all possible samples. After you collect the data and compute one interval, that particular interval either contains $\theta$ or it does not; there is no probability left to speak of. "95% confidence" is a frequency attached to the recipe, like the claim that a factory's bolts are within tolerance 95% of the time. It is not a belief about the bolt in your hand.

The picture below makes this concrete. We fix a true $\theta$, simulate a hundred independent datasets from it, and for each one compute a 95% Wald interval. Each interval is drawn as a horizontal bar, coloured blue if it covers the gold truth line and pink if it misses. About five should miss — not exactly five, because the count is itself random, but the misses cluster around the nominal failure rate. Press Re-simulate and the pattern redraws; the fraction that covers stays near 95% while the individual bars jump around.

Each row is one simulated dataset and its 95% interval for $\theta$. The gold dashed line is the true value. Bars that miss it are pink.

covers the truth misses true parameter

On the last dataset only, the posterior density of $\theta$ under a Beta(2,2) prior. The shaded band is the equal-tailed credible interval; the pink dashed lines are that same dataset's frequentist interval.

posterior density frequentist interval true parameter

Two honest caveats are visible in the animation if you look for them. First, the Wald interval is an approximation, and its real coverage can sit a little below nominal when n is small or $\theta$ is near the edges, because the normal approximation to the binomial is imperfect and $\hat p$ can hit exactly 0 or 1. Drag the trials slider down to 10 and the coverage count drops below the nominal line; that is the method failing, not the simulation. The Agresti–Coull interval, which adds two successes and two failures before computing the Wald interval, repairs most of this at almost no cost.

Second, the number of misses is random. If the true coverage is 95%, the count of misses in a hundred trials follows a binomial distribution with mean 5 and standard deviation about 2.2, so seeing 3 or 8 misses is not evidence that anything is broken. The confidence level is a property of the long-run frequency of the procedure, and a hundred intervals is only a small window onto that long run.

4

Bayesian credible intervals

Read the interval off the posterior

The Bayesian route starts from a prior and updates it with the likelihood, exactly as Part 2 did. For a proportion the conjugate pairing is the Beta prior with the binomial likelihood: place $\theta \sim \mathrm{Beta}(a_0,b_0)$, observe k successes in n trials, and the posterior is again a Beta distribution with its parameters bumped by the data.

$$\theta\sim\mathrm{Beta}(a_0,b_0),\qquad k\sim\mathrm{Binomial}(n,\theta)\quad\Longrightarrow\quad \theta\mid k \sim \mathrm{Beta}(a_0+k,\;b_0+n-k).$$

The prior Beta(2,2) is a gentle nudge away from the extremes: it behaves like having already seen two successes and two failures, which matters most when n is small and washes out as the data accumulate. The posterior mean is (a_0+k)/(a_0+b_0+n), a weighted compromise between the prior mean of one half and the sample proportion. For the last dataset in the simulation, with k successes in n trials, that is the number printed in the readout, and the density curve below the bars is the whole posterior, not just its centre.

A credible interval is any interval that contains 95% of the posterior probability. The simplest choice is the equal-tailed interval, cut at the 2.5% and 97.5% quantiles of the posterior: [q_{0.025}, q_{0.975}]. It is what the shaded band shows, and it is easy to compute because the Beta quantile function is a couple of lines of code. The alternative is the highest posterior density interval, the shortest interval containing 95% of the mass, which for a skewed posterior is not equal-tailed at all. When the posterior is roughly symmetric the two coincide, and for the bell-shaped Beta posteriors in this demo the difference is negligible.

$$\text{Credible:}\quad P\big(\theta\in[a,b]\mid x\big)=1-\alpha,\qquad [a,b]=[q_{\alpha/2},\,q_{1-\alpha/2}]\ \text{ of the posterior}.$$

Notice what that sentence says. The probability is over $\theta$, conditioning on the observed data x. There is no long run, no repeated sampling, no reference to a procedure. You can say, without translation, "given these data there is a 95% chance the true rate lies in this interval." That is exactly the sentence most people think a confidence interval licenses, and it is the reason credible intervals are so attractive to practitioners.

The same structure works for a mean. With a normal prior $\theta\sim\mathcal N(\mu_0,\tau^2)$ and data whose sample mean has standard error $\sigma/\sqrt n$, conjugacy gives a normal posterior whose precision is the sum of the prior and data precisions.

$$\theta\mid x \sim \mathcal N\!\left(\frac{\tau^{-2}\mu_0+\sigma^{-2}\bar x}{\tau^{-2}+\sigma^{-2}},\;\frac{1}{\tau^{-2}+\sigma^{-2}}\right).$$

This is the normal–normal model that Part 2 introduced and that Part 17's Kalman filter runs in a loop. Its posterior mean is a precision-weighted average of the prior mean and the data mean: more data, or a tighter prior, pulls the answer toward itself. The prior and posterior part builds this update from scratch, one observation at a time.

5

The comparison

Same numbers, different sentences

Put the two intervals from the demo next to each other. They are close, often overlapping almost completely, and for large n with a vague prior they are nearly identical. The frequentist interval is centred on $\hat p$ and has half-width $z\operatorname{se}(\hat p)$. The credible interval is centred near the posterior mean and has half-width set by the posterior standard deviation, which shrinks as prior precision plus data precision. When the prior is flat and n is large, the prior contributes nothing, the posterior standard deviation approaches the standard error, and the two intervals converge. Numerically similar does not mean logically similar.

The frequentist statement is about repetition: $P(\text{procedure covers }\theta)=0.95$. The credible statement is about this dataset: $P(\theta\in[a,b]\mid x)=0.95$. One quantifies the reliability of a method across hypothetical reruns; the other quantifies where the parameter plausibly sits given what you saw. If you have to make a decision now, with the data you have, the second is the statement you actually want, which is why Bayesian intervals feel so natural.

The price is the prior. The credible interval depends on a_0 and b_0, and a reader who disagrees with your prior should discount your interval accordingly. Drag the prior in Part 2 and you can watch the posterior shift; the confidence interval has no such dial, which is simultaneously its virtue (no prior to defend) and its limitation (no way to inject real prior knowledge). Frequentist intervals are also guaranteed to be calibrated: across repeated samples they cover at the advertised rate by construction, whereas a credible interval's frequentist coverage is only as good as the prior.

In practice the two agree whenever the data are informative enough to dominate the prior, and diverge when they are not. A trial with ten patients and a strong prior will produce a credible interval pulled toward the prior and a confidence interval that ignores it entirely. This is the same tension that evaluation faces when it asks whether a model's stated uncertainty is calibrated: calibration is a frequentist property, coherence is a Bayesian one, and a system can be excellent at one while failing the other. Report both if you can, and always say which sentence you mean.

$$\underbrace{P_\theta(\text{CI covers }\theta)=1-\alpha}_{\text{about the procedure}}\qquad\text{vs.}\qquad \underbrace{P(\theta\in[a,b]\mid x)=1-\alpha}_{\text{about the parameter}}$$
6

Where this shows up

Intervals under the machinery

Robotics

Covariance as a credible region

A robot's pose estimate comes with a covariance, and the region it implies is a credible set under the filter's posterior. In SLAM the loop closures either shrink that region or, if a closure is wrong, drag the truth outside it — an inconsistency you can only diagnose if you know the region is a belief, not a guarantee.

AI / ML

Error bars on benchmarks

A benchmark score is a mean over a finite sample, so its confidence interval decides whether a two-point win is real. A model that reports a probability should also be checked for calibration, which is the frequentist reading applied to its outputs.

Statistics

Estimator quality

The width of a confidence interval is a direct readout of an estimator's variance, and the estimators part showed that Fisher information sets a floor beneath it. Choosing a better estimator is the only way to shrink an interval without more data.

Bayesian inference

Posterior summaries

Every posterior distribution can be summarised by an interval, and the prior and posterior part is where those posteriors are built. When the posterior has no closed form, MCMC or a particle filter draws samples from it and the interval becomes a quantile of the sample.

7

Cheat sheet

Every formula in one place

IdeaFormulaIntuition
Confidence interval$\hat\theta \pm z_{1-\alpha/2}\operatorname{se}(\hat\theta)$Estimate plus a multiple of its standard error.
Standard error of a proportion$\sqrt{\hat p(1-\hat p)/n}$Width shrinks like $1/\sqrt n$.
Standard error of a mean$s/\sqrt n$, or t_{n-1} multiplier for small nUse the t quantile when the scale is estimated.
Common multipliersz_{0.95}=1.645, z_{0.975}=1.96, z_{0.995}=2.576Two-sided, so the tail is $\alpha/2$ on each side.
Frequentist statement$P(\text{CI covers }\theta)=0.95$The procedure covers in 95% of repeated samples.
Credible statement$P(\theta\in[a,b]\mid x)=0.95$Given this data, 95% of the posterior mass is in the interval.
Beta–binomial posterior$\theta\mid k\sim\mathrm{Beta}(a_0+k,\,b_0+n-k)$Prior successes and failures add to the data counts.
Equal-tailed interval$[q_{\alpha/2},\,q_{1-\alpha/2}]$ of the posteriorCut 2.5% from each tail; use HPD when skewed.
Normal–normal posterior$\theta\mid x\sim\mathcal N\!\big(\frac{\tau^{-2}\mu_0+\sigma^{-2}\bar x}{\tau^{-2}+\sigma^{-2}},\frac{1}{\tau^{-2}+\sigma^{-2}}\big)$Precision-weighted average of prior and data.
Coveragecount of intervals containing $\theta$ / totalRead it live in the simulator; it wobbles around the nominal level.
8

Further reading

Where to go deeper

9

Check your understanding

0/6 answered