Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

What is the 95% a probability of?

You draw a sample, compute a mean, and attach a margin of error. The result is a confidence interval. It has a lower bound and an upper bound, both of them numbers you can write down, and the recipe promises "95% confidence". Now ask the obvious question: 95% of what? The population mean is a fixed number — it is either inside your interval or it is not, with no probability about it. The interval, on the other hand, was computed from random data, so it is the interval that is random. Before you draw the sample, the interval is a random interval, and 95% is the chance that this random interval will contain the fixed mean. After you draw it, that chance has collapsed to a fact you cannot see. The percentage belongs to the recipe, not to the dish.

That single sentence is the whole subject of this part. Everything else — the formula, the t multiplier, the square-root-of-n width — is machinery that makes the sentence concrete. Confidence intervals also sit between the point estimate and the hypothesis test, carrying the same information dressed as a range instead of a verdict, and usually making the more honest presentation: a test says whether an effect is distinguishable from zero, while an interval says how large it might be and how much precision you bought.

💡 By the end of this part you'll see why a confidence interval is a random interval built by a fixed rule, why 95% counts how often the rule's output captures a fixed parameter over repeated experiments, and why "there is a 95% probability the mean is in this interval" is not what the number means.
2

A random interval

The recipe, and the one distribution it needs

Start with the simplest setting: independent observations drawn from a population with unknown mean μ. The sample mean x̄ estimates it, and the sample standard deviation s estimates the spread. The estimate has a standard error, the standard deviation of its sampling distribution, estimated as s/√n. If the population were exactly normal and s were the true σ, the standardised estimate would be standard normal and the multiplier for 95% would be 1.96. Because s is itself estimated from the same small sample, the multiplier is a little larger and the correct reference distribution is Student's t with n − 1 degrees of freedom:

$$\bar{x} \;\pm\; t^{*}_{\,n-1,\,1-\alpha/2}\;\frac{s}{\sqrt{n}}$$

Read the two pieces. The centre x̄ moves when the data move; the half-width t* s/√n moves too, because s is computed from the same draws. So the whole interval slides and stretches from sample to sample. What stays fixed is the rule: given a sample of size n and a chosen confidence level, produce this interval. That rule, applied over and over to fresh samples, produces a cloud of different intervals. The claim behind the confidence level is a claim about that cloud.

The demo below draws one experiment. The ticks are the sampled values, the solid line is the sample mean, the bar is the confidence interval, and the dashed line is the true population mean — the number the interval is trying to catch and which, in real life, you would never see. Press the button and the interval moves. Sometimes it brackets the dashed line, sometimes it does not, and you cannot tell from looking which case you are in. The next step resolves that by running the experiment many times.

One sample, its mean (solid), its confidence interval (bar), and the fixed population mean (dashed).

⚠ A wide interval is not a failure. It is an honest report that the data cannot pin the parameter down tightly. A narrow interval from a biased sample is far more dangerous than a wide one from a clean sample, because the width measures only random error. If the data are not representative, the interval can be narrow and wrong, and no confidence level will tell you.
3

A hundred experiments

Coverage, counted in front of you

Here is the experiment that makes the meaning of 95% undeniable. Instead of one sample, draw a hundred independent samples. For each, apply the same rule and draw its interval as a horizontal row. The vertical dashed line is the fixed population mean. Now simply count: how many rows cross that line? If the procedure deserves to be called 95%, then in a large batch of experiments the fraction crossing should be close to 95% — not because each interval has a 95% chance after the fact, but because the procedure was calibrated to hit that rate over repetitions.

Watch the count, not any single row. Individual intervals miss in both directions: one can land entirely above the truth or entirely below it. Nothing about a missed interval is contradictory, since a 95% procedure is supposed to miss about one in twenty. If the count were 100 out of 100 every time, the level would be lying, and a procedure that always covers is usually just absurdly wide.

Each row is one experiment: a sample mean with its confidence interval. Red rows are the ones that missed the true μ (dashed).

Change the level and run again. At 90% the bands are narrower and the misses are more frequent; at 99% the bands are wide and almost everything covers. The point is not that 99% is "better". A wider interval is a less informative one. Choosing a level is a trade between how often you want to be right and how much you are willing to say, the same trade that hypothesis testing makes with its significance threshold, in a later part.

Change the sample size too. Larger n pulls the intervals inward toward the truth, so the rows cluster tightly around the dashed line. Even with tight clustering the occasional row still misses: shrinking the width does not change the coverage rate, it changes how precisely each interval locates the same fixed target. Coverage stays at the nominal level by construction; precision is what improves.

4

More data, narrower interval

Width grows like 1/√n, and t* is the small-sample tax

The width of the interval is 2 t* s/√n, and almost all of its behaviour is the 1/√n. Quadruple the sample and the interval halves. To halve it again you must quadruple once more. This is the same square-root law that governs the standard error, and it is the single most useful fact for planning an experiment: you can predict, before collecting anything, roughly how wide the answer will be, and therefore how much data you need to make the interval narrow enough to be worth reporting.

The multiplier t* is the small-sample tax. For large n it settles toward the normal value of 1.96 at 95%; for n of five it is closer to 2.78, inflating the interval by more than forty percent. The reason is that a small sample can produce an unusually small s by luck, and an interval built from an underestimate of the spread would miss too often. The t distribution corrects for that by demanding a larger multiplier, an admission that the estimated spread is itself uncertain.

Sample size is not the only lever. The other is the population spread σ, which you cannot change: a noisier measurement process simply produces wider intervals. When you cannot collect more data and cannot reduce the noise, the only honest moves are to report the wide interval or to report a quantity that is measured more precisely.

The confidence level enters through the same multiplier: moving from 95% to 99% widens the interval by the ratio of the two t quantiles, on the order of thirty percent for moderate n. The higher level is not more truthful, just a different trade. Most fields settle on 95% by convention, and it is worth remembering that the number is a convention you can choose differently as long as you say so.

5

What it does not mean

The most common misreading, and the alternative it points to

The sentence to retire is: "there is a 95% probability that the true mean lies in this interval." It is tempting, it is how most people read a confidence interval, and it is not what the procedure delivers. The probability statement is about the interval, before it is observed. Once the numbers are on the page, the mean is either in the range or it is not; there is no remaining randomness to attach a probability to. A confidence interval is a statement about the long-run behaviour of a method, not a posterior belief about the parameter.

A confidence interval does not let you say "the effect is probably between these bounds". It lets you say that the method which produced these bounds covers the truth in 95% of experiments. If you want the other reading — a probability distribution over the parameter itself, so that "the effect is probably in this range" becomes legitimate — you need a different construction. A credible interval places a prior on the parameter, updates it with the data through Bayes' rule, and reports a range of the resulting posterior. Then the probability really is about the parameter, and the interval really is a statement of belief.

The two intervals can look identical, and often do when the data are plentiful and the prior is weak. That resemblance is a trap: it makes the frequentist interval look like it carries a Bayesian meaning that it does not. The honest position is that they answer different questions with different guarantees. A confidence interval controls how often a procedure captures a fixed unknown; a credible interval describes what you should believe about a random-seeming unknown given your prior and the data. A later part develops the Bayesian machinery and shows where the two coincide and where they part ways.

There is a second misreading worth naming: treating the interval as a range of "plausible values for the next observation". It is not. It is a statement about the mean, and a single new observation would scatter far more widely — that is a prediction interval, and its width means something different.

6

Where this shows up

Error bars are intervals, whether or not they say so

Robotics

Error bounds on a pose estimate

A robot estimating its pose reports a covariance, and the associated error ellipse is the multi-dimensional version of the interval on this page. When an odometry estimate accumulates drift and quotes a bound on position error, it is quoting the spread of a random estimate around a fixed true pose. The same coverage question applies: a claimed 95% bound is a statement about the estimator's behaviour over many runs, not a guarantee for the one drive you just completed.

ML / AI

Intervals on a benchmark score

A model's accuracy on a finite test set is an estimate of its error on a population of inputs. The evaluation chapter reports a single number; a confidence interval would report the number plus how much it would move on a different test set. That interval is the precision of the benchmark, and without it a difference of a fraction of a percent between two models is unreadable — it may be smaller than the interval the benchmark itself carries.

The pattern generalises. Every bar in an error-bar plot and every shaded band around a regression fit is an interval estimate, and its width is only meaningful once you know what it is an interval for and how it was calibrated. The bootstrap part builds intervals without a formula by resampling the data at hand; the regression part builds them around a fitted line; the Bayesian parts build the credible version. They differ in machinery, but all inherit the lesson of the hundred rows: the interval is random, the parameter is not.

Further reading

The references below treat the confidence interval as a statement about a procedure rather than a single range, which is the reading that survives contact with real data. If you take away one thing, take away the picture of a hundred intervals and a count of the misses.

Cheat sheet

TermMeaning here
Confidence intervalA random interval produced by a fixed rule from a random sample
CoverageThe fraction of repeated experiments whose interval contains the fixed parameter
The 95%A property of the rule over experiments, not of the single interval on the page
Half-width$t^{*}\,s/\sqrt{n}$ — the distance from the estimate to either end
WidthScales like $1/\sqrt{n}$; four times the data halves it
$t^{*}$The small-sample multiplier; larger than 1.96 and shrinking with $n$
Confidence levelBought with width: higher level, wider interval, same estimate
Credible intervalA Bayesian range with a genuine probability statement about the parameter
7

Check your understanding

0/4 answered