Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

What does a cloud of estimates look like?

In probability you are handed a distribution and asked what a sample will look like. Here the direction reverses. A population has a distribution over individual values, summarised by numbers such as its mean and standard deviation. You draw n values from it and compute a single number from them — a mean, a median, a proportion. That number is a guess about a fixed parameter you cannot see.

But the guess depends on which values happened to arrive. Before the data are drawn, the estimate is a random variable: a function of random inputs, and therefore itself random. Its distribution is called the sampling distribution of the estimator, and it is a second, higher level of description. The population describes a single observation; the sampling distribution describes a summary of an entire sample. Confusing the two levels is the classic first mistake in statistics, and the stacked canvases below are there to keep them apart.

The sampling distribution is never observed directly. You have one sample, one estimate, one point. It is a thought experiment — what the estimate would do if the whole study were rerun — and all of inference is the art of reasoning about that unobserved distribution from the single point you can actually see. This part makes it visible by rerunning the study, over and over, on screen.

💡 By the end of this part you'll see why an estimator is a random variable with a distribution of its own, why that distribution is centred on the parameter but has a spread that is not the spread of the data, and why that spread — the standard error — is the quantity every interval and test is built from.
2

The sampling distribution

One estimate, over and over

The demo below runs the experiment continuously. About twelve times a second it draws a fresh sample of n values from the population and drops the sample mean into the growing histogram on the lower canvas. Early on the histogram is a lumpy stack of bars; after a few hundred draws it settles into a smooth shape centred near the population mean. That shape is the sampling distribution of the mean, built empirically rather than assumed.

Notice what it is not: it is not the population. The pdf on the upper canvas can be flat, skewed, or two-humped, but the histogram of the estimates is narrower and far closer to symmetric. Averaging cancels the independent surprises in the individual values, and what survives is concentration around the truth. Use the shape buttons to switch the population. The four shapes are chosen to share roughly the same mean and spread, so that the comparison is fair: the population's own silhouette changes completely, while the histogram of the estimator barely changes at all. That is a first look at the central limit theorem, which Part 5 turns into a precise statement.

Now change the sample size with the slider. Each sample's mean is built from more values, and the histogram of estimates tightens. Note carefully which width moves: not the population's, which is fixed, but the width of the distribution of the estimate. The dashed curve overlaid on the histogram is not fitted to the data. It is a normal curve whose width is fixed at the theoretical standard error $\sigma/\sqrt{n}$, where σ is the known population standard deviation. As estimates accumulate, the histogram grows into that curve. The readout compares the running sd of the collected estimates with $\sigma/\sqrt{n}$: two numbers obtained by completely different routes, converging in front of you.

The fainter histogram is the sample median. Each statistic has its own sampling distribution: here the medians are centred in the same place as the means but spread wider, because the median throws away how far the extreme values lie. Two rules of thumb about the same data can have two different reliabilities, which is why a standard error is always attached to a particular estimator and not to a dataset.

Top: the population pdf, the current sample, and the true mean. Bottom: every estimate drawn so far — sample means (solid) and sample medians (faint) — with a normal of width $\sigma/\sqrt{n}$ overlaid.

⚠ The estimator is random; the parameter is not. The dashed line marks the fixed truth. Every bar in the histogram is one perfectly legitimate answer computed from one perfectly legitimate sample, and they disagree only because the samples disagree. Watching the estimates scatter while the truth stands still is the whole picture of estimation in one frame.
3

The standard error

A name and a rate for the spread

The standard deviation of a sampling distribution has its own name: the standard error. It is the typical distance between an estimate and the parameter it is trying to recover. "Standard deviation" describes the spread of data; "standard error" describes the spread of an estimate. Same idea, one level up, and the distinction is what the previous section's two canvases were built to make.

For the sample mean of independent observations the standard error has a closed form. Variances of independent quantities add, so the variance of a sum of n observations is nσ2. Dividing by n to form the mean divides the variance by n2, leaving σ2/n, and taking the square root gives the rule:

$$\mathrm{SE}(\bar{x}) = \frac{\sigma}{\sqrt{n}}$$

The √n is the part that matters in practice. Halving the typical error requires four times the data; a hundred observations buy only ten times the precision of one. This is not a limitation of a clever method but a property of averaging, and nothing in the rest of the subject escapes it.

The curve is $\sigma/\sqrt{n}$. The dots are the sd of 400 simulated sample means at each sample size.

The demo measures the sampling distribution directly: at each sample size it runs the experiment 400 times and takes the standard deviation of the resulting means, then plots that against the curve. The dots land on the formula, and neither the formula nor the simulation was told anything about the other. In real work σ is unknown, so the sample standard deviation s is substituted and the extra uncertainty that creates is handled by the t distribution later in the series. When no formula exists at all — for a median, a correlation, a regression coefficient — the sampling distribution is estimated by resampling the data you have, which is the bootstrap.

4

Two spreads, easy to confuse

The sd of the data is not the sd of the estimate

A population with σ = 1 has individual values that typically land about 1 away from the mean. The mean of 100 such values typically lands about 0.1 away. The same word, "spread", is describing two numbers a factor of ten apart, and the two are connected only by the √n. The sd of the data answers "how variable is one observation?" The standard error answers "how variable is my summary?" Confusing them is the most common error in reading a result, because a large standard deviation sounds alarming while it may say nothing at all about the precision of the mean.

The two spreads also differ from estimator to estimator. For normal data the sample median's standard error is about 1.25σ/√n, a quarter wider than the mean's, so on that model the mean is the more efficient estimator. But efficiency is measured against a model, and real data are rarely so obliging. In a skewed population, or one punctuated by occasional wild outliers, the median is pulled around far less: its sampling distribution may be wider at the centre and yet have much lighter tails, which is frequently the better trade. The two histograms in the demo are exactly this tradeoff, drawn side by side.

There is a condition attached to every formula in this section. The σ/√n rule assumes the observations are independent. When they are clustered — several measurements from the same sensor, several frames from the same drive, several responses from the same person — the effective sample is smaller than n, the true standard error is larger than the formula predicts, and intervals built from it are too narrow. The sampling distribution still exists; only the shortcut of dividing by √n fails.

⚠ More data shrinks the random error, not the systematic one. A larger sample narrows the histogram of estimates around whatever the estimator is honestly targeting. If the data are a biased slice of the world, that narrower histogram is simply a more confident march toward the wrong number. Bias lives in the design; the standard error only measures the wobble that remains.
5

What it is for

Everything downstream is built on it

The sampling distribution is the centre of the subject because every tool that expresses uncertainty is a statement about it. A confidence interval is a range that covers the parameter in a stated fraction of imagined reruns; that fraction is a probability computed from the sampling distribution. A test compares an estimate with what the sampling distribution would predict under a null hypothesis, measured in units of the standard error.

Concretely, if the estimate is roughly normal with standard error SE, then the interval estimate ± 1.96 × SE covers the truth in about 95% of reruns, and the statistic (estimate − hypothesised value)/SE reports the surprise in standard units. The familiar t and z machinery is nothing more than the sampling distribution of a particular estimator with its shape and width written down. Without the sampling distribution there is no scale on which to say that an estimate is far from a claim, because "far" is meaningless until you know how much scatter to expect.

When no formula is available, the sampling distribution can be measured directly by resampling the data you have: draw new samples from the sample, compute the estimate each time, and use the resulting spread as the standard error. That is the bootstrap, and it is this same picture with the population replaced by the best stand-in available. It is why the picture matters even where the algebra runs out.

The probability course's chapter on distributions supplies the raw material — normal, exponential, gamma — that the sampling distribution is assembled from. The same reasoning appears in least squares, where the standard errors of fitted coefficients are read off the diagonal of a covariance matrix, and in odometry, where a robot accumulates many noisy measurements and needs to know how precisely it has localised.

6

Where this shows up

A random sample of methods that run this loop

Robotics & Vision

Estimators inside RANSAC

RANSAC in the multi-view geometry guide is an estimator whose randomness is explicit: it draws a random minimal sample of correspondences, fits a model, and counts inliers. Different draws give different models, and a different random seed gives a different answer from the same data. Its guarantees — how many iterations are needed, how likely a clean sample is — are statements about the sampling distribution of that estimate, which is why the method reasons about inlier probabilities rather than a single fitted line.

ML / AI

Latency percentiles over repeated runs

A reported p99 latency is a statistic computed from a finite set of requests, and a second run over a different set of requests returns a different p99. The serving metrics chapter treats those numbers as estimates with their own sampling variability, which is why comparing two deployments on a single measurement is unreliable: the difference may be smaller than the standard error of the percentile itself. The same reasoning applies when comparing two benchmark scores.

Once you can see the parameter/estimate split, the formulas downstream stop being arbitrary. A filtered robot pose is an estimate whose sampling distribution the Kalman update is tracking; a held-out benchmark score is an estimate over a population of inputs; a fitted slope is an estimate whose standard error appears right next to it in every summary table. The sampling distribution is not one technique among many — it is the object all of them are describing.

Further reading

The references below treat the sampling distribution as the central object of inference rather than a preliminary to formulas. If you take away one thing, take away the picture of a fixed line and a histogram of estimates growing around it, its width shrinking like 1/√n.

Wasserman is the compact modern treatment this series follows in spirit, and his chapter on the bootstrap is the natural next step once the standard error has a name. Casella and Berger derive the sampling distributions of the standard estimators at the level a first-year graduate course expects. For the idea that the estimate is itself a random variable, the visual treatments repay a second look.

Cheat sheet

TermMeaning here
Population distributionThe spread of individual observations; fixed and unknown
Sampling distributionThe distribution of an estimate over repeated samples of size n
EstimatorThe rule applied to a sample; a random variable before the data arrive
Standard errorThe standard deviation of the sampling distribution
σ/√nThe standard error of a sample mean of independent observations
sd of the data vs sd of the estimateσ is how variable one value is; σ/√n is how variable the mean is
Quadruple nHalves the standard error, because of the √n
Standard error of the medianAbout 1.25 σ/√n for normal data — wider than the mean, but robust to outliers
7

Check your understanding

0/4 answered