What statistics actually asks
Statistics begins with an asymmetry that probability does not have. In probability you are handed the distribution and asked what a sample will look like. In statistics you are handed one sample and asked what the distribution was — or at least some number that summarises it. You never get to see the answer, and if you draw a different sample you will get a different estimate. This part makes that asymmetry something you can watch: a population you cannot inspect, samples you can redraw at a button press, and the slow, honest fact that the estimate jumps around while the quantity it is chasing does not move at all.
The question
What can one sample tell you about everything?
Imagine a process that produces numbers: the height of a person drawn from a city, the time between failures of a machine, the token length of a request arriving at a server. There is a population — a probability distribution over all the values the process could produce — and a parameter of it, some single number such as its mean, that you would like to know. The parameter is fixed. It does not change when you collect data. But you never observe it directly. You observe a finite sample: a handful of values the process happened to produce, and no more.
The task of statistics is to go from the sample to a claim about the parameter. The simplest such claim is a point estimate: use the sample mean to estimate the population mean, use the sample maximum to estimate the largest value the process can produce, use the fraction of successes in the sample to estimate the rate. The estimate is a function of the data, computed by a rule you chose. Because the data are random, the estimate is random too: redraw the sample and the estimate changes, even though the parameter has not.
Everything else in this series is a refinement of that observation. The next few parts ask how good such a rule can be, how to quantify its wobble, how to attach an interval and a decision to it, and how to fold in prior knowledge. But they all rest on one picture, and this part is that picture.
Draw a sample
The estimate moves; the parameter does not
Press the button and a fresh sample is drawn. The grey ticks along the baseline are the values in the sample; the solid marker on the axis is the sample mean; the dashed line is the population mean. The estimate lands near the truth on average, but it is never exactly on it, and a second press of the button gives a different answer. That scatter is not a defect of the method — it is the irreducible consequence of having seen only part of the population.
The readout keeps the last several estimates so you can see them differ. Every one of them is a legitimate answer to the same question, computed from an equally legitimate sample. They disagree because the data disagree.
Top: the hidden population, the sample ticks, the sample mean (solid) and the true mean (dashed). Bottom: the estimates from every sample drawn so far.
The estimate wobbles, the truth does not
A random variable made of data
Before the data are drawn, the sample mean is a random variable: it is a function of random values, so it has a distribution of its own. Statisticians call it the sampling distribution of the estimator, and it is the central object of this series. The bottom canvas builds it one sample at a time — each press of the button adds one estimate to a growing histogram. Come back after a hundred presses and the shape is unmistakable: a hump centred near the true mean, with a width that measures how far an estimate typically strays.
The sampling distribution lives at a different level from the population. The population describes individual observations; the sampling distribution describes a summary of a whole sample. They have different means (the same, for an unbiased estimator) and very different spreads. Keeping the two levels straight is most of what makes statistics confusing the first time, and the two stacked canvases are here to keep them apart.
Nothing in this picture requires the population to be normal or symmetric. The estimates pile up around the truth because averaging cancels the independent errors of the individual observations, not because the data are well behaved. Part 5 tells that story properly.
More data, less wobble
The one lever everybody reaches for
Drag the sample-size slider up and press the button again. With more values in each sample, the estimates cluster more tightly around the truth; the histogram narrows. Drag it down to a handful and the estimates scatter widely. This is the single most reliable fact in the subject: averaging more independent observations reduces the spread of the average, at a rate the next parts make precise.
That rate is not free lunch. To halve the typical error you need four times the data, because the spread of an average shrinks with the square root of the sample size. Part 5 turns that sentence into the standard error, $ \mathrm{SE} = \sigma/\sqrt{n} $, as a curve you can read off the picture rather than a formula you memorise.
There is also a caution buried here. Increasing n shrinks the random part of the error but does nothing about the systematic part. If the sample is not really drawn from the population you care about — if the survey reaches only people with phones, if the sensor is miscalibrated, if the data are a biased slice of the world — then a bigger sample makes you more confident about the wrong number. Bias lives in the design, not in the sample size, and Part 6 is where the two error sources are separated and added.
Where this shows up
The loop the rest of the site runs
Robust fitting from noisy points
RANSAC in the multi-view geometry guide is an estimator with exactly this structure: it draws a random minimal sample of correspondences, fits a model, and counts inliers. Different draws give different models. Its guarantees are statements about the sampling distribution of that estimate, which is why the guide reasons about inlier probabilities rather than a single fitted line.
Evaluation is estimation
A held-out benchmark score is an estimate of a model's true error on a population of inputs. The evaluation chapter reports accuracy on a finite test set; a different test set would give a different number, just as a different sample gives a different mean. Confidence intervals on benchmark scores are this part's wobble, made quantitative.
The pattern repeats everywhere on the site. A robot estimating its position from a noisy measurement is running the loop in the bottom canvas; the estimate is a random variable, and the filter is a rule for its distribution. A language model sampled at a temperature is drawing from a conditional population, and any statistic computed from its outputs is an estimate with this same sampling wobble. Once you can see the parameter/estimate split, the formulas that follow are bookkeeping rather than new ideas.
Further reading
The references below treat estimation as the central act of statistics rather than a preliminary to formulas. If you take away one thing, take away the picture of a fixed line and a cloud of estimates scattered around it.
Wasserman's All of Statistics is the compact modern treatment this series follows in spirit; Casella and Berger derive the properties of estimators at the level a first-year graduate course expects; and the standard-error chapter of any introductory text is the next step. For the conceptual trap of "significant by sample size alone", Ioannidis is worth reading once.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapters 1–6 — probability, random variables and the basic estimators.
- George Casella and Roger Berger, Statistical Inference, chapter 7 — point estimation, bias, variance and MSE.
- 3Blue1Brown, "Why do we use the mean?" and the broader Essence of Statistics series — the visual style this guide follows.
- John Ioannidis, "Why Most Published Research Findings Are False", 2005 — what goes wrong when bias, not variance, dominates.
Cheat sheet
| Term | Meaning here |
|---|---|
| Population | The distribution all observations are drawn from; fixed, unobserved |
| Parameter | A number describing the population, such as its mean; fixed, unknown |
| Sample | The observed values; random, and the only thing you actually have |
| Estimate | A rule applied to a sample, such as $\bar{x}$; random before the data arrive |
| Sampling distribution | The distribution of the estimate over repeated samples |
| i.i.d. | Independent and identically distributed — the assumption every method here starts from |
| Standard error | The standard deviation of the sampling distribution; shrinks like $\sigma/\sqrt{n}$ |