Null distributions and the p-value
A confidence interval in the last part answered “how far off could this estimate plausibly be?” A hypothesis test answers a different question: “if there were really no effect, how surprising would this data be?” To answer that you need a reference distribution — the null distribution, the values the statistic takes when the effect is genuinely absent. This part builds that reference by the most direct method available: keep the data, throw away the labels, and shuffle them. Under the hypothesis that the labels mean nothing, every shuffle is as plausible as the one you started with. Do that a thousand times and a distribution appears; the fraction of it at least as extreme as what you observed is the p-value. No formula is required, and none will be used.
The question
Is this difference real, or just wobble?
Suppose you run an experiment on two groups — a control group and a treatment group — and the two group means come out different. The last several parts taught you to expect differences even when nothing is going on, because every sample is random. So the observed gap has two possible explanations. Either the treatment really moved the mean, or the labels are arbitrary and the gap is the ordinary scatter of two samples drawn from one population. You cannot tell which by staring at the numbers. You need a way to ask how large a gap this process produces on its own.
That is the job of a hypothesis test. It formalises the boring explanation as the null hypothesis: the labels carry no information, the two groups are drawn from the same distribution, and the population means are equal. Against it stands the alternative hypothesis, that the treatment changed something. You then compute a test statistic — here, the difference in group means — and ask where that number falls relative to what the null hypothesis would generate. The reference distribution for that comparison is the null distribution.
Part 11 built intervals by resampling the data you have. This part resamples the labels instead, a small change with a large consequence: it produces an exact reference distribution under one clean assumption, with no formula for a density and no appeal to the central limit theorem.
Two groups and one number
Reducing a table of data to a single quantity
A test begins by collapsing the data into one number, the test statistic. The natural choice for two groups is the difference in sample means,
A statistic is just a function of the data chosen to be sensitive to the alternative you care about. If the treatment raises the mean, T will be large and positive; if it lowers the mean, large and negative; if it does nothing, T will hover near zero but will rarely be exactly zero, because the two samples are different draws. That hover is not noise to be explained away — it is the entire reference distribution we are about to build.
You will also meet standardised versions of T: divide the difference by its standard error and you get something close to a t statistic, the classical route to a p-value and the subject of Part 15. Standardising imports extra assumptions — roughly normal data, or enough data that the central limit theorem rescues you. The permutation approach below takes the other road: it derives the reference distribution from the data by simulation, and asks only that the labels be exchangeable under the null. Keep the observed value of the statistic in view, call it t_{\text{obs}}, and the rest is one question: how unusual is it among the values the same statistic takes when the null is true?
The null, built by shuffling
Simulating the world in which nothing happened
To answer that question we need the distribution of T under the null hypothesis. There are two ways to get it. The classical way assumes a shape for the data and derives the distribution of T mathematically, which is where the t, z, F and χ² curves come from. The other way is to simulate the null.
The key observation is what the null hypothesis says about the labels. If the treatment does nothing, then it does not matter which observation was labelled “treatment” and which “control”. The two sets of numbers are just one pooled set of numbers, arbitrarily split into two piles. Statisticians call this exchangeability: under the null, any reassignment of the labels is equally plausible. So instead of assuming a shape and deriving a curve, we can physically perform the reassignment. Join the two groups into one pool, shuffle it, split it back into piles of the same sizes, and recompute T. That shuffled value is one draw from the null distribution.
Repeat the shuffle many times and the histogram of those values is the null distribution — not an approximation to a formula, but the thing itself, up to the finite number of shuffles. This is a permutation test. It is remarkably assumption-free: it does not care whether the data are normal, skewed, or lumpy, because it never uses a formula for the data's distribution. It uses only the fact that, under the null, the labels are meaningless. That single premise is also its price: if the labels are not exchangeable — paired measurements, clustered samples, a treatment applied to whole groups — the naive shuffle is wrong and the test must be adapted.
The demo below makes the argument visible. Two groups are generated, one with a real effect you control. The pooled shuffle builds the null distribution from those same numbers, and the observed difference is dropped onto it as a vertical line.
Permute and watch the null grow
The signature interaction
Press Permute once and a single reshuffle is appended to the histogram, one bar at a time. Press Permute 1000 and a thousand shuffles stream in until the shape settles. The solid vertical line is the observed difference; the dashed line mirrors it on the other side. Bars beyond either line are the reshuffles that matched or beat the real result — those bars are shaded, and the p-value is simply their share of the whole histogram.
Drag the true-effect slider to zero and the two groups are drawn from the same distribution. Now the observed difference is itself just another reshuffle, so it lands in the middle of the pile and almost nothing is shaded: a large p-value, exactly as the null hypothesis deserves. Push the slider up and the real difference moves out toward the edge of the null distribution, the shaded tail shrinks, and the p-value falls. Nothing about the procedure changed between the two settings; only the world did.
Histogram of the difference in means over label shuffles. Shaded bars are at least as extreme as the observed difference (solid line); the dashed line is its mirror.
The p-value is a tail area
P(statistic at least this extreme | null)
With the null distribution in hand, the p-value is a counting exercise. Count the shuffles whose statistic is at least as extreme as the one you observed, and divide by the number of shuffles:
In practice one uses (c + 1)/(B + 1) rather than c/B, so that the p-value is never exactly zero and the test stays honest at very small counts. That is the entire definition. The p-value is the probability, computed under the null hypothesis, of a statistic at least as extreme as the observed one. Read that sentence again and notice the conditional. It is P(data this extreme | null), not P(null | data). The first is what the test computes; the second is what most people hear when they read a p-value, and it is not available without a prior. A small p-value says the data would be surprising if the null were true. It does not say the null is probably false, nor that the effect is large, nor that the finding is important.
“At least as extreme” needs a direction, and that is the role of the two-sided version used in the demo: we count permutations whose absolute difference is at least the absolute observed difference, so a large effect in either direction counts. Two-sidedness is the cautious default, because in advance of the experiment you often do not know which way an effect would go. A one-sided test counts only one tail and therefore produces smaller p-values; it is legitimate only when the other direction is genuinely impossible or uninteresting, decided before seeing the data.
A p-value is not a property of the data alone. It is a statement about the data relative to a chosen statistic, a chosen null, and a chosen number of shuffles. Change the statistic and the p-value changes even though the data have not. That is not a flaw; it is the structure of the question, and it is why the question has to be posed before the data arrive.
One-sided, two-sided, and α
Turning a number into a decision
To convert a p-value into a yes/no decision, a threshold — the significance level α — is fixed in advance. If p \le \alpha the result is called statistically significant and the null is rejected; otherwise it is not rejected. The conventional choice, α = 0.05, means that if the null were true and the experiment repeated many times, about one run in twenty would produce a p-value this small by chance. The threshold is a policy about how often you are willing to be fooled, not a law of nature, and it belongs to the experimenter, decided before the data arrive.
The threshold is also where the two error types live, and Part 13 is devoted to them. Rejecting a true null is a Type I error, whose rate is bounded by α. Failing to reject a false null is a Type II error, and one minus its rate is the power of the test. Power depends on the size of the effect, the noise, and the sample size, which is why a non-significant result is never proof that nothing happened — it may just be a small study.
This is also the bridge to the classical zoo. Standardise the difference in means and assume normal data, and the permutation null approaches a t distribution while the permutation p-value approaches the t-test p-value. Part 15 lays out that family — t, z, χ², F and the likelihood-ratio test — as one idea seen through different assumptions. The permutation test is the same idea with the assumptions stripped away, which is why it is worth understanding first.
Where this shows up
A test by any other name
Testing whether a fit is real
RANSAC in the multi-view geometry guide repeatedly samples a minimal set of correspondences, fits a model, and counts inliers. The natural question — is this many inliers more than chance would give? — is a hypothesis test. The null is a random correspondence set, the statistic is the inlier count, and the threshold on that count plays the role of α. The geometry page reasons about those probabilities exactly as this part does.
A/B model comparisons
Comparing two models on the same benchmark is an A/B test. With paired inputs a permutation over the paired differences gives an exact null without assuming normality, which matters because benchmark scores are bounded and often skewed. The evaluation chapter reports the scores; this part is how you decide whether a gap between them is real or is the wobble of a finite test set.
The same shape appears across the site. An online experiment comparing two serving configurations is a two-group test against a null of no difference, and so is a robotics pipeline that swaps one sensor for another and measures error. A language-model evaluation reporting an improvement after fine-tuning is claiming a difference in means, and a permutation test is one of the few tools that does not demand the scores be well behaved. Wherever a number is compared to what chance would have produced, a null distribution is being imagined — even if nobody draws it.
Further reading
The references below treat the permutation test as the intuitive foundation and the classical tests as shortcuts available when extra assumptions hold. If you take away one thing, take away the picture of a histogram built from shuffled labels with a shaded tail beyond the observed value.
Wasserman covers hypothesis testing compactly and proves the properties the permutation test inherits. Efron and Tibshirani's bootstrap book has the cleanest short treatment of permutation and rank tests side by side, and Good's book is the dedicated reference if you want the theory. The American Statistical Association's statement is worth reading once for what a p-value does and does not mean.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapter 10 — hypothesis testing, p-values and the Neyman–Pearson framework.
- Bradley Efron and Robert Tibshirani, An Introduction to the Bootstrap, chapter 15 — permutation tests and why they are exact under exchangeability.
- Phillip Good, Permutation, Parametric, and Bootstrap Tests of Hypotheses — the dedicated treatment, including the counting correction.
- Ronald Wasserstein and Nicole Lazar, "The ASA Statement on p-Values: Context, Process, and Purpose", 2016 — what a p-value does and does not mean.
Cheat sheet
| Term | Meaning here |
|---|---|
| Null hypothesis H0 | Labels carry no information; both groups share one distribution |
| Alternative H1 | The treatment changed the distribution — usually its mean |
| Test statistic | The single number summarising the data; here, the difference in means |
| Null distribution | The statistic's distribution when H0 is true; built here by shuffling |
| Exchangeability | Under the null, any relabelling is equally plausible — what makes the shuffle valid |
| p-value | P(statistic at least this extreme | null), estimated as (c+1)/(B+1) |
| α | Significance level fixed in advance; reject when p \le \alpha |
| Two-sided | Counts |permuted| ≥ |observed|, either direction; the cautious default |