Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

If you test enough things, something will look real

Part 12 gave you the p-value: the probability, assuming the null hypothesis is true, of seeing a result at least as extreme as the one in front of you. Part 13 gave you its partner, the threshold. Pick $\alpha = 0.05$, declare "significant" whenever the p-value falls below it, and you have committed to a procedure that fools you about one time in twenty. For a single test that is a reasonable deal. You cannot buy certainty, so you rent it at five percent.

The deal is priced for one test. Run twenty independent tests when the null is true for all twenty, and the chance that at least one of them crosses the line is $1 - 0.95^{20} \approx 0.64$. Run forty and it is above eighty percent. Nothing has changed about the tests or the data: there is still no effect anywhere. What changed is how many chances you gave noise to impersonate an effect. The property you thought you had bought — a five percent chance of a false positive — was a property of the single test, and you spent it twenty times.

This is where the word significant turns out to be doing more work than it can support. In Fisher's original framing the p-value was a loose flag, evidence to be weighed alongside everything else. Under the textbook decision rule it became a switch, and a switch can be flipped by watching it long enough. The sandbox in this part makes that failure mechanical: the truth is fixed and known to be nothing, and yet a green flag appears on the third try.

💡 By the end of this part you'll see why the chance of at least one false positive grows with the number of comparisons, why "significant" names a decision rather than a discovery, and how Bonferroni and Benjamini–Hochberg trade one kind of error control for another.
2

The garden of forking paths

Researcher degrees of freedom

Real p-hacking, the deliberate kind, is comparatively rare, and it is not what makes the problem dangerous. Andrew Gelman and Eric Loken gave the sharper diagnosis: a researcher walks through a garden of forking paths. At every step there is a choice — which outcome to analyse, which covariates to adjust for, which subgroup to report, whether to log-transform the response, whether to drop the two outlying runs, whether to stop collecting data on Monday. Each choice is defensible on its own. Each path leads to a slightly different test. And the reported p-value is computed as though the path were the only one, when in fact it was selected after seeing where the paths led.

The formal name for the hazard is the multiple-comparisons problem. Once you have a family of tests rather than one, the relevant error is no longer the per-test rate but the chance that the family contains at least one false positive. Statisticians call it the family-wise error rate, FWER. Under independence it is $1 - (1-\alpha)^m$ for a family of $m$ tests, and it climbs toward one as the family grows. A p-value of $0.04$ in a family of fifty is not evidence of anything in particular; it is one card in a large hand.

The uncomfortable part is that the data cannot tell you how big the family was. They record the comparison you finally reported, not the thirty you tried and discarded, and not the forks you took for reasons that felt obvious at the time. A p-value is a statement about a procedure — about the whole plan of testing — so a procedure assembled after the fact from the best-looking branch describes something that never actually ran. That is the precise sense in which forking paths invalidate the number, without anyone lying.

Gelman's image also explains why replication fails more often than p-values promise. If the reported analysis was the best of many branches, the effect it measured was partly noise, and a fresh run has nothing to reproduce. The sandbox below lets you feel that maximum-of-many effect before the corrections arrive.

3

A p-hacking sandbox

Add comparisons until something turns up

The demo below is a null world. Every comparison is generated by drawing two independent samples from the same distribution, so there is no effect to find — not a small one, not a hidden one, none. Each press of the button is a fresh, independent comparison: a random subgroup, a covariate, an outcome, a transformation. The t-test is computed honestly from the data. The only thing dishonest is the plan, because there isn't one.

Each dot on the canvas is one p-value, plotted against the order in which it was computed. The dashed line is $\alpha = 0.05$; dots below it are flagged significant, and the smallest p-value seen so far is ringed. The list beside the plot shows every comparison with its raw p-value and, when a correction is switched on, its adjusted p-value. The readout reports the two numbers that matter: how many comparisons you have made, and how many you have flagged.

Press the button a few times and watch for the first flag. In this run it arrives on the third comparison, which is early but entirely ordinary: any single comparison has a one-in-twenty chance. Press Run 20 comparisons and the count settles near one, exactly as the arithmetic predicts — twenty tests times five percent. Nothing you did made an effect real. You simply checked enough places for noise to show up, and noise obliged.

One dot per comparison, drawn under a true null. The dashed line is α = 0.05; the ring marks the smallest p-value so far.

⚠ "Significant" is a decision, not a discovery. The threshold converts a continuous p-value into a yes-or-no act, and acts can be repeated. In a null world the raw count grows roughly linearly with the number of comparisons, while the corrected count stays near zero. The gap between them is the price of looking in many places.
4

Bonferroni, Benjamini–Hochberg, and the design fix

Two error rates, two corrections, one habit

The correction toggle switches between three ways of turning each comparison's p-value into a decision. On raw every comparison is judged alone, which is the p-hacked behaviour the sandbox was built to expose. The other two judge the family as a whole, and they exist because there are two different things you might want to control.

The first is the family-wise error rate: the probability of making at least one false positive anywhere in the family. Bonferroni controls it by the crudest possible means, dividing the threshold by the number of tests. Equivalently, it multiplies each p-value by $m$ and compares the result to $\alpha$:

$$p_i^{\text{adj}} = \min\!\left(1,\; m\,p_i\right), \qquad \text{reject when } p_i^{\text{adj}} \le \alpha$$

This is exactly the $\alpha/m$ threshold in disguise, and it is conservative: it controls FWER under any dependence structure, at the cost of demanding very small p-values from every test. That stiffness is a feature when the family is a handful of confirmatory questions where even one false positive would be costly, and a liability when the family is a genome-wide scan or a sweep of benchmark variants where you are screening, not concluding. In the sandbox, Bonferroni at twenty comparisons and $\alpha = 0.05$ demands a raw p-value below $0.0025$; the single p-hacked flag quietly disappears.

The second target is the false discovery rate: among the comparisons you call significant, the expected proportion that are false. Controlling FDR rather than FWER is the right ambition when you are willing to carry a small fraction of mistakes in exchange for not throwing away real signals. Benjamini–Hochberg does it with a ranking: sort the p-values from smallest to largest, find the largest index $k$ satisfying

$$p_{(k)} \le \frac{k}{m}\,\alpha,$$

and reject every comparison up to $k$. The adjusted p-values it reports are monotone — once one comparison is significant, every smaller p-value is too — and they are always at least as small as Bonferroni's, which is why BH rejects more. Under a global null, where every hypothesis is true, the two corrections have to make nearly the same demand at the top of the list, so the sandbox looks similar under both. The difference appears as soon as some real effects are present: BH spends its error budget on the strongest signals and lets the weaker ones ride along, while Bonferroni protects each individual test equally and pays for it in power. In the LLM setting, that is the difference between a confirmatory safety claim and an exploratory sweep over forty prompts.

Both corrections are arithmetic bandages on a design wound, and it is worth saying so plainly. The deeper fix is to decide the analysis before seeing the data, which is what a pre-registered analysis plan does: it collapses the garden of forking paths to one path by paying the cost of commitment up front. Where pre-registration is impossible — most exploratory work, and most machine learning — the discipline is to separate the data you explore on from the data you confirm on. Generate hypotheses on one split, then test the handful that survived on a fresh split that played no part in choosing them. That is a held-out set and a multiple-comparison problem wearing a statistician's clothes, and the probability course makes the underlying "many independent chances" arithmetic explicit in the probability part.

5

Where this shows up

Two places where many tests are run quietly

ML / AI

Leaderboards are families

A model release advertises the benchmarks it wins, and the choice of which benchmarks to headline is made after the numbers are in. The evaluation chapter makes the same point from the other side: every held-out score is an estimate with sampling error, and scoring a model on a dozen suites and reporting the best is a garden of forking paths. The honest report is the pre-declared suite, or an adjustment that counts how many were tried. The serving metrics chapter meets the same arithmetic when it compares many latency configurations at once.

Vision & Geometry

RANSAC tests many hypotheses

RANSAC in the multi-view geometry guide draws one minimal sample of correspondences at a time, fits a model, and keeps it if enough points agree. Each draw is a hypothesis test against noise, and the confidence that at least one all-inlier sample is found obeys the same "many independent chances" arithmetic as FWER. Choosing the number of iterations is exactly choosing how small a family failure probability you can tolerate.

The pattern is that any procedure which selects the best of many candidates — the best subgroup, the best benchmark, the best random sample, the best hyperparameter, the best pipeline — has quietly converted a per-test error rate into a family-wise one. Feature selection in the scaling chapter, reward-model comparison in RLHF, and the calibration of a sensor in odometry all run into it in one form or another. Seeing it once, as a count that climbs, makes the corrections feel less like arbitrary penalties and more like bookkeeping that was always owed.

Further reading

The references below are the canonical statements of the problem and its fixes. If you take away one thing, take away that a p-value is a property of a procedure, so a procedure chosen after looking at the data was never really run.

Gelman and Loken's essay is the clearest account of why honest researchers still generate false positives, and Simmons and colleagues give the experimental demonstration. Benjamini and Hochberg is the original paper behind the correction implemented in the sandbox above, and reading it once is worth the effort for the precise definition of the false discovery rate.

Cheat sheet

TermMeaning here
p-valueProbability, under the null, of a result at least this extreme; a property of the whole procedure
Family-wise error rateProbability of at least one false positive anywhere in the family; grows like $1-(1-\alpha)^m$
False discovery rateExpected proportion of the flagged comparisons that are false
BonferroniReject when $p_i \le \alpha/m$; controls FWER, conservative, no dependence assumptions
Benjamini–HochbergLargest $k$ with $p_{(k)} \le (k/m)\alpha$; controls FDR, rejects more than Bonferroni
Garden of forking pathsDefensible analysis choices made after seeing the data, which quietly enlarge the family
Researcher degrees of freedomThe unreported choices available between the data and the reported test
Pre-registrationFixing the analysis before the data arrive; the design fix the corrections approximate
6

Check your understanding

0/4 answered