p-hacking and the garden of forking paths
A single test is allowed to be wrong five percent of the time. That is the bargain you accept when you run one test. The bargain collapses the moment you run many: try enough subgroups, covariates, outcomes and reasonable-looking analyses, and a false positive stops being bad luck and becomes a near certainty. This part builds a sandbox in which the null hypothesis is true for every comparison you make, so every significant result you find is a false one — and then hands you the corrections that put the count back under control.
The question
If you test enough things, something will look real
Part 12 gave you the p-value: the probability, assuming the null hypothesis is true, of seeing a result at least as extreme as the one in front of you. Part 13 gave you its partner, the threshold. Pick $\alpha = 0.05$, declare "significant" whenever the p-value falls below it, and you have committed to a procedure that fools you about one time in twenty. For a single test that is a reasonable deal. You cannot buy certainty, so you rent it at five percent.
The deal is priced for one test. Run twenty independent tests when the null is true for all twenty, and the chance that at least one of them crosses the line is $1 - 0.95^{20} \approx 0.64$. Run forty and it is above eighty percent. Nothing has changed about the tests or the data: there is still no effect anywhere. What changed is how many chances you gave noise to impersonate an effect. The property you thought you had bought — a five percent chance of a false positive — was a property of the single test, and you spent it twenty times.
This is where the word significant turns out to be doing more work than it can support. In Fisher's original framing the p-value was a loose flag, evidence to be weighed alongside everything else. Under the textbook decision rule it became a switch, and a switch can be flipped by watching it long enough. The sandbox in this part makes that failure mechanical: the truth is fixed and known to be nothing, and yet a green flag appears on the third try.
The garden of forking paths
Researcher degrees of freedom
Real p-hacking, the deliberate kind, is comparatively rare, and it is not what makes the problem dangerous. Andrew Gelman and Eric Loken gave the sharper diagnosis: a researcher walks through a garden of forking paths. At every step there is a choice — which outcome to analyse, which covariates to adjust for, which subgroup to report, whether to log-transform the response, whether to drop the two outlying runs, whether to stop collecting data on Monday. Each choice is defensible on its own. Each path leads to a slightly different test. And the reported p-value is computed as though the path were the only one, when in fact it was selected after seeing where the paths led.
The formal name for the hazard is the multiple-comparisons problem. Once you have a family of tests rather than one, the relevant error is no longer the per-test rate but the chance that the family contains at least one false positive. Statisticians call it the family-wise error rate, FWER. Under independence it is $1 - (1-\alpha)^m$ for a family of $m$ tests, and it climbs toward one as the family grows. A p-value of $0.04$ in a family of fifty is not evidence of anything in particular; it is one card in a large hand.
The uncomfortable part is that the data cannot tell you how big the family was. They record the comparison you finally reported, not the thirty you tried and discarded, and not the forks you took for reasons that felt obvious at the time. A p-value is a statement about a procedure — about the whole plan of testing — so a procedure assembled after the fact from the best-looking branch describes something that never actually ran. That is the precise sense in which forking paths invalidate the number, without anyone lying.
Gelman's image also explains why replication fails more often than p-values promise. If the reported analysis was the best of many branches, the effect it measured was partly noise, and a fresh run has nothing to reproduce. The sandbox below lets you feel that maximum-of-many effect before the corrections arrive.
A p-hacking sandbox
Add comparisons until something turns up
The demo below is a null world. Every comparison is generated by drawing two independent samples from the same distribution, so there is no effect to find — not a small one, not a hidden one, none. Each press of the button is a fresh, independent comparison: a random subgroup, a covariate, an outcome, a transformation. The t-test is computed honestly from the data. The only thing dishonest is the plan, because there isn't one.
Each dot on the canvas is one p-value, plotted against the order in which it was computed. The dashed line is $\alpha = 0.05$; dots below it are flagged significant, and the smallest p-value seen so far is ringed. The list beside the plot shows every comparison with its raw p-value and, when a correction is switched on, its adjusted p-value. The readout reports the two numbers that matter: how many comparisons you have made, and how many you have flagged.
Press the button a few times and watch for the first flag. In this run it arrives on the third comparison, which is early but entirely ordinary: any single comparison has a one-in-twenty chance. Press Run 20 comparisons and the count settles near one, exactly as the arithmetic predicts — twenty tests times five percent. Nothing you did made an effect real. You simply checked enough places for noise to show up, and noise obliged.
One dot per comparison, drawn under a true null. The dashed line is α = 0.05; the ring marks the smallest p-value so far.
Bonferroni, Benjamini–Hochberg, and the design fix
Two error rates, two corrections, one habit
The correction toggle switches between three ways of turning each comparison's p-value into a decision. On raw every comparison is judged alone, which is the p-hacked behaviour the sandbox was built to expose. The other two judge the family as a whole, and they exist because there are two different things you might want to control.
The first is the family-wise error rate: the probability of making at least one false positive anywhere in the family. Bonferroni controls it by the crudest possible means, dividing the threshold by the number of tests. Equivalently, it multiplies each p-value by $m$ and compares the result to $\alpha$:
This is exactly the $\alpha/m$ threshold in disguise, and it is conservative: it controls FWER under any dependence structure, at the cost of demanding very small p-values from every test. That stiffness is a feature when the family is a handful of confirmatory questions where even one false positive would be costly, and a liability when the family is a genome-wide scan or a sweep of benchmark variants where you are screening, not concluding. In the sandbox, Bonferroni at twenty comparisons and $\alpha = 0.05$ demands a raw p-value below $0.0025$; the single p-hacked flag quietly disappears.
The second target is the false discovery rate: among the comparisons you call significant, the expected proportion that are false. Controlling FDR rather than FWER is the right ambition when you are willing to carry a small fraction of mistakes in exchange for not throwing away real signals. Benjamini–Hochberg does it with a ranking: sort the p-values from smallest to largest, find the largest index $k$ satisfying
and reject every comparison up to $k$. The adjusted p-values it reports are monotone — once one comparison is significant, every smaller p-value is too — and they are always at least as small as Bonferroni's, which is why BH rejects more. Under a global null, where every hypothesis is true, the two corrections have to make nearly the same demand at the top of the list, so the sandbox looks similar under both. The difference appears as soon as some real effects are present: BH spends its error budget on the strongest signals and lets the weaker ones ride along, while Bonferroni protects each individual test equally and pays for it in power. In the LLM setting, that is the difference between a confirmatory safety claim and an exploratory sweep over forty prompts.
Both corrections are arithmetic bandages on a design wound, and it is worth saying so plainly. The deeper fix is to decide the analysis before seeing the data, which is what a pre-registered analysis plan does: it collapses the garden of forking paths to one path by paying the cost of commitment up front. Where pre-registration is impossible — most exploratory work, and most machine learning — the discipline is to separate the data you explore on from the data you confirm on. Generate hypotheses on one split, then test the handful that survived on a fresh split that played no part in choosing them. That is a held-out set and a multiple-comparison problem wearing a statistician's clothes, and the probability course makes the underlying "many independent chances" arithmetic explicit in the probability part.
Where this shows up
Two places where many tests are run quietly
Leaderboards are families
A model release advertises the benchmarks it wins, and the choice of which benchmarks to headline is made after the numbers are in. The evaluation chapter makes the same point from the other side: every held-out score is an estimate with sampling error, and scoring a model on a dozen suites and reporting the best is a garden of forking paths. The honest report is the pre-declared suite, or an adjustment that counts how many were tried. The serving metrics chapter meets the same arithmetic when it compares many latency configurations at once.
RANSAC tests many hypotheses
RANSAC in the multi-view geometry guide draws one minimal sample of correspondences at a time, fits a model, and keeps it if enough points agree. Each draw is a hypothesis test against noise, and the confidence that at least one all-inlier sample is found obeys the same "many independent chances" arithmetic as FWER. Choosing the number of iterations is exactly choosing how small a family failure probability you can tolerate.
The pattern is that any procedure which selects the best of many candidates — the best subgroup, the best benchmark, the best random sample, the best hyperparameter, the best pipeline — has quietly converted a per-test error rate into a family-wise one. Feature selection in the scaling chapter, reward-model comparison in RLHF, and the calibration of a sensor in odometry all run into it in one form or another. Seeing it once, as a count that climbs, makes the corrections feel less like arbitrary penalties and more like bookkeeping that was always owed.
Further reading
The references below are the canonical statements of the problem and its fixes. If you take away one thing, take away that a p-value is a property of a procedure, so a procedure chosen after looking at the data was never really run.
Gelman and Loken's essay is the clearest account of why honest researchers still generate false positives, and Simmons and colleagues give the experimental demonstration. Benjamini and Hochberg is the original paper behind the correction implemented in the sandbox above, and reading it once is worth the effort for the precise definition of the false discovery rate.
- Andrew Gelman and Eric Loken, "The garden of forking paths", 2013 — why multiple comparisons arise without any deliberate p-hacking.
- Joseph Simmons, Leif Nelson and Uri Simonsohn, "False-Positive Psychology", 2011 — researcher degrees of freedom demonstrated experimentally.
- Yoav Benjamini and Yosef Hochberg, "Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing", JRSS-B, 1995 — the original BH procedure the demo uses.
- John Ioannidis, "Why Most Published Research Findings Are False", 2005 — the consequences at the scale of a literature.
- Open Science Collaboration, "Estimating the Reproducibility of Psychological Science", 2015 — what replication looks like when the garden is left unmapped.
Cheat sheet
| Term | Meaning here |
|---|---|
| p-value | Probability, under the null, of a result at least this extreme; a property of the whole procedure |
| Family-wise error rate | Probability of at least one false positive anywhere in the family; grows like $1-(1-\alpha)^m$ |
| False discovery rate | Expected proportion of the flagged comparisons that are false |
| Bonferroni | Reject when $p_i \le \alpha/m$; controls FWER, conservative, no dependence assumptions |
| Benjamini–Hochberg | Largest $k$ with $p_{(k)} \le (k/m)\alpha$; controls FDR, rejects more than Bonferroni |
| Garden of forking paths | Defensible analysis choices made after seeing the data, which quietly enlarge the family |
| Researcher degrees of freedom | The unreported choices available between the data and the reported test |
| Pre-registration | Fixing the analysis before the data arrive; the design fix the corrections approximate |