Hypothesis testing, and how it goes wrong
A test is a decision procedure with two ways to be wrong, and it is designed to control only one of them. You fix a rule, you look at the data, and you either reject a default story or fail to. The machinery is simple — one test statistic, two distributions, a threshold — but almost every statistical scandal in the sciences is a story about people who used the simple machinery while quietly doing something else: testing many hypotheses and reporting one, choosing the analysis after seeing the data, or reading a small p-value as a probability that the null is true. This part builds the test from its two error rates, then simulates exactly the ways it is abused, so you can recognize them in the wild.
The question
Choosing between stories, with a known cost of being wrong
Suppose you run a new recommendation model against the old one and the new one wins on your holdout by a hair. Did you improve the system, or did you watch a coin land heads? The data are random, so both explanations are live, and no amount of staring at the number will separate them. A hypothesis test makes the choice explicit: name a default story, name the alternative you care about, pick a statistic that behaves differently under the two, and reject the default only when the statistic lands somewhere the default almost never produces.
The default story is the null hypothesis H_0, usually "no effect", "no difference", "the coin is fair". The alternative H_1 is the world you hope to detect. The test is deliberately asymmetric. You assume H_0 while you compute, you control how often you would wrongly reject it, and the alternative only enters when you ask how often the test succeeds. That asymmetry is the source of most of the confusion that follows, and it is not a bug — it is the design.
A concrete running example: a classifier that should be right one half the time if its features carry nothing. You collect n labelled examples and count correct predictions. Under the null the count is binomial with success probability 1/2. If you observe far more correct than one half, you reject the null and claim the classifier learned something. If you observe near one half, you fail to reject — which is not the same as proving the classifier learned nothing, and the gap between those two statements is another recurring trap.
Four quantities govern everything. The chance of rejecting a true null is $\alpha$; the chance of failing to reject a false null is $\beta$; the chance of catching a real effect is the power $1-\beta$; and the quantity that decides whether there is any hope of catching it is the effect size. Fix $\alpha$ small, spend more data, and power rises; ignore effect size, and you will design experiments that cannot possibly succeed while still producing publishable noise.
Null, alternative, α and β
One statistic, two distributions, two ways to be wrong
Standardise the running example. Let $\bar X$ be the average score of your classifier over n independent examples, with a known spread $\sigma$ per example. The test statistic is the standardised distance from the null's idea of "no skill", $Z=(\bar X-\mu_0)/(\sigma/\sqrt n)$. Under H_0 the central limit theorem makes Z approximately standard normal. Under a specific alternative with effect $\delta=(\mu_1-\mu_0)/\sigma$, the same statistic is normal with mean $\delta\sqrt n$ and variance one. The whole test is a picture of two bell curves.
Choose a critical value c and the rule "reject H_0 when $Z\ge c$". Everything follows. The area of the H_0 curve to the right of c is the false-positive rate $\alpha$, the size of the test. The area of the H_1 curve to the left of c is $\beta$, the false-negative rate. Slide $\alpha$ down and the critical value moves right: fewer false alarms, but the shaded slab of missed effects grows if the alternative has not moved.
The canvas makes the trade visible. The two bells are the sampling distributions of Z; the vertical line is the critical value; the pink tail on the right belongs to H_0 and is $\alpha$; the blue slab on the left belongs to H_1 and is $\beta$. Drag the handle on the blue peak to change the effect size, and use the sliders for the sample size and $\alpha$. When the blue bell slides right, $\beta$ shrinks and power climbs. When it overlaps the pink bell, no threshold can separate them. That overlap, not the arithmetic, is what a weak study looks like.
Two sampling distributions of Z. Pink = the $\alpha$ tail of H_0 beyond the critical value; blue = the $\beta$ slab of H_1 below it. Drag the handle on the blue peak to set the effect size.
Two details the picture clarifies. First, the critical value depends only on $\alpha$, not on the alternative: the test is specified before you see the data, and that is what makes $\alpha$ a rate rather than a description. Second, the distance between the curve centres is $\delta\sqrt n$, the noncentrality. Quadrupling the sample size doubles that distance, which is why data buys power only as its square root, and why an effect of size zero can never be detected no matter how long you run.
Power and effect size
The probability that the experiment works
Power is the chance the test rejects when there really is something to find. It is the single most informative number about a study, and it is almost never reported. An underpowered study is not merely unlucky: it fails to detect real effects most of the time, and the effects it does detect are, by a selection effect, systematically larger than the truth. Among experiments that barely clear the threshold, the only way to clear it with random noise is for the noise to push upward, so published estimates are inflated. This is the winner's curse, and it is the same mechanism that makes a backtest look brilliant until it is run forward.
For the one-sided normal test you can invert the picture and solve for the sample size that reaches a target power. If $z_{1-\alpha}$ and $z_{\text{power}}$ are the corresponding normal quantiles, the requirement $\delta\sqrt n\ge z_{1-\alpha}+z_{\text{power}}$ gives a compact formula. It is the design equation: tell it the smallest effect you care about and the power you want, and it tells you how much data you need before you have any right to run the test.
The formula is brutal in the way it punishes small effects. Halving $\delta$ quadruples n. With the conventional $\alpha=0.05$ and 80% power, the parenthetical factor is about 2.49, so a weak effect of $\delta=0.1$ needs $n\approx 620$ per group, while a strong $\delta=0.8$ needs only about ten. Most of the replications that fail in the behavioural sciences are replications of studies that were never powered for the effect they claimed.
There is a second lesson in the curve: power is not a property of the test alone. Change the effect size and the same n moves from hopeless to certain. That is why reporting a p-value without an effect size and an interval is incomplete — a tiny p-value from a huge sample can correspond to an effect too small to matter, and the confidence-interval part is the honest companion to any test. Drag the effect-size slider and watch the whole power curve shift; the vertical line marks the sample size that reaches the 80% target.
Power of a one-sided test at $\alpha=0.05$ as a function of n for the effect size you choose. The dashed line is the 80% target; the vertical line is the n that reaches it.
The curve is the design sledgehammer, but note what it assumes: a known $\sigma$, a one-sided alternative, one test. Real analyses choose among several endpoints, look at the data midway, or drop a group that is not behaving. Each of those choices quietly changes the null distribution of the statistic, and the nominal $\alpha$ stops meaning what it says. That is the subject of the last two sections.
The p-value, and what it is not
A tail area, not a posterior probability
The p-value is the chance of seeing a statistic at least as extreme as the one you observed, assuming the null is exactly true: $p=P(T\ge t_{\text{obs}}\mid H_0)$. It is a statement about the data given the null. It is emphatically not the probability that the null is true given the data, which is the quantity everyone wants and which requires a prior. The two are related only through Bayes' rule, and when the prior on a real effect is small — say most tested genes do nothing — a p-value of 0.01 can coexist with a posterior probability of the null above one half.
The clean fact that exposes the misunderstanding is the distribution of the p-value itself. When the null is true and the test is valid, the p-value is uniformly distributed on [0,1]. Every value is equally likely, including 0.04. So under a true null, five percent of experiments will produce p<0.05, one percent will produce p<0.01, and a single "significant" result is exactly what the null predicts at the advertised rate. A small p-value is not evidence that something is there; it is evidence that the data are surprising under a hypothesis you already doubted.
When the null is false, the p-value distribution shifts toward zero, and the size of the shift measures power. With a large effect and plenty of data, p-values pile up near zero and a single experiment is convincing. With a small effect and little data, the distribution barely tilts and most experiments produce p-values indistinguishable from uniform. The simulation below draws thousands of idealized experiments and histograms their p-values. Under H_0 the bars are flat. Under H_1 they lean left. Toggle between the two and watch the lean grow with the effect size.
Histogram of p-values from thousands of simulated one-sided z-tests. Flat under H_0; piled toward zero under H_1. The dashed line is what a uniform distribution would give.
Two consequences follow immediately. First, a p-value cannot be compared across studies with different sample sizes or different models; it is not an effect size, and the same physical effect produces smaller p-values from larger samples. Second, the logic of "reject or not" throws away the magnitude. Reporting p=0.049 as a discovery and p=0.051 as a failure is a ritual, not an inference; the two experiments carry almost identical information. What a test gives you is a calibrated false-alarm rate under a specified null, and that is all it gives you.
Multiplicity and p-hacking
Why the best of many tests is always significant
The false-alarm rate $\alpha$ is a promise about one pre-specified test. Run several tests and the promise lapses. If you test m independent true nulls, each at level $\alpha$, the chance that at least one comes back significant is $1-(1-\alpha)^m$. At $\alpha=0.05$ that probability passes one half by m=14, reaches 64% at twenty tests, and approaches one as m grows. The expected number of false positives is exactly $m\alpha$, so twenty null tests buy you one spurious discovery on average, free of charge.
The fix is to move the goalposts for each individual test. The Bonferroni correction tests each hypothesis at $\alpha/m$; by the union bound the chance of any false positive across the family is then at most $\alpha$. It is conservative when the tests are positively correlated, which they usually are, but it never fails to control the family-wise error rate. The Šidák correction, testing each at $1-(1-\alpha)^{1/m}$, is exact for independent tests and slightly less conservative; both answer the same question, "what is my chance of at least one false claim across the whole family?"
P-hacking is multiplicity performed without acknowledging it. If you try several model specifications and report the one that cleared the threshold, you have run the family and hidden the losers. If you add data points until significance and stop, you have given the test many chances to fire. If you choose the endpoint after seeing the data, you have purchased a tail area with degrees of freedom the null distribution never knew about. None of this requires dishonesty; the garden of forking paths is wide, and every fork is a test that nobody wrote down. The remedy is pre-specification: decide the test, the endpoint and the stopping rule before the data arrive, or disclose the family and correct for it.
Family-wise error rate against the number of null tests m. The rising curve is uncorrected; the Bonferroni and Šidák curves are pinned at $\alpha=0.05$; the dots are a simulation of many families.
Look at the simulated dots against the rising curve: the theory is not pessimistic, it is exact. Then look at the two flat lines. Bonferroni and Šidák control the family-wise error at $\alpha$ no matter how large the family grows, at the cost of making each individual test harder to pass. There is an unavoidable exchange rate here. If you want the freedom to search, you must pay for it either in stricter thresholds, in a held-out confirmation set, or in a prior that shrinks small effects toward zero. The one thing you cannot do is run the family and report the winner as if it were the only test.
Where this shows up
The test is everywhere data are compared
Leaderboards are families of tests
Comparing many models on many tasks is a multiplicity problem with dozens of correlated hypotheses. A single best-on-the-board result is the maximum of a noisy family, so benchmark claims need confidence intervals and corrections, not just a winning score.
Significance in model evaluation
A/B tests and evaluation suites decide whether an update ships, and they run at scale. Bootstrapped intervals, paired tests and multiplicity control are the difference between a real gain and a lucky slice of evaluation noise.
Tests and intervals are duals
Rejecting the null at level $\alpha$ is the same as the $1-\alpha$ confidence interval excluding the null value. That duality is built in Confidence vs credible, and it is why an interval always tells you more than a binary verdict.
Selecting the best model is a test
Keeping whichever configuration scored best on a validation set is a selection over a family, and the winner's curse inflates its apparent accuracy. Robust pipelines account for this, much as nonlinear optimization accounts for noise in the objective it minimizes.
Cheat sheet
Every formula in one place
| Idea | Formula | Reading |
|---|---|---|
| Null hypothesis | H_0 | The default: no effect, no difference. Assumed while you compute. |
| Alternative | H_1 | The effect you hope to detect; sets the power, not the threshold. |
| Type I error | $\alpha=P(\text{reject}\mid H_0)$ | False alarm rate. You choose it; conventionally 0.05. |
| Type II error | $\beta=P(\text{fail to reject}\mid H_1)$ | Missed effect rate. Inherited from data and effect size. |
| Power | $1-\beta$ | Chance the experiment succeeds when the effect is real. |
| Noncentrality | $\delta\sqrt n$ | Distance between the two distributions; grows like $\sqrt n$. |
| p-value | $p=P(T\ge t_{\text{obs}}\mid H_0)$ | A tail area under the null, not $P(H_0\mid\text{data})$. |
| p under H_0 | uniform on [0,1] | Every value equally likely; 5% land below 0.05 by construction. |
| Sample size | $n\approx((z_{1-\alpha}+z_{\text{power}})/\delta)^2$ | Design equation: small effects cost quadratically more data. |
| Family-wise error | $1-(1-\alpha)^m$ | Chance of at least one false positive across m null tests. |
| Bonferroni | $\alpha/m$ | Per-test threshold that caps family-wise error at $\alpha$. |
| Šidák | $1-(1-\alpha)^{1/m}$ | Exact per-test threshold for independent tests; slightly looser. |
Further reading
Where to go deeper
- Ronald A. Fisher, Statistical Methods for Research Workers, 1925 — where the p-value and the notion of a significance level were introduced.
- Jerzy Neyman and Egon Pearson, "On the Problem of the Most Efficient Tests of Statistical Hypotheses", 1933 — the two-error framework, power and the alternative.
- Jacob Cohen, Statistical Power Analysis for the Behavioral Sciences, 1988 — effect sizes and the design equation in practice.
- Andrew Gelman and Eric Loken, "The Statistical Crisis in Science", 2014 — the garden of forking paths and why p-hacking needs no fraud.
- Ronald L. Wasserstein and Nicole A. Lazar, "The ASA Statement on p-Values", 2016 — the six things a p-value is not, written by the professional society.
- Bradley Efron and Trevor Hastie, Computer Age Statistical Inference, chapters 3 and 4 — testing and false discovery, with the same simulations as this part.