Errors, power, and sample size
A test is a rule that looks at data and returns one of two answers: there is an effect, or there is not enough evidence for one. Both answers can be wrong, and the two ways of being wrong pull against each other. Set the bar low and you will cry wolf on data that are only noise; set it high and real effects will walk past undetected. The whole trade is visible in one picture of two overlapping densities, and the quantity that summarises the second kind of error — the probability that you actually catch an effect when it is there — is called power. Power is not a property of the test alone. It depends on how big the effect really is, how much data you collected, how much noise sits in each measurement, and where you chose to put the bar.
The question
If there is an effect, how often will you see it?
The previous part built a null distribution by shuffling labels and asked how surprising the observed data would be if nothing were going on. That gives a p-value and a decision. But a decision procedure has a second job a p-value does not describe. Rejecting when the null is true is only one of the two mistakes available; the other — the quieter one — is failing to reject when an effect genuinely exists. A study that can only rarely detect a real effect is one whose negative results mean almost nothing.
Both mistakes have names. A type I error is a false positive: the null is true and the test rejects it, with probability the significance level $\alpha$. A type II error is a miss: the alternative is true and the test fails to reject, with probability written $\beta$. Power is the complement, $1-\beta$: the probability of detecting an effect that is really there. Neither error rate is visible in a single result; both are properties of the procedure before the data arrive.
This part makes the two error rates something you can drag. Two normal densities share one axis — one for the world with no effect, one for the world with one — and a critical value cuts the axis into accept and reject. The null's mass past the cut is $\alpha$, the alternative's mass before it is $\beta$, and the rest of the alternative is power. Move the four levers below and watch the outcomes re-balance.
Two densities, four outcomes
The same data, read under two worlds
Picture a simple question: does a treatment change the mean of a measurement? Under the null hypothesis the sample mean is distributed around zero; under the alternative it is centred at the true effect $d$. Both distributions are normal with the same spread, the standard error $\mathrm{SE} = \sigma/\sqrt{n}$; here we set $\sigma = 1$, so the spread is $1/\sqrt{n}$.
The test places a vertical line on this axis, the critical value, and rejects the null whenever the sample mean lands to its right. For a one-sided test at level $\alpha$, that line sits $z_{1-\alpha}$ standard errors above zero. The line defines a rejection region, and the two densities pour mass across it in different amounts: the null's right tail beyond the line is $\alpha$, the false-positive rate, while the alternative's mass below the line is $\beta$, the miss rate. The rest of the alternative is power.
Drag the sliders. The effect-size slider slides the alternative density left and right. The sample-size slider squeezes both densities narrower or fatter, because the standard error shrinks like $1/\sqrt{n}$. The $\alpha$ slider moves the critical value, trading false positives for misses. The shaded areas and readout follow immediately, and the simulation button checks the claimed power by running thousands of imaginary experiments and counting the rejections.
Dashed: the null density, centred at zero. Solid: the alternative, centred at the true effect. The vertical line is the critical value; its right side is the rejection region.
The four levers
Effect, sample size, threshold, noise
Power is not a single knob. It is the product of four facts about the design, and each one can be moved, at a cost.
The effect size is the first. A large effect separates the two densities, so little of the alternative falls below the critical value. A tiny effect leaves the humps almost on top of each other, and almost all of the alternative sits in the acceptance region. So a study can be capable of detecting a large effect and hopeless for a small one: power is always power against a specific effect.
The sample size is the lever you actually control. More observations shrink the standard error, narrowing both densities around their centres, so the alternative's overlap with the acceptance region falls and power climbs. The climb is slow: the standard error shrinks like $1/\sqrt{n}$, so quadrupling the sample size only halves it, and going from a barely detectable effect to a comfortable one can take an order of magnitude more data.
The significance level is the third, and the one with a direct trade. Loosening $\alpha$ moves the critical value left, growing the rejection region, so the false-positive rate and the power rise together; tightening $\alpha$ does the reverse. There is no free setting: a study that insists on $\alpha = 0.001$ will miss effects one at $0.05$ would catch, and fewer false alarms means more silent misses.
The noise is the fourth. The standard error is $\sigma/\sqrt{n}$, so reducing the measurement variance is as powerful as adding observations — and often cheaper. A better instrument, a tighter protocol, or a paired design that cancels subject-to-subject variation multiplies power without a single extra participant. In machine learning this lever is the one under the most control: a cleaner evaluation set or a lower-variance metric buys detection ability directly.
Power as a function of sample size for the current effect and threshold. The dashed line is the conventional 80% target; the marker is the n you selected above.
The curve rises and then flattens — the classic diminishing return. Early observations buy power quickly because the standard error falls fast when it is large, but squeezing the last few points of power out of a tiny effect may demand more data than the experiment can ever justify.
Planning a study means choosing three of the four and solving for the last. Conventionally the target is 80% power: fix the smallest effect you care about, your tolerable false-positive rate and your noise, then compute the sample size that gives an 80% chance of detecting it. The readout above reports that number. It is a minimum, not a guarantee, and the smaller the effect you insist on detecting the more brutal the requirement becomes.
The formal relation ties the pieces together. For a one-sided test the standardized effect $d/\mathrm{SE} = d\sqrt{n}$ must clear the critical value while leaving room for the target power:
Read it as a sentence: the sample size is the square of the required separation divided by the effect, so doubling the effect quarters the data you need, and halving it quadruples the bill.
Why underpowered studies mislead
The damage is not only the missed effect
It is tempting to treat low power as a mild inconvenience: you ran a study, found nothing, and conclude the effect is small or absent. That inference is invalid. A non-significant result from an underpowered test is barely evidence of anything, because the test would have missed even a substantial effect most of the time. Absence of evidence is not evidence of absence, and without knowing the power you cannot tell the two apart.
There is a subtler and more damaging consequence. When power is low, the only studies that reach significance are those in which noise happened to push the estimate far from zero. Conditional on a significant result, the reported effect is systematically exaggerated — the winner's curse. The literature then fills with effect sizes larger than the truth, not because anyone cheated but because selection did the distorting. Meta-analyses inherit the inflation, and replication fails against effects that never existed at the advertised size.
This is why a power calculation is an ethical as well as a statistical matter. An underpowered experiment spends real participants, time and money on a question it cannot answer. Worse, its positive results are unreliable and its negative results uninformative, so it adds noise to the record rather than knowledge. The remedy is to decide the smallest effect worth detecting, then size the study to detect it.
Power is also the mirror image of the confidence interval. A test rejects at level $\alpha$ exactly when the $1-\alpha$ interval excludes the null value. So "80% power" also means "if the effect had been $d$, there was an 80% chance the interval would have excluded zero." Power is the expected reliability of that interval, fixed before the data are seen; a very wide interval is the signature of a study that was never going to be decisive.
Where this shows up
Two thresholds, two costs
Detecting a small benchmark improvement
A new model improves a benchmark by a fraction of a point, and the question is whether that is real. The evaluation chapter estimates error on a finite test set, so the measured delta is a sample mean with a standard error. A small test set has low power against a small improvement, which is how teams ship "no change" verdicts that later reverse. Sizing the evaluation set is a power calculation.
Sensor fault detection thresholds
A robot comparing a sensor reading against an expected value is running a hypothesis test every cycle. A loose threshold lets a drifting sensor pass as healthy; a tight one makes ordinary noise trigger false alarms that waste recovery behaviour. The pose filter fuses measurements into a pose estimate, and fault detection is the residual check around it. The threshold is the critical value, and choosing it is the $\alpha$/$\beta$ trade under real cost.
The same accounting appears wherever a decision is made from noisy evidence. A serving dashboard watching for latency regressions in the metrics chapter trades alert fatigue against missed outages, which is exactly the $\alpha$/$\beta$ trade. In probability, the normal quantile that locates the critical value is the same object that sets the width of every interval. And a threshold on a reprojection error decides whether a correspondence is an inlier — low power there quietly discards good data. Deciding what counts as signal and what counts as noise is the same problem in every one of them.
Further reading
The references below treat power as a design obligation rather than an afterthought. If you take away one thing, take away the picture of two overlapping densities and the three areas a critical value carves out of them.
Cohen's book is the standard practical treatment of power analysis and the origin of the 80% convention; it is also honest that the convention is a convention. Wasserman places the test, its two errors and its power inside the formal decision framework this series follows, and the experimental-design literature is where the four levers get their real prices.
- Jacob Cohen, Statistical Power Analysis for the Behavioral Sciences — effect sizes, the 80% convention, and sample-size planning.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, chapters 10–11 — hypothesis testing, the two error rates and power.
- Erich Lehmann and Joseph Romano, Testing Statistical Hypotheses — the likelihood-ratio and Neyman–Pearson view of the size and power trade.
- Katherine Button and colleagues, "Power failure: why small sample size undermines the reliability of neuroscience", 2013 — the inflated-effect and replication problem in the wild.
Cheat sheet
| Term | Meaning here |
|---|---|
| Type I error ($\alpha$) | Rejecting a true null; the false-positive rate, set by the critical value |
| Type II error ($\beta$) | Failing to reject a true alternative; the miss rate |
| Power ($1-\beta$) | Probability of detecting a real effect of a given size |
| Critical value | The threshold on the test statistic; $z_{1-\alpha}$ standard errors from the null |
| Effect size $d$ | The true separation of the two densities, in measurement units |
| Standard error | $\sigma/\sqrt{n}$; the shared spread of both sampling distributions |
| Sample size for power | $n \ge ((z_{1-\alpha}+z_{1-\beta})/d)^2$ for a one-sided normal test |
| CI connection | Rejecting at level $\alpha$ is the same event as the $1-\alpha$ interval excluding the null value |