Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

If there is an effect, how often will you see it?

The previous part built a null distribution by shuffling labels and asked how surprising the observed data would be if nothing were going on. That gives a p-value and a decision. But a decision procedure has a second job a p-value does not describe. Rejecting when the null is true is only one of the two mistakes available; the other — the quieter one — is failing to reject when an effect genuinely exists. A study that can only rarely detect a real effect is one whose negative results mean almost nothing.

Both mistakes have names. A type I error is a false positive: the null is true and the test rejects it, with probability the significance level $\alpha$. A type II error is a miss: the alternative is true and the test fails to reject, with probability written $\beta$. Power is the complement, $1-\beta$: the probability of detecting an effect that is really there. Neither error rate is visible in a single result; both are properties of the procedure before the data arrive.

This part makes the two error rates something you can drag. Two normal densities share one axis — one for the world with no effect, one for the world with one — and a critical value cuts the axis into accept and reject. The null's mass past the cut is $\alpha$, the alternative's mass before it is $\beta$, and the rest of the alternative is power. Move the four levers below and watch the outcomes re-balance.

💡 By the end of this part you'll see why a test has two error rates rather than one, why power depends on the effect size, the sample size, the threshold and the noise, why an underpowered study distorts the literature even when it reports nothing, and why a confidence interval excluding zero is the same event as a significant test.
2

Two densities, four outcomes

The same data, read under two worlds

Picture a simple question: does a treatment change the mean of a measurement? Under the null hypothesis the sample mean is distributed around zero; under the alternative it is centred at the true effect $d$. Both distributions are normal with the same spread, the standard error $\mathrm{SE} = \sigma/\sqrt{n}$; here we set $\sigma = 1$, so the spread is $1/\sqrt{n}$.

The test places a vertical line on this axis, the critical value, and rejects the null whenever the sample mean lands to its right. For a one-sided test at level $\alpha$, that line sits $z_{1-\alpha}$ standard errors above zero. The line defines a rejection region, and the two densities pour mass across it in different amounts: the null's right tail beyond the line is $\alpha$, the false-positive rate, while the alternative's mass below the line is $\beta$, the miss rate. The rest of the alternative is power.

Drag the sliders. The effect-size slider slides the alternative density left and right. The sample-size slider squeezes both densities narrower or fatter, because the standard error shrinks like $1/\sqrt{n}$. The $\alpha$ slider moves the critical value, trading false positives for misses. The shaded areas and readout follow immediately, and the simulation button checks the claimed power by running thousands of imaginary experiments and counting the rejections.

Dashed: the null density, centred at zero. Solid: the alternative, centred at the true effect. The vertical line is the critical value; its right side is the rejection region.

⚠ Power is computed under the alternative, not from the data. It is the long-run rate of correct rejections if the true effect were exactly $d$. You can never observe it in one study, and it is a property you choose at design time — which is why it belongs in the planning, not the post-mortem.
3

The four levers

Effect, sample size, threshold, noise

Power is not a single knob. It is the product of four facts about the design, and each one can be moved, at a cost.

The effect size is the first. A large effect separates the two densities, so little of the alternative falls below the critical value. A tiny effect leaves the humps almost on top of each other, and almost all of the alternative sits in the acceptance region. So a study can be capable of detecting a large effect and hopeless for a small one: power is always power against a specific effect.

The sample size is the lever you actually control. More observations shrink the standard error, narrowing both densities around their centres, so the alternative's overlap with the acceptance region falls and power climbs. The climb is slow: the standard error shrinks like $1/\sqrt{n}$, so quadrupling the sample size only halves it, and going from a barely detectable effect to a comfortable one can take an order of magnitude more data.

The significance level is the third, and the one with a direct trade. Loosening $\alpha$ moves the critical value left, growing the rejection region, so the false-positive rate and the power rise together; tightening $\alpha$ does the reverse. There is no free setting: a study that insists on $\alpha = 0.001$ will miss effects one at $0.05$ would catch, and fewer false alarms means more silent misses.

The noise is the fourth. The standard error is $\sigma/\sqrt{n}$, so reducing the measurement variance is as powerful as adding observations — and often cheaper. A better instrument, a tighter protocol, or a paired design that cancels subject-to-subject variation multiplies power without a single extra participant. In machine learning this lever is the one under the most control: a cleaner evaluation set or a lower-variance metric buys detection ability directly.

Power as a function of sample size for the current effect and threshold. The dashed line is the conventional 80% target; the marker is the n you selected above.

The curve rises and then flattens — the classic diminishing return. Early observations buy power quickly because the standard error falls fast when it is large, but squeezing the last few points of power out of a tiny effect may demand more data than the experiment can ever justify.

Planning a study means choosing three of the four and solving for the last. Conventionally the target is 80% power: fix the smallest effect you care about, your tolerable false-positive rate and your noise, then compute the sample size that gives an 80% chance of detecting it. The readout above reports that number. It is a minimum, not a guarantee, and the smaller the effect you insist on detecting the more brutal the requirement becomes.

The formal relation ties the pieces together. For a one-sided test the standardized effect $d/\mathrm{SE} = d\sqrt{n}$ must clear the critical value while leaving room for the target power:

$$n \;\ge\; \left(\frac{z_{1-\alpha} + z_{1-\beta}}{d}\right)^2$$

Read it as a sentence: the sample size is the square of the required separation divided by the effect, so doubling the effect quarters the data you need, and halving it quadruples the bill.

4

Why underpowered studies mislead

The damage is not only the missed effect

It is tempting to treat low power as a mild inconvenience: you ran a study, found nothing, and conclude the effect is small or absent. That inference is invalid. A non-significant result from an underpowered test is barely evidence of anything, because the test would have missed even a substantial effect most of the time. Absence of evidence is not evidence of absence, and without knowing the power you cannot tell the two apart.

There is a subtler and more damaging consequence. When power is low, the only studies that reach significance are those in which noise happened to push the estimate far from zero. Conditional on a significant result, the reported effect is systematically exaggerated — the winner's curse. The literature then fills with effect sizes larger than the truth, not because anyone cheated but because selection did the distorting. Meta-analyses inherit the inflation, and replication fails against effects that never existed at the advertised size.

This is why a power calculation is an ethical as well as a statistical matter. An underpowered experiment spends real participants, time and money on a question it cannot answer. Worse, its positive results are unreliable and its negative results uninformative, so it adds noise to the record rather than knowledge. The remedy is to decide the smallest effect worth detecting, then size the study to detect it.

Power is also the mirror image of the confidence interval. A test rejects at level $\alpha$ exactly when the $1-\alpha$ interval excludes the null value. So "80% power" also means "if the effect had been $d$, there was an 80% chance the interval would have excluded zero." Power is the expected reliability of that interval, fixed before the data are seen; a very wide interval is the signature of a study that was never going to be decisive.

5

Where this shows up

Two thresholds, two costs

ML / AI

Detecting a small benchmark improvement

A new model improves a benchmark by a fraction of a point, and the question is whether that is real. The evaluation chapter estimates error on a finite test set, so the measured delta is a sample mean with a standard error. A small test set has low power against a small improvement, which is how teams ship "no change" verdicts that later reverse. Sizing the evaluation set is a power calculation.

Robotics

Sensor fault detection thresholds

A robot comparing a sensor reading against an expected value is running a hypothesis test every cycle. A loose threshold lets a drifting sensor pass as healthy; a tight one makes ordinary noise trigger false alarms that waste recovery behaviour. The pose filter fuses measurements into a pose estimate, and fault detection is the residual check around it. The threshold is the critical value, and choosing it is the $\alpha$/$\beta$ trade under real cost.

The same accounting appears wherever a decision is made from noisy evidence. A serving dashboard watching for latency regressions in the metrics chapter trades alert fatigue against missed outages, which is exactly the $\alpha$/$\beta$ trade. In probability, the normal quantile that locates the critical value is the same object that sets the width of every interval. And a threshold on a reprojection error decides whether a correspondence is an inlier — low power there quietly discards good data. Deciding what counts as signal and what counts as noise is the same problem in every one of them.

Further reading

The references below treat power as a design obligation rather than an afterthought. If you take away one thing, take away the picture of two overlapping densities and the three areas a critical value carves out of them.

Cohen's book is the standard practical treatment of power analysis and the origin of the 80% convention; it is also honest that the convention is a convention. Wasserman places the test, its two errors and its power inside the formal decision framework this series follows, and the experimental-design literature is where the four levers get their real prices.

Cheat sheet

TermMeaning here
Type I error ($\alpha$)Rejecting a true null; the false-positive rate, set by the critical value
Type II error ($\beta$)Failing to reject a true alternative; the miss rate
Power ($1-\beta$)Probability of detecting a real effect of a given size
Critical valueThe threshold on the test statistic; $z_{1-\alpha}$ standard errors from the null
Effect size $d$The true separation of the two densities, in measurement units
Standard error$\sigma/\sqrt{n}$; the shared spread of both sampling distributions
Sample size for power$n \ge ((z_{1-\alpha}+z_{1-\beta})/d)^2$ for a one-sided normal test
CI connectionRejecting at level $\alpha$ is the same event as the $1-\alpha$ interval excluding the null value
7

Check your understanding

0/4 answered