Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Does B beat A — and would you believe a p-value you stopped on?

You have a baseline experience A and a proposed change B, and you want to know whether B causes more conversions or whether any apparent difference is just the ordinary noise of a finite sample. The standard answer is the A/B test: split traffic at random, measure each arm, and run a two-sample test on the difference.

The interesting failure is not in the test but in the watching. A test is designed as a single, pre-declared look at a pre-declared amount of data, but dashboards update continuously and teams check them. Peek every day, stop the moment the p-value dips below your threshold, and report that flattering number — and the error rate you thought you controlled is no longer the error rate you have. Under the null it does not drift; it climbs fast.

This part explains why randomisation, not the formula, licenses a causal claim, treats the A/B test as the two-sample test it is, and then lets you peek. The simulator runs thousands of experiments in which A and B are truly identical and stops as soon as the test is significant, measuring how often it declares a winner that is not there.

💡 By the end of this part you'll see why random assignment is the source of causal warrant, why a fixed-horizon p-value is defined only for a horizon fixed in advance, and how optional stopping inflates the false-positive rate — along with the alpha-spending and always-valid ideas that buy the guarantee back.
2

Randomization and causal claims

Why the coin, not the calculator, does the work

Compare users who happened to see the new page with users who happened to see the old one, and any difference is entangled with every other way the groups differ: the page might be served to returning visitors, on faster devices, in one country. Those are confounders, variables that affect the outcome and differ between the groups, and no amount of data cleans them out because most were never measured.

Random assignment breaks the entanglement. With each user's arm decided by an independent coin flip, the groups are identical in expectation in every respect except the treatment — the variables you recorded and the ones you never thought to. That is why randomisation licenses a causal reading: the assignment is the only systematic difference left. It is a statement about the procedure, not about any one user, whose outcome is still a flip in the dark. Randomisation balances confounders on average, not perfectly in every realised split, and that residual imbalance is the sampling wobble the p-value quantifies. Each user also has two potential outcomes, only one of which is observed.

3

The A/B test as a two-sample test

One comparison, one horizon

Once the groups are comparable, the comparison is a two-sample test. The null says both arms come from the same distribution; the alternative says they differ. For a conversion metric each user contributes a zero or a one, and the natural statistic compares two sample means. The two-sample Welch t-statistic does the job: with means $ \bar x_A, \bar x_B $, variances $ s_A^2, s_B^2 $ and sizes $ n_A, n_B $,

$$t = \frac{\bar x_A - \bar x_B}{\sqrt{\dfrac{s_A^2}{n_A} + \dfrac{s_B^2}{n_B}}}$$

Under the null the statistic follows a t distribution with Welch's approximate degrees of freedom, and the two-sided p-value is the probability of a difference at least this extreme if A and B were the same. A small p means the data are surprising under the null, and $ \alpha $ is the rate at which you agreed to be surprised when nothing is going on.

All of this is defined for a test fixed in advance: the sample size, the metric, and the single point at which you look are part of the design. A p-value is the tail probability of a statistic whose distribution is known only once the rule for reading it is fixed. Power and sample size (Part 13) are the other half of that design, and choosing them beforehand is what makes a significant result mean something.

4

Peeking and optional stopping

The simulator: watch the false-positive rate climb

The temptation is obvious: data arrive continuously, and you open the dashboard every morning. A single look at $ \alpha = 0.05 $ is wrong one time in twenty under the null, but you are taking a sequence of looks and stopping at the first that dips below the line. Were the looks independent, the chance of at least one false positive over $ L $ looks would be $ 1-(1-\alpha)^L $ — about 40% for ten looks at $ \alpha = 0.05 $. They share data rather than being independent, but the direction is the same, and the simulator measures the real number.

Press Run simulation and the canvas does the experiment you cannot do on your dashboard. It generates thousands of worlds in which A and B are identical, streams the data over the number of looks you chose, and declares significance the moment the running p-value falls below $ \alpha $. Nobody is peeking opportunistically; the rule is fixed and mechanical. Optional stopping needs no bad intent, only a stopping time that depends on the data.

Top: the running p-value for six simulated experiments under the null, against the stopping threshold. Bottom: how often the procedure has declared a winner by the time that many looks have been taken, versus the level you chose.

Read the lines together. The dashed line is $ \alpha $, the rate you think you have. The marker at the right is a single test taken at the very end, with no peeking: it lands near $ \alpha $, the guarantee a fixed-horizon test gives. The rising curve is the peek-every-look procedure — several times $ \alpha $ at ten looks, and higher the more you look. At any fixed look the p-value is uniform on $ [0,1] $, so each look has chance $ \alpha $ of looking small; but the sequence is a random walk and you are waiting for its minimum to cross a threshold, and the minimum of a random walk is biased downward. The number you report is not the probability of data this extreme under the null; it is the probability of reaching that stopping time with data at least this extreme.

⚠ A p-value is only as good as the horizon it was computed for. A number that would be a valid tail probability at a pre-declared look is merely the running minimum of a process once the stopping time is data-dependent. Reporting it as if the horizon were fixed is the error, and it is invisible in the output: the screen shows 0.03, just as it would for an honest test.
5

Always-valid inference

Buying the guarantee back, at a price

The lesson is not that peeking is forbidden: data stream in whether you look or not, and refusing to look wastes information. The error budget must be spent across the looks rather than re-spent at each one. That is alpha spending — choose thresholds $ \alpha_1,\dots,\alpha_L $ whose sum is bounded by the level you chose,

$$P(\text{any false positive}) \;\le\; \sum_{k=1}^{L} \alpha_k \;\le\; \alpha .$$

The crude version is the Bonferroni split $ \alpha_k \approx \alpha / L $, valid but harsher than necessary because the looks are correlated. A spending function spends little early, when evidence is thin, and more later: Pocock uses a roughly constant boundary, while O'Brien–Fleming is conservative early and approaches the fixed-horizon $ \alpha $ at the end. A group-sequential test builds these boundaries into a two-sample test, holding the family-wise rate at $ \alpha $ while allowing an early stop for a large effect.

Always-valid inference goes further, giving a confidence sequence whose whole path covers the true effect with probability $ 1-\alpha $ uniformly over time; you may peek whenever you like and the coverage survives. The price is a wider band than a fixed-horizon interval at the same sample size. The connection to multiple comparisons (Part 14) is direct: ten looks is a family of ten tests, inflating the error rate for the same reason testing ten metrics does. Decide the stopping rule before you look, and if you cannot resist peeking, use a method whose guarantee survives it. The red curve's inflation is this same bill seen from the other side.

6

Where this shows up

The same stopping rule, two different dashboards

ML / AI Serving

Online A/B tests of serving changes

A new quantization level, cache policy, or routing rule is rolled out to a fraction of traffic and judged on latency and error rate. These metrics are watched live, which makes the peeking problem the default rather than the exception. The serving metrics chapter defines the quantities; when they may be read is this part's stopping rule. If the gate stops at the first significant latency win, its false-positive rate is not the level it was configured with.

ML / AI Training

Evaluation is a stopping decision

Choosing a checkpoint, a prompt, or a reward model by its score on a held-out set is a sequence of comparisons, and taking the best one is optional stopping in another costume. The evaluation chapter reports scores on finite test sets; selecting the winner of many trials inflates the apparent gain, exactly as looking many times inflates the false-positive rate. The fix is the same: pre-declare the comparison, or use a method that prices the search.

The pattern is general. Anywhere a decision rule adapts to accumulating evidence — a bandit shifting traffic toward the leader, a dashboard alerting on the first anomaly, a model-selection loop promoting the current best — the naive reading of the evidence is optimistic, because the rule that chose it was itself chosen from it. Separate the design that generates the data from the rule that stops on it.

Further reading

The references below cover the classical sequential-analysis line and the modern always-valid one. Wald founded sequential analysis; Armitage and colleagues showed that repeated significance testing inflates the error rate; Pocock and O'Brien–Fleming supplied the boundaries trials still use; always-valid inference frees you from fixing the number of looks at all.

Cheat sheet

TermMeaning here
RandomizationIndependent coin flips assigning units to arms; balances confounders in expectation
A/B testA two-sample test comparing the arms; for a binary metric, a two-sample Welch $ t $ test
Fixed horizonA sample size and a single look declared in advance; the setting a classical p-value is defined in
Optional stoppingChoosing when to stop from the data; makes the reported p-value the minimum of a random walk
PeekingRe-testing as data accumulate; each look re-spends error budget it did not have to spend
Alpha spendingA schedule $ \alpha_1,\dots,\alpha_L $ with $ \sum_k \alpha_k \le \alpha $, so the family-wise rate stays at $ \alpha $
Always-valid intervalA confidence sequence covering the effect uniformly over time, so any stopping rule is safe
Family-wise error rateChance of at least one false positive across all looks; what inflates under peeking
7

Check your understanding

0/4 answered