Randomization, A/B tests, and peeking
An A/B test looks like the simplest experiment in the world: show one version to some users, another to others, and see which converts better. What makes the result trustworthy is not the arithmetic but the assignment — randomisation turns a difference between two groups into a claim about cause. And the moment you watch the dashboard and stop the test when it looks good, that guarantee quietly stops holding. This part builds a sequential-testing simulator: under a true null, with no real difference between A and B, you watch the false-positive rate climb far past the level you chose.
The question
Does B beat A — and would you believe a p-value you stopped on?
You have a baseline experience A and a proposed change B, and you want to know whether B causes more conversions or whether any apparent difference is just the ordinary noise of a finite sample. The standard answer is the A/B test: split traffic at random, measure each arm, and run a two-sample test on the difference.
The interesting failure is not in the test but in the watching. A test is designed as a single, pre-declared look at a pre-declared amount of data, but dashboards update continuously and teams check them. Peek every day, stop the moment the p-value dips below your threshold, and report that flattering number — and the error rate you thought you controlled is no longer the error rate you have. Under the null it does not drift; it climbs fast.
This part explains why randomisation, not the formula, licenses a causal claim, treats the A/B test as the two-sample test it is, and then lets you peek. The simulator runs thousands of experiments in which A and B are truly identical and stops as soon as the test is significant, measuring how often it declares a winner that is not there.
Randomization and causal claims
Why the coin, not the calculator, does the work
Compare users who happened to see the new page with users who happened to see the old one, and any difference is entangled with every other way the groups differ: the page might be served to returning visitors, on faster devices, in one country. Those are confounders, variables that affect the outcome and differ between the groups, and no amount of data cleans them out because most were never measured.
Random assignment breaks the entanglement. With each user's arm decided by an independent coin flip, the groups are identical in expectation in every respect except the treatment — the variables you recorded and the ones you never thought to. That is why randomisation licenses a causal reading: the assignment is the only systematic difference left. It is a statement about the procedure, not about any one user, whose outcome is still a flip in the dark. Randomisation balances confounders on average, not perfectly in every realised split, and that residual imbalance is the sampling wobble the p-value quantifies. Each user also has two potential outcomes, only one of which is observed.
The A/B test as a two-sample test
One comparison, one horizon
Once the groups are comparable, the comparison is a two-sample test. The null says both arms come from the same distribution; the alternative says they differ. For a conversion metric each user contributes a zero or a one, and the natural statistic compares two sample means. The two-sample Welch t-statistic does the job: with means $ \bar x_A, \bar x_B $, variances $ s_A^2, s_B^2 $ and sizes $ n_A, n_B $,
Under the null the statistic follows a t distribution with Welch's approximate degrees of freedom, and the two-sided p-value is the probability of a difference at least this extreme if A and B were the same. A small p means the data are surprising under the null, and $ \alpha $ is the rate at which you agreed to be surprised when nothing is going on.
All of this is defined for a test fixed in advance: the sample size, the metric, and the single point at which you look are part of the design. A p-value is the tail probability of a statistic whose distribution is known only once the rule for reading it is fixed. Power and sample size (Part 13) are the other half of that design, and choosing them beforehand is what makes a significant result mean something.
Peeking and optional stopping
The simulator: watch the false-positive rate climb
The temptation is obvious: data arrive continuously, and you open the dashboard every morning. A single look at $ \alpha = 0.05 $ is wrong one time in twenty under the null, but you are taking a sequence of looks and stopping at the first that dips below the line. Were the looks independent, the chance of at least one false positive over $ L $ looks would be $ 1-(1-\alpha)^L $ — about 40% for ten looks at $ \alpha = 0.05 $. They share data rather than being independent, but the direction is the same, and the simulator measures the real number.
Press Run simulation and the canvas does the experiment you cannot do on your dashboard. It generates thousands of worlds in which A and B are identical, streams the data over the number of looks you chose, and declares significance the moment the running p-value falls below $ \alpha $. Nobody is peeking opportunistically; the rule is fixed and mechanical. Optional stopping needs no bad intent, only a stopping time that depends on the data.
Top: the running p-value for six simulated experiments under the null, against the stopping threshold. Bottom: how often the procedure has declared a winner by the time that many looks have been taken, versus the level you chose.
Read the lines together. The dashed line is $ \alpha $, the rate you think you have. The marker at the right is a single test taken at the very end, with no peeking: it lands near $ \alpha $, the guarantee a fixed-horizon test gives. The rising curve is the peek-every-look procedure — several times $ \alpha $ at ten looks, and higher the more you look. At any fixed look the p-value is uniform on $ [0,1] $, so each look has chance $ \alpha $ of looking small; but the sequence is a random walk and you are waiting for its minimum to cross a threshold, and the minimum of a random walk is biased downward. The number you report is not the probability of data this extreme under the null; it is the probability of reaching that stopping time with data at least this extreme.
Always-valid inference
Buying the guarantee back, at a price
The lesson is not that peeking is forbidden: data stream in whether you look or not, and refusing to look wastes information. The error budget must be spent across the looks rather than re-spent at each one. That is alpha spending — choose thresholds $ \alpha_1,\dots,\alpha_L $ whose sum is bounded by the level you chose,
The crude version is the Bonferroni split $ \alpha_k \approx \alpha / L $, valid but harsher than necessary because the looks are correlated. A spending function spends little early, when evidence is thin, and more later: Pocock uses a roughly constant boundary, while O'Brien–Fleming is conservative early and approaches the fixed-horizon $ \alpha $ at the end. A group-sequential test builds these boundaries into a two-sample test, holding the family-wise rate at $ \alpha $ while allowing an early stop for a large effect.
Always-valid inference goes further, giving a confidence sequence whose whole path covers the true effect with probability $ 1-\alpha $ uniformly over time; you may peek whenever you like and the coverage survives. The price is a wider band than a fixed-horizon interval at the same sample size. The connection to multiple comparisons (Part 14) is direct: ten looks is a family of ten tests, inflating the error rate for the same reason testing ten metrics does. Decide the stopping rule before you look, and if you cannot resist peeking, use a method whose guarantee survives it. The red curve's inflation is this same bill seen from the other side.
Where this shows up
The same stopping rule, two different dashboards
Online A/B tests of serving changes
A new quantization level, cache policy, or routing rule is rolled out to a fraction of traffic and judged on latency and error rate. These metrics are watched live, which makes the peeking problem the default rather than the exception. The serving metrics chapter defines the quantities; when they may be read is this part's stopping rule. If the gate stops at the first significant latency win, its false-positive rate is not the level it was configured with.
Evaluation is a stopping decision
Choosing a checkpoint, a prompt, or a reward model by its score on a held-out set is a sequence of comparisons, and taking the best one is optional stopping in another costume. The evaluation chapter reports scores on finite test sets; selecting the winner of many trials inflates the apparent gain, exactly as looking many times inflates the false-positive rate. The fix is the same: pre-declare the comparison, or use a method that prices the search.
The pattern is general. Anywhere a decision rule adapts to accumulating evidence — a bandit shifting traffic toward the leader, a dashboard alerting on the first anomaly, a model-selection loop promoting the current best — the naive reading of the evidence is optimistic, because the rule that chose it was itself chosen from it. Separate the design that generates the data from the rule that stops on it.
Further reading
The references below cover the classical sequential-analysis line and the modern always-valid one. Wald founded sequential analysis; Armitage and colleagues showed that repeated significance testing inflates the error rate; Pocock and O'Brien–Fleming supplied the boundaries trials still use; always-valid inference frees you from fixing the number of looks at all.
- Abraham Wald, Sequential Analysis, 1947 — the original theory of tests that stop on the data.
- Peter Armitage et al., "Repeated Significance Tests on Accumulating Data", JRSS A, 1969 — how peeking inflates the error rate.
- Stuart Pocock, "Group sequential methods in the design and analysis of clinical trials", Biometrika, 1977 — the constant-boundary design.
- Peter O'Brien and Thomas Fleming, "A multiple testing procedure for clinical trials", Biometrics, 1979 — conservative early, near-nominal at the end.
- Ron Kohavi, Diane Tang, and Ya Xu, Trustworthy Online Controlled Experiments, 2020 — A/B testing as practised, peeking trap included.
- Ramesh Johari et al., "Always Valid Inference: Continuous Monitoring of A/B Tests", Operations Research, 2022 — confidence sequences you can peek at freely.
Cheat sheet
| Term | Meaning here |
|---|---|
| Randomization | Independent coin flips assigning units to arms; balances confounders in expectation |
| A/B test | A two-sample test comparing the arms; for a binary metric, a two-sample Welch $ t $ test |
| Fixed horizon | A sample size and a single look declared in advance; the setting a classical p-value is defined in |
| Optional stopping | Choosing when to stop from the data; makes the reported p-value the minimum of a random walk |
| Peeking | Re-testing as data accumulate; each look re-spends error budget it did not have to spend |
| Alpha spending | A schedule $ \alpha_1,\dots,\alpha_L $ with $ \sum_k \alpha_k \le \alpha $, so the family-wise rate stays at $ \alpha $ |
| Always-valid interval | A confidence sequence covering the effect uniformly over time, so any stopping rule is safe |
| Family-wise error rate | Chance of at least one false positive across all looks; what inflates under peeking |