Confounding, Simpson's paradox, and DAGs
Two variables move together. Whether the movement is caused by one of them, by a third variable pushing both, or by nothing at all is not a question the data can answer on their own — it is answered by an assumption about how the world is wired. A directed acyclic graph, or DAG, writes that assumption down: which variables carry real effects, which merely correlate, and what happens when you condition on a variable. This part makes the three canonical structures watchable. Condition on the right node and a spurious association melts away; condition on the wrong node and a real one disappears, or a fake one appears that was never there.
The question
What does it mean for a number to be an effect?
Correlation is a fact about the sample you hold. Causation is a fact about the world, and the world is not in the sample. When a regression coefficient of $0.8$ comes back from a model of $Y$ on $X$, the number is real — but it is a summary of a joint distribution, not a measurement of what would happen if you intervened and set $X$ yourself. Two mechanisms can produce the same coefficient, and only one means that changing $X$ changes $Y$.
The first mechanism is direct: $X$ causes $Y$. The second is a shared cause: a third variable $Z$ influences both, so they move together without either acting on the other. The third is selection: both are consequences of a common effect, and attention is restricted to cases where that effect is extreme. All three produce correlation, and telling them apart is not solved by more data or a cleverer estimator — only by writing down what you believe and checking your analysis against it.
A DAG does exactly that. Each node is a variable; each arrow means "is a direct cause of"; acyclic means you cannot follow arrows around a loop. The graph encodes assumptions, not facts, and those assumptions license the adjustments. Once it is on the page, two questions are answerable by eye: is the association confounded, and which variables should be adjusted for — and which should not be.
Association is not causation
DAGs are the assumptions you were already making
Every regression already carries a causal claim. Fitting $Y$ on $X$ implicitly asserts that no other variable opens a back-channel between them, and a DAG makes that assertion explicit as a picture. The picture does not know whether it is true; it tells you what follows if it is, separating the mathematics — is the effect identified? — from the domain knowledge where the assumptions come from.
The vocabulary is small. A path is any route between two nodes, following arrows in either direction. A directed path follows arrows forward and represents a possible causal influence. A back-door path is any path from $X$ to $Y$ that starts by following an arrow into $X$ — one that enters through $X$'s causes rather than leaving through its effects. Confounding is the presence of an open back-door path, and adjustment aims to block all of them while leaving the causal paths alone.
Conditioning means holding a variable fixed: stratifying by it, regressing it out, matching on it. In graph language, conditioning on a node blocks the paths that run through it — but only under a rule you have to be careful about, because a node with two incoming arrows behaves the opposite way. That rule is the whole of this part.
Fork, chain, collider
Three ways three variables can be wired
Three nodes and two arrows give exactly three shapes, and they behave in three different ways. Click between them below. The same $240$ observations are generated by the graph you see, the scatter plots $X$ against $Y$, and the points are coloured by whether $Z$ is above or below its median. The toggle marks $Z$ as conditioned; the readout then reports the partial correlation of $X$ and $Y$ given $Z$, obtained by regressing each on $Z$ and correlating the residuals.
Top: the assumed graph, with Z filled when conditioned. Bottom: the $X$–$Y$ scatter, coloured by $Z$'s median split, with the unadjusted fit (dashed) and, when conditioned, the fit adjusted for $Z$ (solid).
The fork is the classic confounder: $Z$ is a common cause, $X \leftarrow Z \rightarrow Y$. Because $Z$ pushes both, $X$ and $Y$ correlate even if neither causes the other, and the unconditional correlation is large. Conditioning on $Z$ blocks the only path, so the partial correlation collapses toward zero. Whatever $Z$ is — a demographic, a time of day, an environmental factor — leaving it out attributes its work to $X$.
The chain is a mediator: $X \rightarrow Z \rightarrow Y$. Here $Z$ is not a nuisance but the mechanism. The causal influence travels through $Z$, so the marginal correlation is a real effect and conditioning on $Z$ blocks the path that carries it. The partial correlation drops to zero not because confounding was removed but because the effect was deleted. This is over-control, and it is why "control for everything you measured" erases the very thing you are estimating.
The collider is the surprise: $X \rightarrow Z \leftarrow Y$. Marginally, $X$ and $Y$ are independent and the scatter is a formless cloud. But $Z$ is a common consequence, and conditioning on it — selecting cases where $Z$ is high, or holding it fixed in a regression — induces a negative association. Knowing $Z$ is high, an unusually large $X$ implies a smaller $Y$ that still sums to the same total. Statisticians call it explaining away; the graph calls it a collision, and it is the one node you must never adjust for. Use the strength slider to see the manufactured correlation grow.
Simpson's paradox
A reversal you can watch appear
The three shapes explain the most famous embarrassment in applied statistics. Suppose every group in a dataset shows one trend, but the pooled dataset shows the opposite. That is Simpson's paradox, and it looks like a paradox only until the graph is drawn: group membership is a common cause of both the predictor and the outcome. Within a group the real relationship is visible; pool the groups without adjusting and the between-group difference swamps it.
Below, two groups are generated with an identical negative within-group slope of $-1$. Group 1 simply sits higher and to the right — it has both more predictor and more outcome — so the dashed pooled line slopes upward. The slider moves the groups apart; at zero separation there is no paradox, and as the separation grows the pooled slope flips while the group slopes do not budge.
Group 0 and group 1 with their within-group fits, plus the pooled fit. Drag the gap and watch the dashed line reverse.
The lesson is not that one analysis is right and the other wrong by arithmetic. Both lines are correctly computed from the same numbers. The question is which corresponds to an intervention: if you could change the predictor within a group, the group slope is what you would get, and the pooled slope is a mixture of that effect with the group difference. The pooled line answers "given an observation with a high predictor, what outcome should I expect?" — prediction, not cause. Mixing the two is how a real treatment gets blamed for a difference that is entirely due to who was selected into it.
The back-door criterion
Adjust for confounders, not for everything
Combining the three cases gives a usable rule. To estimate the causal effect of $X$ on $Y$, find a set $S$ that blocks every back-door path and contains no descendant of $X$. If such a set exists the effect is identified, and it can be estimated by adjusting for $S$; the directed causal path survives because none of its interior nodes is in $S$.
Two failure modes make the rule worth stating precisely. Under-adjustment leaves a common cause out of $S$, so a back-door path stays open and the estimate is confounded. Over-adjustment is subtler and more common, because including every column is so tempting. Adding a mediator deletes part of the causal path. Adding a collider, or a descendant of one, opens a path that was closed and actively creates bias. The two errors look identical in a coefficient table and have opposite causes.
The practical consequence is that variable selection is a causal act. "Coefficients after controlling for everything" is not safer than a small, principled adjustment set; it is a different and usually worse one. The graph decides which is which before you look at the numbers. Part 32 takes the adjustment set for granted and asks how to estimate the effect once you know which variables belong in it.
Where this shows up
The same trap in two very different places
Confounded benchmark comparisons
A model that scores higher may simply be evaluated on an easier slice. Benchmark, prompt distribution and decoding settings are common causes of both model choice and score — a fork. The evaluation chapter treats a score as an estimate over a population; this part adds that the population itself may be selected. Matching models on prompt difficulty is what turns a ranking into a comparison.
Biased sensor fusion
A robot fusing wheel odometry with a vision estimate can be fooled by a variable that drives both errors: a slippery floor raises wheel slip and also degrades visual features, so the two measurements agree for reasons unrelated to the true motion. A fused pose estimator falls for exactly the same reason; the DAG says the shared cause must be modelled, and that conditioning on a downstream consistency check can induce spurious agreement rather than remove it.
The pattern generalises across the site. Anywhere a system estimates an effect from observational data, the same three shapes are present and the same rule decides which variables to hold fixed. The estimation machinery of this series is agnostic about what you hand it; the graph is where the causal content lives.
Further reading
The references below approach causation as a modelling problem rather than a regression trick. If you take away one thing, take away the collider: conditioning is not a neutral act of "controlling," and the variables you leave alone matter as much as the ones you adjust for.
Pearl's The Book of Why is the accessible version of the graph-first view; Hernán and Robins give the same material the rigour of a textbook; and Simpson's original note is short and worth seeing in its own words.
- Pearl and Mackenzie, The Book of Why, 2018 — forks, chains and colliders for a general reader.
- Hernán and Robins, Causal Inference: What If, 2020 — graphs, the back-door criterion and the adjustment formula in full detail.
- Judea Pearl, Causality: Models, Reasoning, and Inference, 2nd ed., 2009 — do-calculus and identification.
- Edward Simpson, "The Interpretation of Interaction in Contingency Tables", 1951 — the reversal that started the discussion.
Cheat sheet
| Term | Meaning here |
|---|---|
| DAG | Nodes and directed arrows encoding assumed direct causes; acyclic |
| Back-door path | A path from $X$ to $Y$ that enters $X$ through one of its causes; carries confounding |
| Fork / confounder | $X \leftarrow Z \rightarrow Y$; conditioning on $Z$ removes the spurious association |
| Chain / mediator | $X \rightarrow Z \rightarrow Y$; conditioning on $Z$ blocks the causal path |
| Collider | $X \rightarrow Z \leftarrow Y$; conditioning on $Z$ creates an association |
| Partial correlation | Correlation of the residuals after regressing both variables on the conditioned set |
| Over-adjustment | Conditioning on a mediator or collider, adding bias instead of removing it |
| Simpson's paradox | Pooled trend opposite to every within-stratum trend, caused by a fork |