Estimating causal effects
Everything so far has described what the data look like. This part asks a different kind of question: what would have happened if we had done something else? On one simulated dataset where the true effect is known, a covariate drives both who gets treated and how they turn out, so the plain difference in means is wrong. Three estimators — nearest-neighbour propensity matching, inverse-probability weighting, and difference-in-differences on a before-and-after panel — attack the same bias from different directions, and their estimates move toward the truth as you adjust for the covariate. Watching which method recovers the effect, and which does not, is the clearest way to see what a causal claim actually rests on.
The question
What would have happened otherwise?
Prediction asks what goes with what. A causal question asks what would change if you intervened — if this patient took the drug, if this reader saw the new layout, if this policy generated the response. The trouble is that you never get to see both worlds. For any single unit you observe the outcome under the treatment it actually received, and the outcome under the treatment it did not receive is a counterfactual that is missing by construction.
That missing half is the whole subject. Causal inference is a collection of strategies for replacing it with something you can estimate, and every strategy buys its answer with an assumption you have to be willing to defend. When the treatment is assigned by a coin flip the assumption is almost free, which is why randomized experiments are the gold standard. When the treatment is assigned by people, by circumstance, or by a model, the assumption is doing real work and can be wrong.
This part makes the machinery concrete. A simulator builds one dataset from a known causal model, so the true effect sits on the screen as a dashed line that no real analyst would ever see. Around it sit four estimates: the naive difference in means, and three adjustments. Adjust nothing and the estimate misses; adjust correctly and it moves back to the truth. The distance between the two is confounding, made visible.
Potential outcomes and one dataset
The Rubin causal model, and the problem it names
Give every unit two potential outcomes. Let $Y_i(1)$ be the outcome unit $i$ would have if treated and $Y_i(0)$ the outcome if not treated. The individual treatment effect is their difference, $Y_i(1) - Y_i(0)$, and the average treatment effect — the ATE — is the mean of that difference over the population:
This notation is called the Rubin causal model, or the potential-outcomes framework, and it immediately exposes the fundamental problem of causal inference: for each unit you observe only $Y_i = T_i\,Y_i(1) + (1-T_i)\,Y_i(0)$, where $T_i$ marks the treatment. One of the two potential outcomes is always missing, so the individual effect is never observed. You cannot compare a unit to itself; you can only compare the group that was treated to the group that was not, and hope the two groups are alike in every way that matters.
In a randomized experiment they are alike, in expectation, because the coin decides $T$ independently of everything else. Then the difference in group means estimates the ATE honestly. In observational data the treatment is chosen, and the units differ in ways that affect the outcome. The demo below builds that situation deliberately. A covariate $X$ raises the outcome and also raises the probability of treatment, so treated units would have had higher outcomes even without treatment. That is a confounder, and it is the reason the naive comparison is biased.
The generator is short enough to state: $X \sim \mathcal{N}(0,1)$; treatment is a logistic coin, $P(T=1 \mid X) = \sigma(\beta_0 + \beta\,X)$; and the potential outcomes are $Y(0) = \mu_0 + \tau X + \varepsilon$ and $Y(1) = Y(0) + \text{ATE}$, with the same ATE for every unit. Because we build the world, the ATE is a known number, here $2.0$. The estimate you see is computed from the data alone; only the dashed line knows the answer.
Every simulated unit is a dot: covariate X across, fitted propensity $P(T=1\mid X)$ down. Treated units cluster to the right because X drives both the treatment and the outcome.
Slide β to zero and the treatment becomes a fair coin: the two groups overlap completely, propensity is flat near one half, and the naive gap lands on the truth. Push β up and the groups separate along X. The fitted propensity curve steepens, the treated dots slide up and to the right, and the naive gap drifts away from the dashed truth. Nothing about the outcome model changed; only who got treated changed, and that was enough to break the comparison.
Matching and weighting
Three estimates, one dashed truth
The naive estimate is just the difference in group means, $\bar Y_{T=1} - \bar Y_{T=0}$. Its bias is exactly the difference in the non-treatment potential outcomes between the groups: the treated were already different, and the naive comparison credits treatment with a gap that existed before it. Both adjustments below try to reconstruct the missing counterfactual by making the treated and control groups comparable on $X$.
Matching. For each treated unit, find the control unit with the closest propensity score and use its outcome as the counterfactual. This estimates the average effect on the treated, the ATT, because the treated group is the one being matched. The caliper is the largest propensity gap you will accept; tighten it and you throw away treated units whose counterfactual would have been a poor match, trading a smaller, better-matched sample for coverage. Matching is transparent and assumption-light, but it uses only part of the control pool and degrades when propensities overlap poorly.
Inverse-probability weighting. Rather than pairing units, reweight them. Weight a treated unit by $1/p(X)$ and a control by $1/(1-p(X))$, so that a unit which was unlikely to receive its treatment counts for more. The weighted groups then mimic the whole population and the weighted difference estimates the ATE:
IPW is efficient because it uses every unit, and fragile because it does: a unit with a propensity near zero or one gets an enormous weight and can dominate the estimate. That is the positivity problem, and it shows up directly in the readout as the largest weight. When overlap is good, matching and IPW agree; when it is poor, they disagree, and neither is trustworthy. Combining them gives the doubly robust estimators, which stay consistent if either the outcome model or the propensity model is correct — a useful safety net, though not a licence to skip the assumptions.
Each row is one estimator on the same dataset; the dashed line is the truth. Points are estimates and bars are approximate standard errors.
Toggle Adjust for X off and only the naive row remains, sitting far from the dashed truth. Toggle it on and three adjusted estimates appear and move back toward the line: matching recovers the effect by pairing like-with-like, IPW recovers it by reweighting, and difference-in-differences — the next step — recovers it by subtraction. Push the confounding slider in the previous demo and watch the naive estimate walk away while the adjusted ones stay put. That stability is what it means for an estimator to be robust to a particular confounder.
Difference-in-differences
When you have a before and an after
Sometimes you do not need to model $X$ at all, because you observe the same units before and after. Suppose a treated group and a control group are each measured in two periods. Each group's change from before to after combines a common time trend with the treatment effect, if any. Subtract the control group's change from the treated group's change and the common trend cancels:
The identifying assumption is parallel trends: absent treatment, the two groups would have moved in parallel. It is not the same as the groups being identical — a fixed difference in $X$ between them is fine, because differencing removes any confounder that does not change over time. What breaks the method is a confounder that moves differently in the two groups, which is precisely the assumption you cannot test from the data you have. The demo's control group is the untreated version of the treated group, so the trends really are parallel and DiD lands on the truth.
On the same simulated dataset, the pre-period outcome is the untreated potential outcome $Y(0)$ and the post-period outcome adds a common trend plus the effect for treated units. Because every unit is measured twice, the fixed elevation that $X$ gives the treated group appears in both periods and cancels in the subtraction — the same confounding that defeated the naive cross-section is harmless here.
Group means before and after. The dashed continuation is what the treated group would have done without treatment; the gap at the right is the DiD estimate.
DiD is the workhorse of policy evaluation for a reason: it is credible, it is easy to explain, and it needs no propensity model. Its weakness is equally easy to state. If the treated group was already on a different trajectory — if a new technology was spreading faster among the treated, or a concurrent policy hit them harder — the parallel-trends assumption fails and the estimate absorbs that divergence as if it were the effect. The honest response is to plot the pre-trends, look for a placebo effect before treatment, and treat a parallel-looking history as evidence rather than proof.
What identification needs
The price of the counterfactual
Every adjusted estimator in this part is trading one assumption for another, and it is worth naming them once. Consistency (with the related SUTVA) says a unit's observed outcome is the potential outcome for the treatment it actually got, and that one unit's treatment does not change another's outcome — no interference, no hidden versions of the treatment. Ignorability, or unconfoundedness, says that conditional on the covariates you measured, treatment is as good as randomly assigned: $\{Y(0), Y(1)\} \perp T \mid X$. That is what lets matching and weighting work, and it is untestable — it is a claim about the confounders you did not measure. Positivity says every unit had a nonzero chance of either treatment, which is what keeps the weights finite and the matches available.
There is also the question of which effect you want. The ATE averages over everyone; the ATT averages over the treated only. They coincide when the effect is constant, as in this simulation, but in general a program can help the people who choose it more than it would help the population, or the reverse. Matching naturally targets the ATT, IPW can target either by choice of weights, and DiD targets the treated group's post-period change. Reporting "the effect" without saying which estimand is a common and consequential slippage.
Randomization is the cleanest identification strategy because it satisfies ignorability and positivity by construction: the coin makes treatment independent of the potential outcomes, and with any reasonable allocation probability everybody can go either way. That is the subject of the randomized-experiment part, where the remaining danger is not confounding but the human urge to peek at a running experiment. The honest comparison of two policies is an experiment; when you cannot run one, you are estimating a causal effect from observational data with a model of the assignment mechanism standing in for the coin. Reading a causal claim is therefore reading its assumptions: what was conditioned on, what was assumed to be parallel, and what could have moved quietly underneath the estimate.
Where this shows up
The same question, two more worlds
Did the model change help?
Shipping a new model and comparing its logged scores to the old model's is an observational comparison: traffic, time of day, and prompt mix all differ between the two. The evaluation chapter is where this becomes a causal question, and the answer is an experiment — randomized traffic, a fixed analysis, and no peeking. The same logic governs comparing two reward models trained on logged human preferences, which are observational by nature.
Comparing controllers from logs
A robot's recorded runs are a biased sample: harder terrain, a different operator, or a calibration drift can coincide with whichever controller was installed. The fix has the same shape as everything here — match runs on terrain and speed, weight by how likely each run was under each controller, or difference before-and-after segments to cancel a fixed bias. Metrics that ignore assignment are confounded by it, as the serving metrics chapter warns from the systems side.
The pattern recurs wherever numbers are compared across groups that were not assigned by chance. Regression adjustment is the simplest special case: if the outcome is linear in the confounders, adding those confounders to the model estimates the effect holding them fixed, which is what the propensity score does nonparametrically. And because causal claims are ultimately claims about a probability model, they inherit the discipline of the whole series: a parameter, an estimate, and an interval that says how much the estimate would move if the data were redrawn — the same wobble the guide opened with, now attached to a counterfactual.
Further reading
If you take away one thing, take away the potential-outcomes frame: two outcomes per unit, one of them always missing, and every method an assumption about how to fill it in. The references below are the standard entry points, ordered from the gentle to the formal.
Imbens and Rubin is the modern textbook; Hernán and Robins builds the whole subject on a single tool, the g-formula, and is unusually clear about the assumptions; Angrist and Pischke is the econometric counterpart where DiD and instrumental variables live. Morgan and Winship is a good bridge for social scientists.
- Guido Imbens and Donald Rubin, Causal Inference for Statistics, Social, and Biomedical Sciences, 2015 — potential outcomes, ignorability, the propensity score, and matching.
- Miguel Hernán and James Robins, Causal Inference: What If, 2020 — free online; the g-formula, IPW, doubly robust estimation, and target trials.
- Joshua Angrist and Jörn-Steffen Pischke, Mostly Harmless Econometrics, 2009 — difference-in-differences, parallel trends, and instrumental variables.
- Paul Rosenbaum and Donald Rubin, "The central role of the propensity score in observational studies for causal effects", Biometrika, 1983 — the paper that introduced propensity matching.
- Judea Pearl, Causality: Models, Reasoning, and Inference, 2009 — the graphical view of confounding and the back-door criterion.
Cheat sheet
| Term | Meaning here |
|---|---|
| Potential outcomes | $Y_i(1)$ and $Y_i(0)$: the outcomes under each treatment, one of them always unobserved |
| Fundamental problem | Only one potential outcome per unit is ever observed, so individual effects are unidentifiable |
| ATE / ATT | Average effect over everyone / over the treated only; they differ unless the effect is constant |
| Ignorability | $\{Y(0),Y(1)\} \perp T \mid X$ — treatment is as good as random given the covariates |
| Positivity | Every unit has a nonzero chance of each treatment; keeps weights finite and matches available |
| Consistency / SUTVA | The observed outcome is the potential outcome for the treatment received, with no interference |
| Matching | Pair each treated unit with a similar control; estimates the ATT; the caliper bounds the match quality |
| IPW | Weight by $1/p(X)$ or $1/(1-p(X))$ so the groups resemble the population; estimates the ATE |
| Difference-in-differences | Treated change minus control change; needs parallel trends, not identical groups |
| Doubly robust | Consistent if either the outcome model or the propensity model is right |