Independence, and the paradoxes
Independence is the one assumption that makes probability tractable, and the one that causes the most trouble when it is applied by reflex. It says that learning one fact tells you nothing about another — not that the two facts have nothing to do with each other. A shared cause can leave two events independent while a shared effect makes them dependent the moment you look at it, which is why doctors, pollsters and vision engineers all have to know when to condition and when not to. The same machinery produces Simpson's paradox, where a treatment wins in every subgroup and loses in the pool. This part builds both effects, draws them, and shows that they are the same algebra wearing different clothes.
The question
What independence is, and what it is not
Two events are independent when the probability of both happening is the product of their probabilities, $P(A\cap B)=P(A)P(B)$. Read it as an information statement rather than a causal one: if you already know that B happened, your probability for A does not budge, $P(A\mid B)=P(A)$. The word "independent" is doing quiet work here. It does not mean unrelated, uncorrelated in every sense, or physically disconnected. It means conditionally irrelevant to each other under the current model, and it is always relative to a probability measure and a background of other information.
The first trap is confusing independence with mutual exclusivity. Two disjoint events are as dependent as events can be: if one happens, the other is impossible. When both have positive probability, $P(A\cap B)=0$ while P(A)P(B)>0, so the multiplicative rule fails spectacularly. Mutual exclusivity is a statement about the sample space; independence is a statement about information. They are almost opposites.
The second trap is assuming independence travels. It does not. Events can be pairwise independent yet not jointly independent — the classic construction flips two fair coins and looks at "first is heads", "second is heads" and "the two agree", three events that are pairwise independent but whose joint probability is 1/4, not 1/8. And events that are independent can become dependent the moment you condition on something else, including something caused by both of them. That reversal is the whole subject of this part.
The third trap is treating conditional independence as a technicality. It is the load-bearing assumption of every graphical model, every naive Bayes classifier, every hidden Markov model and every Kalman filter. Without it, the joint distribution of n variables needs a table of size exponential in n; with it, the same distribution collapses into a product of small local tables. Learning when you may assume it — and when conditioning creates dependence instead of removing it — is the difference between a model that works and one that is confidently wrong.
So we do this in three movements. First the definitions and the chain rule, because the conditional version is the one you actually use. Then a three-node network where two independent causes share an effect, and you watch explaining-away happen live. Then Simpson's paradox, where the same algebra of conditioning and averaging reverses an association when two subgroups are pooled. By the end, the two paradoxes should feel like the same fact seen from two angles.
Independence and conditional independence
One product rule, two directions
Everything begins with the multiplication rule, which is true always and needs no assumption: $P(A\cap B)=P(A\mid B)P(B)$ when P(B)>0. It says that the chance of both events is the chance of B times the chance of A given B. Independence is the special case where the second factor simplifies, because conditioning on B changed nothing: $P(A\mid B)=P(A)$. Substituting gives the definition you can check on a single table.
Conditional independence is the same product rule, with the entire calculation carried out inside the slice of the sample space where C holds. Two events are independent given C when
and the colon "given C" is doing real work. Conditioning reweights the sample space: outcomes inconsistent with C get zero probability and everything else is renormalised. Two events that ignore each other globally can become entangled inside that slice, and two events that are entangled globally can ignore each other inside it. Neither implication runs in the direction your intuition expects, which is exactly why the paradoxes in this part surprise people.
The reason conditional independence is the version that matters is the chain rule. For any n variables, repeated application of the multiplication rule gives $P(x_1,\dots,x_n)=\prod_i P(x_i\mid x_1,\dots,x_{i-1})$, an identity with no assumptions at all. The problem is that the last factor can depend on all n-1 earlier variables, so writing the joint down explicitly needs a table with exponentially many entries. A Bayesian network replaces the arbitrary collection of earlier variables by a small set of parents and asserts that each variable is conditionally independent of its non-descendants given its parents. That assertion is what turns the identity into a compact factorisation: $P(x_1,\dots,x_n)=\prod_i P(x_i\mid \mathrm{pa}(x_i))$. Strip out conditional independence and the network buys you nothing.
A concrete example makes the compression vivid. In the three-node network A → C ← B, the joint distribution is $P(A,B,C)=P(A)P(B)P(C\mid A,B)$. A general joint over three binary variables needs 2^3-1=7 free numbers; the factorisation uses 1+1+3=5, and for the noisy version where C is a cause-and-effect combination it uses fewer still. Scale to twenty variables and the saving goes from a factor of one-and-a-bit to a factor of hundreds of thousands. This is why the independence assumption, spelled out as a graph, is the backbone of probabilistic machine learning rather than a footnote in a probability book.
One warning before the demonstrations. Conditional independence is a property of a model, and it can hold for one conditioning set and fail for another. It is not implied by correlation being low, because correlation measures only linear association. Two variables can have zero correlation and be strongly dependent — think of a point on a circle and its square. So when a method claims measurements are conditionally independent given the state, that claim is a modelling commitment to be defended, not a mathematical consequence of the data being random.
Explaining-away
Two independent causes, one shared effect
Build a network with two independent causes and a common effect: A and B each point into C. Before looking at anything, A and B are independent by construction. Now imagine the effect fires. If you already believe A is the explanation, the probability that B is also responsible drops, because the effect is already accounted for. The two causes become negatively dependent once you condition on their shared effect — each one "explains away" the need for the other. This is the collider pattern, and it is the mirror image of the more familiar case where conditioning on a common cause makes two events independent.
The clean example is a burglary alarm. Let A be "there was a burglary" and B be "there was an earthquake"; suppose these are independent a priori. Let C be "the alarm rang", which either cause can trigger. The alarm rings and you are told there was no burglary. Your belief in an earthquake jumps — not because earthquakes and burglaries influence each other, but because the alarm needs an explanation and the burglar is ruled out. Conditioning on the effect has connected the causes.
The arithmetic is worth doing once by hand. Take P(A)=P(B)=0.5, and let the effect be a noisy OR: $P(C=1\mid A=1,B=0)=P(C=1\mid A=0,B=1)=0.9$ and $P(C=1\mid A=0,B=0)=0$. The joint over the four cause configurations is uniform, and conditioning on C=1 reweights each configuration by its likelihood. The corner (A=0,B=0) is wiped out, the two single-cause corners survive at full weight, and the double-cause corner is multiplied by 1-(1-0.9)^2=0.99. The posterior mass therefore tilts toward one cause at a time: given the alarm, $P(A=1,B=1\mid C=1)$ is smaller than $P(A=1\mid C=1)\,P(B=1\mid C=1)$. The difference is exactly the negative dependence you can read off the strip below.
A and B are independent causes; C is their shared effect. Toggle whether C is observed and watch the readout: conditioning on the effect makes the two causes repel. The sliders set the prior chances and the two single-cause likelihoods; the combined likelihood $P(C\mid A,B)$ is the noisy OR of them.
Two practical consequences follow. The first is a warning about regression and "controlling for" variables. If you condition on a common effect — a collider — you create a spurious association between its causes, and the association is real in the conditional distribution even though no causal link exists. This is the selection-bias mechanism behind the famous result that a school's admission rates look biased when you stratify by department, and behind the way that conditioning on "is a hospital patient" links unrelated diseases. The second is constructive: explaining-away is the engine of diagnostic reasoning, from fault trees in engineering to differential diagnosis in medicine, because observing a symptom shifts probability mass between its possible causes.
There is a tidy graphical rule, covered fully in the causal-inference literature: two variables are dependent given a set if every path between them is blocked, and conditioning on a collider opens a path rather than blocking it, unless you also condition on one of its descendants. That single asymmetry — common causes block, common effects unblock — is why you cannot decide what to adjust for from the data alone. You need the graph, or an experiment.
Simpson's paradox
A trend in every subgroup, reversed in the pool
Simpson's paradox is the arithmetic of averaging meeting the arithmetic of conditioning. Suppose a treatment beats the control in every subgroup of patients, but loses when the subgroups are pooled. Nothing is wrong with the numbers. The treatment groups had different mixes of subgroups, and the pooled rate is a weighted average whose weights differ between arms. An association that holds inside each stratum can reverse when the strata are combined, because the combination adds a second, confounding association between the stratum and the treatment.
The classic illustration is kidney-stone treatment. Two procedures are compared, and one wins for small stones and wins again for large stones, yet the other wins overall — because the first procedure was used mostly on the large stones, which are harder to treat. The modern textbook case is graduate admissions at Berkeley in 1973, where the overall admission rate looked lower for women but most departments admitted women at equal or higher rates, because women applied disproportionately to departments with low admission rates. In both cases the pooled table is not lying; it is answering a different question from the stratified tables.
The demo below makes the mechanism draggable. There are two subgroups, each with a control rate and a treatment rate, and each subgroup favours treatment. Two mix sliders decide what fraction of the control arm and what fraction of the treatment arm fall in subgroup 1. Push the controls toward the high-baseline subgroup and the treated toward the low-baseline subgroup, and the pooled bars cross: treatment now looks worse overall even though it is better everywhere. That crossing is the paradox, and the mix sliders are the confounder made visible.
Blue is control, magenta is treatment. Each subgroup panel shows treatment ahead; the pooled panel shows the weighted average, which can flip. Drag the mix sliders to move patients between subgroups, and the rate sliders to change the success chances.
Which table is correct? That is not a probability question; it is a causal question, and probability can only tell you what the reversal costs if you answer it wrongly. If the subgroups are defined by a variable that the treatment does not affect and that genuinely changes the outcome — severity of disease, say — then the pattern behind the reversal is confounding, and the stratified comparison is the causally meaningful one. If instead the subgroups are defined by a variable on the causal path from treatment to outcome, then stratifying would block part of the treatment's effect and the pooled table could be closer to right. Pearl's backdoor criterion is the formal test: adjust for a set of variables if it blocks every non-causal path and none of the causal ones.
There is a second, more unsettling reading. Simpson's paradox shows that "the data" cannot settle the direction of a trend without a model of how the data were produced. Two analysts with the same table and different graphs will rationally report opposite signs, and both can be honest. The defence is not to pick the table you like but to write down the causal assumptions before looking, and to prefer an experiment that randomises the confounder away. When you cannot experiment, the reversal is a signal that the pooled average is mixing populations you should not be averaging over.
What ties this back to independence is that both pheomena are statements about which conditional distributions are stable. Independence says a conditional equals a marginal; Simpson's paradox says a conditional association can point one way while the marginal points the other. Neither is a contradiction, because they are conditioning on different things. Confusing the two is how a correct number becomes a wrong conclusion.
Where this shows up
Independence assumptions under the hood
Factorised language models
An autoregressive language model writes the joint probability of a sequence as a product of next-token conditionals. That is the chain rule, and the practical compression comes from assuming each prediction depends only on a bounded context — a conditional-independence claim about the past. Naive Bayes and mixture models make the same move more crudely, and their failures are usually failures of that assumption.
Conditional independence in geometry
In multi-view geometry the measurements are modelled as conditionally independent given the camera poses and structure, which is what makes bundle adjustment a clean least-squares problem rather than a coupled one. When that assumption is wrong — correlated correspondences, rolling shutter — the optimiser still converges, but its uncertainty estimates are optimistic.
Filters and correlated noise
A Kalman filter assumes the measurement noise is independent across steps and independent of the process noise. Odometry violates the first part badly: integration errors persist and correlate, so a naive filter becomes overconfident. Tracking that correlation, usually with a state-augmentation or factor-graph formulation, is what keeps the covariance honest.
Covariance and linear dependence
Two random variables are uncorrelated when their covariance is zero, which is weaker than independence and is exactly what linear algebra sees. The covariance matrix encodes second moments only; a diagonal covariance means no linear dependence, not no dependence. The difference between the two is the same gap that Simpson's reversal exploits when it hides a confounding variable inside an average.
Cheat sheet
Every formula in one place
| Idea | Formula | Reading |
|---|---|---|
| Independence | $P(A\cap B)=P(A)P(B)$ | Learning one event tells you nothing about the other. |
| Equivalent forms | $P(A\mid B)=P(A)$ and $P(B\mid A)=P(B)$ | Valid when the conditioning event has positive probability. |
| Mutually exclusive | $A\cap B=\varnothing \Rightarrow P(A\cap B)=0$ | Disjoint events are dependent, not independent, when both have positive probability. |
| Multiplication rule | $P(A\cap B)=P(A\mid B)P(B)$ | True always; independence is the case that removes the conditioning. |
| Conditional independence | $P(A\cap B\mid C)=P(A\mid C)P(B\mid C)$ | Inside the slice where C holds, the two events ignore each other. |
| Chain rule | $P(x_1,\dots,x_n)=\prod_i P(x_i\mid x_1,\dots,x_{i-1})$ | An identity; conditional independence is what makes it small. |
| Bayesian network | $P(x_1,\dots,x_n)=\prod_i P(x_i\mid \mathrm{pa}(x_i))$ | Each variable depends only on its parents, by assumption. |
| Explaining-away | Conditioning on a collider unblocks a path | Two independent causes become negatively dependent once their shared effect is seen. |
| Simpson reversal | Pooled rate $=\sum_s w_s \cdot \text{rate}_s$, weights w_s differ by arm | Stratified association can flip sign when subgroups are pooled. |
| Backdoor criterion | Adjust for a set that blocks non-causal paths only | Deciding the right table is a causal question, not an arithmetic one. |
Further reading
Where to go deeper
- Judea Pearl, Causality: Models, Reasoning, and Inference, 2nd edition, 2009 — d-separation, colliders, and the backdoor criterion in full.
- Judea Pearl and Dana Mackenzie, The Book of Why, 2018 — the accessible account of explaining-away and Simpson's paradox, with the Berkeley admissions case.
- E. H. Simpson, "The Interpretation of Interaction in Contingency Tables", JRSS B 13(2), 1951 — the original paper that gave the paradox its name.
- Joseph Blitzstein and Jessica Hwang, Introduction to Probability, 2nd edition, 2019, chapters 2–3 — independence, conditional independence and the two-envelope-style counterexamples done carefully.
- Daphne Koller and Nir Friedman, Probabilistic Graphical Models: Principles and Techniques, 2009, chapter 3 — conditional independence and factorisation as a graph-theoretic language.
- Stanford Encyclopedia of Philosophy, "Simpson's Paradox" — a careful survey of when the pooled table can be defended and when it cannot.