Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

What independence is, and what it is not

Two events are independent when the probability of both happening is the product of their probabilities, $P(A\cap B)=P(A)P(B)$. Read it as an information statement rather than a causal one: if you already know that B happened, your probability for A does not budge, $P(A\mid B)=P(A)$. The word "independent" is doing quiet work here. It does not mean unrelated, uncorrelated in every sense, or physically disconnected. It means conditionally irrelevant to each other under the current model, and it is always relative to a probability measure and a background of other information.

The first trap is confusing independence with mutual exclusivity. Two disjoint events are as dependent as events can be: if one happens, the other is impossible. When both have positive probability, $P(A\cap B)=0$ while P(A)P(B)>0, so the multiplicative rule fails spectacularly. Mutual exclusivity is a statement about the sample space; independence is a statement about information. They are almost opposites.

The second trap is assuming independence travels. It does not. Events can be pairwise independent yet not jointly independent — the classic construction flips two fair coins and looks at "first is heads", "second is heads" and "the two agree", three events that are pairwise independent but whose joint probability is 1/4, not 1/8. And events that are independent can become dependent the moment you condition on something else, including something caused by both of them. That reversal is the whole subject of this part.

The third trap is treating conditional independence as a technicality. It is the load-bearing assumption of every graphical model, every naive Bayes classifier, every hidden Markov model and every Kalman filter. Without it, the joint distribution of n variables needs a table of size exponential in n; with it, the same distribution collapses into a product of small local tables. Learning when you may assume it — and when conditioning creates dependence instead of removing it — is the difference between a model that works and one that is confidently wrong.

So we do this in three movements. First the definitions and the chain rule, because the conditional version is the one you actually use. Then a three-node network where two independent causes share an effect, and you watch explaining-away happen live. Then Simpson's paradox, where the same algebra of conditioning and averaging reverses an association when two subgroups are pooled. By the end, the two paradoxes should feel like the same fact seen from two angles.

💡 By the end of this part you'll see why $P(A\cap B)=P(A)P(B)$ is an information claim, why a Bayesian network needs $P(A\cap B\mid C)=P(A\mid C)P(B\mid C)$ to factorise a joint distribution, and how the same conditioning that removes dependence in one direction can manufacture it in another.
2

Independence and conditional independence

One product rule, two directions

Everything begins with the multiplication rule, which is true always and needs no assumption: $P(A\cap B)=P(A\mid B)P(B)$ when P(B)>0. It says that the chance of both events is the chance of B times the chance of A given B. Independence is the special case where the second factor simplifies, because conditioning on B changed nothing: $P(A\mid B)=P(A)$. Substituting gives the definition you can check on a single table.

$$A\perp B \iff P(A\cap B)=P(A)\,P(B) \iff P(A\mid B)=P(A)\ \text{and}\ P(B\mid A)=P(B).$$

Conditional independence is the same product rule, with the entire calculation carried out inside the slice of the sample space where C holds. Two events are independent given C when

$$A\perp B \mid C \iff P(A\cap B\mid C)=P(A\mid C)\,P(B\mid C),$$

and the colon "given C" is doing real work. Conditioning reweights the sample space: outcomes inconsistent with C get zero probability and everything else is renormalised. Two events that ignore each other globally can become entangled inside that slice, and two events that are entangled globally can ignore each other inside it. Neither implication runs in the direction your intuition expects, which is exactly why the paradoxes in this part surprise people.

The reason conditional independence is the version that matters is the chain rule. For any n variables, repeated application of the multiplication rule gives $P(x_1,\dots,x_n)=\prod_i P(x_i\mid x_1,\dots,x_{i-1})$, an identity with no assumptions at all. The problem is that the last factor can depend on all n-1 earlier variables, so writing the joint down explicitly needs a table with exponentially many entries. A Bayesian network replaces the arbitrary collection of earlier variables by a small set of parents and asserts that each variable is conditionally independent of its non-descendants given its parents. That assertion is what turns the identity into a compact factorisation: $P(x_1,\dots,x_n)=\prod_i P(x_i\mid \mathrm{pa}(x_i))$. Strip out conditional independence and the network buys you nothing.

A concrete example makes the compression vivid. In the three-node network A → C ← B, the joint distribution is $P(A,B,C)=P(A)P(B)P(C\mid A,B)$. A general joint over three binary variables needs 2^3-1=7 free numbers; the factorisation uses 1+1+3=5, and for the noisy version where C is a cause-and-effect combination it uses fewer still. Scale to twenty variables and the saving goes from a factor of one-and-a-bit to a factor of hundreds of thousands. This is why the independence assumption, spelled out as a graph, is the backbone of probabilistic machine learning rather than a footnote in a probability book.

One warning before the demonstrations. Conditional independence is a property of a model, and it can hold for one conditioning set and fail for another. It is not implied by correlation being low, because correlation measures only linear association. Two variables can have zero correlation and be strongly dependent — think of a point on a circle and its square. So when a method claims measurements are conditionally independent given the state, that claim is a modelling commitment to be defended, not a mathematical consequence of the data being random.

3

Explaining-away

Two independent causes, one shared effect

Build a network with two independent causes and a common effect: A and B each point into C. Before looking at anything, A and B are independent by construction. Now imagine the effect fires. If you already believe A is the explanation, the probability that B is also responsible drops, because the effect is already accounted for. The two causes become negatively dependent once you condition on their shared effect — each one "explains away" the need for the other. This is the collider pattern, and it is the mirror image of the more familiar case where conditioning on a common cause makes two events independent.

The clean example is a burglary alarm. Let A be "there was a burglary" and B be "there was an earthquake"; suppose these are independent a priori. Let C be "the alarm rang", which either cause can trigger. The alarm rings and you are told there was no burglary. Your belief in an earthquake jumps — not because earthquakes and burglaries influence each other, but because the alarm needs an explanation and the burglar is ruled out. Conditioning on the effect has connected the causes.

The arithmetic is worth doing once by hand. Take P(A)=P(B)=0.5, and let the effect be a noisy OR: $P(C=1\mid A=1,B=0)=P(C=1\mid A=0,B=1)=0.9$ and $P(C=1\mid A=0,B=0)=0$. The joint over the four cause configurations is uniform, and conditioning on C=1 reweights each configuration by its likelihood. The corner (A=0,B=0) is wiped out, the two single-cause corners survive at full weight, and the double-cause corner is multiplied by 1-(1-0.9)^2=0.99. The posterior mass therefore tilts toward one cause at a time: given the alarm, $P(A=1,B=1\mid C=1)$ is smaller than $P(A=1\mid C=1)\,P(B=1\mid C=1)$. The difference is exactly the negative dependence you can read off the strip below.

A and B are independent causes; C is their shared effect. Toggle whether C is observed and watch the readout: conditioning on the effect makes the two causes repel. The sliders set the prior chances and the two single-cause likelihoods; the combined likelihood $P(C\mid A,B)$ is the noisy OR of them.

Two practical consequences follow. The first is a warning about regression and "controlling for" variables. If you condition on a common effect — a collider — you create a spurious association between its causes, and the association is real in the conditional distribution even though no causal link exists. This is the selection-bias mechanism behind the famous result that a school's admission rates look biased when you stratify by department, and behind the way that conditioning on "is a hospital patient" links unrelated diseases. The second is constructive: explaining-away is the engine of diagnostic reasoning, from fault trees in engineering to differential diagnosis in medicine, because observing a symptom shifts probability mass between its possible causes.

There is a tidy graphical rule, covered fully in the causal-inference literature: two variables are dependent given a set if every path between them is blocked, and conditioning on a collider opens a path rather than blocking it, unless you also condition on one of its descendants. That single asymmetry — common causes block, common effects unblock — is why you cannot decide what to adjust for from the data alone. You need the graph, or an experiment.

4

Simpson's paradox

A trend in every subgroup, reversed in the pool

Simpson's paradox is the arithmetic of averaging meeting the arithmetic of conditioning. Suppose a treatment beats the control in every subgroup of patients, but loses when the subgroups are pooled. Nothing is wrong with the numbers. The treatment groups had different mixes of subgroups, and the pooled rate is a weighted average whose weights differ between arms. An association that holds inside each stratum can reverse when the strata are combined, because the combination adds a second, confounding association between the stratum and the treatment.

The classic illustration is kidney-stone treatment. Two procedures are compared, and one wins for small stones and wins again for large stones, yet the other wins overall — because the first procedure was used mostly on the large stones, which are harder to treat. The modern textbook case is graduate admissions at Berkeley in 1973, where the overall admission rate looked lower for women but most departments admitted women at equal or higher rates, because women applied disproportionately to departments with low admission rates. In both cases the pooled table is not lying; it is answering a different question from the stratified tables.

The demo below makes the mechanism draggable. There are two subgroups, each with a control rate and a treatment rate, and each subgroup favours treatment. Two mix sliders decide what fraction of the control arm and what fraction of the treatment arm fall in subgroup 1. Push the controls toward the high-baseline subgroup and the treated toward the low-baseline subgroup, and the pooled bars cross: treatment now looks worse overall even though it is better everywhere. That crossing is the paradox, and the mix sliders are the confounder made visible.

Blue is control, magenta is treatment. Each subgroup panel shows treatment ahead; the pooled panel shows the weighted average, which can flip. Drag the mix sliders to move patients between subgroups, and the rate sliders to change the success chances.

Which table is correct? That is not a probability question; it is a causal question, and probability can only tell you what the reversal costs if you answer it wrongly. If the subgroups are defined by a variable that the treatment does not affect and that genuinely changes the outcome — severity of disease, say — then the pattern behind the reversal is confounding, and the stratified comparison is the causally meaningful one. If instead the subgroups are defined by a variable on the causal path from treatment to outcome, then stratifying would block part of the treatment's effect and the pooled table could be closer to right. Pearl's backdoor criterion is the formal test: adjust for a set of variables if it blocks every non-causal path and none of the causal ones.

There is a second, more unsettling reading. Simpson's paradox shows that "the data" cannot settle the direction of a trend without a model of how the data were produced. Two analysts with the same table and different graphs will rationally report opposite signs, and both can be honest. The defence is not to pick the table you like but to write down the causal assumptions before looking, and to prefer an experiment that randomises the confounder away. When you cannot experiment, the reversal is a signal that the pooled average is mixing populations you should not be averaging over.

What ties this back to independence is that both pheomena are statements about which conditional distributions are stable. Independence says a conditional equals a marginal; Simpson's paradox says a conditional association can point one way while the marginal points the other. Neither is a contradiction, because they are conditioning on different things. Confusing the two is how a correct number becomes a wrong conclusion.

5

Where this shows up

Independence assumptions under the hood

AI / ML

Factorised language models

An autoregressive language model writes the joint probability of a sequence as a product of next-token conditionals. That is the chain rule, and the practical compression comes from assuming each prediction depends only on a bounded context — a conditional-independence claim about the past. Naive Bayes and mixture models make the same move more crudely, and their failures are usually failures of that assumption.

Vision

Conditional independence in geometry

In multi-view geometry the measurements are modelled as conditionally independent given the camera poses and structure, which is what makes bundle adjustment a clean least-squares problem rather than a coupled one. When that assumption is wrong — correlated correspondences, rolling shutter — the optimiser still converges, but its uncertainty estimates are optimistic.

Robotics

Filters and correlated noise

A Kalman filter assumes the measurement noise is independent across steps and independent of the process noise. Odometry violates the first part badly: integration errors persist and correlate, so a naive filter becomes overconfident. Tracking that correlation, usually with a state-augmentation or factor-graph formulation, is what keeps the covariance honest.

Math

Covariance and linear dependence

Two random variables are uncorrelated when their covariance is zero, which is weaker than independence and is exactly what linear algebra sees. The covariance matrix encodes second moments only; a diagonal covariance means no linear dependence, not no dependence. The difference between the two is the same gap that Simpson's reversal exploits when it hides a confounding variable inside an average.

6

Cheat sheet

Every formula in one place

IdeaFormulaReading
Independence$P(A\cap B)=P(A)P(B)$Learning one event tells you nothing about the other.
Equivalent forms$P(A\mid B)=P(A)$ and $P(B\mid A)=P(B)$Valid when the conditioning event has positive probability.
Mutually exclusive$A\cap B=\varnothing \Rightarrow P(A\cap B)=0$Disjoint events are dependent, not independent, when both have positive probability.
Multiplication rule$P(A\cap B)=P(A\mid B)P(B)$True always; independence is the case that removes the conditioning.
Conditional independence$P(A\cap B\mid C)=P(A\mid C)P(B\mid C)$Inside the slice where C holds, the two events ignore each other.
Chain rule$P(x_1,\dots,x_n)=\prod_i P(x_i\mid x_1,\dots,x_{i-1})$An identity; conditional independence is what makes it small.
Bayesian network$P(x_1,\dots,x_n)=\prod_i P(x_i\mid \mathrm{pa}(x_i))$Each variable depends only on its parents, by assumption.
Explaining-awayConditioning on a collider unblocks a pathTwo independent causes become negatively dependent once their shared effect is seen.
Simpson reversalPooled rate $=\sum_s w_s \cdot \text{rate}_s$, weights w_s differ by armStratified association can flip sign when subgroups are pooled.
Backdoor criterionAdjust for a set that blocks non-causal paths onlyDeciding the right table is a causal question, not an arithmetic one.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered