Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

What does it mean to know something?

Suppose you roll a fair six-sided die and I tell you the result is even. What is the chance it is a three? Obviously zero, and not because threes became rarer — because the only outcomes now under discussion are two, four and six. The information did not reward you with a free number; it deleted half the possibilities out from under the original distribution. That deletion, and the rescaling that has to follow it, is the whole content of conditional probability.

It is tempting to treat conditioning as an afterthought, a way to use extra facts once the real probability has been computed. That ordering is backwards. In almost every application the unconditional probability is the artificial object: nobody walks into a clinic with the marginal probability of a disease, they walk in with a test result. Machine-learning models are trained and evaluated under conditioning every step of the way — given this prompt, given this image, given this robot pose. The marginal is what you get when you refuse to condition, and refusing to condition throws away information.

The clean way to see the mechanism is geometric. On Part 1 we built the sample space — the set of all outcomes — and treated the probability of an event as a kind of mass spread over it. Conditioning on an event B says: if B did not happen, this run of the experiment is not one we are talking about. So discard everything outside B, keep only the mass that sits inside it, and then scale that mass back up so the new universe has total probability one. Nothing about the relative weights inside B changes. Dividing by P(B) is exactly the rescaling.

Once you internalize that picture, the definition is not something to memorise. It is the only rescaling that makes the truncated world into a legitimate probability distribution, with the three axioms of Part 1 intact. This part builds the definition from the picture, shows how it generates the multiplication rule and the chain rule, and then uses it to dismantle a few of the fallacies that make conditional reasoning feel slippery. The interactives let you resize the two events and watch the numbers move, and they let you run the host's decision in the Monty Hall problem until the conditional answer stops being surprising.

💡 By the end of this part you'll see why conditioning is literally cropping the sample space and renormalising it — so that $P(A\mid B)=P(A\cap B)/P(B)$, why the multiplication rule is just that identity rearranged, how the chain rule factors a whole conjunction, and when learning B leaves A untouched.
2

Conditioning crops the sample space

The unit square, cut down and rescaled

Take the simplest non-trivial sample space there is: the unit square. Each point $(\omega_1,\omega_2)$ with both coordinates in [0,1] is an outcome, and because the square has area one, the probability of a region is simply its area. You can think of the first coordinate as the outcome of one experiment and the second as another, or as two coordinates of a single continuous measurement — the picture is the same. An event is a region, and its probability is what fraction of the square it covers.

Now let A and B be rectangles inside the square. The canvas below draws them, and the four corner handles let you drag them around. The shaded purple region is the intersection $A\cap B$: the outcomes where both events happen. Because areas behave like probabilities here, P(A) is the area of A, P(B) is the area of B, and $P(A\cap B)$ is the area of the overlap. Nothing is hidden and nothing is approximate; you can check every readout against the geometry.

Press Condition on B and watch what happens. Everything outside B dims, because those outcomes are no longer candidates. The square is still there as a frame, but the live probability mass now lives only in the B rectangle. The strip below the square performs the renormalisation explicitly: it stretches B out until it has length one, and shades the portion of it that also lies inside A. That shaded fraction is $P(A\mid B)$, by definition and by picture at the same time.

$$P(A\mid B)=\frac{P(A\cap B)}{P(B)},\qquad P(B)>0.$$

The formula is doing two jobs that are easy to conflate. The numerator $P(A\cap B)$ keeps only the outcomes that survive the information. The denominator P(B) rescales, dividing by the total surviving mass so that the cropped world is a probability space in its own right. If you forget the denominator you are not computing a probability but an unnormalised mass, and it will not obey $P(\Omega)=1$ in the new world. If you forget the numerator you are asking about all of B rather than the part of it inside A.

Drag the handles until the readout says the events are independent. This happens when $P(A\cap B)$ is exactly the product P(A)P(B), which forces $P(A\mid B)=P(A)$. Geometrically the overlap has to be just the right size relative to the two rectangles; the default layout on load is deliberately set up that way, with two half-unit squares offset diagonally, so you can start from the independent case and then break it. As you drag B across A you will see $P(A\mid B)$ climb above and fall below P(A): knowing B can raise the chance of A, lower it, or leave it alone.

Drag the four round handles to move and resize A (blue) and B (pink). Toggle conditioning and the mass outside B dims; the strip below stretches B to unit length and shades P(A|B).

A B A ∩ B

One detail the picture hides is the condition P(B)>0. If B has probability zero there is nothing to renormalise, and $P(A\mid B)$ is left undefined rather than set to some convenient value. In the demo the handles are clamped so a rectangle can never collapse below a small minimum, which is why the readouts stay finite; in a real problem you check that the conditioning event has positive probability before writing the formula. When B is a continuum and has probability zero, the fix is a limit or a density rather than this ratio, and that refinement is treated with the continuous random variables later in this volume.

3

The multiplication rule and the chain rule

Rearranging the definition, one factor at a time

Read the definition of $P(A\mid B)$ and multiply both sides by P(B). What comes out is the multiplication rule, and it is the workhorse for computing the probability that two things happen together.

$$P(A\cap B)=P(A\mid B)\,P(B)=P(B\mid A)\,P(A).$$

The two forms on the right are the same statement written from the two viewpoints, and the equality between them is the seed of Bayes' rule in the next part. Each form decomposes a joint probability into a marginal times a conditional: first let B happen, then, given that, let A happen. For a sequential experiment this is how you actually compute. Draw a card, then draw another without replacing the first. The chance the first is an ace is 4/52; given that, three aces remain among fifty-one cards, so the chance the second is also an ace is 3/51. The joint probability of two aces is the product, $\frac{4}{52}\cdot\frac{3}{51}$.

Notice what the rule does not say. It does not say $P(A\cap B)=P(A)P(B)$; that equality is a special case that holds only for independent events, and assuming it in general is the most common error in elementary probability. The whole point of writing a conditional factor is to correct the marginal of A for the information that B happened. When the correction is no correction at all — $P(A\mid B)=P(A)$ — the events are independent and the two-factor product is legitimate. Until then, keep the conditional.

Apply the rule repeatedly and you get the chain rule, also called the general multiplication rule. The probability that all of $A_1,\dots,A_n$ occur is the probability of the first, times the probability of the second given the first, times the probability of the third given the first two, and so on.

$$P(A_1\cap\cdots\cap A_n)=P(A_1)\,P(A_2\mid A_1)\,P(A_3\mid A_1\cap A_2)\cdots P(A_n\mid A_1\cap\cdots\cap A_{n-1}).$$

The chain rule is not a new fact, just the multiplication rule peeled one factor at a time. Its value is that it turns a question about a whole sequence into a sequence of questions about one step, each conditioned on the history so far. Memoryless models, Markov chains and autoregressive language models are all built on exactly this factorisation: the probability of a whole token sequence is the product of next-token probabilities, each conditioned on the prefix. Without the chain rule you could not write down the likelihood of a sentence at all.

A tree is the natural bookkeeping device for the chain rule. Branch on A_1 or its complement, multiply the branch probabilities along a path, and the probability of an outcome is the product of the edge labels on the way to it. Doing this in reverse — summing the leaf probabilities that are consistent with an observation — is how you condition on the observation, and it is precisly the calculation Monty Hall forces you to make carefully. The next section runs that calculation, and separates it from the intuition that gets it wrong.

4

Independence and common fallacies

When B tells you nothing about A — and when you only think it does

Two events are independent when the information that one occurred does not move the probability of the other. Writing that out with the conditional definition gives three equivalent statements, and it is worth keeping all three in view because you will use them in different moods.

$$A\perp B\quad\Longleftrightarrow\quad P(A\cap B)=P(A)P(B)\quad\Longleftrightarrow\quad P(A\mid B)=P(A)\quad(\text{when }P(B)>0).$$

Independence is a property of the probability measure, not of the sets. Two events can be independent under one distribution and dependent under another, which is why the word is so often misused. It is also not the same as being disjoint: disjoint events are maximally dependent in the sense that knowing one happened tells you the other definitely did not. Students who first meet the multiplication rule often conflate "no outcome in common" with "no influence on each other"; the canvas makes the difference visible, because disjoint rectangles are the easiest way to break independence, not to achieve it.

The Monty Hall problem is the standard test of whether the conditioning has actually sunk in. Three doors hide one car and two goats. You pick a door; the host, who knows where the car is, opens one of the other doors to reveal a goat, and offers you the remaining closed door in exchange for your original pick. Should you switch? The naive answer is that two doors remain, so each is equally likely, and it does not matter. The naive answer conditions on the wrong event: it treats the host's action as if it carried no information, when in fact the host had to open a goat door and the choice of which one is constrained by where the car is.

Compute it with the multiplication rule. Your first pick is right with probability 1/3; then switching loses, because the car is behind your door and the host has simply revealed one goat at random. Your first pick is wrong with probability 2/3; then the car is behind one of the other two doors, the host is forced to reveal the only goat among them, and the remaining closed door is the car — switching wins. So switching wins with probability 2/3, staying with 1/3. With N doors and a single revealed goat, the same conditioning gives $\frac{N-1}{N}\cdot\frac{1}{N-2}=\frac{N-1}{N(N-2)}$, which is 2/3 when N=3 and approaches 1/N only slowly as the number of doors grows.

Blue bars are a simulation of the host's game; pale bars are the exact conditional probabilities. They should agree, which is the point.

Closely related is the two-child problem. A family has two children, and you learn that at least one is a boy. What is the probability that both are boys? The sample space of ordered births is $\{BB,BG,GB,GG\}$, each equally likely; conditioning on "at least one boy" removes only GG, leaving three equally likely outcomes, of which one is BB. The answer is 1/3, not 1/2, because the information is about the family rather than a randomly chosen child. If instead you meet one child at random and that child is a boy, the answer really is 1/2: the conditioning event is different, even though the English sentence sounds nearly identical. The lesson generalises: the answer to a conditional question is determined by the exact conditioning event, not by the paraphrase you happen to use.

Two more fallacies are worth naming because they cost real money. The base-rate fallacy ignores P(B) in the denominator: a 99%-accurate test for a one-in-ten-thousand condition yields mostly false positives, because the small sensitivity loss on a huge healthy population swamps the true positives. The prosecutor's fallacy then swaps the conditional in the other direction, quoting $P(\text{evidence}\mid\text{innocent})$ as though it were $P(\text{innocent}\mid\text{evidence})$. Both errors are the same mistake — treating two conditionals as interchangeable. The multiplication rule makes the asymmetry explicit: $P(E\mid I)P(I)=P(I\mid E)P(E)$, and the two sides are equal only because both are the same joint probability, not because the two conditionals are equal to each other.

5

Where this shows up

Conditioning under every model

AI / ML

Likelihoods and evaluation

A language model is trained to maximise $P(\text{token}\mid\text{context})$, and the chain rule turns the probability of a whole document into a product of those conditionals. Measures such as perplexity and log loss are computed inside that conditioning, and evaluation is largely a matter of choosing the right conditional to score.

Vision

Fusing what each measurement knows

A camera pose is estimated given the observed features, and a consistency check is a conditional probability: how likely is this correspondence, given that it is an inlier? In SLAM the whole map is the posterior of the world conditioned on a stream of sensor readings, updated landmark by landmark.

Robotics

Beliefs that update

A robot's odometry is a noisy displacement, so the pose estimate is really a distribution conditioned on every measurement so far. Bayes' rule is the recursion that folds one new reading into that conditional, and the multiplication rule is what justifies the recursion.

Math

The algebra behind the crop

Changing coordinates, projecting, and computing conditional distributions all lean on the same linear-algebra machinery — orthogonal projections and the Schur complement of a covariance matrix — developed in Linear algebra. Conditioning a Gaussian is a projection, which is why the formula looks so clean.

6

Cheat sheet

Every formula in one place

IdeaFormulaIntuition
Definition$P(A\mid B)=\dfrac{P(A\cap B)}{P(B)}$, P(B)>0Crop to B, then rescale so the crop has mass one.
Multiplication rule$P(A\cap B)=P(A\mid B)P(B)$Joint equals marginal times conditional, in either order.
Symmetry$P(A\mid B)P(B)=P(B\mid A)P(A)$The seed of Bayes' rule; both sides equal $P(A\cap B)$.
Chain rule$P(A_1\cap\cdots\cap A_n)=\prod_k P(A_k\mid A_1\cap\cdots\cap A_{k-1})$Peel a joint into one conditional per step.
Independence$P(A\cap B)=P(A)P(B)\iff P(A\mid B)=P(A)$Learning B leaves the odds of A unchanged.
Complement$P(A^c\mid B)=1-P(A\mid B)$Conditioning is still a probability, so the axioms survive.
Partition / total probability$P(A)=\sum_i P(A\mid B_i)P(B_i)$Sum the conditioned pieces over a partition of $\Omega$.
Two-child$P(BB\mid\text{at least one }B)=1/3$The conditioning event is the family, not a random child.
Monty Hall$P(\text{win}\mid\text{switch})=\frac{N-1}{N(N-2)}$The host's forced reveal is the conditioning.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered