Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Two variables, one table

Suppose you survey one hundred students and record two things about each: how much they slept the night before, bucketed into four ranges, and how they did on an exam, bucketed into four grades. Each student is one outcome, but the outcome now has two coordinates. The natural object is no longer a list of probabilities over a single axis; it is a grid. Cell (i,j) holds the probability that a random student landed in sleep bucket i and grade bucket j.

That grid is the joint distribution, written p(x,y). It contains everything the survey knows. Every question you could ask about sleep and grades together is answered by reading a cell, a row, a column, or a ratio of them. The remarkable thing is that all three readings — joint, marginal, conditional — are not three separate distributions to memorise. They are the same table, summed or sliced or divided.

The reason this matters is that most modelling mistakes are reading errors. People sum the wrong axis, divide by the wrong total, or assume two variables are independent because the conditional barely moved. A table small enough to hold in your head makes those mistakes visible. So the running example here is a four-by-four joint on sleep and grades, drawn as a heatmap you can click. The marginals are the strips on its edges. The conditional is a row you pick, divided by its own total. And a toggle replaces the table by an outer product, so you can watch the marginals stay exactly where they were while the conditional flattens out.

One caveat before we start: the table below is hand-authored, not sampled from real students. Its numbers are chosen to be mildly dependent — more sleep, better grades — so that conditioning visibly moves the answer. The arithmetic, though, is the arithmetic of any joint distribution you will ever meet.

💡 By the end of this part you'll see why a joint distribution is a table, a marginal is a sum over one of its axes, a conditional is a slice renormalised by its own marginal, and independence is exactly the statement that every slice is the same — which is what makes the table an outer product.
2

The joint distribution

A probability for every pair

Let X be sleep bucket and Y be grade bucket, each taking four values. The joint probability mass function is the function $p(x,y)=P(X=x,\;Y=y)$ — the probability that both coordinates take the values you name. Because the outcomes are mutually exclusive and cover everything, the entries are non-negative and sum to one. That is the only constraint; the table is otherwise free.

$$p(x,y)=P(X=x,\;Y=y)\ge 0,\qquad \sum_{x}\sum_{y} p(x,y)=1.$$

The heatmap below is exactly this table. Darker means more probability. The running example is a survey: rows are sleep, columns are grades. The grid is asymmetric — the mass sits along the diagonal from top-left to bottom-right, because sleeping more and scoring higher tend to come together. If the two variables had nothing to do with each other, the mass would be spread out in a pattern you will be able to recognise by the end of the page.

Click any cell, or drag the row slider, to pick a row x. The chosen row is outlined; its total p(x) appears in the readout, and the third canvas below shows that row divided by its total. The strips on the right of the grid are the row totals, and the strips underneath are the column totals. They are the marginals, and the next section explains why they live on the margins.

For continuous variables the same idea needs a density. The joint density p(x,y) is a surface over the plane; probabilities are volumes under it, and the sum becomes an integral over both axes. Everything on this page has a continuous twin, and the twin is always the discrete statement with $\Sigma$ replaced by $\int$.

$$\int\!\!\int p(x,y)\,dx\,dy=1,\qquad P\big((X,Y)\in A\big)=\int\!\!\int_A p(x,y)\,dx\,dy.$$

Click a cell or drag the slider to choose a row x. The right strip is p(x), the bottom strip is p(y). The button replaces the table by the outer product p(x)p(y).

Watch what does not change when you press the independence button. The row strip and the column strip stay put, because they are computed from the table you started with, and the outer product was built from those very marginals. What changes is the interior: the diagonal ridge vanishes and the mass redistributes into a checkerboard-free product pattern. The marginals were never the whole story, and the joint was never determined by them — unless you also know the two variables are independent.

3

Marginals by summing

Throw one variable away

Often you do not care about the pair; you want the distribution of sleep alone. To get it, take the joint and forget y: for each value of x, add up the probabilities across the whole row. The result is the marginal distribution of X.

$$p(x)=\sum_{y} p(x,y),\qquad p(y)=\sum_{x} p(x,y),\qquad \sum_x p(x)=1.$$

The name comes from the ledger habit of writing row and column totals in the margins of a table, and that is precisely what they are. Each row total is a marginal probability of X; each column total is a marginal probability of Y. Adding a row is the operation of integrating out y — of summing over every way the second variable could have gone, weighted by how likely each way was. It is a weighted average in disguise, and it is why the marginal of X is a single honest distribution even though X was tangled up with Y.

Two facts about marginals are worth fixing now. First, they are computed from the joint, not the other way round. Many different joints share the same two marginals; the survey table and its outer product are exactly such a pair. The marginals tell you about each variable in isolation, and they are silent about how the variables move together. Second, marginals add: if you partition the values of Y into groups and sum p(x,y) within each group, you recover the same total. This is just associativity of addition, but it is the reason marginalisation is safe to do in stages.

In the language of linear algebra, a joint table is a matrix, and summing its rows is a matrix–vector product against a vector of ones. The row marginal is $p(x) = \sum_y M_{xy}$, and the column marginal is the same operation on the transpose. That is a dot product between a row of the table and the all-ones vector, which is why the machinery of Part 7 of the linear algebra guide keeps reappearing here: a marginal is a projection of the joint onto one axis.

$$p(x)=\int p(x,y)\,dy,\qquad p(y)=\int p(x,y)\,dx.$$

The top group is the row sum p(x); the bottom group is the column sum p(y). These are the numbers on the edges of the heatmap above, and they do not move when the joint changes to its outer product.

The continuous version replaces each sum by an integral. The marginal density of X is the joint density integrated over every value of y. In two dimensions that integral is the area under a vertical slice of the joint surface, and the collection of those areas, plotted as a function of x, is the marginal curve. When you first meet this in densities, it is worth redoing a two-dimensional example by hand once, because the picture of "slice, then integrate the slice" is the whole idea.

4

Conditioning by slicing

What a row looks like after dividing by its total

Now keep x and ask about y. Restrict attention to students who slept in the bucket you named: take that row of the table. Its entries are the joint probabilities p(x,y), and they do not sum to one — they sum to the row total p(x). To turn the row into a distribution over y, divide every entry by that total. The result is the conditional distribution of Y given X=x.

$$p(y\mid x)=\frac{p(x,y)}{p(x)},\qquad \sum_y p(y\mid x)=1\ \text{ for every }x.$$

The operation is worth naming precisely, because it is the one people get wrong. Slicing picks out a set of outcomes and throws away the rest. Renormalising divides by the probability of the set you kept, so that what remains sums to one. The denominator is always the marginal of the variable you conditioned on. Conditioning is therefore not a new kind of probability; it is ordinary probability computed inside a smaller sample space, the one where X=x.

That reading makes Bayes' rule look less like a formula and more like symmetry. Since $p(x,y)=p(y\mid x)p(x)$ and also $p(x,y)=p(x\mid y)p(y)$, the two expressions for the same cell give $p(y\mid x)=p(x\mid y)p(y)/p(x)$. Nothing was added; two ways of factoring one table were equated. The Bayes part of this volume is entirely contained in that observation, plus the fact that the denominator can be recovered by marginalising the numerator.

The canvas below draws the chosen row twice: as bars for the renormalised conditional $p(y\mid x)$, and as a dashed line for the marginal p(y). Where the bars rise above the line, that grade is over-represented among students who slept that much; where they fall below, under-represented. The gap between the bars and the line is the whole content of "these two variables are related", and when the joint is independent the gap closes everywhere at once.

$$p(y\mid x)=\frac{p(x,y)}{p(x)},\qquad p(x\mid y)=\frac{p(x,y)}{p(y)},\qquad p(x,y)=p(y\mid x)\,p(x).$$

Bars: the conditional $p(y\mid x)$ for the selected row. Dashed line: the marginal p(y). The distance between them is what conditioning on x buys you.

Continuous conditioning works the same way, with the density evaluated along a line. Fix x, divide the joint density by the marginal density at that x, and you get a one-dimensional density over y. The only new subtlety is that conditioning on an event of probability zero needs a limit rather than a ratio of finite numbers, and the ratio of densities is exactly that limit. The change-of-variables machinery from the previous part is what makes the ratio well behaved when coordinates are rescaled.

5

Independence as a product

When every slice is the same

Press the independence button on the first canvas and the interior changes but the edges do not. That is the definition in pictures. Two variables are independent when their joint distribution is the product of their marginals.

$$p(x,y)=p(x)\,p(y)\quad\text{for all }x,y\ \Longleftrightarrow\ p(y\mid x)=p(y)\ \text{ for all }x.$$

Read the right-hand side first. Independence says that learning x tells you nothing about y: the conditional is the same for every row, so the rows of the table are all proportional to one another, and renormalising them all produces the same curve. That is the sense in which the conditional goes flat — flat across x. It is not that $p(y\mid x)$ becomes uniform over y; it is that it stops depending on which row you chose. The left-hand side is the algebraic statement of the same thing: the table is an outer product of two vectors.

The outer product is why the marginals cannot change when you press the button. If you build a table from p(x)p(y), its row sum is $p(x)\sum_y p(y)=p(x)$, and its column sum is p(y). The marginals are baked in by construction. So the outer product is the unique joint with those marginals in which the two variables are independent, and the survey table is a different joint with the same marginals in which they are not. A table has far more freedom than its edges, and independence uses up exactly all of that freedom.

Two cautions. Independence is a property of the joint, not of the marginals, so you cannot read it off the edges; you must look inside. And independence is stronger than being uncorrelated: zero correlation only rules out a linear relationship, while independence rules out every relationship. The next part makes the weaker notion precise, and the gap between them is a standard source of confusion. For now, keep the picture: independent means the table is a product, and a product's rows are all the same shape.

For continuous variables the definition is unchanged: p(x,y)=p(x)p(y) for every pair, with the marginals now integrals of the joint density. A joint Gaussian with a diagonal covariance matrix is exactly this — a product of two one-dimensional Gaussians — and the contours of such a density are axis-aligned ellipses rather than tilted ones. That geometric reading returns in Part 15.

$$p(x,y)=p(x)\,p(y),\qquad p(x)=\int p(x,y)\,dy,\qquad p(y)=\int p(x,y)\,dx.$$
6

Where this shows up

Tables, slices and products everywhere

Robotics

Fusing a pose and a measurement

A SLAM or odometry pipeline carries a joint belief over pose and landmark, then conditions it on each new range reading. That conditioning is literally a slice and a renormalisation; the marginal over the pose is what the planner finally consumes, and the two operations are implemented as matrix factorisations of the same joint.

AI / ML

Attention and co-occurrence tables

An attention matrix is a joint distribution over query and key positions before its rows are renormalised; the softmax makes each row sum to one, which is conditioning by construction. Co-occurrence counts in pretraining are joint tables whose row and column marginals drive pointwise mutual information and the embeddings built from it.

Vision

Stereo and conditional depth

In multi-view geometry the disparity at a pixel is a conditional distribution given the pixel's neighbourhood, and the joint over disparity fields is what a matching cost approximates. Marginalising over neighbouring disparities is the slice-and-sum that turns a joint cost volume into a per-pixel estimate.

Math

Independence, and what breaks it

The independence part treats the product condition as a hypothesis to test; the covariance part measures the linear part of the dependence. Between them they cover the two questions a joint table always raises: does one variable inform the other, and if so, in what direction and by how much?

7

Cheat sheet

Every formula in one place

IdeaFormulaReading
Joint pmfp(x,y)=P(X=x,Y=y), $\sum_x\sum_y p(x,y)=1$The cell of the table; probability of the pair.
Marginal of X$p(x)=\sum_y p(x,y)$Sum a row: forget y, keep the total.
Marginal of Y$p(y)=\sum_x p(x,y)$Sum a column: forget x.
Conditional$p(y\mid x)=p(x,y)/p(x)$Slice the row, divide by its total.
Product rule$p(x,y)=p(y\mid x)\,p(x)=p(x\mid y)\,p(y)$Two ways to factor the same cell.
Bayes' rule$p(y\mid x)=p(x\mid y)\,p(y)/p(x)$Equate the two factorisations.
Independence$p(x,y)=p(x)\,p(y)$The table is an outer product; every slice is the same.
Independence test$p(y\mid x)=p(y)$ for all xConditioning changes nothing.
Continuous marginal$p(x)=\int p(x,y)\,dy$Integrate out the other coordinate.
Continuous conditional$p(y\mid x)=p(x,y)/p(x)$Divide the joint density by the marginal density.
Continuous independence$p(x,y)=p(x)\,p(y)$ for all x,yAxis-aligned contours; the surface is a product.
Matrix view$p(x)=\sum_y M_{xy}$, row times all-onesMarginalising is a projection of the table.
8

Further reading

Where to go deeper

9

Check your understanding

0/6 answered