Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Why we map outcomes to numbers

An event is a yes/no question about an outcome: did the coin land heads, did the test come back positive, did the model emit the right token. That framing was enough for the last five parts, where the whole game was to compute the probability of an event by counting or conditioning. But most real questions are quantitative. Not "was the total over the limit?" but "how far over?"; not "did the beam hit the landmark?" but "how many metres off was it?"; not "is the next token correct?" but "what is the probability of each candidate token?" A number is being produced by the experiment, and we want its distribution, not just the chance that it crosses some line.

The move that makes this possible is almost too simple. A random variable is a function X from the sample space to the real numbers, $X:\Omega\to\mathbb{R}$. It takes an outcome $\omega$ and returns a number $X(\omega)$. In the die example the outcome is the face itself, so the most obvious random variable is $X(\omega)=\omega$. In the two-dice example the outcome is the ordered pair $\omega=(i,j)$, and the sum is the function $X(\omega)=i+j$. In a spam filter the outcome is a full email and X is the number of times the word "free" appears.

The terminology is one of the worst in mathematics. A random variable is not random and not a variable. The randomness lives entirely in $\omega$, which is drawn from $\Omega$ according to P; the function X is perfectly deterministic. Given $\omega$, there is no uncertainty about $X(\omega)$ at all. The phrase survived because it is suggestive — the number $X(\omega)$ does change from run to run — but the precise statement, the one that makes every later formula obvious, is that X is a measurement you apply to an outcome.

Functions let us push probability forward. We never observe $\omega$ itself; we observe $X(\omega)$. So we care about events written in terms of the value of X, and each such event is a preimage. The statement X=x is shorthand for the event $\{\omega\in\Omega : X(\omega)=x\}$, and the statement $X\le x$ is shorthand for $\{\omega : X(\omega)\le x\}$. Both are honest events, subsets of $\Omega$, so the axioms from Part 1 apply to them directly. This is the whole bridge: a question about a number becomes a question about a set, and sets already have probabilities.

Here is the running example for the rest of the page. Roll two dice, each independently loaded so that a six comes up with probability q and the other five faces share 1-q equally. The outcome space is the six-by-six table of ordered pairs (i,j), thirty-six cells. The random variable is the sum X=i+j. Notice immediately that the map is many-to-one: the value X=7 is produced by six different cells, while X=2 is produced by only one. That folding — many outcomes, one number — is where all the interesting structure comes from, and it is exactly what the probability mass function records.

💡 By the end of this part you'll see why a random variable is just a function on outcomes, why its PMF is the pushforward of probability onto the number line, why the CDF is the same information written as one nondecreasing staircase, and how to get the distribution of any function of a variable by summing over preimages.
2

Mass, not measure

The bars are the probabilities, moved onto the line

For a discrete random variable — one that takes values in a finite or countable set — the entire distribution is captured by the chance of each individual value. That function is the probability mass function, or PMF, and its definition is nothing more than the preimage idea written in symbols: p_X(x)=P(X=x). Read it aloud: the mass at x is the probability of all the outcomes that the measurement sends to x. For the two-dice sum, p_X(7) is the probability of the six cells on the anti-diagonal, which for fair dice is 6/36.

$$p_X(x)=P(X=x)=P\bigl(\{\omega\in\Omega : X(\omega)=x\}\bigr),\qquad p_X(x)\ge 0,\qquad \sum_{x} p_X(x)=1.$$

Two facts do all the work. Non-negativity is inherited straight from the axiom $P\ge 0$, because a preimage is an event. The sum-to-one identity is countable additivity applied to the partition of $\Omega$ into the level sets $\{\omega: X(\omega)=x\}$: those preimages are disjoint and their union is everything, so their probabilities add to $P(\Omega)=1$. Nothing else is needed. A PMF is simply a nonnegative function on the possible values that totals one, and conversely any such function is the PMF of some random variable you can build by choosing a point in a unit interval.

For the loaded dice the computation is a small convolution. Write p for the single-die PMF, with p(6)=q and p(f)=(1-q)/5 for $f=1,\dots,5$. Independence of the two rolls means the pair probabilities multiply, so the sum's PMF is

$$p_{X}(k)=P(i+j=k)=\sum_{i=1}^{6} p(i)\,p(k-i),\qquad k=2,\dots,12.$$

With fair dice this is the familiar triangle: 1/36 at the ends, rising by 1/36 to 6/36 at seven and falling again. Load the die toward six and the triangle tilts: mass migrates toward twelve, the peak softens, and the distribution becomes skewed. The demo below shows this directly. It has three views of the same object, and they are linked: hover a value anywhere and the same value is highlighted in the outcome table, in the PMF bars, and on the CDF staircase. Press Roll the dice to draw one sample and watch a single outcome light up in the table at the same time as its value lights up on the line below.

The outcome space: the 6×6 table of ordered rolls (i,j). Each cell shows the sum i+j. Hover a cell to select its value; click to record it as the sampled outcome.

The PMF: a bar at every value the sum can take. The selected value's bar is drawn in the second accent colour.

The CDF: the running total $F_X(x)=P(X\le x)$. The height of each vertical jump equals the bar above it.

The reduction in the number of atoms is the thing to notice. The outcome view has thirty-six cells, the PMF has eleven bars, and the CDF is a single staircase. Information has not been lost — every cell still knows which bar it feeds — but the question "what value will the sum take?" only needs the bars. This is why we are allowed to forget the outcome space once the PMF is in hand: for any event phrased in terms of X, its probability is the sum of the masses of the values inside it, and the underlying $\omega$ never has to be mentioned again.

A word on the name "mass". For a discrete variable the probability literally sits as weight on individual points, and the bars are those weights. A continuous random variable has no such atoms: the chance of any single value is zero, and its analogue of the PMF is a density that must be integrated over an interval to give probability. The two objects behave differently — one is summed, the other integrated — but they share the CDF, which is the subject of the next section and the reason the cumulative view is the more fundamental one. Part 10 takes up densities properly; here the discrete picture is deliberately kept whole.

3

The CDF staircase

One function that holds every atom at once

The PMF answers "what is the chance of exactly this value?". The complementary question is cumulative: "what is the chance of this value or anything smaller?" That is the cumulative distribution function, defined for every real number x by $F_X(x)=P(X\le x)$. For a discrete variable it is the running sum of the bars, $F_X(x)=\sum_{k\le x}p_X(k)$, and its graph is a staircase that is flat between values and jumps at each one.

$$F_X(x)=P(X\le x)=\sum_{k\le x}p_X(k), \qquad F_X(-\infty)=0,\qquad F_X(+\infty)=1.$$

Three properties characterise a CDF completely, and each is visible in the picture. First, F is nondecreasing: if $x\le y$ then $\{X\le x\}\subseteq\{X\le y\}$, and probability is monotone, so $F_X(x)\le F_X(y)$. Second, F is right-continuous: it equals the value it approaches from above, because the events $\{X\le x+1/n\}$ shrink down to $\{X\le x\}$ and the probability of a decreasing sequence with a nonempty limit is the limit of the probabilities. Third, its limits at the two ends are 0 and 1, since the events $\{X\le -n\}$ empty out and the events $\{X\le n\}$ fill the space. Conversely, any function with those three properties is the CDF of some random variable.

The staircase form makes the discrete case vivid. Between atoms, F is constant, so its value does not move as x slides through a gap. At an atom k it jumps by exactly p_X(k), the size of the bar there. That jump is the giveaway that the variable is discrete: the left-hand limit $F_X(k^-)=\lim_{x\uparrow k}F_X(x)=P(X<k)$ is strictly below $F_X(k)=P(X\le k)$ whenever p_X(k)>0, and the gap between them is the mass at k. For a continuous variable the staircase smooths out, every jump disappears, and F becomes a continuous curve — that is the whole difference, and it is why right-continuity is stated as a separate property rather than left-continuity.

The CDF also answers interval questions without any subtraction of preimages. For any a, the event $\{X\le b\}$ is the disjoint union of $\{X\le a\}$ and $\{a<X\le b\}$, so additivity gives the very useful identity

$$P(a<X\le b)=F_X(b)-F_X(a),\qquad P(X>a)=1-F_X(a).$$

Watch it on the demo: pick a value, read its height on the staircase, and the readout reports both p_X(k) and F_X(k). Because the highlighting is shared, selecting k=7 shades the block of cells that sum to seven, lifts the matching bar, and lands a marker on the staircase exactly one jump above the mass at six. When you drag the loading slider the shape of F changes — a heavier six makes the top steps climb faster — but the three properties never break. A CDF that dipped anywhere would describe a probability that is negative or fails to normalise, and the axioms forbid both.

4

Functions of a variable

Push a distribution through a map

Random variables compose, and this is where the function viewpoint pays for itself. If X is a random variable and g is any function, then Y=g(X) is again a random variable: it is the composition of the measurement X with the map g, so it too assigns a number to every outcome. The question is how to get the distribution of Y from the distribution of X. Because Y=y happens exactly when X lands somewhere in the preimage $g^{-1}(y)=\{x:g(x)=y\}$, and because those preimages are disjoint as y varies, additivity gives the change-of-variable formula for PMFs:

$$p_Y(y)=P(g(X)=y)=\sum_{x:\,g(x)=y} p_X(x).$$

The two-dice sum is the leading example. There the underlying measurement is the pair (I,J) with its thirty-six equal- or loaded-weight cells, and g is addition. Summing p over the anti-diagonal cells is precisely the convolution from the previous section. Whenever the map is many-to-one, mass is merged: several values of X collapse onto one value of Y, and the merged amount is the sum. When the map is one-to-one, nothing merges and the masses just move to new locations; you can think of that as a relabelling of the line.

The smallest interesting collapse is an indicator. Take the single loaded die and let $Y=\mathbf{1}\{X=6\}$, which is one when the die shows a six and zero otherwise. The preimage of one is the single face six, and the preimage of zero is the other five faces. The formula then gives p_Y(1)=q and p_Y(0)=1-q: a Bernoulli variable carved out of a six-valued one. The demo below draws both rows at once, with colour marking which bucket each face folds into, so you can see five bars pour into one and a single bar stand alone.

Top row: the PMF of the face X. Bottom row: the PMF of $Y=\mathbf{1}\{X=6\}$. The faint lines trace which faces fold into which value of Y.

Two cautions belong here. The first is that the formula needs the PMF of X alone; it does not care how X was built, which is the sense in which the distribution is a complete description. The second is that when several variables meet, independence cannot be waved away. The convolution $p_{X+Y}(k)=\sum_i p_X(i)\,p_Y(k-i)$ holds exactly when X and Y are independent; for dependent variables the sum is governed by the joint distribution, a two-dimensional table whose marginals alone do not suffice. The independence part developed that table, and Part 11 uses it to define covariance. For continuous variables the same preimage idea becomes the change-of-variables formula with a Jacobian, the subject of change of variables.

There is a final thread to pull. Once Y=g(X) has a distribution, we can ask for its average, and the calculation $\sum_y y\,p_Y(y)$ can be rewritten as $\sum_x g(x)\,p_X(x)$ — the "law of the unconscious statistician". That identity, and the whole machinery of expectation, is the next part. For now it is worth noting that the rewrite is exactly the preimage sum for the distribution of Y, rearranged.

5

Where this shows up

Every noisy number is a random variable

Robotics

Poses and range readings as variables

A robot never observes its state directly; it observes functions of it, corrupted by noise. Odometry integrates wheel turns into a pose estimate, and every sensor is a random variable whose PMF or density the filter carries forward. The histogram filter in this guide is literally a discretised PMF over positions, and its update step is the preimage sum applied to a measurement model.

AI / ML

Next-token distributions

A language model's softmax outputs a PMF over the vocabulary, and sampling a token is drawing a random variable from it. The decode loop reshapes that PMF at every step — temperature sharpens or flattens the bars, and nucleus sampling keeps the smallest set of values whose CDF reaches a threshold. Top-p decoding is an inverse-CDF operation, run millions of times a second.

Math

Random vectors and linear maps

A list of random variables is a random vector, a function into $\mathbb{R}^n$, and a linear map sends one random vector to another. Covariance matrices, whitening, and PCA all live here: the second-moment structure of a vector-valued variable. The linear algebra guide provides the eigen-decomposition that diagonalises that structure.

Counting

Why the dice triangle has its shape

The number of cells that sum to k is a counting problem, and its answer is why p_X(7) is six times p_X(2). That cell count is the preimage size from the transform formula. The counting part is where those coefficients are derived, and it is what turns an outcome table into a PMF.

6

Cheat sheet

Every formula in one place

IdeaFormulaReading
Random variable$X:\Omega\to\mathbb{R}$A deterministic measurement applied to a random outcome.
Event as preimage$\{X=x\}=\{\omega: X(\omega)=x\}$A statement about a value is a subset of outcomes.
PMFp_X(x)=P(X=x)Mass sitting on the value x.
Normalisation$p_X(x)\ge 0,\ \sum_x p_X(x)=1$Nonnegative, totals one; inherited from the axioms.
CDF$F_X(x)=P(X\le x)=\sum_{k\le x}p_X(k)$Running total; defined for every real x.
CDF propertiesnondecreasing, right-continuous, $0\to1$The three conditions that characterise a CDF.
Jump sizeF_X(k)-F_X(k^-)=p_X(k)At an atom the staircase rises by exactly the bar height.
Interval$P(a<X\le b)=F_X(b)-F_X(a)$Differences of the staircase give interval probabilities.
Function of X$p_Y(y)=\sum_{x:g(x)=y} p_X(x)$Merge mass over the preimage of the map.
Independent sum$p_{X+Y}(k)=\sum_i p_X(i)p_Y(k-i)$Convolution: the distribution of a sum of independent terms.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered