Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

One number that stands in for a whole distribution

Take a random variable X with a probability mass function p(x)=P(X=x). The mass function is complete: it tells you the chance of every value. But a single number often carries the decision. A casino does not care which hand wins; it cares about the average take per hand. A robot does not care about one noisy range reading; it cares about the pose that best explains a hundred of them. The expectation is that number, and it is defined by weighting each value by its probability and adding.

$$\mathbb{E}[X]=\sum_x x\,p(x),\qquad \mathbb{E}[X]=\int_{-\infty}^{\infty} x\,f_X(x)\,dx.$$

For a fair six-sided die this gives $\frac{1}{6}(1+2+3+4+5+6)=3.5$, a value the die never shows. That is the first surprise and the first hint of what the quantity really is: not an outcome, but the centre of mass of the outcomes. Flip a coin that pays one dollar on heads and nothing on tails and the expectation is fifty cents, a sum you cannot hold. The expectation lives one level up, on the distribution rather than on any single draw.

There is a second reading, and it is the one that justifies the name. Draw $X_1,X_2,\dots,X_n$ independently from the same distribution and average them. As n grows, that sample mean creeps toward $\mathbb{E}[X]$. The expectation is the fixed point of averaging, the value the running mean is chasing. That convergence is the law of large numbers, and it is why an expected value is a practical prediction and not only an algebraic decoration.

But if expectation were only a long-run average, it would be useless whenever the long run does not exist. You bet once on a single election; you run a filter for one trajectory; you train a model on one finite corpus. The miracle is that expectation does not need the long run to be useful, because it is linear. The mean of a sum is the sum of the means, and that identity holds whether or not the terms are independent. This part builds that fact from the physical picture, then spends its time showing the three consequences: linearity, the push-forward shortcut known as LOTUS, and the indicator trick that turns counting problems into expectation problems.

💡 By the end of this part you'll see why $\mathbb{E}[X]$ is the balance point of a distribution, why $\mathbb{E}[X+Y]=\mathbb{E}[X]+\mathbb{E}[Y]$ survives any amount of dependence while $\mathbb{E}[XY]$ does not, how $\mathbb{E}[g(X)]$ is computed without ever finding the distribution of g(X), and how indicators turn an expectation into a probability.
2

The balance point

Torques, the centre of mass, and why the mean weighs by probability

Put a horizontal beam on a fulcrum and load it with weights. Each weight w_i at position x_i pulls the beam down with a torque proportional to w_i times its distance from the pivot. The beam balances when the clockwise pull equals the anticlockwise pull, and that happens at exactly one point: the centre of mass. Probability turns this machine into a definition. Place a point mass p(x) at each value x. The total mass is $\sum_x p(x)=1$, so the beam is always fully loaded, and the pivot that balances it is the expectation.

The balancing condition is $\sum_x (x-\mathbb{E}[X])\,p(x)=0$. Expand the bracket and the condition becomes $\sum_x x\,p(x)=\mathbb{E}[X]\sum_x p(x)=\mathbb{E}[X]$, which is the definition rearranged. Nothing is assumed beyond the masses summing to one; the pivot is forced. A distribution with a long right tail carries a heavy mass far to the right, so the pivot slides toward it, which is why a mean can sit far from where most of the probability lives. The median, by contrast, only splits the mass in half and is insensitive to how far the tail reaches. Expectation is the centre of mass, not the centre of probability.

The wheel below makes the analogy literal. Five masses sit on the axis at x=0,1,2,3,4. Each slider controls one mass; the bars are drawn with height proportional to probability, and the pink fulcrum slides to wherever the torques cancel. Drag a heavy weight to the right and watch the pivot chase it; set the masses equal and the pivot snaps to the middle. The readout keeps the two ledger lines honest: the probabilities still sum to one, and the torques on either side of the fulcrum are equal.

Masses on the beam are probabilities; the fulcrum is $\mathbb{E}[X]=\sum x\,p(x)$. Move the weights and the balance point follows.

Two operations on X have clean pictures on the beam. Scaling by a stretches every distance from the origin by a, so the pivot scales too; shifting by b slides the whole beam, and the pivot slides with it. Together they give the affine rule $\mathbb{E}[aX+b]=a\,\mathbb{E}[X]+b$, which holds for any constants, any distribution, and no independence at all. You can read it straight off the geometry, which is the point: expectation commutes with the linear operations you are most likely to perform, and it is that commutativity, not the averaging story, that makes it powerful.

$$\mathbb{E}[aX+b]=a\,\mathbb{E}[X]+b,\qquad \mathbb{E}[X]\text{ is the value }m\text{ with }\sum_x (x-m)\,p(x)=0.$$
3

Linearity, even under dependence

The mean of a sum is the sum of the means, always

Here is the fact that pays for the rest of the guide. For any two random variables X and Y defined on the same experiment, $\mathbb{E}[X+Y]=\mathbb{E}[X]+\mathbb{E}[Y]$. There is no hypothesis. The variables may be independent, positively correlated, negatively correlated, or joined in some way no word covers; the identity still holds. The proof is one line of bookkeeping: write the joint probability of each pair (x,y) and sum (x+y) against it. The sum splits into two sums, and each collapses by the marginal rule.

$$\mathbb{E}[X+Y]=\sum_x\sum_y (x+y)\,p(x,y)=\sum_x x\,p_X(x)+\sum_y y\,p_Y(y)=\mathbb{E}[X]+\mathbb{E}[Y].$$

Contrast that with the product. The identity $\mathbb{E}[XY]=\mathbb{E}[X]\mathbb{E}[Y]$ is not free; it requires the variables to be uncorrelated, which follows from independence but is weaker than it. The gap between the two sides is the covariance, $\operatorname{Cov}(X,Y)=\mathbb{E}[XY]-\mathbb{E}[X]\mathbb{E}[Y]$, and the next part is devoted to it. So there is an asymmetry at the heart of the subject: addition is always safe, multiplication is conditional. Dragging a slider below is the fastest way to feel the difference.

The demo builds the joint distribution of two Bernoulli variables from a single knob. Both have mean one half, so $\mathbb{E}[X+Y]=1$ no matter what. At one end of the slider the two are perfectly aligned, so their sum is either 0 or 2 and is maximally spread; at the other they are perfectly opposed, so the sum is always exactly 1 and has no spread at all. Watch the readouts: $\mathbb{E}[X]$, $\mathbb{E}[Y]$, and $\mathbb{E}[X+Y]$ are frozen while $\mathbb{E}[XY]$ and $\operatorname{Var}(X+Y)$ sweep the whole range. The mean of the sum does not care how the terms conspire; only the second moment does.

Left: the joint table, shaded by probability. Right: the distribution of S=X+Y. The correlation knob tilts the joint but never moves $\mathbb{E}[S]$.

Linearity extends to any finite weighted sum: $\mathbb{E}\!\left[\sum_i a_iX_i+b\right]=\sum_i a_i\,\mathbb{E}[X_i]+b$, again with no independence. That single line is the engine behind most of what follows. The expected number of successes in a hundred dependent trials is the sum of the hundred success probabilities, even when the trials share information. The expected loss on a corpus is the sum of the per-example losses, even when the examples are correlated. Whenever an average of many small pieces is hard but each piece is easy, linearity is the lever.

4

LOTUS and functions of a variable

Push the values through, leave the probabilities alone

Suppose you feed X into a function and want the mean of the output, $\mathbb{E}[g(X)]$. The honest route is two steps: first find the distribution of Y=g(X) by pushing the mass of each value x onto g(x), merging values that collide; then average Y. That works, but it is usually more work than necessary, because the merging never changes the total contribution of the original masses. The shortcut is called the law of the unconscious statistician, LOTUS, and it says you may weight g(x) by the original p(x) and sum.

$$\mathbb{E}[g(X)]=\sum_x g(x)\,p_X(x),\qquad \mathbb{E}[g(X)]=\int_{-\infty}^{\infty} g(x)\,f_X(x)\,dx.$$

The demo makes the equivalence visible. A binomial variable on $0,\dots,5$ is drawn on the left. Choose a map g — square it, fold it around 2.5, or take it modulo three — and the right panel shows the pushed-forward distribution of Y=g(X), where colliding values have their probabilities added. The same expectation appears twice in the readout: once computed the long way by averaging the pushed-forward masses, and once computed the LOTUS way by summing $g(x)\,p(x)$. They agree to machine precision, which is the entire content of the theorem.

Left: p_X(x) for a binomial. Right: the distribution of Y=g(X). Both routes to $\mathbb{E}[g(X)]$ give the same number.

One warning belongs here because it is the most common expectation mistake. In general $\mathbb{E}[g(X)]\ne g(\mathbb{E}[X])$. Squaring has $\mathbb{E}[X^2]\ge(\mathbb{E}[X])^2$, with equality only when X is constant; that inequality is the definition of variance in disguise. A convex map like a square or an exponential pushes the mean up, a concave map like a logarithm pushes it down, and the size of the gap is governed by the spread of X. Jensen's inequality names the pattern, and it is why the expected loss of a model is not the loss at the expected parameter.

5

The indicator trick

Probability is just the mean of a yes-or-no variable

For an event A, define the indicator $\mathbf{1}_A$ to be one when A occurs and zero otherwise. It is a random variable with two values, so its expectation is immediate: $\mathbb{E}[\mathbf{1}_A]=1\cdot P(A)+0\cdot P(A^c)=P(A)$. That tiny identity is a bridge between the two halves of probability. Anything you can count, you can express as a sum of indicators, and linearity then hands you its expectation — and a sum of indicators is exactly a count.

$$\mathbf{1}_A(\omega)=\begin{cases}1,& \omega\in A\\ 0,& \omega\notin A\end{cases}\qquad \mathbb{E}[\mathbf{1}_A]=P(A).$$

The pattern is always the same. To find the expected number of things that happen, name the events, write the count as $X=\sum_i \mathbf{1}_{A_i}$, and push the expectation inside: $\mathbb{E}[X]=\sum_i P(A_i)$. Because linearity needs no independence, the events may be as tangled as you like. The classic example is the expected number of fixed points of a random permutation of n items: there are n positions, each is fixed with probability 1/n, so the expected count is $n\cdot\frac1n=1$, independent of n. Computing the distribution of the number of fixed points is a hard combinatorial problem; its mean is one line.

The canvas shows the identity doing real work: a Monte Carlo estimate of the area of a disk by sampling points in a square. Each sample is a Bernoulli trial for the event A that the point lands inside, and the average of the indicators, the fraction of points inside, estimates $\mathbb{E}[\mathbf{1}_A]=P(A)=\pi r^2$. That is expectation as integration, and it is the same idea that underlies randomised algorithms and importance sampling.

Uniform points in the square; pink points landed in the disk A. The fraction inside is the sample mean of $\mathbf{1}_A$, an estimate of P(A).

6

Where this shows up

Expectation under everything

Math

Variance and concentration

The next part defines variance as $\mathbb{E}[(X-\mathbb{E}[X])^2]$, a second moment built entirely out of expectation, and then bounds how far a variable can stray using only its mean. Linearity is what makes $\operatorname{Var}(X)=\mathbb{E}[X^2]-\mathbb{E}[X]^2$ and $\operatorname{Var}(X+Y)$ computable. See Spread, moments and concentration.

AI / ML

Training loss as an expectation

A language model is trained to minimise the expected negative log-likelihood of the next token, and the corpus loss is a sample mean of that expectation. Perplexity is its exponential. Linearity is why the per-example gradients can be averaged and why the loss on a batch estimates the loss on the distribution. See language models.

Vision

How many RANSAC iterations?

Robust fitting samples minimal sets and counts inliers, and the expected number of iterations needed to hit an all-inlier sample is a short expectation calculation over indicators. That is why the required iteration count grows so steeply with the inlier ratio. See RANSAC.

Linear algebra

Expectation as an inner product

Write the values as a vector $\mathbf{x}$ and the probabilities as a vector $\mathbf{p}$. Then $\mathbb{E}[X]=\mathbf{x}\cdot\mathbf{p}$: expectation is a dot product against a vector of weights that sums to one. The dot product carries the whole geometry of projections with it.

7

Cheat sheet

Every formula in one place

IdeaFormulaReading
Expectation, discrete$\mathbb{E}[X]=\sum_x x\,p(x)$Probability-weighted average; the balance point.
Expectation, continuous$\mathbb{E}[X]=\int x\,f_X(x)\,dx$The same weighted average with density replacing mass.
Balance point$\sum_x (x-\mathbb{E}[X])p(x)=0$Torques cancel at the fulcrum; total mass is one.
Affine map$\mathbb{E}[aX+b]=a\,\mathbb{E}[X]+b$Scaling and shifting commute with expectation.
Linearity of sums$\mathbb{E}[X+Y]=\mathbb{E}[X]+\mathbb{E}[Y]$Always true, dependent or not.
Linear combinations$\mathbb{E}[\sum_i a_iX_i+b]=\sum_i a_i\mathbb{E}[X_i]+b$The lever for averages of many pieces.
Product rule$\mathbb{E}[XY]=\mathbb{E}[X]\mathbb{E}[Y]$ iff $\operatorname{Cov}(X,Y)=0$Needs uncorrelatedness; independence suffices.
LOTUS$\mathbb{E}[g(X)]=\sum_x g(x)p(x)$No need to find the distribution of g(X).
Second moment$\operatorname{Var}(X)=\mathbb{E}[X^2]-\mathbb{E}[X]^2$Spread is an expectation of a square.
Indicator$\mathbb{E}[\mathbf{1}_A]=P(A)$Probability as the mean of a yes/no variable.
Counting trick$X=\sum_i\mathbf{1}_{A_i}\Rightarrow \mathbb{E}[X]=\sum_i P(A_i)$Expected counts without any independence.
8

Further reading

Where to go deeper

9

Check your understanding

0/6 answered