Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Why does averaging turn noise into order?

A single observation is a random variable with whatever distribution the world handed you: skewed, lumpy, bounded, nothing like a bell. But an average is a different random variable. It is built from many observations at once, and the act of adding and dividing imposes a regularity that the individual draws never had. The question of this part is precise: given the distribution of one draw, what is the distribution of the average of n draws, and how does it depend on n?

At first you expect the answer to inherit every detail of the population. It does not. For any population with a finite variance, the distribution of the average becomes approximately normal as n grows, and the only features of the population that survive into the limit are its mean and its variance. The shape — the skew, the gaps, the multiple humps — gets washed out. That single fact is the central limit theorem, and it is why the bell curve appears in places that have no business being bell-shaped.

The demo path below is deliberately concrete. We start with a small discrete population you can edit by dragging, because a finite set of bars makes convolution something you can watch rather than something you have to trust. Then we add the population to itself, over and over, and overlay the normal curve that the result should approach. Finally we measure the width of that result as a function of the sample size, which is the standard error.

💡 By the end of this part you'll see why adding independent random variables is convolution, why averaging makes a distribution concentrate, why its shape approaches a normal regardless of the population's shape, and why the standard error is the width of that bell and falls like $\sigma/\sqrt{n}$.
2

Adding random variables

Convolution is what addition does to distributions

Start with a random variable X that takes one of a handful of values, each with a probability. The demo below is exactly that: nine outcomes, one bar each, all the mass adding to one. Drag any handle and you move probability mass into or out of that outcome; the rest of the bars rescale so the total stays at one. The readout tracks the mean $\mu = \sum_i p_i x_i$ and the standard deviation $\sigma = \sqrt{\sum_i p_i (x_i-\mu)^2}$, the two numbers that will turn out to be the only ones that matter.

Now draw a second, independent value Y from the same population and add them. What is the distribution of the sum? For each possible total s, the pairs that produce it are (x, s-x), and because the draws are independent their probabilities multiply. So

$$P(X+Y=s) = \sum_x P(X=x)\,P(Y=s-x)$$

That sum is a convolution of the two probability mass functions. It is not an obscure operation: it is just "enumerate the ways the total could have happened and add their probabilities". If X can take nine values, then X+Y can take up to seventeen, and the convolution smears the original mass into a wider, smoother shape. This is the same convolution you meet as a change of variables in the probability chapter, specialised to a discrete world where everything can be written down exactly.

The demo lets you choose the population, and the choice matters for the story. A skewed population starts with its mass piled at low outcomes and a long thin tail; a two-humped one has a gap in the middle; a flat one is as unlike a bell as a distribution can be. Convolve any of them with itself a few times and the differences begin to fade. That is the whole surprise, and it is worth setting up the most awkward population you can before moving on.

Drag a handle to move probability mass between outcomes. Bars renormalise to total 1; the dashed line is the mean $\mu$.

⚠ The convolution formula needs independence. Multiplying $P(X=x)$ by $P(Y=s-x)$ is only valid because one draw tells you nothing about the other. If the draws are correlated — repeated measurements from a drifting sensor, tokens from the same document — the sum does not merely have a different shape, it can have a different variance, and every $\sigma/\sqrt{n}$ statement below needs rethinking. Part 4 made the assumption explicit; here it is doing real work.
3

Convolve the population with itself

The shape relaxes to a bell

Averaging k draws is two operations: add them, then divide by k. We already know how to add, by convolving the pmf with itself k times — each new draw is another convolution, so the cost grows as $O(k \cdot \text{bins}^2)$ and is trivial at this scale. Dividing by k simply rescales the horizontal axis: the total s becomes the average s/k, and every gap between outcomes narrows from one unit to 1/k.

The slider below sets k. At k = 1 the bars are the population, unchanged. At k = 2 the pmf has been convolved once: the shape is smoother, the peak is nearer the centre, and the support has spread. By k = 10 the outline is recognisably a bell; by k = 30 the bars are so fine that they read as a filled curve. Nothing was pushed toward normal — no randomness was injected and no smoothing was applied. The bell emerges purely from repeated convolution, which is the algebraic meaning of "average more things".

The overlay makes the convergence measurable. The curve is the normal density with the same mean $\mu$ and the variance of the average, $\sigma^2/k$, so its standard deviation is $\sigma/\sqrt{k}$. The readout reports the largest gap between the discrete density and that curve. As you drag k upward, watch that number fall: the shape is not merely "getting bell-like", it is getting closer to one specific normal distribution, the one fixed by the population's mean and variance alone.

This is the heart of the central limit theorem stated as a picture. If $X_1,\dots,X_k$ are independent draws from any population with finite mean $\mu$ and finite variance $\sigma^2$, then the distribution of the average $\bar X = (X_1+\cdots+X_k)/k$ approaches $\mathrm{Normal}(\mu, \sigma^2/k)$ as k grows, whatever the population's shape. The population above is discrete, skewed, and finite — the least normal object imaginable — and it obeys. The convergence is in distribution: individual averages still land on the discrete grid of 1/k steps, but the envelope they fill is the bell.

Bars: the exact distribution of the average of k draws, by repeated convolution. Curve: the matching normal, $\mathrm{Normal}(\mu,\sigma/\sqrt{k})$.

Two things are worth noticing before moving on. First, the mean of the average is \mu for every k: convolution slides the distribution around but never biases it, so the bell is always centred on the population mean. Second, the width shrinks with k, and that is what makes the average useful rather than merely pretty. Part 4 drew the sampling distribution one sample at a time; here we have computed it exactly, and the two agree.

4

The standard error is the width

σ/√n as a curve, not a formula

The standard deviation of the average is the number you would report whenever you quote an estimate, and it has its own name: the standard error. For the mean of n i.i.d. draws it is

$$\mathrm{SE} = \frac{\sigma}{\sqrt{n}}$$

where \sigma is the standard deviation of a single draw. The demo plots this curve for the population you shaped, and overlays points obtained by simulation: for each sample size n on the axis, it draws many samples of size n, computes each sample's mean, and plots the standard deviation of those means. The dots land on the curve because they are estimating exactly the quantity the curve describes. If you reshape the population, both the curve and the dots move together — the simulated points never need to know the formula, and they still fall on it.

The shape of the curve is the practical content of the whole part. Error falls like $n^{-1/2}$, so it takes four times the data to halve the standard error, sixteen times to quarter it, and a hundred times to divide it by ten. Precision is bought at a quadratic price. A study that has already collected a thousand observations and wants twice the precision must collect three thousand more; the curve is flat by the time you reach the right-hand side, and effort spent there buys almost nothing. Reading the curve before designing the experiment is cheaper than discovering the slowdown afterwards.

Notice also what the horizontal line at the far left says. At n = 1 the standard error equals \sigma: one observation is exactly as uncertain as the population is variable. The entire gain of statistics comes from moving right along the curve. And because the curve is defined from \sigma alone, a population with a huge variance is tamed by the same \sqrt{n} law — it just starts higher.

Curve: $\sigma/\sqrt{n}$. Dots: simulated standard deviation of the sample means at several n, computed by drawing many samples from your population.

⚠ The central limit theorem is narrower than it sounds. It describes the sampling distribution of the average, not the data: individual observations stay exactly as skewed or lumpy as the population, and the CLT never claims otherwise. It says nothing about bias: a biased sampling frame makes the average concentrate on the wrong number, and $n \to \infty$ only makes that wrongness more certain. It needs a finite variance: for a heavy-tailed distribution such as the Cauchy, the average of n draws has the same width as one draw, and the bell never arrives. And it needs something close to independence; strongly correlated data shrink the standard error far more slowly than $\sigma/\sqrt{n}$.
5

Where this shows up

The bell that arrives uninvited

Robotics

Averaging noisy measurements

A robot that fuses repeated range readings, or that averages the residual error over many trials, is relying on exactly this law. A pose estimator has to quantify how the position estimate's uncertainty grows as noisy increments accumulate; the per-step error does not vanish, but the average over many independent samples does concentrate. Monte Carlo sweeps of a controller report means and error bars that shrink with $\sigma/\sqrt{n}$ precisely because the CLT makes the average's distribution predictable without knowing the sensor's noise shape.

ML / AI

Latency and throughput averages

A serving system reports mean latency over a window of requests, and the metrics chapter is careful to give intervals around that number because a different window of requests would give a different mean. The per-request latency distribution is wildly skewed — long tails, spikes, cache hits — and yet the average over a window is roughly normal with standard error $\sigma/\sqrt{n}$, which is why windowed means are comparable at all. The same logic licenses the confidence interval on a benchmark score.

Once you see it, the bell is everywhere. A pixel's intensity averaged over many exposures, a Monte Carlo estimate of an integral, a bootstrap replicate mean, the gradient noise averaged by a large batch — each is a sum of many small independent contributions, and each is therefore approximately normal around its target with a width controlled by the count. The population being averaged does not need to be normal, and usually is not; the averaging is what makes normality appear.

This is also why the next part can talk about one estimate's quality. If the sampling distribution of the mean is approximately normal, centred at \mu, with standard error $\sigma/\sqrt{n}$, then "how good is this estimate?" becomes "how many standard errors away might it plausibly be?", and that is a question a number can answer. Bias and variance are the two ways an estimator can miss, and the CLT supplies the variance half of the account.

Further reading

The references below treat the central limit theorem as a statement about convolution rather than a magic trick, which is the order that makes it believable. If you take away one thing, take away the picture of a lumpy pmf being convolved with itself until the lumps are too fine to see.

Wasserman gives the clean modern statement and the delta method that follows from it; Casella and Berger prove the theorem through moment generating functions and show the finite-variance hypothesis doing its work; and the classic Feller volumes place it in the longer story of sums of random variables. The 3Blue1Brown treatment is the visual source for the convolution picture used here.

Cheat sheet

TermMeaning here
ConvolutionThe pmf of a sum of independent variables: $P(X+Y=s)=\sum_x P(X=x)P(Y=s-x)$
Sum of k drawsMean $k\mu$, variance $k\sigma^2$ — means and variances add
Average of k drawsMean $\mu$, variance $\sigma^2/k$, so standard deviation $\sigma/\sqrt{k}$
Central limit theoremFor i.i.d. finite-variance draws, the average approaches $\mathrm{Normal}(\mu,\sigma^2/k)$ in shape, whatever the population
Standard errorThe standard deviation of the sampling distribution of the mean: $\mathrm{SE}=\sigma/\sqrt{n}$
The √n lawTo halve the standard error, collect four times the data; precision costs quadratically
What it does not sayThe data are not normal, bias is not removed, infinite variance breaks it, independence is assumed
6

Check your understanding

0/4 answered