Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Why the bell curve is everywhere

Roll one fair die and the outcome is flat: each of $1,\dots,6$ has probability 1/6, with not a bell in sight. Roll two dice and add them, and the shape is a triangle peaking at seven. Roll five and the profile already looks like a rounded hump. Roll fifty and the histogram is a smooth curve that would pass for a normal density drawn by hand. Nothing in the die was normal; the addition manufactured the normality. This is the phenomenon the central limit theorem explains, and it is the single most useful fact in statistics.

The reason the raw sum keeps changing shape is that it also keeps getting wider. What stabilises is the standardised sum: subtract the mean so the centre is pinned at zero, then divide by the standard deviation so the width is pinned at one. Once both the location and the scale are fixed, all that is left is shape, and the theorem says the shape converges. That is why we speak of the bell curve rather than of a family of bells indexed by the base distribution: after standardisation there is only one limit.

Why should a flat die, a lopsided waiting time, and a yes/no coin all funnel into the same curve? Because the normal distribution is the unique fixed point of the operation "add two independent copies and rescale", which is exactly the operation that builds an n-fold sum. Convolution widens a distribution; standardisation narrows it back; the normal is the one shape that comes out the way it went in. Every other distribution is pulled toward it a little more with each draw. The sums and convolution part made that stability concrete; here we use it as the explanation for the whole theorem.

Two questions then remain, and they are distinct. The first is the limit: what exactly does the shape converge to, and under what assumptions? The second is the rate: how large must n be before the bell is a good description, and can we put a number on the error at finite n? The first has a clean answer with almost no hypotheses. The second has an answer that is uniform in x, sharp in its n-dependence, and conservative in its constant — the Berry–Esseen theorem, which the final section draws.

💡 By the end of this part you'll see why averaging any finite-variance distribution standardises into a Gaussian, why the error shrinks like $1/\sqrt n$ regardless of the base, and how to read the sample skew and kurtosis to see the convergence happening in real time.
2

Standardisation and the theorem

Remove the location and the scale; only shape remains

Let $X_1,X_2,\dots$ be independent draws from one distribution with finite mean $\mu$ and finite variance $\sigma^2$. Form the sample mean $\bar X_n=(X_1+\cdots+X_n)/n$. Linearity of expectation gives $\mathbb E[\bar X_n]=\mu$, because the n terms each contribute $\mu/n$. Independence makes the variances add, and the factor 1/n is squared when it comes out, so $\operatorname{Var}(\bar X_n)=\sigma^2/n$. The sample mean is centred on the truth and its spread shrinks like the square root of the sample size — that is the standard error, and it is the same $1/\sqrt n$ law that governed the law of large numbers.

$$\operatorname{Var}(\bar X_n)=\frac{\sigma^2}{n},\qquad \operatorname{SE}(\bar X_n)=\frac{\sigma}{\sqrt n},\qquad \bar X_n\approx N\!\left(\mu,\frac{\sigma^2}{n}\right)\ \text{for large }n.$$

That approximation is the practical face of the theorem, but the clean statement is about the standardised variable. Subtract $\mu$ and divide by $\sigma/\sqrt n$, or equivalently multiply the deviation by $\sqrt n/\sigma$:

$$Z_n=\frac{\bar X_n-\mu}{\sigma/\sqrt n}=\frac{\sqrt n\,(\bar X_n-\mu)}{\sigma}\;\xrightarrow[n\to\infty]{d}\;N(0,1).$$

By construction $\mathbb E[Z_n]=0$ and $\operatorname{Var}(Z_n)=1$ for every n, so the theorem is making a claim about the shape alone. The arrow marked d is convergence in distribution: the cumulative distribution function of Z_n, evaluated at any fixed point x, converges to the standard normal CDF $\Phi(x)$. There is no assumption that the base distribution looks normal, is symmetric, or is continuous. The only price of admission is a finite variance, and even that can be relaxed at the cost of a different limit.

The mechanism behind the convergence is the characteristic function, which turns the sum into a product. Each of the n standardised summands has a Taylor expansion whose first non-trivial term is -t^2/2, because its mean is zero and its variance is one. Multiplying n such factors and letting n grow gives

$$M_{Z_n}(t)=\left(1-\frac{t^2}{2n}+o\!\left(\frac{1}{n}\right)\right)^{n}\;\longrightarrow\;e^{-t^2/2},$$

which is exactly the characteristic function of N(0,1). Notice that the third and higher moments of the base never entered. This is why the theorem is universal: skewness and kurtosis affect how fast the limit is reached, not what the limit is. The Gaussian is the attractor of the averaging map, and the higher moments are the transient that decays.

The demo below shows the raw, unstandardised sample mean, so you can watch the $\sigma/\sqrt n$ narrowing directly. The histogram is the sampling distribution of $\bar X_n$ from many seeded runs; the smooth curve is the normal approximation $N(\mu,\sigma^2/n)$; the two inner dashed lines mark $\mu\pm\sigma/\sqrt n$. At n=1 the histogram is just the base distribution and the normal fit can be poor, especially for the lopsided exponential. Increase n and the histogram tightens onto the curve.

Histogram of the raw sample mean over 4000 seeded runs. The smooth curve is $N(\mu,\sigma^2/n)$; the dashed lines sit at $\mu$ and $\mu\pm\sigma/\sqrt n$. The horizontal scale is fixed per base distribution, so the collapse toward the centre is real.

Two cautions belong here. First, standardisation is an affine map, $x\mapsto (x-\mu)/(\sigma/\sqrt n)$, and an affine map sends a normal to a normal. That is why the theorem only had to identify one limit: once the standardised variable is approximately N(0,1), the raw mean is automatically approximately $N(\mu,\sigma^2/n)$, and the whole family is fixed. Second, "approximately" hides a finite-sample error that depends on the base distribution. A symmetric base with no heavy shoulders can be indistinguishable from the bell at n=5; a skewed one may need hundreds of draws. Quantifying that sentence is the job of the next two sections.

3

The CLT machine

Pick a base, slide n, watch the bell appear

The machine below is the theorem with the algebra hidden. Choose a base distribution from the button row — a flat uniform, a lopsided exponential, a two-humped mixture, a coin, or a die. Slide n from one to fifty. For each setting the machine draws 4000 independent sample means from a seeded stream, standardises each one to $z=\sqrt n(\bar x-\mu)/\sigma$, and histograms the results. The smooth curve is the standard normal density, and it does not move: only the histogram moves toward it.

Set n=1 and the histogram is the base distribution itself, standardised. For the uniform it is a flat-topped block, not a bell. For the exponential it is a sharp spike near zero with a long right tail, and its skewness is large. For the die it is six isolated spikes. None of these is remotely Gaussian, which is the point: the theorem is about the limit, not about a single draw. Now drag n upward in steps. The discrete spikes broaden and merge; the exponential's tail gets pulled in; the uniform's block develops shoulders. By n=50 every one of the five bases is visually pinned to the same curve.

The readout quantifies what your eye is doing. The mean and variance of the standardised sample means hover around 0 and 1 for every n, as the construction guarantees. The skewness starts at whatever the base has and decays toward zero; the kurtosis, defined here as the fourth standardised moment, starts at whatever the base has and decays toward three, the Gaussian value. The line comparing the empirical variance of the sample means with the theoretical $\sigma^2/n$ is the standard-error law of the previous section, checked numerically. Watching those four numbers rather than the picture is the honest way to see convergence, because the eye forgives small departures that the moments do not.

One feature of the display is deliberate. Everything is standardised, so the horizontal axis means "how many standard errors from the mean" rather than "dollars" or "seconds". That is the only way five distributions with wildly different units can share one plot, and it is also the reason the theorem is stated for Z_n rather than for $\bar X_n$. Rescale each base by its own $\mu$ and $\sigma$ and the differences in location and spread vanish; what is left is pure shape, and shape is what converges.

The histogram is built from a finite sample of 4000 means, so it has its own Monte Carlo noise of order $1/\sqrt{4000}\approx 0.016$ in CDF terms. That noise is why the bars are not perfectly smooth even when the underlying convergence is excellent, and why a reseed gives a visibly different but statistically identical picture. Seeded randomness keeps the drawing reproducible; the reseed button advances the seed so you can see the variation is sampling, not a bug.

Histogram of the standardised sample mean $z=\sqrt n(\bar x-\mu)/\sigma$ over 4000 seeded runs, against the fixed N(0,1) density. The base is selected by the buttons; n by the slider.

There is a real limitation on display here. For the die and the coin the base is discrete, so the standardised values at small n land on a lattice and the histogram has gaps. That is not a failure of the theorem; it is a reminder that convergence in distribution is a statement about CDFs, which are defined at every point even when the mass is concentrated. As n grows the lattice of possible means becomes fine enough that the gaps are invisible at this resolution, and the curve takes over. The next section turns the visual "eventually" into an inequality with a number in it.

4

Rate of convergence

The Berry–Esseen bound

Convergence in distribution is a limit statement, and limits say nothing about how fast they are approached. There is no rate in the theorem itself. To get one we must measure the distance between the CDF of Z_n, written F_n, and the standard normal CDF $\Phi$, and ask how that distance shrinks with n. The natural distance is the uniform one: the largest vertical gap between the two curves, taken over all x. The Berry–Esseen theorem bounds that gap using a single extra moment of the base, the third absolute central moment $\rho=\mathbb E|X-\mu|^3$.

$$\sup_{x\in\mathbb R}\big|F_n(x)-\Phi(x)\big|\;\le\;\frac{C\,\rho}{\sigma^3\sqrt n},\qquad \rho=\mathbb E|X-\mu|^3,\qquad C\le 0.4748.$$

Read the bound in three pieces. The $1/\sqrt n$ is the rate, and it is the same square-root rate as the standard error; to halve the worst-case CDF error you must quadruple the sample size. The constant C is universal and known to be at most about 0.4748, though it is not sharp — the true worst-case constant is smaller, and for any particular distribution the actual error is typically much smaller than the bound. The factor $\rho/\sigma^3$ is the part that depends on the shape of the base. It is dimensionless, because both $\rho$ and $\sigma^3$ carry the units of the variable cubed, and it is a kind of normalised skewness: it is large exactly when the distribution has a long, heavy, one-sided tail.

Two numbers make the shape-dependence concrete. For the uniform distribution on an interval, $\rho/\sigma^3\approx 1.30$, its smallest possible value among non-degenerate distributions, which is why uniform summands converge unusually quickly. For the exponential distribution, $\rho/\sigma^3\approx 2.41$, nearly twice as large, which is why the lopsided case in the machine needs many more draws to look Gaussian. Feeding these through the bound gives a worst-case error of about $0.62/\sqrt n$ for the uniform and $1.15/\sqrt n$ for the exponential. At n=100 that is a gap of at most about 0.062 and 0.115 respectively, and at n=10{,}000 about 0.006 and 0.011. The bell is a good description at a few hundred draws and an excellent one at a few thousand, with the exact threshold set by the shape factor.

It is worth separating this from the concentration inequalities of the spread and concentration part. Chebyshev controls the probability that $\bar X_n$ is far from $\mu$ and gives a deviation bound falling like 1/n; the CLT is a much more precise statement about the whole distribution, and its CDF error falls like $1/\sqrt n$. The slower-looking rate is not a weakness — it is the price of controlling every quantile at once, including the tails, rather than one fixed threshold. For tail probabilities far from the centre the normal approximation needs a correction, and the exponential's skew is precisely the reason.

The canvas below measures the Berry–Esseen distance directly. For the base distribution currently selected it recomputes the standardised sample mean at a spread of sample sizes, sorts the 1500 standardised values, and takes the largest gap between their empirical CDF and $\Phi$. Those measured points are the dots; the smooth decreasing curve is the theoretical bound $C\rho/(\sigma^3\sqrt n)$. The measured error must sit below the bound, and for the exponential it comes close enough to make the shape factor visible. At large n the dots flatten out at the Monte Carlo floor of about $1/\sqrt{1500}\approx 0.026$; the true error has kept shrinking, but the empirical estimate can no longer see it.

Dots: the measured sup gap between the standardised empirical CDF and $\Phi$, at selected sample sizes. Curve: the Berry–Esseen bound $C\rho/(\sigma^3\sqrt n)$ for the selected base. The dashed line marks the current n on the slider above.

One honest qualification. The Berry–Esseen constant is a bound, not an equality, and for smooth symmetric bases the real error can be an order of magnitude below it. The value of the theorem is not that it predicts the number to two digits; it is that it certifies the $1/\sqrt n$ rate with no assumption beyond a third moment, so the convergence in the machine cannot be a fluke of the particular seeds or of a lucky base distribution. When the distribution has no third moment — or no second moment at all — the theorem's guarantee evaporates, and the next part shows what takes its place.

5

Where this shows up

The bell curve under the measurements

AI / ML

Scaling laws and averaged gradients

A training step averages a mini-batch of per-example gradients, each a noisy estimate of the true descent direction. The CLT is what makes the averaged gradient approximately Gaussian with covariance shrinking like 1/|B|, which is exactly the noise model behind the scaling laws for learning rate and batch size. Loss curves and benchmark spreads inherit their bell shape from sums over many independent tokens.

Robotics / Vision

Averaging noisy measurements

A robot fuses thousands of range readings, and a vision pipeline averages residuals across many correspondences. Each individual measurement has an awkward, non-Gaussian error; the aggregate behaves like Gaussian noise with standard deviation $\sigma/\sqrt n$. That is the licence for the least-squares and gradient-based estimators in nonlinear optimisation to report parameter uncertainties as ellipses.

Probability

Why the Gaussian is the default

The continuous families part introduced the Gaussian twice: as the distribution of maximum entropy for a fixed variance, and as the limit selected by convolution. The CLT is the bridge between those two stories, and it is the reason the normal shows up as a modelling default long before any data has been checked against it.

Limits

From averages to fluctuations

The law of large numbers says the sample mean settles on $\mu$; the CLT magnifies the residual by $\sqrt n$ to reveal the shape of the settling. Together they are the two halves of the asymptotic story: where the average goes, and how it wobbles on the way there.

6

Cheat sheet

Every formula in one place

IdeaFormulaReading
Sample mean$\bar X_n=\frac{1}{n}\sum_{i=1}^n X_i$Average of n i.i.d. draws.
Mean and variance$\mathbb E[\bar X_n]=\mu,\quad \operatorname{Var}(\bar X_n)=\frac{\sigma^2}{n}$Unbiased, with spread shrinking like $1/\sqrt n$.
Standard error$\operatorname{SE}=\frac{\sigma}{\sqrt n}$Width of the sampling distribution of the mean.
CLT (standardised)$Z_n=\frac{\sqrt n(\bar X_n-\mu)}{\sigma}\xrightarrow{d}N(0,1)$The shape converges, location and scale already fixed.
CLT (raw)$\bar X_n\approx N\!\left(\mu,\frac{\sigma^2}{n}\right)$The practical approximation for large n.
Sum form$\frac{\sum X_i-n\mu}{\sigma\sqrt n}\approx N(0,1)$Same statement, written for totals rather than means.
Assumptionsindependent, identically distributed, $0<\sigma^2<\infty$No normality of the base is required.
Berry–Esseen$\sup_x|F_n(x)-\Phi(x)|\le \frac{C\rho}{\sigma^3\sqrt n}$Uniform CDF error; $\rho=\mathbb E|X-\mu|^3$, $C\le 0.4748$.
Ratebound $\propto 1/\sqrt n$Quadruple n to halve the worst-case error.
Shape factor$\rho/\sigma^3$Dimensionless skewness; $\approx 1.30$ uniform, $\approx 2.41$ exponential.
Moment targetsskew $\to 0$, kurtosis $\to 3$Third and fourth standardised moments of Z_n.
Fails when$\operatorname{Var}(X)=\infty$Cauchy sums stay Cauchy; no Gaussian limit.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered