Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

When does an average converge?

Start with the simplest stochastic process there is: independent draws $X_1,X_2,\dots$ from one distribution with mean $\mu$ and finite variance $\sigma^2$. Form the running mean $\bar X_n=(X_1+\cdots+X_n)/n$. At n=1 it is the first draw, which can be anywhere. At n=2 it is the midpoint of two draws. As n grows, each new observation matters less, and it becomes harder for any single wild draw to drag the average far away. The question of this part is sharp and concrete: in what sense, exactly, does $\bar X_n$ approach $\mu$?

The first thing to notice is that $\bar X_n$ is random at every n. You cannot write $\bar X_n\to\mu$ and mean it for each fixed outcome, because different outcomes give different sequences and some of them wander. So "converges" needs a qualifier, and the qualifier is the whole content of the law. The weak law quantifies the average behaviour at a fixed time: for any tolerance $\varepsilon$, the probability that $\bar X_n$ is further than $\varepsilon$ from $\mu$ goes to zero as n grows. The strong law quantifies behaviour of the entire path: with probability one, the sequence $\bar X_n$ has $\mu$ as its limit in the ordinary calculus sense.

Those two sentences do not say the same thing, and the difference is not pedantry. Weak convergence in probability only promises that at each time a large fraction of the ensemble is doing well; it permits a shrinking minority of runs to be badly wrong at every time, and it permits individual runs to keep jumping out of the band forever. Strong convergence promises that the exceptional set — the runs that fail to settle — has probability zero, so almost every path is eventually well behaved forever. The weak law is a statement about a crowd at a moment; the strong law is a promise to each member of the crowd.

The machinery that separates them is a small taxonomy called the modes of convergence. Convergence in probability is the weak one. Almost sure convergence is the strong one. Convergence in mean square, also written L^2, is another path to the same limit, and convergence in distribution is the weakest notion of all. Knowing which implies which — and which implications fail, with a counterexample for each failure — is what lets you read a theorem and know exactly what was promised. The plan is the funnel picture first, then the two laws, then the map.

💡 By the end of this part you'll see why a running average settles at all, why the weak and strong laws are genuinely different claims, and how the four modes of convergence sit in a hierarchy with a concrete counterexample guarding every missing arrow.
2

The weak law

A crowd of averages, funnelling inward

Here is the entire proof, and it uses nothing but the variance facts you already have. The mean of $\bar X_n$ is $\mu$, because expectation is linear and the n terms each contribute $\mu/n$. The variance of $\bar X_n$ is $\sigma^2/n$, because the draws are independent so the variances add and then the factor 1/n is squared. Chebyshev then says that for any $\varepsilon>0$,

$$P\big(|\bar X_n-\mu|>\varepsilon\big)\;\le\;\frac{\operatorname{Var}(\bar X_n)}{\varepsilon^2}\;=\;\frac{\sigma^2}{n\,\varepsilon^2}\;\xrightarrow[n\to\infty]{}\;0.$$

That inequality is the weak law for finite variance. It is crude — the true tail is usually far smaller — but it needs almost nothing, and the shape of the bound is what matters. Fix a tolerance $\varepsilon$ and the failure probability falls off like 1/n. Ask instead for a fixed failure probability and the tolerance you can guarantee shrinks like $\sigma/\sqrt{n}$. Either way the barrier against a large deviation grows with the number of draws, which is the formal reason a longer experiment is a better experiment.

The canvas below makes the $1/\sqrt n$ visible. It draws thirty independent runs at once, each one the cumulative mean of its own seeded stream of draws, all on the same axes around the true mean. Early on the paths are scattered across the whole range, because after one or two draws the average is whatever those draws happened to be. As n grows the paths squeeze into a funnel whose walls narrow like $2\sigma/\sqrt n$, drawn as the dashed envelope. Nothing pulls a single path toward the middle — each step is still a fresh random draw — but a single new draw can only move the average by a factor 1/n, so the jitter dies out on its own.

Change the distribution and the picture changes only in its scale: the centre moves to that distribution's mean and the funnel opens or closes with its standard deviation, but the shape is always the same, because the variance of an average is always $\sigma^2/n$. This is why the law of large numbers is a theorem about all finite-variance distributions rather than a fact about coins. The expectation part built $\mu$ as the balance point; the law says the sample balance point homes in on it.

Thirty running averages, each from an independent seeded stream. The dashed funnel is $\mu\pm 2\sigma/\sqrt{n}$; the dashed centre line is the true mean $\mu$.

Watch the readout as you shorten the window. At the right edge, the empirical spread of the thirty running means is close to the theoretical $\sigma/\sqrt n$, and the two numbers track each other as the distribution slider moves. The paths never stop moving — a running average is a random walk of the mean, not a settling stone — but the size of its moves keeps shrinking. That is the weak law seen from the side: most of the time most of the paths are close, and the fraction that is far away is what goes to zero.

It is worth being honest about what this simulation cannot show. Thirty paths are a finite sample; the band contains almost all of them because $2\sigma$ is generous, not because a theorem guarantees it. The weak law is a statement about the limit of a probability, and a picture can only suggest the limit. The next section supplies the sharper statement and shows why it is genuinely stronger, not just a restatement of the same fact with more confidence.

3

The strong law

Almost every path settles, forever

Almost sure convergence looks at the whole sequence at once. It says that the set of outcomes $\omega$ for which the numerical sequence $\bar X_n(\omega)$ fails to converge to $\mu$ has probability zero. Equivalently, for each tolerance $\varepsilon$, the event "the average exceeds $\varepsilon$ from $\mu$ only finitely often" has probability one. So there is, for almost every run, a last time at which the average strays beyond the band; after that last crossing it stays inside forever. That is a much stronger promise than the weak law, which never rules out repeated excursions, only makes them unlikely at pre-specified times.

$$\bar X_n=\frac{X_1+\cdots+X_n}{n}\;\longrightarrow\;\mu\quad\text{almost surely},\qquad\text{i.e.}\quad P\!\left(\lim_{n\to\infty}\bar X_n=\mu\right)=1.$$

The proof is not Chebyshev alone — Chebyshev bounds a probability at one time, not the probability of infinitely many bad times. The classical route restricts to a subsequence n_k growing fast enough that $\sum_k P(|\bar X_{n_k}-\mu|>\varepsilon)$ converges, applies the first Borel–Cantelli lemma to conclude that the subsequence converges almost surely, then fills in the gaps between the n_k using the fact that the average cannot move much between neighbouring points. A cleaner route assumes a finite fourth moment and puts Borel–Cantelli directly on the events $|\bar X_n-\mu|>\varepsilon$. Either way the engine is the same: summable failure probabilities force finitely many failures, and finitely many failures is exactly almost sure convergence.

The demo below contrasts the two claims on the same simulated data. For the weak law we ask, at a checkpoint n_0, what fraction of the runs have a running mean outside the tolerance right now. For the strong law we ask a pathwise question: what fraction of the runs ever stray outside the tolerance at some time $k\ge n_0$, counting the whole future of the path rather than a single instant. The two curves are honestly different objects. The pathwise fraction is always at least the snapshot fraction, because the future includes the present, yet both fall to zero, and the runs still at risk at n_0 are precisely the ones whose settling has not happened yet.

Drag the tolerance and the checkpoint. Tighten $\varepsilon$ and both curves rise, because closeness is harder to demand; this is not a failure of the theorem, only a reminder that the tolerance is a free parameter. Move the checkpoint to the right and the at-risk count keeps falling, sometimes to zero out of two hundred — no run has a future excursion left. That is almost sure convergence in a single glance: not that every path is close now, but that the list of paths that will ever be far again is thinning to nothing. The central limit theorem in the next part answers the complementary question of how the remaining error is distributed.

Solid: fraction outside $\varepsilon$ at time n (weak). Dashed: fraction with an excursion at some time $k\ge n$ (strong, pathwise). The vertical line is the checkpoint n_0.

One caveat keeps the picture honest. With a finite ensemble the pathwise fraction can only be estimated, and it is bounded below by the snapshot fraction rather than being smaller; the weaker-looking curve is not a tighter theorem but a different event. What makes the strong law stronger is not a smaller number at a fixed time, it is the quantifier: for almost every run there exists a time after which it stays close. That quantifier is what turns a statistical tendency into a statement about individual sequences, and it is why the law of large numbers is the bridge from probability to statistics.

The hypothesis that the mean exists is not decoration. Feed a $\text{Cauchy}$ stream into the running average and it does not settle at all: the sample mean of n Cauchy draws has the same Cauchy distribution as one draw, no matter how large n is. The law fails because $\mu$ is undefined, and with it the condition that makes both proofs run. The heavy-tails part explores what replaces the law when the variance is infinite.

4

Modes of convergence

Four notions, and the arrows that can be broken

A sequence of random variables can approach a limit in four inequivalent ways, and the law of large numbers is stated in whichever one the theorem can afford. Convergence in probability is the weak law: for every $\varepsilon>0$, $P(|X_n-X|>\varepsilon)\to 0$. Convergence almost surely is the strong law: $P(X_n\to X)=1$. Convergence in mean square, or L^2, demands that the mean squared error vanish: $\mathbb{E}[(X_n-X)^2]\to0$. Convergence in distribution asks only that the cumulative distribution functions converge at every point where the limit is continuous — it compares the shapes of the laws and ignores the joint behaviour of the sequence entirely.

$$\mathbb{E}\big[(\bar X_n-\mu)^2\big]=\frac{\sigma^2}{n}\;\longrightarrow\;0,\qquad\text{so}\quad \bar X_n\xrightarrow{\;L^2\;}\mu\;\Longrightarrow\;\bar X_n\xrightarrow{\;P\;}\mu.$$

The implications form a small hierarchy, and the fastest way to remember it is the map on the canvas. Almost sure convergence and L^2 convergence are each strong enough to force convergence in probability; convergence in probability forces convergence in distribution. But almost sure and L^2 do not imply each other, and convergence in distribution — the weakest mode — forces nothing stronger. Each of those broken arrows has a standard counterexample, and each is worth keeping in your pocket, because a surprising theorem is usually surprising precisely because one of these arrows was assumed wrongly.

Here is the reflex to build. When a statement says a sequence converges, ask in which mode, and ask whether the proof used a single time or a whole path. A bound on $\mathbb{E}[(X_n-X)^2]$ that lends itself to a summable series gives almost sure convergence by Borel–Cantelli, and that is the usual route in learning theory, because the probabilistic bound is computed at one time and then upgraded. Conversely, when you only need the limiting distribution of a statistic, convergence in distribution is the cheapest tool and is all the central limit theorem ever claims.

The demo below is that hierarchy, drawn. The large region is convergence in distribution; nested inside it is convergence in probability; overlapping inside that are almost sure convergence and L^2 convergence. Pick a mode and the readout gives the counterexample that guards its outward arrows. Notice that the overlaps are genuine: a sequence can converge both almost surely and in mean square, and the laws of large numbers usually prove one and then inherit the other.

Pick a mode to highlight its region and read the counterexample for the implication that fails from it.

Two of the counterexamples are simulations you can picture. The moving spike: place a narrow bump of width 1/n on the unit interval and let it march across the interval forever. The probability that a uniformly chosen point sits under the bump is 1/n, so the variables go to zero in probability — yet every single point is covered infinitely often, so nothing converges almost surely. The tall skinny spike: let X_n be n on an interval of width 1/n and zero elsewhere. Now $\mathbb{E}[X_n^2]=1$ for every n, so no L^2 convergence, but each point is hit only finitely often, so the sequence does converge to zero almost surely. Same interval machinery, opposite failure.

The L^2 mode deserves its place for a practical reason. It is the only one of the four defined by an expectation you can compute, so it is the one whose hypotheses are checkable with the algebra of moments. The route from an L^2 bound to almost sure convergence — sum the squared errors over a subsequence and apply Borel–Cantelli — is the standard machinery behind consistency proofs, and it is why the variance calculation $\sigma^2/n$ does double duty: it proves the weak law directly and, taken along a sparse subsequence, builds the strong law.

5

Where this shows up

Why the average is the default estimator

Robotics

Averaging out sensor noise

A robot integrates thousands of noisy odometry and range readings. The law of large numbers is why the running estimate of a constant bias converges, and why a filter can treat the average residual as a trustworthy calibration. The strong law is the version that justifies treating a single long run as representative.

AI / ML

Concentration as the engine of scaling

Training and evaluation are averages over data. The law of large numbers says a larger sample gives a mean closer to the population value, and the concentration inequalities sharpen the rate. That is the statistical spine of the scaling story: more data narrows the gap between measured and expected loss.

Vision

How many random samples are enough

Robust fitting in multi-view geometry draws random minimal samples until enough of them are all-inlier. The inlier fraction is itself estimated by an average, and the law of large numbers is what lets a RANSAC loop trust the running estimate enough to stop.

Math

The bridge from probability to statistics

The expectation of a distribution is an abstract balance point; the sample mean is a computable object. The law of large numbers is the theorem that connects them, and it is the licence for every estimator that quotes a sample average as if it were the population value.

6

Cheat sheet

Every formula in one place

IdeaFormulaReading
Sample mean$\bar X_n=\frac{1}{n}\sum_{i=1}^n X_i$The running average of i.i.d. draws.
Mean and spread$\mathbb{E}[\bar X_n]=\mu,\quad \operatorname{Var}(\bar X_n)=\sigma^2/n$Unbiased, and spread shrinks like $1/\sqrt n$.
Weak law$P(|\bar X_n-\mu|>\varepsilon)\to 0$ for every $\varepsilon>0$At each fixed time, most runs are close.
Chebyshev bound$P(|\bar X_n-\mu|>\varepsilon)\le \sigma^2/(n\varepsilon^2)$The finite-variance proof of the weak law.
Strong law$\bar X_n\to\mu$ almost surelyAlmost every path settles and stays settled.
Mean square$\mathbb{E}[(\bar X_n-\mu)^2]=\sigma^2/n\to 0$The L^2 statement, checkable with moments.
In probability$P(|X_n-X|>\varepsilon)\to 0$Implied by L^2 and by almost sure convergence.
Almost sure$P(X_n\to X)=1$Pathwise convergence; not implied by convergence in probability.
In distribution$F_n(x)\to F(x)$ at continuity pointsThe weakest mode; implied by convergence in probability.
Hierarchy$\text{a.s.}\Rightarrow P,\quad L^2\Rightarrow P,\quad P\Rightarrow D$No arrow runs the other way without extra hypotheses.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered