Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

A number with a distribution

Fix a statistical model: a family of distributions $p(x;\theta)$ indexed by an unknown parameter $\theta$, together with an assumption that the data were drawn from one member of the family. An estimator is any rule that turns a sample $X_1,\dots,X_n$ into a guess $\hat\theta(X_1,\dots,X_n)$. The rule is a function, and a function of random variables is a random variable. So the estimate inherits randomness from the data, and the object we must study is its distribution.

That distribution is called the sampling distribution of $\hat\theta$. It answers the question you actually care about: if I repeated this whole experiment many times, where would my estimates land? Its centre tells you about bias, the systematic distance between the average estimate and the truth. Its spread tells you about variance, the wobble that remains even when the centre is right. Bias and variance pull in opposite directions more often than you would like, and their sum is the expected squared error, a single number that grades the estimator.

There is a second, sharper question. Suppose you have banished bias entirely. Is there a limit to how small the variance can go? Information theory says yes, and the quantity that sets the limit is the Fisher information $I(\theta)$, the expected curvature of the log-likelihood at the truth. The more sharply the likelihood peaks, the more the data say about $\theta$, and the lower the variance any unbiased estimator can achieve. The Cramér–Rao bound $\operatorname{Var}(\hat\theta)\ge 1/I(\theta)$ is that statement made exact.

These three ideas are the same idea seen at three distances. Up close, an estimator is a formula. One step back, it is a distribution with a bias and a standard error. Two steps back, that standard error is bounded by the model's geometry. Everything in the applied half of this volume — fitting, intervals, testing, filtering — is built on this last view, so it is worth getting right.

💡 By the end of this part you'll see why an estimate is a random variable, how bias and variance trade off inside the mean squared error, why the sample variance with n in the denominator is biased while the n-1 version is not, and why Fisher information draws a hard line under the variance of any unbiased estimator.
2

Sampling distributions and standard error

Draw the experiment a thousand times and watch the estimates pile up

Start with a model we know completely: $X_1,\dots,X_n$ are independent draws from a normal distribution with mean $\mu$ and variance $\sigma^2$. Pick an estimator. The sample mean averages the draws, the sample variance squares their deviations from the mean, the median takes the middle order statistic. Each of these produces one number per dataset. Generate a thousand datasets from the same model and you get a thousand numbers, and a histogram of those numbers is the sampling distribution.

For the mean there is a clean theory to check the histogram against. Each draw has mean $\mu$ and variance $\sigma^2$, so the average has mean $\mu$ and variance $\sigma^2/n$. Its standard deviation, the standard error $\operatorname{se}(\hat\theta)=\sqrt{\operatorname{Var}(\hat\theta)}$, is $\sigma/\sqrt n$. Notice what that $1/\sqrt n$ does: quadrupling the data only halves the error, which is why precise estimates are expensive. The animation below adds datasets one batch at a time and lets the histogram fill, with the true value and the average estimate drawn on top.

$$\hat\theta \sim p(\,\cdot\mid\theta),\qquad \operatorname{se}(\hat\theta)=\sqrt{\operatorname{Var}(\hat\theta)},\qquad \operatorname{Var}(\bar X)=\frac{\sigma^2}{n}.$$

Switch the estimator and the picture changes character. The sample mean and the median both centre on $\mu$, but the median's histogram is fatter, because throwing away the magnitudes of the data wastes information. The variance estimator is more interesting: divide the sum of squared deviations by n and the histogram sits visibly to the left of $\sigma^2$, because you measured deviations from the sample mean, which is closer to the data than the truth is. Divide by n-1 and the histogram centres on $\sigma^2$. That one-unit correction is the whole content of "unbiased".

An estimator with a sampling distribution centred on the truth is unbiased; otherwise it is biased, and its bias is $\mathbb{E}[\hat\theta]-\theta$. For the population variance the bias is exactly $-\sigma^2/n$, a formula the readout will let you confirm against the simulation. For a small sample that bias is large in relative terms — with n=5 the biased variance is twenty percent short on average — and it is the reason the n-1 version exists at all.

Each bar counts how many simulated datasets produced an estimate in that interval. The pink line is the true value; the dark line is the average of the estimates, and its distance from the truth is the bias.

One habit is worth forming here. When someone reports an estimate without a standard error, they have reported a draw from a distribution while hiding the distribution. The standard error is not decoration; it is the scale on which the estimate should be read. The next section splits that scale into the part you could in principle remove and the part you cannot.

3

Bias and variance

Two ways to miss, and only one number that adds them

An estimate can miss the target in two distinct ways. It can miss the same way every time, because the rule is systematically off — that is bias. Or it can miss in a different direction each time, because the data are noisy — that is variance. The target picture is a dartboard: bias is the offset of the whole cloud of darts from the bullseye, variance is the size of the cloud around its own centre. A tight cluster in the wrong place and a wide scatter around the right place both lose the game.

What grades an estimator is the expected squared distance from the truth, the mean squared error. A one-line algebra identity decomposes it. Add and subtract the mean estimate inside the square, expand, and the cross term vanishes because deviations from the mean average to zero. What is left is the square of the bias plus the variance, so the error splits cleanly into a systematic and a random part.

$$\operatorname{MSE}(\hat\theta)=\mathbb{E}\big[(\hat\theta-\theta)^2\big]=\operatorname{Bias}(\hat\theta)^2+\operatorname{Var}(\hat\theta).$$

The decomposition explains an apparently irrational choice. Suppose you must estimate a quantity and you have a candidate that is slightly biased but much less variable. Its MSE can be smaller than that of an unbiased rival, and MSE is what you actually pay. The classic example is estimating a normal mean: the ordinary sample mean is unbiased, but shrinking it slightly toward zero gives up a little bias to buy a lot of variance, and across many problems the shrunken estimate wins on average. Machine learning makes this trade with a dial — regularisation is bias bought with variance — and the whole point of tuning is to find the bend in the curve.

Drag the bias and spread sliders below. Moving the bias drags the whole cloud off centre; moving the spread inflates it. The readout keeps the identity $\operatorname{MSE}=\operatorname{Bias}^2+\operatorname{Variance}$ exact, and the empirical values computed from the darts on screen should track it. Watch the two knobs push the error up by completely different routes.

The bullseye is the true value. The cross is the centre of the cloud of darts; its distance from the bullseye is the bias, and the cloud's size is the variance.

There is a catch that makes the trade bearable. You cannot have zero bias and zero variance at once unless the estimator is the truth itself, so for any finite sample the bias-variance decomposition is a conservation law with real teeth. It is also the source of the most useful diagnostic in applied work: a model that is too simple is biased and stable, one that is too flexible is unbiased and jittery, and the validation curve that bottoms out is the point where the two contributions cross.

4

Fisher information and the Cramér–Rao bound

How sharply the likelihood peaks sets the best any estimator can do

The bias-variance split says what an estimator costs. The next question is what it could cost at best. The answer comes from the shape of the log-likelihood $\ell(\theta)=\log p(X_1,\dots,X_n;\theta)$. Near the true parameter the log-likelihood is roughly a downward parabola, and its curvature measures how fast the data rule out neighbouring values. A sharp peak means the data are informative; a flat peak means they are not. Fisher information is that curvature, averaged over the randomness in the data.

$$I(\theta)=\mathbb{E}\!\left[\left(\partial_\theta\log p(X;\theta)\right)^2\right]=-\mathbb{E}\!\left[\partial_\theta^2\log p(X;\theta)\right],\qquad \operatorname{Var}(\hat\theta)\ge\frac{1}{I(\theta)}.$$

The two expressions are equal by a short integration by parts, and the second is the one to remember: information is minus the expected second derivative of the log-likelihood. For n independent observations the log-likelihood is a sum, so the information is a sum too: $I_n(\theta)=n\,I_1(\theta)$. Information accumulates linearly in the sample size, so the variance floor $1/I_n(\theta)$ falls like 1/n and the standard error falls like $1/\sqrt n$. This is the deep reason for the square-root law we saw in the histogram — it is not a quirk of the mean, it is a property of information.

For the normal model with known variance the information is especially clean. If $X\sim N(\theta,\sigma^2)$, then $I_1(\theta)=1/\sigma^2$ and $I_n(\theta)=n/\sigma^2$, so the floor on the variance of any unbiased estimator is $\sigma^2/n$. The sample mean has exactly that variance. It attains the bound; no unbiased estimator of the normal mean can beat it for any sample size, so it is called efficient. The sample median, by contrast, has asymptotic variance $\pi\sigma^2/(2n)\approx 1.5708\,\sigma^2/n$, about fifty-seven percent worse. Both are unbiased, and one is provably optimal.

The left canvas plots $\ell(\theta)$ for the first n observations of a fixed dataset; raising n sharpens the peak, which is information made visible. The right canvas plots the variance of the mean and the median against n and draws the Cramér–Rao floor as a horizontal line for the current sample size. The mean sits on the floor; the median floats above it. Only estimators whose score function is a linear function of the estimator can touch the floor, which is why the bound is attained so rarely outside the exponential families.

Top: the log-likelihood with its constant dropped, so the peak sits at zero. Bottom: variance against sample size, with the Cramér–Rao floor drawn for the current n. The sample mean lands on the floor; the median cannot.

Two warnings keep the bound honest. First, it constrains unbiased estimators only; a biased estimator can have smaller variance, which is the bias-variance trade reappearing. Second, it is local: the information describes the likelihood near the true parameter, so the bound assumes you are already in the right neighbourhood. Away from the truth the log-likelihood can curve in ways the quadratic approximation misses, which is exactly why nonlinear least squares needs a good initial guess.

5

Where this shows up

Error floors everywhere something is fitted

AI / ML

Uncertainty in evaluation

A benchmark score is an estimate from a sample, so its standard error is $\sigma/\sqrt n$-shaped and two systems can differ less than the noise. Reporting intervals rather than point scores is the whole argument of evaluation, and it is this part's sampling distribution applied to a leaderboard.

Vision

The floor under a fit

Least-squares bundle adjustment is a maximum-likelihood estimator under Gaussian noise, so its covariance is bounded by the inverse Fisher information — the normal matrix of nonlinear optimization. The Hessian that a solver assembles is, up to scale, the information matrix that sets the error floor.

Math

Why the variance formula has an n-1

The estimator built in variance and standard deviation divides by n-1 for exactly the reason the histogram above moves when you switch denominators: deviations measured from the sample mean are too small, and the correction removes the bias.

Math

Curvature is a step size

Newton's method multiplies the gradient by the inverse curvature, so a sharp likelihood means a small confident step. The optimizers in optimizers are walking the same parabola whose second derivative is Fisher information.

6

Cheat sheet

Every formula in one place

IdeaFormulaReading
Estimator$\hat\theta=\hat\theta(X_1,\dots,X_n)$A rule from data to a guess; a random variable.
Sampling distribution$\hat\theta\sim p(\,\cdot\mid\theta)$Where the estimates land if the experiment repeats.
Standard error$\operatorname{se}(\hat\theta)=\sqrt{\operatorname{Var}(\hat\theta)}$The scale on which to read the estimate.
Bias$\operatorname{Bias}(\hat\theta)=\mathbb{E}[\hat\theta]-\theta$Systematic offset of the estimate cloud.
Mean squared error$\operatorname{MSE}=\operatorname{Bias}^2+\operatorname{Var}$Expected squared miss, split into two causes.
Unbiased variance$\frac{1}{n-1}\sum_i(X_i-\bar X)^2$Divides by n-1; the n version is short by $\sigma^2/n$.
Fisher information$I(\theta)=\mathbb{E}[(\partial_\theta\log p)^2]=-\mathbb{E}[\partial_\theta^2\log p]$Expected curvature of the log-likelihood.
Information adds$I_n(\theta)=n\,I_1(\theta)$Independent observations add information.
Cramér–Rao bound$\operatorname{Var}(\hat\theta)\ge 1/I(\theta)$Floor for any unbiased estimator.
Gaussian mean$I(\theta)=n/\sigma^2,\quad \operatorname{Var}(\bar X)=\sigma^2/n$The sample mean attains the bound; it is efficient.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered