Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Two ways an estimate can miss

Part 4 built the sampling distribution of an estimator: the histogram you would get if you could repeat the whole experiment forever. Its centre tells you where the estimates pile up, and its width tells you how far a typical estimate strays. Those are two separate properties, and an estimator can be good at one while being bad at the other. An estimate can be centred exactly on the truth yet spread so widely that it is rarely close. Or it can be tightly clustered in the wrong place, missing the truth every single time by almost the same amount.

The two failure modes have names. Bias is the distance between the centre of the sampling distribution and the parameter: the average miss, if you could average over infinitely many experiments. Variance is the squared spread around that centre: the wobble. Either alone is only half the story, so the subject folds them into one number, the mean squared error, which is the average squared distance from a single estimate to the truth. The whole part turns on a small piece of algebra that says these three quantities are related by Pythagoras in disguise,

$$\mathrm{MSE}(\hat\theta) \;=\; \mathrm{Bias}(\hat\theta)^2 \;+\; \mathrm{Var}(\hat\theta).$$

Read it as a budget. Total error is a sum of a systematic part that survives every repetition and a random part that cancels when you average. Reducing one does not automatically reduce the other, and the tension between them is the most useful idea in estimation.

💡 By the end of this part you'll see why MSE separates cleanly into bias and variance, why the familiar divisor $n-1$ exists and what it buys, why the sample maximum of a uniform is too small on average yet still converges to the truth, and why the estimator with the smallest error is sometimes deliberately biased.
2

A race between three rules

Same data, three answers, one truth

Below is the signature experiment. A hidden parameter is chosen and fixed. Thousands of independent samples are drawn from a population that depends on it, and three rival rules convert each sample into an estimate. Every estimate is compared with the truth, and the squared misses are accumulated. The top canvas overlays the three sampling distributions so you can see where the rules place their mass; the bottom canvas stacks each rule's error into a variance part and a bias-squared part. The bars have the same height as total MSE, and the identity above is not a slogan here — it is what the two segments add up to, computed from the simulation.

Two scenarios are on offer, because the lesson repeats in different clothes. In the population variance scenario the data are normal with variance one, and the three rules differ only in what they divide the sum of squared deviations by: $n$, $n-1$, or $n+1$. In the uniform endpoint scenario the data are uniform on $(0,\theta)$ and the rules are the sample maximum, twice the sample mean, and the maximum scaled up by $(n+1)/n$. The parameter is one number in both, the rules are all defensible, and the simulation decides nothing — it only makes the contest visible.

Top: overlaid sampling distributions of the three rules, with the true parameter dashed. Bottom: each rule's error split into variance (light) and bias² (dark); the total is MSE.

Watch what happens as you slide $n$ upward. The histograms slide over toward the dashed truth and narrow. The bias-squared segments shrink toward zero while the variance segments shrink more slowly. That is the qualitative difference the next two sections make precise: some errors go away as you collect more data and some are manufactured by the rule itself.

⚠ An estimator is a rule, not an estimate. The simulation is not judging one number; it is judging the procedure that produced thousands of numbers. When the readout says a rule is biased, it means the procedure is biased, which is a statement about the rule and the population. A single estimate from a single sample is neither biased nor unbiased — it is just close or far, and you cannot tell which without the truth this demo hands you for free.
3

Unbiased, consistent, and the difference

Centred on the truth versus converging to it

An estimator is unbiased when its expected value equals the parameter, $\mathbb{E}[\hat\theta] = \theta$. Unbiasedness is a statement about the centre of the sampling distribution and nothing else. It says that if you could run the experiment infinitely many times and average the answers, the average would be exactly right. It does not promise that any answer you actually see is close, and it does not promise that the spread is small. The sample mean is the canonical unbiased estimator; the unbiased variance rule in the race is another.

An estimator is consistent when it converges to the parameter as the sample grows, so that $\hat\theta_n \to \theta$ in probability as $n \to \infty$. Consistency is the weaker and, for practical purposes, the more important property. An estimator can be biased for every finite $n$ and still be perfectly consistent, provided the bias shrinks to zero. That is exactly the sample-maximum story. For $n$ observations uniform on $(0,\theta)$ the maximum is below $\theta$ with probability one, so it is biased low for every sample size. But as $n$ grows the maximum crowds up against the endpoint: the probability that it falls short of $\theta$ by more than any fixed margin goes to zero. The rule is wrong on average forever and still gets the right answer in the limit.

Slide $n$ up in the uniform scenario and watch both things at once. The bias-squared segment for the plain maximum shrinks toward nothing, and its histogram presses against the dashed line from the left. Twice the sample mean is unbiased at every $n$, but its distribution is wide and narrows only like $1/\sqrt{n}$. The scaled maximum is unbiased and much tighter, because it uses the one observation that carries the most information about an endpoint. All three are consistent; only two are unbiased; and, as the next section shows, the race for smallest error is not won by the unbiased pair alone.

The distinction matters because bias and variance respond to sample size differently. Variance of a well-behaved estimator typically falls like $1/n$, so it can be made as small as you like by collecting more data. Bias is a property of the rule and the population, and increasing $n$ does not touch it at all. A biased estimator whose bias does not vanish is broken no matter how much data you feed it; a biased estimator whose bias vanishes is merely awkward, and often worth the awkwardness.

4

Where ÷(n−1) comes from

The correction hiding inside the sample mean

The population variance is the average squared deviation from the population mean: $\sigma^2 = \mathbb{E}[(X-\mu)^2]$. If you knew $\mu$, you would estimate it by averaging $(X_i-\mu)^2$ over the sample, and that rule would be unbiased. You do not know $\mu$. You use the sample mean $\bar x$ instead, and the sample mean is the value that makes the sum of squared deviations as small as possible. That is the trouble: $\bar x$ is fitted to the very data you are measuring the spread of, so it sits closer to the data than $\mu$ does. Measuring deviations from $\bar x$ systematically understates how far the data really scatter.

The understatement has a clean size. The sum of squared deviations from the sample mean has expectation $(n-1)\sigma^2$, not $n\sigma^2$. Dividing by $n$ therefore produces an estimator whose average value is $\frac{n-1}{n}\sigma^2$ — a bias of $-\sigma^2/n$, always downward. Divide by $n-1$ instead, and the expected value is exactly $\sigma^2$. That single change is the Bessel correction, and it explains the divisor that everyone learns by rote. The intuition is that the $n$ deviations are constrained to sum to zero once $\bar x$ is subtracted, so only $n-1$ of them are free to vary. The data contain $n-1$ independent pieces of information about the spread, so the honest denominator is $n-1$.

$$S^2_{n-1} = \frac{1}{n-1}\sum_{i=1}^{n}(x_i-\bar x)^2, \qquad \mathbb{E}\!\left[S^2_{n-1}\right] = \sigma^2.$$

The correction is not free, and Part 5 hinted why. Dividing by the smaller number $n-1$ makes each estimate a little larger, which adds spread. The unbiased rule has strictly higher variance than the maximum-likelihood rule that divides by $n$, and for normal data that extra variance is not repaid by the removal of bias. Run the variance scenario in the demo and compare the two bars: for every $n$ shown, the divisor $n$ has the smaller MSE even though the divisor $n-1$ is the one that is exactly centred. Unbiasedness is a virtue, but it is not the only virtue and it is not always the winning one.

5

The bias–variance trade-off

Sometimes a little bias is exactly what you want

Because MSE is bias squared plus variance, an estimator can lower its total error by accepting a small, deliberate bias in exchange for a larger reduction in spread. That is the trade-off. The deterministic rules in the variance scenario make it unusually crisp, because every one of them is just the sum of squared deviations multiplied by a constant $c$. Choosing $c$ is choosing how much to shrink the estimate toward zero. Shrink too little and the estimate is noisy; shrink too much and it is biased low. The total error as a function of $c$ is a parabola, and its minimum sits at $c = 1/(n+1)$, not at the unbiased value $1/(n-1)$ and not at the maximum-likelihood value $1/n$.

$$\mathrm{MSE}\!\left(\frac{\textstyle\sum(x_i-\bar x)^2}{c}\right) = \frac{2(n-1)}{c^2} + \left(\frac{n-1}{c}-1\right)^2,$$

which is smallest at $c = n+1$. So the best of the three rules in the race is the one that divides by a number larger than either the sample size or the degrees of freedom. It is deliberately biased toward zero, and that bias is more than paid for by the variance it removes. The demo shows this as a shorter total bar for the divisor $n+1$ even though its bias-squared segment is the largest of the three. This is not a statistical trick; it is the same principle that makes ridge regression preferable to ordinary least squares when a design is unstable, and it is why the least-squares chapter treats regularization as part of the estimation story rather than an afterthought.

Between two unbiased estimators there is still a contest, and the currency is efficiency. When both are unbiased their MSEs are just their variances, so the one with the smaller variance is strictly better, and the ratio of variances measures how many times more data the looser one would need to match. In the uniform scenario, twice the sample mean and the scaled maximum are both unbiased, but the scaled maximum has variance $\\theta^2 / (n(n+2))$ while twice the mean has variance $\\theta^2/(3n)$. The maximum is more efficient by a factor of $(n+2)/3$: for a sample of twenty, it extracts several times more information from the same data. An endpoint estimator should look at the endpoint, not at the average.

None of this makes unbiasedness worthless. A vanishing bias is a guarantee that the errors do not all point the same way, which is what makes averaging multiple estimates safe. The point is only that "unbiased" is one property among several, and a rule should be judged on the error it actually incurs. The language of MSE is what lets you compare a biased rule with an unbiased one on the same scale, and it is the bridge to the maximum-likelihood principle in the next part, where the rule is chosen by optimizing a criterion rather than by imposing a property.

6

Where this shows up

The same contest, wearing other clothes

Vision & Geometry

Estimator trade-offs under outliers

The RANSAC chapter is a study in exactly this trade-off. Fitting a line to all the correspondences is unbiased when the noise is well behaved but has enormous variance when a few matches are grossly wrong; fitting to a minimal random sample and counting inliers accepts a little bias in exchange for robustness. Choosing the inlier threshold and the number of iterations is a bias–variance decision, and the guide reasons about it in the same terms the bars use here. The calibration part makes the same choice when it decides how much to trust a noisy correspondence.

ML / AI

Benchmarks and regularisation

A held-out score is an estimator of a model's true error, and its spread across test sets is variance you can measure. The evaluation chapter is careful about finite test sets for precisely this reason. Regularisation in training is the bias–variance trade-off chosen on purpose: a model that fits the training data too tightly buys a small bias with a large variance. The optimizers chapter sees the same tug when it decides how far to move on noisy gradients.

The theme repeats whenever a number must be guessed from data that contain both signal and noise. A robot fusing an odometry estimate with a measurement is blending a low-variance, drifting estimate with a high-variance, unbiased one, which is the trade-off in miniature; a pose filter and the SLAM chapter both rely on it. A sampling-based optimizer trading step size for stability is moving along the same curve. Wherever you can compute an average squared error, you can separate it into the part that data will erase and the part that only a better rule will fix, and that separation is what makes the choice of estimator an engineering decision rather than a matter of taste.

Further reading

The references below treat bias, variance and MSE as the organising quantities of point estimation rather than as formulas to memorise. If you take away one thing, take away the picture of three histograms over one fixed truth and the two stacked segments that explain why they differ.

Casella and Berger derive the MSE decomposition and the Cramér–Rao bound in the same chapter, which is the shortest honest route from bias to efficiency. Wasserman gives the same material in compressed form alongside the plug-in principle. For the idea that deliberately biased estimators can win, Efron's account of shrinkage and the James–Stein phenomenon is the classic surprise, and the trade-off is the same one that governs modern regularised fitting.

Cheat sheet

QuantityMeaning here
Bias$\mathbb{E}[\hat\theta]-\theta$: the average miss, a property of the rule
Variance$\mathrm{Var}(\hat\theta)$: squared wobble of the estimates around their own mean
MSE$\mathbb{E}[(\hat\theta-\theta)^2]=\mathrm{Bias}^2+\mathrm{Var}$: total error
UnbiasedBias zero: right on average, not necessarily close on any sample
ConsistentConverges to the truth as $n\to\infty$; can be biased at every finite $n$
EfficiencyAmong unbiased rules, the smaller variance wins; the ratio says how many times more data the loser needs
÷(n−1)Bessel correction: the sample mean already cost one degree of freedom
÷(n+1)Minimum-MSE shrinkage of the variance estimate: biased low on purpose
Uniform maximumBiased low by $\theta/(n+1)$, consistent; $(n+1)/n$ times it is unbiased and efficient
7

Check your understanding

0/4 answered