Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Why does a model that fits the data better get worse on new data?

Suppose a true curve generates your data and you see only noisy samples from it. A more flexible model follows the training points more closely, so its error on the data you have falls almost monotonically as flexibility grows. Yet its error on fresh data falls, bottoms out, then climbs. Something must explain a quantity that improves on training data while worsening on new data.

The model was never trained on one fixed dataset; it was trained on whichever dataset you happened to collect. Change the noise and a rigid model barely moves while a flexible one swings wildly. So the error on a new sample has two sources: how far the model's average answer sits from the truth, and how far a single fit scatters around that average. The first is bias, the second is variance, and they add. Part 6 decomposed this for a single estimator; a fitted curve is just many parameters estimated at once, and the decomposition survives the promotion to a whole function.

💡 By the end of this part you'll see why the expected test error splits exactly into bias squared plus variance plus noise, why flexibility trades bias for variance, why more data lowers the variance term and leaves the bias term untouched, and where regularisation and the double-descent caveat fit into the picture.
2

The decomposition

Three terms that add to the error

Fix a point $x$ where you want to predict. The data are $y = f(x) + \varepsilon$, with $f$ the true curve and $\varepsilon$ noise of mean zero and variance $\sigma_\varepsilon^2$. Run your fitting procedure on a random dataset and it produces an estimate $\hat f(x)$; before you see the data that estimate is a random variable. Its expected squared error at $x$, averaged over all datasets you might have drawn, splits into three pieces exactly:

$$E\big[(y-\hat f(x))^2\big] = \underbrace{\big(E[\hat f(x)]-f(x)\big)^2}_{\text{bias}^2} + \underbrace{E\big[(\hat f(x)-E[\hat f(x)])^2\big]}_{\text{variance}} + \underbrace{\sigma_\varepsilon^2}_{\text{noise}}$$

Read the terms in order. Bias squared is the gap between the average fit and the truth: collect many datasets, fit each, average the curves, and compare that average to $f$. It belongs to the model family, not the sample — a line fitted to a wave stays biased however much data it sees. Variance is how much one fit scatters around that average, and it grows with how strongly the fit reacts to noise. Noise is the part nobody can remove; even a perfect model pays $\sigma_\varepsilon^2$ per new observation.

The identity is algebra, not approximation. Expand $(y-\hat f)^2$, split $y-f$ into the noise and $f-\hat f$, and the cross terms vanish because the noise is mean-zero and independent of the fit. You can trade bias against variance, but their sum cannot fall below the noise floor, and pushing one term down usually pushes the other up.

Bias is reported as its square because that is the term in the sum; an expectation over datasets is what no single number can display, which is why the next step watches many fits.

3

An ensemble of fits

Many datasets, one model family

The demo below hides a true curve and generates data from it. The grey points are one noisy sample; behind the scenes the page draws many fresh samples of the same size from the same curve and fits the same polynomial degree to each. The faint curves are those fits, the dashed curve is the truth, and the solid curve is their average. The gap between dashed and solid is the bias; the spread of the faint curves around the solid one is the variance.

Drag the degree slider. At low degree the polynomial is too rigid: the faint fits lie on top of one another, so variance is tiny, but the average misses the wave because no polynomial that simple can bend enough — bias dominates. Raise the degree and the average moulds itself to the truth, so bias falls, but the faint curves separate as each chases its own noise, so variance rises. At high degree some fits swing up between data points and others down, and the average is no better than before even though every individual fit chased its training data more closely.

Faint lines: one fit per resampled dataset. Dashed: the true curve. Solid: the average fit. The gap is bias; the spread is variance.

⚠ A single fit tells you nothing about bias. Bias is a statement about the average of curves you never drew, so the average fit is honest only when the ensemble is large. That is why the B slider exists: at $B = 5$ the solid curve is itself noisy, and the terms in the next chart wobble from redraw to redraw.
4

Underfitting, overfitting, and the U

The two failure modes on one axis

The second canvas turns the picture into a chart. For every degree from zero to the maximum, the page runs the whole ensemble and computes the three terms averaged over a grid of test points. Stacked at each degree they add up to the expected test error, and the total height traces the U. The left arm is underfitting, where bias is large and variance small; the right arm is overfitting, where bias has been squeezed down but variance has exploded. The bottom is the degree that trades the two best at this sample size and noise level.

The stacked bars make the addition literal: bias and variance sit on top of each other, with the noise block unchanged on top of both. That noise block is why the U never reaches zero — error on a new point cannot fall below $\sigma_\varepsilon^2$, because the point carries its own noise.

Stacked expected error versus degree: bias² + variance + noise. The total is the height of each bar.

Move the degree slider in the previous demo; it is the same control, and the highlight here follows it.

Underfitting and overfitting are the same trade seen from two ends. An underfit model's bias is so large that the variance it avoids does not compensate; it is stable and consistently wrong. An overfit model's variance is so large that the bias it removes is not worth it; it is faithful to the training points and untrustworthy elsewhere. Neither label is intrinsic to a degree — the same polynomial that underfits a wavy truth may be right for a straighter one. Complexity only means something relative to how much data you have and how wiggly the truth is.

It is worth naming what bias cannot see. The decomposition assumes data are tested on the distribution they were drawn from, and that noise is random rather than a systematic distortion in collection. A training set that is a biased slice of the world biases the average fit toward that slice, and no flexibility repairs it.

5

Data, complexity, and the caveat

Which term does what

The sliders separate the two levers cleanly. Raise the noise and every term worsens. Enlarge the training sample, holding the degree fixed, and variance falls while bias is essentially unchanged. This is the most useful fact in the picture: more data reduces variance, not bias. A rigid family is confidently wrong at any sample size; a family that can represent the truth simply has its parameters pinned down better.

The same asymmetry explains regularisation, the subject of the next part. A penalty on coefficient size makes the fitted curve less reactive to the sample, reducing variance at the cost of pulling the average fit from the truth, which raises bias. It is a dial on the trade, tuned by estimating test error through cross-validation, an information criterion, or a held-out set. The next part reframes the penalty as a prior over coefficients — the same dial in the language of the previous chapter.

One caveat keeps the picture honest. The clean U assumes you move smoothly through a fixed family. Modern over-parameterised models break that assumption: at the interpolation threshold, where parameters equal data points, variance can blow up, yet push past it and error can fall again — the "double descent" curve. The decomposition still holds; what changes is that bias and variance are no longer monotone in complexity. Test-set error remains the only honest arbiter.

That is why machine learning watches a validation set so obsessively. Training error is a biased estimate of test error, computed on data the model has seen, so it understates the error — and more flexible models understate it more. The apparatus of evaluation exists to estimate the held-out error the decomposition is about.

6

Where this shows up

Capacity, projection, and the rest of the site

ML / AI

Capacity versus data

The scaling chapter argues about model size, dataset size, and compute, which is this trade writ large: a larger model has lower bias and higher variance, and a larger dataset buys back the variance. Empirical scaling laws are a budget line drawn on a bias–variance chart you cannot read directly, which is also why training far past the point where the training loss stops improving often still helps the test loss.

Linear algebra

Fitting as projection

The least-squares part shows the fitted polynomial as the orthogonal projection of the data onto the column space of the design matrix. Higher degree means a larger subspace, which can only reduce the training residual — and it also means more directions for the noise to project into, which is the variance. Projection onto a bigger subspace explains why training error falls while test error can rise.

The two cards point at the two halves of the story: capacity enlarges the subspace and lowers bias, while data constrains the projection and lowers variance. Nothing else on the site escapes the arithmetic — a Kalman filter trading process noise against measurement noise, a pose graph weighting residuals, a model that is regularised or early-stopped are all being moved along the same axis.

There is even a numerical echo. A high-degree Vandermonde design matrix is ill-conditioned, so tiny changes in the data produce large changes in the coefficients — variance, computed by a machine. The numerics part treats that condition number as the sensitivity of a linear solve, the same sensitivity that appears here as the spread of the ensemble. The backpropagation part cares because it governs how gradients propagate through a deep model, and the Jacobian and Hessian part gives the local language for that sensitivity.

Further reading

If you take away one thing, take away the ensemble: bias is the gap between the average fit and the truth, variance is the spread of the fits around their average, and their sum is what you pay on new data.

Hastie, Tibshirani and Friedman derive the decomposition for general estimators and make it the organising principle of their book; Geman, Bienenstock and Doursat give the classic statement of the trade-off.

Cheat sheet

TermMeaning here
BiasHow far the average fit sits from the truth; a property of the model family
VarianceHow much a single fit scatters around that average; a property of the sample's noise
Noise$\sigma_\varepsilon^2$, the irreducible error in each new observation
Expected test error$\text{bias}^2 + \text{variance} + \sigma_\varepsilon^2$, and never below the last term
UnderfittingToo rigid: high bias, low variance, error stuck on the left arm of the U
OverfittingToo flexible: low bias, high variance, error climbing the right arm
More dataLowers variance; leaves bias essentially untouched
RegularisationTrades bias up to bring variance down; its tuning is a choice along the U
7

Check your understanding

0/4 answered