The bias–variance decomposition
Every model you fit is a guess made from data, and every guess can be wrong in two ways: the family you chose may be unable to express the truth, which is bias, or the particular sample you got may have pulled the fit somewhere unusual, which is variance. These two errors do not merely coexist — they add, and they add to the one quantity you actually care about, the expected squared error on new data. This part makes the sum visible: fit a polynomial of a degree you choose to a noisy dataset, repeat the exercise on many resampled datasets, and watch a cloud of fits fan out as bias falls and variance rises, their sum tracing the familiar U.
The question
Why does a model that fits the data better get worse on new data?
Suppose a true curve generates your data and you see only noisy samples from it. A more flexible model follows the training points more closely, so its error on the data you have falls almost monotonically as flexibility grows. Yet its error on fresh data falls, bottoms out, then climbs. Something must explain a quantity that improves on training data while worsening on new data.
The model was never trained on one fixed dataset; it was trained on whichever dataset you happened to collect. Change the noise and a rigid model barely moves while a flexible one swings wildly. So the error on a new sample has two sources: how far the model's average answer sits from the truth, and how far a single fit scatters around that average. The first is bias, the second is variance, and they add. Part 6 decomposed this for a single estimator; a fitted curve is just many parameters estimated at once, and the decomposition survives the promotion to a whole function.
The decomposition
Three terms that add to the error
Fix a point $x$ where you want to predict. The data are $y = f(x) + \varepsilon$, with $f$ the true curve and $\varepsilon$ noise of mean zero and variance $\sigma_\varepsilon^2$. Run your fitting procedure on a random dataset and it produces an estimate $\hat f(x)$; before you see the data that estimate is a random variable. Its expected squared error at $x$, averaged over all datasets you might have drawn, splits into three pieces exactly:
Read the terms in order. Bias squared is the gap between the average fit and the truth: collect many datasets, fit each, average the curves, and compare that average to $f$. It belongs to the model family, not the sample — a line fitted to a wave stays biased however much data it sees. Variance is how much one fit scatters around that average, and it grows with how strongly the fit reacts to noise. Noise is the part nobody can remove; even a perfect model pays $\sigma_\varepsilon^2$ per new observation.
The identity is algebra, not approximation. Expand $(y-\hat f)^2$, split $y-f$ into the noise and $f-\hat f$, and the cross terms vanish because the noise is mean-zero and independent of the fit. You can trade bias against variance, but their sum cannot fall below the noise floor, and pushing one term down usually pushes the other up.
Bias is reported as its square because that is the term in the sum; an expectation over datasets is what no single number can display, which is why the next step watches many fits.
An ensemble of fits
Many datasets, one model family
The demo below hides a true curve and generates data from it. The grey points are one noisy sample; behind the scenes the page draws many fresh samples of the same size from the same curve and fits the same polynomial degree to each. The faint curves are those fits, the dashed curve is the truth, and the solid curve is their average. The gap between dashed and solid is the bias; the spread of the faint curves around the solid one is the variance.
Drag the degree slider. At low degree the polynomial is too rigid: the faint fits lie on top of one another, so variance is tiny, but the average misses the wave because no polynomial that simple can bend enough — bias dominates. Raise the degree and the average moulds itself to the truth, so bias falls, but the faint curves separate as each chases its own noise, so variance rises. At high degree some fits swing up between data points and others down, and the average is no better than before even though every individual fit chased its training data more closely.
Faint lines: one fit per resampled dataset. Dashed: the true curve. Solid: the average fit. The gap is bias; the spread is variance.
Underfitting, overfitting, and the U
The two failure modes on one axis
The second canvas turns the picture into a chart. For every degree from zero to the maximum, the page runs the whole ensemble and computes the three terms averaged over a grid of test points. Stacked at each degree they add up to the expected test error, and the total height traces the U. The left arm is underfitting, where bias is large and variance small; the right arm is overfitting, where bias has been squeezed down but variance has exploded. The bottom is the degree that trades the two best at this sample size and noise level.
The stacked bars make the addition literal: bias and variance sit on top of each other, with the noise block unchanged on top of both. That noise block is why the U never reaches zero — error on a new point cannot fall below $\sigma_\varepsilon^2$, because the point carries its own noise.
Stacked expected error versus degree: bias² + variance + noise. The total is the height of each bar.
Underfitting and overfitting are the same trade seen from two ends. An underfit model's bias is so large that the variance it avoids does not compensate; it is stable and consistently wrong. An overfit model's variance is so large that the bias it removes is not worth it; it is faithful to the training points and untrustworthy elsewhere. Neither label is intrinsic to a degree — the same polynomial that underfits a wavy truth may be right for a straighter one. Complexity only means something relative to how much data you have and how wiggly the truth is.
It is worth naming what bias cannot see. The decomposition assumes data are tested on the distribution they were drawn from, and that noise is random rather than a systematic distortion in collection. A training set that is a biased slice of the world biases the average fit toward that slice, and no flexibility repairs it.
Data, complexity, and the caveat
Which term does what
The sliders separate the two levers cleanly. Raise the noise and every term worsens. Enlarge the training sample, holding the degree fixed, and variance falls while bias is essentially unchanged. This is the most useful fact in the picture: more data reduces variance, not bias. A rigid family is confidently wrong at any sample size; a family that can represent the truth simply has its parameters pinned down better.
The same asymmetry explains regularisation, the subject of the next part. A penalty on coefficient size makes the fitted curve less reactive to the sample, reducing variance at the cost of pulling the average fit from the truth, which raises bias. It is a dial on the trade, tuned by estimating test error through cross-validation, an information criterion, or a held-out set. The next part reframes the penalty as a prior over coefficients — the same dial in the language of the previous chapter.
One caveat keeps the picture honest. The clean U assumes you move smoothly through a fixed family. Modern over-parameterised models break that assumption: at the interpolation threshold, where parameters equal data points, variance can blow up, yet push past it and error can fall again — the "double descent" curve. The decomposition still holds; what changes is that bias and variance are no longer monotone in complexity. Test-set error remains the only honest arbiter.
That is why machine learning watches a validation set so obsessively. Training error is a biased estimate of test error, computed on data the model has seen, so it understates the error — and more flexible models understate it more. The apparatus of evaluation exists to estimate the held-out error the decomposition is about.
Where this shows up
Capacity, projection, and the rest of the site
Capacity versus data
The scaling chapter argues about model size, dataset size, and compute, which is this trade writ large: a larger model has lower bias and higher variance, and a larger dataset buys back the variance. Empirical scaling laws are a budget line drawn on a bias–variance chart you cannot read directly, which is also why training far past the point where the training loss stops improving often still helps the test loss.
Fitting as projection
The least-squares part shows the fitted polynomial as the orthogonal projection of the data onto the column space of the design matrix. Higher degree means a larger subspace, which can only reduce the training residual — and it also means more directions for the noise to project into, which is the variance. Projection onto a bigger subspace explains why training error falls while test error can rise.
The two cards point at the two halves of the story: capacity enlarges the subspace and lowers bias, while data constrains the projection and lowers variance. Nothing else on the site escapes the arithmetic — a Kalman filter trading process noise against measurement noise, a pose graph weighting residuals, a model that is regularised or early-stopped are all being moved along the same axis.
There is even a numerical echo. A high-degree Vandermonde design matrix is ill-conditioned, so tiny changes in the data produce large changes in the coefficients — variance, computed by a machine. The numerics part treats that condition number as the sensitivity of a linear solve, the same sensitivity that appears here as the spread of the ensemble. The backpropagation part cares because it governs how gradients propagate through a deep model, and the Jacobian and Hessian part gives the local language for that sensitivity.
Further reading
If you take away one thing, take away the ensemble: bias is the gap between the average fit and the truth, variance is the spread of the fits around their average, and their sum is what you pay on new data.
Hastie, Tibshirani and Friedman derive the decomposition for general estimators and make it the organising principle of their book; Geman, Bienenstock and Doursat give the classic statement of the trade-off.
- Trevor Hastie, Robert Tibshirani and Jerome Friedman, The Elements of Statistical Learning, chapter 7 — model assessment, selection, and the bias–variance decomposition.
- Stuart Geman, Elie Bienenstock and René Doursat, "Neural Networks and the Bias/Variance Dilemma", Neural Computation, 1992 — the classical analysis of the trade-off.
- Bradley Efron and Trevor Hastie, Computer Age Statistical Inference, chapter 7 — estimation error, prediction error, and the optimism of training error.
- Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal, "Reconciling modern machine-learning practice and the classical bias–variance trade-off", PNAS, 2019 — the double-descent caveat.
Cheat sheet
| Term | Meaning here |
|---|---|
| Bias | How far the average fit sits from the truth; a property of the model family |
| Variance | How much a single fit scatters around that average; a property of the sample's noise |
| Noise | $\sigma_\varepsilon^2$, the irreducible error in each new observation |
| Expected test error | $\text{bias}^2 + \text{variance} + \sigma_\varepsilon^2$, and never below the last term |
| Underfitting | Too rigid: high bias, low variance, error stuck on the left arm of the U |
| Overfitting | Too flexible: low bias, high variance, error climbing the right arm |
| More data | Lowers variance; leaves bias essentially untouched |
| Regularisation | Trades bias up to bring variance down; its tuning is a choice along the U |