Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

How well will this model do on data it has never seen?

A model turns inputs into predictions, and fitting chooses its parameters so that its predictions match the data. The trouble is that the data used for choosing are the same data used for judging. A flexible model can reduce its apparent error by learning the pattern or by memorising the noise in this sample; only the first transfers to a fresh sample.

The quantity you want is the test error: the expected loss on new observations from the same population. You never observe it directly. What you have is a training error on the sample you fitted, and some way of estimating the test error from it — which is the whole of model selection.

This part builds the two standard answers. One holds data out: fit on part of the sample, score on the rest, and rotate the held-out block. The other stays on the training data and adds a penalty for complexity, which is what AIC and BIC do.

By the end of this part you'll see why training error is optimistic and by how much, how k-fold cross-validation turns held-out data into a stable estimate, and how AIC and BIC replace the held-out set with a penalty on the likelihood — with different appetites for complexity.
2

Training error is optimistic

The gap between the number you see and the number you get

Write $\mathrm{Err}_{\text{train}}$ for the average loss over the fitted sample and $\mathrm{Err}_{\text{test}}$ for the expected loss on a fresh draw from the same population. Both are random, because the sample is random. But $\mathrm{Err}_{\text{train}}$ is systematically smaller, and the difference is called the optimism:

$$\text{optimism} = \mathbb{E}\big[\mathrm{Err}_{\text{test}}\big] - \mathbb{E}\big[\mathrm{Err}_{\text{train}}\big] \;\ge\; 0$$

The reason is almost tautological. Fitting minimises the training loss, so the fitted parameters sit at the bottom of that particular valley. Some of the reduction came from fitting quirks of this sample that will not reappear in a fresh draw. With many parameters there are many directions in which to chase them, so training error keeps falling after the real structure has been captured, while held-out points never received that flattering treatment.

The simplest honest estimate is a held-out set: split the sample, fit on one part, and compute the loss on the other. Because the validation points were not used to choose the parameters, their loss is unbiased for the test error of that fitted model. The price is variance: a small validation set is noisy, while a large one leaves less data for training.

A second subtlety matters enormously. The moment you use the validation score to choose among models, the score stops being unbiased for the winner, because the winner was selected partly for how well it happened to do on those points. That selection bias is why a single split is dangerous, and it is what cross-validation controls.

3

K-fold cross-validation

Every point is held out exactly once

One split wastes data and wobbles with the luck of the split. Cross-validation fixes both by rotating the held-out block. Partition the sample into $k$ roughly equal folds. For each fold in turn, fit the model on the other $k-1$ folds and score it on the fold that was left out. Average the $k$ held-out scores, so that every point is used for validation exactly once and for training $k-1$ times.

In the demo below the points are partitioned into folds, shown as the coloured blocks. The current validation fold is highlighted, and its points are drawn in warning colour on the fit canvas — the model is fitted without them and then asked to predict them. Press Step fold to move the held-out block along and watch the error accumulate; press Run all folds for a complete pass. The readout pools the held-out errors into the cross-validation estimate.

Above: the current polynomial fit (solid), all points (light), and the held-out fold (warning colour). Below: the fold partition and each fold's held-out error.

The choice of $k$ is itself a bias–variance trade-off. Each model is trained on a fraction $(k-1)/k$ of the data, so a small $k$ trains on much less data than the final model will see and therefore overestimates the error — it is pessimistic. A large $k$ trains on nearly the whole sample, so the bias shrinks; but the training sets overlap heavily, the fitted models become similar, and the average inherits their correlation. Computation grows with $k$ too.

The extreme is leave-one-out cross-validation, $k = n$: each fit uses all but one point, so it is almost unbiased for the full-data model and involves no random split, at the cost of $n$ fits. For ordinary least squares there is a shortcut that avoids refitting: the leave-one-out residual for point $i$ is the ordinary residual divided by $1 - h_{ii}$, its leverage. In practice $k = 5$ or $k = 10$ is the usual compromise.

⚠ Cross-validation estimates the error of a procedure, not the final model. Each fold's model trained on less data than you will eventually use. For choosing among models that is fine; for the winner's error, re-estimate on fresh data or an outer loop — the nesting in nested cross-validation.
4

The optimism curve

Two curves, one gap that grows

Let capacity grow while the data stay fixed. For each polynomial degree $d = 1$ to $12$, fit on all points and record the training error, then estimate the same model by cross-validation. Training error falls monotonically, because a flexible polynomial can always get closer to the points it was fitted to. Validation error is U-shaped: it falls while capacity captures real structure, then turns up once capacity starts chasing noise.

The vertical gap between the curves is the optimism, and it widens with capacity. The degree where the validation curve bottoms out is the degree cross-validation would choose — and training error would happily pick a far more flexible model.

Training error (solid) and cross-validation error (warning colour) against polynomial degree. The dashed line marks the CV minimum; the faint line marks the degree you have selected.

There is an analytic version of the same picture. For least squares with $p$ parameters and noise variance $\sigma^2$, the expected optimism is about $2\sigma^2 p / n$: each extra parameter buys roughly two units of flattering. That factor of two reappears in Mallows' $C_p$ and in AIC below.

Two cautions about reading the curve. First, the cross-validation curve is itself an estimate with noise, and on a small sample that noise can move the apparent minimum; the common remedy is the one-standard-error rule, choosing the simplest model within one standard error of the minimum. Second, the curve is specific to the training size: change $k$ or $n$ and the shape shifts.

5

AIC and BIC

Complexity as a line item on the likelihood

Cross-validation needs many fits. An alternative needs one: correct the training criterion itself. Fit each candidate by maximum likelihood, let $L$ be the maximised likelihood and $k$ the number of free parameters, and compute

$$\mathrm{AIC} = 2k - 2\ln L, \qquad \mathrm{BIC} = k\ln n - 2\ln L$$

Here $-2\ln L$ is the deviance, a measure of how badly the model fits; larger is worse. The second term is a toll charged per parameter, and smaller values are better. The toll is what stops a flexible model from winning simply by fitting harder: the likelihood improves as parameters are added, and the penalty must be worth it.

The criteria differ only in the price of a parameter. AIC charges a flat $2$ regardless of sample size; BIC charges $\ln n$, which exceeds $2$ once $n \ge 8$ and grows slowly forever. BIC is therefore the stricter accountant and tends to pick simpler models. The table scores every degree on the same data, with the best AIC and BIC highlighted; compare them with the cross-validation minimum.

Where do these formulae come from? AIC estimates the expected Kullback–Leibler divergence between the fitted model and the truth, with a bias correction of one unit per parameter, doubled to $2k$ by convention. BIC approximates the marginal likelihood under a default prior, which is why it behaves like Bayesian model selection: as $n$ grows, BIC is consistent and selects the true model with probability tending to one, while AIC targets predictive accuracy and can keep some excess parameters forever.

For Gaussian regression, with $\hat\sigma^2 = \mathrm{SSE}/n$, the deviance is $n\ln(\mathrm{SSE}/n) + \text{const}$, so AIC is the residual sum of squares plus $2k$ — Mallows' $C_p$ in different notation.

Neither criterion is strictly better. AIC aims at prediction; BIC aims at parsimony and at identifying the true model, which is why it is consistent. Both assume maximum-likelihood fits to the same data with the same response and a meaningful parameter count; for regularised models, use effective degrees of freedom. Use cross-validation when the loss is not a likelihood or the model is a pipeline; use AIC or BIC when one fit must do — AIC for prediction, BIC for a sparser answer. When they disagree the extra complexity is marginal, and the simpler model is safer.

⚠ The nested cross-validation caveat. If you use cross-validation both to choose a hyperparameter and to report the chosen model's error, the reported number is optimistic, because the held-out points helped select the winner. Nested cross-validation separates the jobs: an inner loop chooses on each outer training set, and an outer loop scores the chosen procedure on data it never saw. Believe the outer estimate.
6

Where this shows up

The same gap, at two scales

ML / AI

Choosing a model size

The scaling chapter is this part's picture at industrial scale: training loss keeps falling with parameters and compute while held-out loss bottoms out and then rises. Scaling laws fit the held-out curve, and picking a size depends on never confusing the two.

ML / AI

Held-out evaluation

The evaluation chapter is held-out validation written large. A benchmark score is a validation estimate, and the instant you tune against it — a checkpoint, a prompt, a decoding setting — it acquires the selection bias cross-validation was invented to avoid. Contaminated test sets are optimism that escaped the lab.

The same discipline appears whenever many models share a single score. In the least-squares part the fitted coefficients minimise the residual on the data at hand, and the leave-one-out identity is a one-line consequence of the hat matrix; in a language-model run the held-out loss is the only number that survives contact with new text. The rule never changes: separate the data that scores from the data that chooses.

Further reading

If you take away one thing, take the picture of two curves — one that always improves, one that turns — and the discipline of never letting the same data choose and judge. The references below treat model assessment as central rather than a step tacked onto fitting.

Cheat sheet

TermMeaning here
Training errorLoss on the data used to fit the model; systematically too small
Test errorExpected loss on a fresh draw from the same population; the target
Optimism$\mathbb{E}[\mathrm{Err}_{\text{test}}] - \mathbb{E}[\mathrm{Err}_{\text{train}}]$, about $2\sigma^2 p/n$ for OLS
Validation errorLoss on held-out points; honest for one model
$k$-fold CVAverage held-out error over $k$ disjoint folds; every point held out once
Choice of $k$Small $k$ is pessimistic and cheap; large $k$ is near-unbiased and correlated; $5$–$10$ is usual
Leave-one-out$k = n$: no split randomness, $n$ fits, low bias and correlated errors
AIC$2k - 2\ln L$: flat penalty per parameter; tuned for prediction
BIC$k\ln n - 2\ln L$: penalty grows with $n$; favours parsimony and is consistent
Nested CVOuter loop scores, inner loop selects; the only honest report after tuning
7

Check your understanding

0/4 answered