Cross-validation, AIC and BIC
Every fitted model reports a number for how well it did: the loss on the data you just fitted. That number is too small, and systematically so. This part is about the gap between the error a model shows you and the error it will actually make, the two honest ways to close it — holding data back, and paying for complexity analytically — and the moment more capacity buys a better-looking fit and a worse model.
The question
How well will this model do on data it has never seen?
A model turns inputs into predictions, and fitting chooses its parameters so that its predictions match the data. The trouble is that the data used for choosing are the same data used for judging. A flexible model can reduce its apparent error by learning the pattern or by memorising the noise in this sample; only the first transfers to a fresh sample.
The quantity you want is the test error: the expected loss on new observations from the same population. You never observe it directly. What you have is a training error on the sample you fitted, and some way of estimating the test error from it — which is the whole of model selection.
This part builds the two standard answers. One holds data out: fit on part of the sample, score on the rest, and rotate the held-out block. The other stays on the training data and adds a penalty for complexity, which is what AIC and BIC do.
Training error is optimistic
The gap between the number you see and the number you get
Write $\mathrm{Err}_{\text{train}}$ for the average loss over the fitted sample and $\mathrm{Err}_{\text{test}}$ for the expected loss on a fresh draw from the same population. Both are random, because the sample is random. But $\mathrm{Err}_{\text{train}}$ is systematically smaller, and the difference is called the optimism:
The reason is almost tautological. Fitting minimises the training loss, so the fitted parameters sit at the bottom of that particular valley. Some of the reduction came from fitting quirks of this sample that will not reappear in a fresh draw. With many parameters there are many directions in which to chase them, so training error keeps falling after the real structure has been captured, while held-out points never received that flattering treatment.
The simplest honest estimate is a held-out set: split the sample, fit on one part, and compute the loss on the other. Because the validation points were not used to choose the parameters, their loss is unbiased for the test error of that fitted model. The price is variance: a small validation set is noisy, while a large one leaves less data for training.
A second subtlety matters enormously. The moment you use the validation score to choose among models, the score stops being unbiased for the winner, because the winner was selected partly for how well it happened to do on those points. That selection bias is why a single split is dangerous, and it is what cross-validation controls.
K-fold cross-validation
Every point is held out exactly once
One split wastes data and wobbles with the luck of the split. Cross-validation fixes both by rotating the held-out block. Partition the sample into $k$ roughly equal folds. For each fold in turn, fit the model on the other $k-1$ folds and score it on the fold that was left out. Average the $k$ held-out scores, so that every point is used for validation exactly once and for training $k-1$ times.
In the demo below the points are partitioned into folds, shown as the coloured blocks. The current validation fold is highlighted, and its points are drawn in warning colour on the fit canvas — the model is fitted without them and then asked to predict them. Press Step fold to move the held-out block along and watch the error accumulate; press Run all folds for a complete pass. The readout pools the held-out errors into the cross-validation estimate.
Above: the current polynomial fit (solid), all points (light), and the held-out fold (warning colour). Below: the fold partition and each fold's held-out error.
The choice of $k$ is itself a bias–variance trade-off. Each model is trained on a fraction $(k-1)/k$ of the data, so a small $k$ trains on much less data than the final model will see and therefore overestimates the error — it is pessimistic. A large $k$ trains on nearly the whole sample, so the bias shrinks; but the training sets overlap heavily, the fitted models become similar, and the average inherits their correlation. Computation grows with $k$ too.
The extreme is leave-one-out cross-validation, $k = n$: each fit uses all but one point, so it is almost unbiased for the full-data model and involves no random split, at the cost of $n$ fits. For ordinary least squares there is a shortcut that avoids refitting: the leave-one-out residual for point $i$ is the ordinary residual divided by $1 - h_{ii}$, its leverage. In practice $k = 5$ or $k = 10$ is the usual compromise.
The optimism curve
Two curves, one gap that grows
Let capacity grow while the data stay fixed. For each polynomial degree $d = 1$ to $12$, fit on all points and record the training error, then estimate the same model by cross-validation. Training error falls monotonically, because a flexible polynomial can always get closer to the points it was fitted to. Validation error is U-shaped: it falls while capacity captures real structure, then turns up once capacity starts chasing noise.
The vertical gap between the curves is the optimism, and it widens with capacity. The degree where the validation curve bottoms out is the degree cross-validation would choose — and training error would happily pick a far more flexible model.
Training error (solid) and cross-validation error (warning colour) against polynomial degree. The dashed line marks the CV minimum; the faint line marks the degree you have selected.
There is an analytic version of the same picture. For least squares with $p$ parameters and noise variance $\sigma^2$, the expected optimism is about $2\sigma^2 p / n$: each extra parameter buys roughly two units of flattering. That factor of two reappears in Mallows' $C_p$ and in AIC below.
Two cautions about reading the curve. First, the cross-validation curve is itself an estimate with noise, and on a small sample that noise can move the apparent minimum; the common remedy is the one-standard-error rule, choosing the simplest model within one standard error of the minimum. Second, the curve is specific to the training size: change $k$ or $n$ and the shape shifts.
AIC and BIC
Complexity as a line item on the likelihood
Cross-validation needs many fits. An alternative needs one: correct the training criterion itself. Fit each candidate by maximum likelihood, let $L$ be the maximised likelihood and $k$ the number of free parameters, and compute
Here $-2\ln L$ is the deviance, a measure of how badly the model fits; larger is worse. The second term is a toll charged per parameter, and smaller values are better. The toll is what stops a flexible model from winning simply by fitting harder: the likelihood improves as parameters are added, and the penalty must be worth it.
The criteria differ only in the price of a parameter. AIC charges a flat $2$ regardless of sample size; BIC charges $\ln n$, which exceeds $2$ once $n \ge 8$ and grows slowly forever. BIC is therefore the stricter accountant and tends to pick simpler models. The table scores every degree on the same data, with the best AIC and BIC highlighted; compare them with the cross-validation minimum.
Where do these formulae come from? AIC estimates the expected Kullback–Leibler divergence between the fitted model and the truth, with a bias correction of one unit per parameter, doubled to $2k$ by convention. BIC approximates the marginal likelihood under a default prior, which is why it behaves like Bayesian model selection: as $n$ grows, BIC is consistent and selects the true model with probability tending to one, while AIC targets predictive accuracy and can keep some excess parameters forever.
For Gaussian regression, with $\hat\sigma^2 = \mathrm{SSE}/n$, the deviance is $n\ln(\mathrm{SSE}/n) + \text{const}$, so AIC is the residual sum of squares plus $2k$ — Mallows' $C_p$ in different notation.
Neither criterion is strictly better. AIC aims at prediction; BIC aims at parsimony and at identifying the true model, which is why it is consistent. Both assume maximum-likelihood fits to the same data with the same response and a meaningful parameter count; for regularised models, use effective degrees of freedom. Use cross-validation when the loss is not a likelihood or the model is a pipeline; use AIC or BIC when one fit must do — AIC for prediction, BIC for a sparser answer. When they disagree the extra complexity is marginal, and the simpler model is safer.
Where this shows up
The same gap, at two scales
Choosing a model size
The scaling chapter is this part's picture at industrial scale: training loss keeps falling with parameters and compute while held-out loss bottoms out and then rises. Scaling laws fit the held-out curve, and picking a size depends on never confusing the two.
Held-out evaluation
The evaluation chapter is held-out validation written large. A benchmark score is a validation estimate, and the instant you tune against it — a checkpoint, a prompt, a decoding setting — it acquires the selection bias cross-validation was invented to avoid. Contaminated test sets are optimism that escaped the lab.
The same discipline appears whenever many models share a single score. In the least-squares part the fitted coefficients minimise the residual on the data at hand, and the leave-one-out identity is a one-line consequence of the hat matrix; in a language-model run the held-out loss is the only number that survives contact with new text. The rule never changes: separate the data that scores from the data that chooses.
Further reading
If you take away one thing, take the picture of two curves — one that always improves, one that turns — and the discipline of never letting the same data choose and judge. The references below treat model assessment as central rather than a step tacked onto fitting.
- Trevor Hastie, Robert Tibshirani and Jerome Friedman, The Elements of Statistical Learning, 2nd ed., chapter 7 — model assessment and selection, the $1$-SE rule, and the LOO/leverage identity.
- Sylvain Arlot and Alain Celisse, "A survey of cross-validation procedures for model selection", Statistics Surveys, 2010 — the bias–variance of $k$ in detail.
- Hirotugu Akaike, "A new look at the statistical model identification", 1974 — where AIC comes from.
- Gideon Schwarz, "Estimating the dimension of a model", 1978 — the BIC penalty and its Bayesian reading.
Cheat sheet
| Term | Meaning here |
|---|---|
| Training error | Loss on the data used to fit the model; systematically too small |
| Test error | Expected loss on a fresh draw from the same population; the target |
| Optimism | $\mathbb{E}[\mathrm{Err}_{\text{test}}] - \mathbb{E}[\mathrm{Err}_{\text{train}}]$, about $2\sigma^2 p/n$ for OLS |
| Validation error | Loss on held-out points; honest for one model |
| $k$-fold CV | Average held-out error over $k$ disjoint folds; every point held out once |
| Choice of $k$ | Small $k$ is pessimistic and cheap; large $k$ is near-unbiased and correlated; $5$–$10$ is usual |
| Leave-one-out | $k = n$: no split randomness, $n$ fits, low bias and correlated errors |
| AIC | $2k - 2\ln L$: flat penalty per parameter; tuned for prediction |
| BIC | $k\ln n - 2\ln L$: penalty grows with $n$; favours parsimony and is consistent |
| Nested CV | Outer loop scores, inner loop selects; the only honest report after tuning |