Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

How uncertain is this prediction, and can you prove it?

Regression and classification models are trained to be right on average, and a point prediction is the whole of their output. But a prediction of $2.4$ for some outcome is worthless without an idea of how far off it might be: a band of $\pm 0.1$ and a band of $\pm 3$ are entirely different claims, even though both are centred on the same number. The task of this part is to attach such a band, and then to attach a guarantee — a statement about how often new outcomes actually land inside it.

There are two genuinely different things that can make a prediction uncertain, and mixing them up is the most common mistake in the subject. One is the irreducible randomness of the world: even if you knew the true relationship exactly, new observations would still scatter around it. The other is your ignorance of that relationship: you have seen finitely many data points, so you do not know the function, and where data are sparse that ignorance is large. The first does not shrink when you collect more careful data; the second does. The first is called aleatoric uncertainty, the second epistemic, and the two demos on this page are built to keep them apart.

Alongside them runs a separate idea that is easy to confuse with uncertainty itself: confidence. A confidence statement is a property of a procedure, evaluated over repeated samples, and says nothing about any particular case. Coverage of a prediction interval is in that family. Keeping "how uncertain is this outcome" and "how often does my recipe cover" in separate mental boxes is most of what makes the second half of this page clear.

💡 By the end of this part you'll see why an ensemble's spread estimates epistemic uncertainty, why a conformal interval calibrated on held-out data covers at the rate you asked for, and why that promise is conditional on exchangeability — the assumption a shifted test set quietly violates.
2

Aleatoric and epistemic

Two sources, two behaviours

Write the data-generating process as

$$y = f(x) + \varepsilon, \qquad \varepsilon \sim \mathcal{N}\!\left(0, \sigma^2(x)\right).$$

The function $f$ is the signal you are trying to learn; the noise $\varepsilon$ is everything the input $x$ does not explain. In most real problems the noise is heteroscedastic: its spread depends on $x$. A sensor reads position precisely near its rest point and sloppily at range; a language model's next token is nearly deterministic after a formula and wide open after an ambiguous prompt. The demo data here have exactly that property: the noise grows with $x$, so the points fan out as you read left to right.

That noise is the aleatoric part, and no amount of extra data at the same $x$ removes it. The epistemic part is different. You never observe $f$ itself; you observe a finite sample and fit a model $\hat f$. Two different samples give two different fits, so $\hat f$ has a distribution of its own. Where data are dense every reasonable fit agrees and that distribution is narrow; where data are sparse — a gap in the design, or the region beyond the last observation — the fits diverge. Adding data in the sparse region shrinks it: epistemic uncertainty is reducible, aleatoric is not.

For a squared-error model the two add up in the predictive variance:

$$\operatorname{Var}(y \mid x) \;\approx\; \underbrace{\sigma^2(x)}_{\text{aleatoric}} \;+\; \underbrace{\operatorname{Var}\!\left(\hat f(x)\right)}_{\text{epistemic}}.$$

You almost never know $\sigma^2(x)$, so in practice both terms are estimated. Quantile regression attacks the first directly, an ensemble or a Bayesian posterior estimates the second, and the two can be combined. Part 16 built the Bayesian version — a posterior over the parameters that induces a posterior over $f$ — and Part 24 worried about whether predicted probabilities match observed frequencies. This part is the decision-facing companion: not just "how sure is the model", but "what range will contain the next outcome, and how often".

One caution before the demos. Epistemic uncertainty measured by a model's own spread is an artefact of that model class: a single cubic fit has no spread at all, and an ensemble can only be as diverse as the perturbations you feed it. If the true relationship lies outside every member's reach, the ensemble will agree confidently on the wrong answer. Conformal prediction avoids this by measuring the model's limits from held-out data instead of asking the model to know them.

3

An ensemble of fits

A band you get for free from resampling

The simplest way to see epistemic uncertainty is to stop pretending there is only one fit. Take the training data, make many slightly different training sets from it, fit one model on each, and treat the spread of their predictions as the uncertainty. The classical construction is the bootstrap: each member is fit on a resample of the original points, drawn with replacement, so every member sees a slightly different dataset. Where the data pin the curve down all the members agree; where they do not, the members fan apart. MC dropout is the same idea inside a neural network — each pass drops a random subset of units, giving a different function — and deep ensembles, trained from different initialisations, are a third variant.

The demo below fits a cubic polynomial to a deliberately awkward sample: dense at both ends, sparse in the middle, with noise that grows to the right. Each faint curve is one bootstrap member; the solid curve is their average, and the shaded region spans the middle ninety percent of the members' predictions at each $x$. The band is not the noise: the true aleatoric band, drawn faintly from the true noise level, tracks the fan of the data, while the ensemble band pinches where points crowd and balloons across the gap in the middle, precisely where the model has been given nothing to go on.

Faint curves: bootstrap members. Solid: ensemble mean. Shaded: middle 90% of members (epistemic). Dashed: the true function, with its aleatoric band behind.

⚠ An ensemble's spread is not a guarantee. It measures disagreement within the model class you chose, under the perturbations you gave it. It can be too narrow (members agree on a shared mistake) or too wide (members differ for irrelevant reasons), and nothing in the construction bounds how often a new observation falls inside the band. Quantifying that is a different job, and it is the next section's.
4

Conformal prediction

A coverage promise you can count

Split conformal prediction takes any point predictor — here, the average of the ensemble — and wraps it in an interval with a guarantee, using only a held-out set of data it is allowed to look at. The recipe has four steps and no assumptions about the model. Split the data into three parts. Fit the predictor on the training set. On a separate calibration set, compute a nonconformity score for each point, most simply the absolute residual $s_i = |y_i - \hat f(x_i)|$. Sort those scores and pick the $k$-th smallest, where

$$k = \left\lceil (n+1)(1-\alpha) \right\rceil, \qquad q = s_{(k)}.$$

Then for any new input the interval is $C(x) = [\hat f(x) - q,\; \hat f(x) + q]$, a band of constant width $q$. The surprising part is what that buy purchase: if the calibration points and the new point are exchangeable — roughly, if their joint distribution is symmetric so no point is special — then

$$1-\alpha \;\le\; \mathbb{P}\!\left(y_{\text{new}} \in C(x_{\text{new}})\right) \;\le\; 1-\alpha + \frac{1}{n+1}.$$

The upper bound is the finite-sample exactness. Coverage is not approximately $1-\alpha$ in some large-sample limit; it is pinned to that value from above and below by the size of the calibration set, and the interval is valid for any predictor, however badly specified. A biased model simply gets a wider $q$; the width absorbs the model, so the coverage stays honest. That is a very different kind of promise from the ensemble spread, which had none.

The demo shows the promise being kept. Calibration and test sets are drawn from the same distribution as the training data, the interval is calibrated on the former, and the empirical coverage on the latter — the fraction of test points inside the band — tracks the nominal level as you move the slider. Notice too that the interval is the same width everywhere. Conditional on $x$, coverage is poor: where the noise is small the band is wider than it needs to be, and where the noise is large it is too narrow. The guarantee is marginal, averaged over the distribution of inputs, not conditional on each one. That gap is the most important thing to understand about conformal prediction, and it matters in the next section.

Solid: predictor. Shaded band: conformal interval $\hat f(x) \pm q$. Dots: test points, green inside the band and pink outside.

The order statistic is what makes the guarantee exact rather than asymptotic. With $n = 99$ calibration points and a nominal $90\%$, $k = \lceil 100 \times 0.9 \rceil = 90$, so $q$ is the ninetieth smallest of ninety-nine residuals: you deliberately keep the largest $10\%$ of residuals outside the band, and the finite-sample correction nudges the count to be slightly conservative. A larger calibration set tightens the guarantee toward exactly $1-\alpha$.

5

When the guarantee breaks

Exchangeability is doing all the work

The conformal guarantee rests on a single assumption, and it is not about the model at all. It is exchangeability: the calibration points and the new test point must be drawn in a way that makes their order irrelevant. If the data are i.i.d., they are exchangeable. If the test set has quietly drifted — a new sensor, a new user population, a season later, a different prompt distribution — then calibration and test are no longer interchangeable, the order statistic $q$ was computed for a distribution that no longer applies, and the guarantee evaporates. Nothing in the formula warns you. The interval is still printed, still looks like an interval, and still covers less often than it promised.

Drag the shift slider in the demo above. It slides the whole test set to the right, out of the region the training data covered and into a region where both the model's epistemic uncertainty and the world's noise are larger. At zero shift the empirical coverage sits on the nominal line, as the guarantee says it should. As the shift grows, coverage falls away from nominal — at the extreme, far below it. The calibrated width $q$ is frozen, because it was measured on a world that no longer exists, while the residuals it is supposed to bound have grown. This is the precise sense in which a coverage guarantee can be simultaneously true and useless: true under the exchangeability you assumed, void under the shift you did not check for.

There are principled repairs. Weighted conformal prediction reweights the calibration scores by how likely each point was under the test distribution, restoring a guarantee when that distribution is known. Adaptive methods recompute the score quantile on a rolling window of recent data, which works if drift is slow and labels eventually arrive. Conformal intervals can also be made conditionally sharper, with widths that vary with $x$, though full conditional coverage is impossible in general without further assumptions. None escape the fundamental point: an uncertainty statement is only as good as the assumption about where tomorrow's data come from, and distribution shift is the assumption that fails most often and most silently.

This is the uncertainty/confidence distinction made concrete. The model's own spread said one thing, the calibration guarantee said another, and the shifted data settled it. A prediction system that reports an interval must therefore say which world it is calibrated for — and monitor whether that world is still the one arriving at the door.

6

Where this shows up

Two places a band changes the decision

ML / AI

Abstaining when the band is wide

A deployed assistant should say "I don't know" rather than guess, and the trigger for abstention is a predictive interval that is too wide to act on. The guardrails chapter builds exactly these thresholds. Strong calibration on the evaluation set is also where the evaluation chapter's held-out numbers meet this part's machinery: a benchmark score is a point estimate, and a conformal interval is how you report its uncertainty honestly.

Robotics

Covariance that drives the filter

Odometry and visual-inertial navigation fuse noisy measurements into a state estimate whose covariance is the epistemic uncertainty — the same object as the ensemble band here. The numerics part is the cautionary half: an overconfident covariance, like a conformal interval calibrated on the wrong world, is worse than an honest wide one because the planner trusts it.

The same picture recurs wherever a number drives a decision. Serving systems report latency quantiles, which are conformal-style statements about a tail; the serving metrics chapter lives on them, and a workload shift that changes the request mix is exactly the distribution shift that invalidates a previously calibrated threshold. The probability course's random-variable chapter supplies the distributions these bands are built from. Whether uncertainty is modelled explicitly, as in an ensemble or a Kalman covariance, or measured empirically, as in conformal calibration, a prediction without its uncertainty is an incomplete answer, and an uncertainty without its assumptions is a trap.

Further reading

The references below are ordered from the guarantee to its application. If you take away one thing, take away the split-conformal recipe and the exchangeability assumption it stands on — the rest is elaboration.

Vovk and coauthors introduced conformal prediction and prove the finite-sample coverage bound; Angelopoulos and Bates write the readable modern tutorial the page follows. Lakshminarayanan and coauthors give the deep-ensembles recipe, and Gal and Ghahramani show that dropout at test time is an approximate Bayesian model, which is the cleanest bridge from ensembles to the Bayesian part of this series.

Cheat sheet

TermMeaning here
Aleatoric uncertaintyThe noise $\varepsilon$ in $y = f(x) + \varepsilon$; irreducible, often heteroscedastic
Epistemic uncertaintyUncertainty about $f$ itself; shrinks as data accumulate, grows in sparse regions
Ensemble / MC dropoutA spread of fits whose disagreement estimates the epistemic term
Nonconformity scoreHow badly a calibration point fits; here the absolute residual $|y_i - \hat f(x_i)|$
Conformal quantile$s_{(k)}$ with $k = \lceil (n+1)(1-\alpha)\rceil$, the finite-sample-corrected order statistic
Coverage guarantee$1-\alpha \le \mathbb{P}(y_{\text{new}} \in C) \le 1-\alpha + 1/(n+1)$ under exchangeability
ExchangeabilityThe assumption that calibration and test points are order-interchangeable; i.i.d. implies it
Marginal vs conditionalCoverage averaged over inputs, not guaranteed per input — hence constant width can be locally wrong
7

Check your understanding

0/4 answered