Predictive uncertainty
A model that returns a single number is telling you only half the story. It has not said how much to trust that number, whether a different training sample would have moved it, or how wide a range of outcomes is still consistent with what it knows. This part is about turning a point prediction into a statement you can act on: a mean, a band around it, and a coverage number you can check. We build the band two ways — from the spread of an ensemble of fits, and from split conformal prediction, which comes with a finite-sample guarantee — and then we do the instructive thing: break that guarantee by shifting the data and watch the coverage fall away from the level it promised.
The question
How uncertain is this prediction, and can you prove it?
Regression and classification models are trained to be right on average, and a point prediction is the whole of their output. But a prediction of $2.4$ for some outcome is worthless without an idea of how far off it might be: a band of $\pm 0.1$ and a band of $\pm 3$ are entirely different claims, even though both are centred on the same number. The task of this part is to attach such a band, and then to attach a guarantee — a statement about how often new outcomes actually land inside it.
There are two genuinely different things that can make a prediction uncertain, and mixing them up is the most common mistake in the subject. One is the irreducible randomness of the world: even if you knew the true relationship exactly, new observations would still scatter around it. The other is your ignorance of that relationship: you have seen finitely many data points, so you do not know the function, and where data are sparse that ignorance is large. The first does not shrink when you collect more careful data; the second does. The first is called aleatoric uncertainty, the second epistemic, and the two demos on this page are built to keep them apart.
Alongside them runs a separate idea that is easy to confuse with uncertainty itself: confidence. A confidence statement is a property of a procedure, evaluated over repeated samples, and says nothing about any particular case. Coverage of a prediction interval is in that family. Keeping "how uncertain is this outcome" and "how often does my recipe cover" in separate mental boxes is most of what makes the second half of this page clear.
Aleatoric and epistemic
Two sources, two behaviours
Write the data-generating process as
The function $f$ is the signal you are trying to learn; the noise $\varepsilon$ is everything the input $x$ does not explain. In most real problems the noise is heteroscedastic: its spread depends on $x$. A sensor reads position precisely near its rest point and sloppily at range; a language model's next token is nearly deterministic after a formula and wide open after an ambiguous prompt. The demo data here have exactly that property: the noise grows with $x$, so the points fan out as you read left to right.
That noise is the aleatoric part, and no amount of extra data at the same $x$ removes it. The epistemic part is different. You never observe $f$ itself; you observe a finite sample and fit a model $\hat f$. Two different samples give two different fits, so $\hat f$ has a distribution of its own. Where data are dense every reasonable fit agrees and that distribution is narrow; where data are sparse — a gap in the design, or the region beyond the last observation — the fits diverge. Adding data in the sparse region shrinks it: epistemic uncertainty is reducible, aleatoric is not.
For a squared-error model the two add up in the predictive variance:
You almost never know $\sigma^2(x)$, so in practice both terms are estimated. Quantile regression attacks the first directly, an ensemble or a Bayesian posterior estimates the second, and the two can be combined. Part 16 built the Bayesian version — a posterior over the parameters that induces a posterior over $f$ — and Part 24 worried about whether predicted probabilities match observed frequencies. This part is the decision-facing companion: not just "how sure is the model", but "what range will contain the next outcome, and how often".
One caution before the demos. Epistemic uncertainty measured by a model's own spread is an artefact of that model class: a single cubic fit has no spread at all, and an ensemble can only be as diverse as the perturbations you feed it. If the true relationship lies outside every member's reach, the ensemble will agree confidently on the wrong answer. Conformal prediction avoids this by measuring the model's limits from held-out data instead of asking the model to know them.
An ensemble of fits
A band you get for free from resampling
The simplest way to see epistemic uncertainty is to stop pretending there is only one fit. Take the training data, make many slightly different training sets from it, fit one model on each, and treat the spread of their predictions as the uncertainty. The classical construction is the bootstrap: each member is fit on a resample of the original points, drawn with replacement, so every member sees a slightly different dataset. Where the data pin the curve down all the members agree; where they do not, the members fan apart. MC dropout is the same idea inside a neural network — each pass drops a random subset of units, giving a different function — and deep ensembles, trained from different initialisations, are a third variant.
The demo below fits a cubic polynomial to a deliberately awkward sample: dense at both ends, sparse in the middle, with noise that grows to the right. Each faint curve is one bootstrap member; the solid curve is their average, and the shaded region spans the middle ninety percent of the members' predictions at each $x$. The band is not the noise: the true aleatoric band, drawn faintly from the true noise level, tracks the fan of the data, while the ensemble band pinches where points crowd and balloons across the gap in the middle, precisely where the model has been given nothing to go on.
Faint curves: bootstrap members. Solid: ensemble mean. Shaded: middle 90% of members (epistemic). Dashed: the true function, with its aleatoric band behind.
Conformal prediction
A coverage promise you can count
Split conformal prediction takes any point predictor — here, the average of the ensemble — and wraps it in an interval with a guarantee, using only a held-out set of data it is allowed to look at. The recipe has four steps and no assumptions about the model. Split the data into three parts. Fit the predictor on the training set. On a separate calibration set, compute a nonconformity score for each point, most simply the absolute residual $s_i = |y_i - \hat f(x_i)|$. Sort those scores and pick the $k$-th smallest, where
Then for any new input the interval is $C(x) = [\hat f(x) - q,\; \hat f(x) + q]$, a band of constant width $q$. The surprising part is what that buy purchase: if the calibration points and the new point are exchangeable — roughly, if their joint distribution is symmetric so no point is special — then
The upper bound is the finite-sample exactness. Coverage is not approximately $1-\alpha$ in some large-sample limit; it is pinned to that value from above and below by the size of the calibration set, and the interval is valid for any predictor, however badly specified. A biased model simply gets a wider $q$; the width absorbs the model, so the coverage stays honest. That is a very different kind of promise from the ensemble spread, which had none.
The demo shows the promise being kept. Calibration and test sets are drawn from the same distribution as the training data, the interval is calibrated on the former, and the empirical coverage on the latter — the fraction of test points inside the band — tracks the nominal level as you move the slider. Notice too that the interval is the same width everywhere. Conditional on $x$, coverage is poor: where the noise is small the band is wider than it needs to be, and where the noise is large it is too narrow. The guarantee is marginal, averaged over the distribution of inputs, not conditional on each one. That gap is the most important thing to understand about conformal prediction, and it matters in the next section.
Solid: predictor. Shaded band: conformal interval $\hat f(x) \pm q$. Dots: test points, green inside the band and pink outside.
The order statistic is what makes the guarantee exact rather than asymptotic. With $n = 99$ calibration points and a nominal $90\%$, $k = \lceil 100 \times 0.9 \rceil = 90$, so $q$ is the ninetieth smallest of ninety-nine residuals: you deliberately keep the largest $10\%$ of residuals outside the band, and the finite-sample correction nudges the count to be slightly conservative. A larger calibration set tightens the guarantee toward exactly $1-\alpha$.
When the guarantee breaks
Exchangeability is doing all the work
The conformal guarantee rests on a single assumption, and it is not about the model at all. It is exchangeability: the calibration points and the new test point must be drawn in a way that makes their order irrelevant. If the data are i.i.d., they are exchangeable. If the test set has quietly drifted — a new sensor, a new user population, a season later, a different prompt distribution — then calibration and test are no longer interchangeable, the order statistic $q$ was computed for a distribution that no longer applies, and the guarantee evaporates. Nothing in the formula warns you. The interval is still printed, still looks like an interval, and still covers less often than it promised.
Drag the shift slider in the demo above. It slides the whole test set to the right, out of the region the training data covered and into a region where both the model's epistemic uncertainty and the world's noise are larger. At zero shift the empirical coverage sits on the nominal line, as the guarantee says it should. As the shift grows, coverage falls away from nominal — at the extreme, far below it. The calibrated width $q$ is frozen, because it was measured on a world that no longer exists, while the residuals it is supposed to bound have grown. This is the precise sense in which a coverage guarantee can be simultaneously true and useless: true under the exchangeability you assumed, void under the shift you did not check for.
There are principled repairs. Weighted conformal prediction reweights the calibration scores by how likely each point was under the test distribution, restoring a guarantee when that distribution is known. Adaptive methods recompute the score quantile on a rolling window of recent data, which works if drift is slow and labels eventually arrive. Conformal intervals can also be made conditionally sharper, with widths that vary with $x$, though full conditional coverage is impossible in general without further assumptions. None escape the fundamental point: an uncertainty statement is only as good as the assumption about where tomorrow's data come from, and distribution shift is the assumption that fails most often and most silently.
This is the uncertainty/confidence distinction made concrete. The model's own spread said one thing, the calibration guarantee said another, and the shifted data settled it. A prediction system that reports an interval must therefore say which world it is calibrated for — and monitor whether that world is still the one arriving at the door.
Where this shows up
Two places a band changes the decision
Abstaining when the band is wide
A deployed assistant should say "I don't know" rather than guess, and the trigger for abstention is a predictive interval that is too wide to act on. The guardrails chapter builds exactly these thresholds. Strong calibration on the evaluation set is also where the evaluation chapter's held-out numbers meet this part's machinery: a benchmark score is a point estimate, and a conformal interval is how you report its uncertainty honestly.
Covariance that drives the filter
Odometry and visual-inertial navigation fuse noisy measurements into a state estimate whose covariance is the epistemic uncertainty — the same object as the ensemble band here. The numerics part is the cautionary half: an overconfident covariance, like a conformal interval calibrated on the wrong world, is worse than an honest wide one because the planner trusts it.
The same picture recurs wherever a number drives a decision. Serving systems report latency quantiles, which are conformal-style statements about a tail; the serving metrics chapter lives on them, and a workload shift that changes the request mix is exactly the distribution shift that invalidates a previously calibrated threshold. The probability course's random-variable chapter supplies the distributions these bands are built from. Whether uncertainty is modelled explicitly, as in an ensemble or a Kalman covariance, or measured empirically, as in conformal calibration, a prediction without its uncertainty is an incomplete answer, and an uncertainty without its assumptions is a trap.
Further reading
The references below are ordered from the guarantee to its application. If you take away one thing, take away the split-conformal recipe and the exchangeability assumption it stands on — the rest is elaboration.
Vovk and coauthors introduced conformal prediction and prove the finite-sample coverage bound; Angelopoulos and Bates write the readable modern tutorial the page follows. Lakshminarayanan and coauthors give the deep-ensembles recipe, and Gal and Ghahramani show that dropout at test time is an approximate Bayesian model, which is the cleanest bridge from ensembles to the Bayesian part of this series.
- Vladimir Vovk, Alex Gammerman and Glenn Shafer, Algorithmic Learning in a Random World, 2005 — the origin of conformal prediction and its exchangeability-based guarantee.
- Anastasios Angelopoulos and Stephen Bates, "A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification", 2021 — the tutorial this page's recipe follows.
- Balaji Lakshminarayanan, Alexander Pritzel and Charles Blundell, "Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles", 2017 — ensembling as a practical epistemic estimator.
- Yarin Gal and Zoubin Ghahramani, "Dropout as a Bayesian Approximation", 2016 — MC dropout as approximate posterior uncertainty.
- Larry Wasserman, All of Statistics, chapters 11–13 — the confidence-interval and resampling background this part builds on.
Cheat sheet
| Term | Meaning here |
|---|---|
| Aleatoric uncertainty | The noise $\varepsilon$ in $y = f(x) + \varepsilon$; irreducible, often heteroscedastic |
| Epistemic uncertainty | Uncertainty about $f$ itself; shrinks as data accumulate, grows in sparse regions |
| Ensemble / MC dropout | A spread of fits whose disagreement estimates the epistemic term |
| Nonconformity score | How badly a calibration point fits; here the absolute residual $|y_i - \hat f(x_i)|$ |
| Conformal quantile | $s_{(k)}$ with $k = \lceil (n+1)(1-\alpha)\rceil$, the finite-sample-corrected order statistic |
| Coverage guarantee | $1-\alpha \le \mathbb{P}(y_{\text{new}} \in C) \le 1-\alpha + 1/(n+1)$ under exchangeability |
| Exchangeability | The assumption that calibration and test points are order-interchangeable; i.i.d. implies it |
| Marginal vs conditional | Coverage averaged over inputs, not guaranteed per input — hence constant width can be locally wrong |