Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

How do you rerun a study you cannot rerun?

An estimate is a rule applied to a sample, and its reliability is the spread of the rule over repeated samples. For the sample mean, and only for the sample mean of independent observations under tidy conditions, that spread has a formula, $\mathrm{SE} = \sigma/\sqrt{n}$. For a median, a correlation, a regression coefficient, a ratio of two means, or the maximum of a sample, the formula is either missing or so fragile that quoting it is an act of faith. The bootstrap replaces the formula with a simulation, and the simulation uses only the data on screen.

The move is a substitution. A parameter is a function of the population distribution, written $\theta = T(F)$. Its estimate is the same function of the data, $\hat\theta = T(\hat F)$, where $\hat F$ is the empirical distribution that puts probability $1/n$ on each observed value. Having made that substitution once, you can make it everywhere: the standard error is a function of $F$, so estimate it by the same function of $\hat F$; a confidence interval is a property of the sampling distribution under $F$, so build it from the sampling distribution under $\hat F$. The bootstrap is the machinery that evaluates those functions by sampling.

This part makes the substitution visible. You will watch a resample being drawn in front of you, see which original values came up more than once, and watch a histogram of the bootstrap statistic grow from nothing. Two kinds of interval will be laid over that histogram, and a third will be drawn for comparison. By the end you will be able to say what the bootstrap is assuming, when the assumption holds, and the cases where trusting it will quietly mislead you.

💡 By the end of this part you'll see why drawing with replacement from the sample mimics drawing from the population, why the histogram of resampled statistics plays the role of the sampling distribution, why percentile and BCa intervals differ, and why the bootstrap fails at the edges — heavy tails, extrema, and dependent data.
2

The plug-in principle

Replace the unknown by its best estimate

Suppose the population has CDF $F$, and you want a feature of it such as the mean, the median, or the fraction below a threshold. Every such feature is a functional $T(F)$: a recipe that reads a distribution and returns a number. The computation might be a sum, an integral, or the solution of an equation. In the population version the recipe is exact and the input is unknown. In the sample version the input is known — it is the empirical CDF $\hat F$ — and the recipe is applied to it unchanged.

The empirical CDF is the step function that jumps by $1/n$ at each observation. It is a perfectly ordinary distribution: the distribution you get by choosing one of the $n$ observed values uniformly at random. Its mean is the sample mean, its median is the sample median, and its quantiles are the sample quantiles. So plugging in is not an approximation grafted onto the problem; it is the statement that the sample is itself a small, lumpy population, and that we will treat it as the world.

$$\theta = T(F) \qquad\Longrightarrow\qquad \hat\theta = T(\hat F)$$

The picture below shows the substitution for the CDF itself. The dashed curve is the population CDF, which in real life you never see; the staircase is $\hat F$, built from the one sample. With few observations the staircase is coarse and wanders around the truth. As $n$ grows it hugs the curve. That convergence, measured by the Glivenko–Cantelli theorem, is why the substitution is more than a slogan: the stand-in eventually becomes indistinguishable from the thing it replaces.

Dashed: the true population CDF. Solid staircase: the empirical CDF of the sample, the only population the bootstrap ever gets to use.

The empirical distribution places mass $1/n$ on each observation, and that is the object the next section samples from. Nothing about the original parametric family survives the substitution — only the observed values remain, and the bootstrap never needs to know they came from a gamma curve.

The plug-in idea is not unique to resampling. Substituting the sample estimate for an unknown quantity in a formula is what every applied statistician does when they replace $\sigma$ by $s$ or a true rate by an observed rate. The bootstrap's contribution is to push the substitution all the way into the sampling distribution, so that quantities with no closed form can still be estimated by the same act of faith.

3

Resample in front of you

Draw with replacement, again and again

A bootstrap resample is a new sample of size $n$ drawn from the empirical distribution. Concretely: pick one of the original observations at random, write it down, put it back, and repeat until you have $n$ values. Because you replace each pick, the same original observation can appear several times, and some original observations may not appear at all. The expected number of times a given observation is left out is about $e^{-1} \approx 0.37$, so roughly a third of the sample is missing from each resample and a third of the slots are duplicates. That is not a flaw; it is exactly the variability that drawing from a population produces.

Why does this mimic sampling from the population? If the sample is your best picture of the population, then drawing a value uniformly from the sample is your best imitation of drawing a value from the population. Drawing $n$ such values with replacement imitates $n$ independent draws. The formal reason is conditional independence: given the observed data, the $n$ bootstrap draws are i.i.d. from $\hat F$, which is the same algebraic position the original observations occupied with respect to $F$. The plug-in principle says to use $\hat F$ wherever $F$ appeared, and this is what that means operationally.

The upper canvas shows the original sample as a row of dots. Press Draw a resample and a second row appears: the values that were drawn, sorted, with duplicates stacked. Dots in the warning colour were drawn more than once. Press the button repeatedly and the same values keep surfacing in different multiplicities, which is the whole engine of the method. The lower canvas is the payoff: each resample contributes one number, the statistic recomputed on it, and the histogram of those numbers is the bootstrap distribution. It is the sampling distribution from Part 4, with the population swapped for the sample and no algebra required.

Top: the original sample (dark) and one resample (light), repeated values highlighted. Bottom: the growing histogram of the resampled statistic, with intervals overlaid.

The bootstrap standard error is the standard deviation of this histogram, and it is the plug-in answer to a question that usually has no formula. For the mean it should land near $s/\sqrt{n}$, which is a useful sanity check rather than a definition. For the median it does something no elementary formula can, and for a regression coefficient it does the same job that the diagonal of a covariance matrix does in least squares, but without assuming the errors are normal or the model is correctly specified.

⚠ The bootstrap resamples observations, not information. It cannot manufacture precision that the sample does not contain, and it inherits every bias in how the data were collected. If the sample is a distorted slice of the world, the bootstrap distribution is a precise description of the wrong estimate. It quantifies random wobble, not systematic error.
4

Percentile, BCa, and the normal interval

Turning the histogram into an interval

Once the bootstrap distribution exists, a confidence interval is a matter of reading off two quantiles. The percentile interval takes the $\alpha/2$ and $1-\alpha/2$ quantiles of the resampled statistics as its ends. It is the most direct expression of the plug-in principle — the sampling distribution is estimated by the bootstrap histogram, so its central $1-\alpha$ range is estimated by the central range of the histogram. It is simple, transformation-respecting, and remarkably effective on well-behaved problems.

It has a blind spot. When the bootstrap distribution is skewed, the percentile interval shifts the estimate's value but does not correct for the skewness, so it can miss on one side more often than the other. The bias-corrected and accelerated interval, BCa, repairs this with two numbers. The bias correction $z_0$ measures how far the median of the bootstrap distribution sits from the original estimate, and it slides the interval endpoints in the right direction. The acceleration $a$ measures how the statistic's standard error changes as the estimate moves, and it is read off a jackknife — the sample recomputed $n$ times, once with each observation deleted. The endpoints are the plain quantiles of the bootstrap distribution evaluated at adjusted probabilities:

$$\alpha_{\text{adj}} = \Phi\!\left(z_0 + \frac{z_0 + z_{\alpha}}{1 - a\,(z_0 + z_{\alpha})}\right)$$

where $z_\alpha = \Phi^{-1}(\alpha)$ and $\Phi$ is the standard normal CDF. When the distribution is symmetric and the standard error is steady, $z_0 \approx 0$ and $a \approx 0$, and BCa collapses back onto the percentile interval. When the statistic is skewed, its file is very different. The demo reports both $z_0$ and $a$, and the interval canvas draws the three candidate intervals side by side for comparison.

The third interval is the familiar one: estimate $\pm z_{1-\alpha/2}\,\mathrm{SE}$, with $\mathrm{SE}$ from the normal-theory formula where available. It is the fastest to compute and the narrowest of the three when the sampling distribution is skewed, because it ignores the asymmetry entirely. On the skewed population used here it is the interval that is quietly overconfident. The bootstrap intervals widen and shift to admit the skew, and it is that honesty, not the arithmetic, that makes them worth the computation.

The same 95% interval three ways. The dot is the estimate; the bar is the interval. Where they disagree, the assumption is usually the reason.

Percentile and BCa are computed from the bootstrap histogram above; the normal interval uses the sample standard deviation and $z_{1-\alpha/2}$. The jackknife supplies the acceleration term and an independent standard-error estimate, and it is the bootstrap's close cousin: delete one observation at a time instead of resampling.

The jackknife is worth a sentence of its own. Deleting observation $i$ and recomputing the statistic gives $n$ values; their spread estimates the standard error, and their asymmetry gives the acceleration $a$. It is a linear approximation to the bootstrap and it is the tool of choice for small smooth statistics, but it fails badly for non-smooth ones such as a median or a maximum, where deleting one point barely moves the answer. The bootstrap is the more general instrument, and the two are usually reported together.

5

When the bootstrap fails

The cases where resampling misleads

The bootstrap is not magic, and its failure modes share a single cause: the empirical distribution must resemble the population in the region the statistic cares about. A sample of size $n$ has no values beyond its own maximum, so if the statistic is an extreme value, the plug-in world is missing exactly the tail that drives the answer. The canvas below makes this concrete. The statistic is the maximum of an exponential sample, the population is known, and both the true sampling distribution and the bootstrap distribution are simulated. The bootstrap cloud sits to the left of the truth, because a resample can never exceed the sample maximum, while a fresh sample from the population routinely does. The interval built from it is biased and too narrow.

Faint: the true sampling distribution of the maximum, simulated from the known population. Highlighted: the bootstrap distribution of the maximum. The bootstrap is shifted low and cannot see the tail.

The same failure appears for any statistic that depends on presence in the tail: the minimum, the range, a quantile above the largest observation, and a probability of an event that did not occur. Heavy-tailed distributions make it worse, because a single unseen extreme value can dominate the estimate.

Three more failure modes are worth naming. First, dependent data: if observations are clustered, as in repeated measurements from one sensor or several frames from one drive, resampling individual observations destroys the dependence structure and the resulting intervals are too narrow. Blocked or cluster bootstrap variants exist for exactly this situation. Second, small $n$: with a handful of observations the empirical distribution is a poor stand-in, and the bootstrap histogram is too coarse and too lumpy to trust. Third, non-smooth statistics such as the median with heavy ties, where the resampled value jumps between a few discrete levels and the normal approximation behind BCa degrades.

When the bootstrap is inappropriate, a permutation test can sometimes take its place. It also resamples, but under a different question: instead of asking how variable the estimate is, it asks how surprising the observed value is if a label could be shuffled. That reframing is the subject of the testing part of this series, and the two methods are not substitutes: the bootstrap estimates the shape of an estimate's distribution, while a permutation test calibrates a null hypothesis. Knowing which question you are asking picks the tool.

6

Where this shows up

Resampling beyond the textbook

Vision & Geometry

Robust fitting by resampling

RANSAC in the multi-view geometry guide is resampling used for a different purpose: it draws a random minimal subset of correspondences, fits a model, and counts inliers, repeating until a consensus emerges. The bootstrap asks how variable a statistic is; RANSAC asks which model survives contamination. Both rest on the same insight — that repeated random subsets of your data can answer questions no single fit can.

ML / AI

Error bars on a metric

A benchmark score computed on a finite evaluation set is an estimate, and a different set would give a different number. The serving metrics chapter tracks latency and accuracy under load; the honest version of each number carries an interval. Bootstrapping the evaluation items — resampling queries, not individual tokens — produces confidence intervals on metrics such as accuracy or median latency without assuming a distribution for them.

The same arithmetic appears wherever a quantity is estimated from a finite, noisy set. In odometry, a robot accumulates measurements and needs a believable covariance, not just a point estimate, and resampling is one way to get one when the noise model is uncertain. In probability, the exponential and gamma distributions that generate this page's data are the same objects whose tail behaviour determines whether the bootstrap can be trusted. The plug-in principle is the thread: whenever a formula needs a distribution it does not have, substituting the empirical one is the move, and the bootstrap is that move carried out by simulation.

Further reading

The references below treat the bootstrap as a general-purpose inference tool and are honest about its limits. If you take away one thing, take away the picture of a histogram built by resampling, with the population replaced by the empirical distribution.

Efron's original paper introduced the method and the jackknife connection; Efron and Tibshirani's book is the standard gentle reference, and Davison and Hinkley is the one to reach for when a problem is awkward. Wasserman's chapter places the bootstrap inside the plug-in framework this part follows.

Cheat sheet

TermMeaning here
Plug-in principleReplace the unknown $F$ by the empirical $\hat F$ and recompute the same functional
Empirical distributionMass $1/n$ on each observation; the population the bootstrap samples from
Bootstrap resample$n$ values drawn with replacement from the sample; some repeated, some omitted
Bootstrap distributionHistogram of the statistic over many resamples; stands in for the sampling distribution
Bootstrap SEStandard deviation of that histogram; the plug-in standard error
Percentile intervalCentral $1-\alpha$ range of the bootstrap distribution
BCaPercentile interval with bias correction $z_0$ and acceleration $a$ from the jackknife
JackknifeRecompute the statistic $n$ times, deleting one observation each time
7

Check your understanding

0/4 answered