The bootstrap
You have one sample and one number to report. Somebody asks how uncertain that number is, and for the sample mean there is a tidy answer: the standard error is $s/\sqrt{n}$. But the moment the quantity you care about is the median, a correlation, a ratio of two fitted coefficients, or the ninety-fifth percentile of a latency distribution, the tidy answer is gone. The classical escape is to assume a model and grind out an asymptotic variance — often hard, always approximate, and sometimes wrong in a way nobody notices. The bootstrap takes a different route: it treats the sample itself as a stand-in for the population, draws new samples from it, and watches the statistic move. It needs almost no algebra. This part builds that resampling picture from scratch, turns the resulting cloud of statistics into an error bar, and then spends real effort on the places where the trick quietly lies.
The question
One dataset, and no error formula
Suppose you run an experiment thirty times and record thirty numbers. The quantity you report is some function $\hat\theta$ of those numbers: a mean, a median, a slope, an interquartile range. A single number is never the whole story — repeated experiments would have produced different numbers, and the reader wants to know how much $\hat\theta$ would move if the experiment were run again. That movement is the sampling distribution of $\hat\theta$, and its standard deviation is the standard error. Everything in this part is about getting at that sampling distribution when you have only one sample.
The classical route goes through theory. If the observations are independent draws from F, then the sample mean has variance $\sigma^2/n$, estimated by s^2/n; the sample proportion has variance p(1-p)/n; and for anything smooth, the delta method linearises the statistic and the Fisher information gives an asymptotic variance. That is a beautiful machine, and this volume built it in Part 3. But it has costs. It requires a model for F, it requires the statistic to be differentiable in the data, and each new statistic needs its own derivation. The median of a skewed distribution, the ratio of two sample means, the correlation between two noisy series: each is a small research project.
The bootstrap sidesteps the derivation. It says: you do not know F, but you have an estimate of it sitting right in front of you. The thirty observations, taken together, define the empirical distribution $\hat F$, which puts probability 1/n on each observation. That distribution is a crude but honest model of the population. Draw fresh samples of size n from $\hat F$, compute $\hat\theta$ on each, and you get a cloud of values. The spread of that cloud estimates the spread of $\hat\theta$ under the real F. No formula for the median was needed; the data did the work.
This is the plug-in principle, applied not to a single number but to a whole distribution: wherever the true sampling distribution would use F, substitute $\hat F$. The only real assumption is that the sample is representative, which is the same thing as saying the observations are independent draws from a common distribution. When that assumption holds, the substitution is remarkably effective. When it fails, the failure is usually loud and diagnosable, and the last section of this part is devoted to making it loud.
Resampling from the empirical distribution
The sample as a miniature population
Write the data as $x_1,\dots,x_n$. The empirical distribution function counts how much of the sample lies at or below x, divided by n: a step function that rises by 1/n at each observation and reaches one at the largest one. It is the distribution that would result if each data point were equally likely, and it is the nonparametric maximum-likelihood estimate of F — the distribution that makes the observed sample as probable as possible without assuming a shape.
A bootstrap resample is a fresh sample of size n drawn from $\hat F_n$. Drawing from a distribution that has mass 1/n on each observed value means picking one of the n observations uniformly, n times over, with replacement. That last phrase carries all the weight. Because we replace each chosen value before choosing again, one original point can appear several times in a resample, and another can be left out entirely. The resample is not a permutation of the data. It is a new sample of the same size from a world in which the data are the only possible values.
The expected number of times a particular observation is chosen is exactly one, but that is an average. The count is Binomial(n,1/n), so about $37\%$ of the data points are absent from any given resample and some appear two, three or four times. The statistic $\hat\theta^*$ computed on the resample therefore differs from $\hat\theta$ computed on the original data, and the pattern of those differences is the thing we are trying to learn.
The demo below makes the mechanism visible. There is one fixed dataset of thirty values, generated from a seeded stream so it is the same on every reload. The grey rug along the bottom is that dataset. Press Draw one resample and watch the blue dots land on the grey points: each dot is one draw, and dots stack up on whichever original value was picked. Some grey points collect a small tower; others stay bare. The dashed line is the observed statistic, and the moving marker is the statistic of the resample being built. Each completed resample drops a single tally into the histogram underneath, and after a few hundred of them the histogram has become a picture of the sampling distribution — reconstructed from the one dataset and nothing else.
Grey rug: the observed dataset of thirty values. Blue dots: copies drawn so far in the current resample, stacked on the original value they came from. The dashed line is the observed statistic.
Every finished resample adds one count to this histogram. The shaded band is the 95% percentile interval of the bootstrap distribution.
Two habits are worth forming at this point. First, the resample size must equal the original sample size; drawing fewer points would make the resampled statistic noisier than the real one and inflate the error bar. Second, the resampling is done with replacement, full stop. Sampling without replacement would return the original dataset every time, which tells you nothing about how the statistic would move under a new experiment. The entire information content of the bootstrap lives in the fact that a resample is genuinely different from the data it came from.
The bootstrap distribution and intervals
Turning a cloud of numbers into an error bar
Repeat the resampling B times, where B is a few thousand for real work, and collect the statistics $\hat\theta^{*1},\dots,\hat\theta^{*B}$. Their distribution is the bootstrap distribution. It is not the true sampling distribution of $\hat\theta$; it is an approximation obtained by replacing F with $\hat F$. The claim of the bootstrap is that these two distributions are close whenever the sample is a fair miniature of the population, and that closeness is what makes the method useful.
The simplest thing to read off is the bootstrap standard error: the sample standard deviation of the resampled statistics. It is an estimate of the standard error of $\hat\theta$, obtained without any formula for the standard error of $\hat\theta$. For the sample mean it will land on the familiar $s/\sqrt n$. For the median it returns a number you could not have written down from memory, and that is precisely the point.
An error bar is not an interval, and the bootstrap can produce intervals too. The percentile interval simply reads ordered quantiles off the bootstrap distribution: for a 95% interval, take the 2.5th and 97.5th percentiles of the resampled statistics. It is the most direct method and the easiest to explain, and it is what the shaded band in the demo shows.
The basic interval exploits the fact that the bootstrap distribution estimates the error distribution of $\hat\theta$ around the truth. If $\hat\theta^*-\hat\theta$ has the same distribution as $\hat\theta-\theta$, then rearranging gives the reflected interval above, which corrects a systematic shift that the percentile interval bakes in. When the bootstrap distribution is badly skewed, neither endpoint rule is quite right, and the BCa interval adds two corrections: a bias term z_0 measuring how far the median of the bootstrap distribution sits from $\hat\theta$, and an acceleration term from a jackknife that adjusts for how the statistic's variance changes with the parameter. Even a crude version with the acceleration set to zero, which is what the readout below computes, moves the endpoints visibly when the distribution leans.
With B=1000 the shape and the standard error are stable to about a percent or so; with B=10 the histogram is a rumour. The relevant noise here is Monte Carlo noise — the error introduced by using finitely many resamples when the ideal has infinitely many. It shrinks like $1/\sqrt B$, exactly like every other simulation in this volume, and it is separate from the modelling error of using $\hat F$ in place of F, which no amount of computing removes.
When the bootstrap fails
The maximum, and other statistics that live at the edge
The bootstrap assumes that resampling from $\hat F$ mimics resampling from F. For an average that assumption is gentle: the sample mean responds to the whole distribution, and the empirical distribution captures the whole distribution well. For a statistic that depends on the extreme value of the sample, the assumption collapses. The sample maximum is the clean example.
Take a small sample and bootstrap its maximum. Every observation in a resample is a copy of an original observation, so the largest value any resample can contain is the largest value the original sample contained. The bootstrap distribution of the maximum therefore lives entirely below the observed maximum, with a hard wall at that point. Its mean sits below the observed max — the bootstrap says the maximum is biased low — and its spread is too narrow, because the real sampling distribution of the maximum has an upper tail that the bootstrap can never reach. The bootstrap is not merely imprecise here; it is structurally incapable of representing the uncertainty, and any confidence interval built from it will be too short.
The demo below makes the failure quantitative. For this one demonstration we allow ourselves to know the true distribution: the sample is drawn from a standard normal. That lets us compute an oracle sampling distribution — thousands of fresh samples of the same size from the true normal — and plot it in pink beside the bootstrap in blue. On the right, for the mean, the two histograms sit almost on top of each other: the bootstrap has recovered the sampling distribution, and the standard errors agree. On the left, for the maximum, the oracle extends into a tail the bootstrap cannot enter, and the bootstrap's standard error is far too small. Drag the sample-size slider down and both effects worsen, because a small sample says even less about the tail it has not seen.
Left: the maximum. Pink is the true sampling distribution from fresh samples, blue is the bootstrap; the bootstrap cannot pass the observed maximum, so it underestimates the spread. Right: the mean, where the two agree.
The maximum is the sharpest case, but it is not the only one. Any statistic whose value is determined by a handful of the most extreme observations — the minimum, the range, the largest order statistic, the 99th percentile of a heavy-tailed distribution — inherits the same defect, because $\hat F$ has no mass beyond the observed extremes. The general diagnosis is that the bootstrap fails when the statistic is non-smooth in the data or when the empirical distribution is a poor approximation in the region that matters. A statistic that is a smooth function of many observations, like a mean, a variance or a regression coefficient under well-behaved noise, is safe; a statistic that hinges on a threshold or a maximum is not.
There is a second, independent failure mode, and it is about the data rather than the statistic. The bootstrap treats the observations as independent draws from a common F. When the data are a time series or a spatial field, nearby observations are correlated, and independent resampling destroys that dependence. The resampled series is smoother than any real series could be, and error bars come out too narrow. The repair is to resample blocks — contiguous stretches long enough to carry the dependence — rather than individual points, which is the block bootstrap; the same logic extends to clustered and hierarchical data through the cluster bootstrap. And the smallest failure of all is the least dramatic: with very few observations, $\hat F$ is a coarse approximation and every bootstrap interval is optimistic, no matter how many resamples you draw.
Where this shows up
Error bars for the statistics nobody has a formula for
Robust fitting and RANSAC
A RANSAC estimate is built from whichever random minimal sample happened to be all-inlier, so its variability is not a textbook standard error. Resampling the correspondences gives that estimate an honest error bar, and the same block-resampling idea repairs the case where correspondences are correlated across a sequence.
Uncertainty on benchmark scores
An evaluation reports accuracy, an F1, a pass rate on a rubric — none of which has a clean standard error. Bootstrapping the held-out examples produces the confidence intervals that evaluation dashboards need, and pairing the resamples across two models gives a test of whether the difference is real.
The estimators it stands beside
The bootstrap is the computational sibling of the sampling-distribution theory in Is this estimator any good?: where that part computes bias and variance from a model, the bootstrap measures them from the data. The two should agree whenever the theory applies, and disagreement is itself diagnostic.
Intervals, without the algebra
A confidence interval is a statement about the sampling distribution, which is exactly what the bootstrap constructs. The percentile and basic intervals here are the resampling counterpoints to the Wald and likelihood intervals of Confidence vs credible, and the BCa correction is what to reach for when the simple ones lean.
Cheat sheet
Every formula in one place
| Idea | Formula | Reading |
|---|---|---|
| Empirical CDF | $\hat F_n(x)=\frac{1}{n}\sum_i \mathbf{1}\{x_i\le x\}$ | Mass 1/n on each observation; the plug-in model for F. |
| Resample | $x_1^*,\dots,x_n^*\sim\hat F_n$ | Draw n values with replacement, uniformly from the data. |
| Bootstrap statistic | $\hat\theta^{*}=T(x_1^*,\dots,x_n^*)$ | Recompute the statistic on each resample; collect the cloud. |
| Bootstrap SE | $\sqrt{\frac{1}{B-1}\sum_b(\hat\theta^{*b}-\bar{\hat\theta}^{*})^2}$ | The spread of the resampled statistics estimates the standard error. |
| Percentile interval | $[\hat\theta^{*}_{(B\alpha/2)},\hat\theta^{*}_{(B(1-\alpha/2))}]$ | Read the endpoints straight off the sorted bootstrap cloud. |
| Basic interval | $[2\hat\theta-\hat\theta^{*}_{(1-\alpha/2)},\,2\hat\theta-\hat\theta^{*}_{(\alpha/2)}]$ | Reflect the bootstrap error distribution around the estimate. |
| BCa | z_0 bias plus jackknife acceleration a | Adjusts the endpoints when the cloud is skewed or the variance moves with the parameter. |
| Monte Carlo noise | shrinks like $1/\sqrt B$ | Finite B adds noise; it never fixes the modelling error of using $\hat F$. |
| Fails for | maxima, minima, extreme order statistics | $\hat F$ has no mass past the observed extremes, so the upper tail is unreachable. |
| Fix for dependence | block or cluster bootstrap | Resample contiguous blocks, not single points, when observations are correlated. |
Further reading
Where to go deeper
- Bradley Efron and Robert Tibshirani, An Introduction to the Bootstrap, 1993 — the standard reference: the plug-in principle, the interval zoo, and the failure cases in one place.
- Bradley Efron and Trevor Hastie, Computer Age Statistical Inference, 2016 — the bootstrap in the context of modern inference, with the accuracy theory behind BCa.
- A. C. Davison and D. V. Hinkley, Bootstrap Methods and Their Application, 1997 — the careful treatment of dependent data and block resampling.
- Persi Diaconis and Bradley Efron, "Computer-Intensive Methods in Statistics", Statistical Science, 2023 — a retrospective on why resampling took hold.
- Larry Wasserman, All of Statistics, chapter 8 — a compact proof sketch that the bootstrap works for smooth statistics, and a crisp statement of when it does not.