Section 3.6

Bootstrap & Resampling

In Stats 1, the sampling distribution required imagining thousands of repeated studies you never ran. The bootstrap is a shortcut: treat the one sample you do have as a stand-in for the population, and resample from it many times to see how much your statistic varies. It needs no formula and no normality assumption.

The trick: resample with replacement

Take your sample of n values. Draw a new sample of n values from it with replacement, so some original points get picked two or three times, others not at all. Compute your statistic (say, the mean) on that resample. Repeat thousands of times. The spread of those resampled statistics estimates the real sampling variability, directly from your data.

🎮 Build a Bootstrap Distribution

Top: your one sample of 15 values. Each resample picks 15 of them with replacement (taller orange stacks = picked more than once; faded dots = not picked). Its mean drops into the distribution below.

① Your sample (the current resample in orange)

② Bootstrap distribution of the mean, with the 95% percentile interval

Resamples0
Original sample mean—
Bootstrap SE—
95% bootstrap CI—

What you get for free

  • The bootstrap standard error is just the standard deviation of all those resampled means, an estimate of how much your statistic varies from sample to sample.
  • The percentile confidence interval is even simpler: sort the bootstrap statistics and read off the 2.5th and 97.5th percentiles. That range is a 95% confidence interval — no t-tables required.

Every package also offers a refinement. The plain percentile interval is slightly off when the bootstrap distribution is skewed or when the statistic is a biased estimate. It assumes the resampling distribution is centered and shaped like the true sampling distribution. The BCa interval (bias-corrected and accelerated) corrects for both and is the usual default recommendation. In most software it takes only a different option. Percentile intervals are fine for a symmetric statistic like a mean, so the difference rarely shows up in a first course, but use BCa when you bootstrap a median, a ratio, or a correlation.

Why it's powerful: the bootstrap works for statistics that have no tidy formula: medians, correlations, ratios, the difference between two trimmed means, almost anything. When the classic formula doesn't exist or its assumptions don't hold, you can almost always bootstrap instead.

The one assumption that remains

The bootstrap relies on one assumption: that your sample is representative of the population. If the sample is biased, resampling it reproduces the bias, and if it is tiny, resampling cannot add information that isn't there. The ordinary bootstrap also fails for dependent data and for very small, very skewed samples. With a reasonably sized, representative sample, it is very reliable.

Why it matters: the bootstrap is one of the most important ideas in modern statistics. It replaces mathematical derivation with computation, which computers make cheap, and it links the formula-based methods of Stats 1 and 2 to simulation-based data analysis. A related resampling idea, testing a model on data it was not fitted to, is the basis of cross-validation.

Common questions

How many bootstrap resamples do I need?

For standard errors, ~1,000 is plenty; for confidence intervals (which depend on the distribution's tails), 5,000–10,000 is the modern norm, and since computation is cheap there's no reason to skimp. Note what B does and doesn't fix: more resamples reduce simulation noise, but the information ceiling is set by your original n. B = 100,000 can't rescue a sample of 12.

My bootstrap interval and my t-interval disagree. Which one do I trust?

First check how far apart they are. For a mean from a reasonably symmetric sample the two should be very close. A clear disagreement usually means the sampling distribution is skewed. The t-interval assumes it is symmetric and the bootstrap does not, so in that case prefer the bootstrap, and prefer a BCa interval over a plain percentile one. Before concluding anything, rule out two things. One is too few resamples: an interval built on 1,000 replicates still varies in its third digit, so rerun with 10,000 and see whether the disagreement remains. The other is a sample so small that neither method is trustworthy. If the two agree, that tells you the parametric assumption was doing no harm here.

What is the difference between bootstrapping and permutation tests?

They answer different questions. The bootstrap resamples with replacement to estimate uncertainty: standard errors and confidence intervals for an estimate. A permutation test reshuffles group labels to build the null distribution, asking "what differences would chance produce if the labels meant nothing?" — yielding an exact p-value. Estimation → bootstrap; hypothesis testing → permutation.