Statistics, Interactively
A step-by-step guide to the statistics the rest of the site assumes - descriptive statistics, estimation, testing, Bayesian inference, the statistics of machine learning, recursive state estimation, and causal inference - in the 3Blue1Brown tradition: geometry and simulation first, formulas second, every idea attached to something you can drag.
Every subject on this site leans on statistics it never teaches. The multi-view geometry guide assumes noise models and RANSAC inlier probabilities; the nonlinear optimization guide assumes that least squares is a maximum-likelihood estimate; the LLM guides talk about sampling, temperature, cross-entropy, calibration and latency percentiles; and every robot that localises itself is running a Bayes filter. This series builds that machinery from pictures: a running sample you can redraw, an estimate you can watch wobble, a null distribution you can reshuffle by hand, and a covariance ellipse you can drag.
Parts build on each other but each one stands alone; the glossary collects every term, the distribution table and the formula cards in one place. Probability is a prerequisite, not part of this series. It assumes you already know sample spaces and conditional probability, Bayes' rule as an identity, random variables and their PMF/PDF/CDF, expectation, variance and covariance, the named distributions (Bernoulli through Beta), joint and conditional distributions, and the statements of the law of large numbers and the central limit theorem. Part 1 recaps only the load-bearing pieces; if the symbols are new, start with the probability course and come back.
The parts
A hidden population, one sample, and an estimate that moves every time you redraw it while the truth stays put.
Histogram, KDE and ECDF of the same data, and how a single outlier pulls the mean away from the median.
Two variables, one ellipse: drag the correlation and the points, and meet the four datasets with identical summaries.
Repeatedly draw a sample, drop its mean into a growing histogram, and watch the spread shrink as the sample grows.
Convolve a draggable population with itself, and read the standard error off a curve instead of a formula.
Race three estimators of the same quantity across thousands of simulated samples, and split the error into the part that shrinks and the part that does not.
Data as ticks on an axis; drag the parameters over a live log-likelihood surface and watch the density slide onto the data.
The curvature of the log-likelihood is the error bar: zoom the peak, read off the width, and compare it to the best any estimator can do.
Resample the data in front of you, build the bootstrap distribution, and get an interval without a formula.
Drag points and watch the fit, its standard error, the residuals and the confidence band move together.
A hundred repeated experiments, a hundred intervals, and a count of how many miss. The interval is random; the parameter is not.
Shuffle the labels, rebuild the null distribution, and shade the tail - a permutation test you can watch, with no formula required.
Two overlapping densities: drag the effect size, the sample size and the threshold, and watch the four outcomes trade against each other.
A p-hacking sandbox: add covariates and subgroups until something is significant, then correct with Bonferroni or Benjamini-Hochberg.
One interactive reference: pick a question, and see which test, which null distribution and which assumption it needs.
Beta-Binomial inference: drag the prior, feed coin flips one at a time, and watch the data overwhelm it.
Two Gaussians merging into a precision-weighted average - deliberately the seed of the Kalman filter.
Credible vs confidence side by side, and a loss function that decides whether the mean, the median or the mode is the number to report.
Live Metropolis-Hastings on a two-dimensional posterior: proposal width, trace plot, acceptance rate, burn-in and R-hat.
Fit a Gaussian to a skewed posterior, and toggle KL(q‖p) against KL(p‖q) to see mode-seeking and mass-covering.
A polynomial-degree slider and an ensemble of fits over resampled data, drawing the bias and variance terms directly.
Ridge is a Gaussian prior and lasso is a Laplace prior: watch the constraint region and the loss contours meet, and see MAP become penalised maximum likelihood.
A k-fold animation, and the optimism of training error made visible as a gap that grows with capacity.
One threshold slider driving the confusion matrix, the ROC point, the PR point and the reliability diagram together.
Ensembles, MC dropout and split conformal prediction: the coverage guarantee demonstrated, then broken by distribution shift.
A draggable covariance, marginals on the margins and a conditioning slice: the multivariate Gaussian made geometric.
A robot in a corridor: a histogram belief that blurs when it moves and multiplies when it sees a door.
Track in 2D with a live covariance ellipse, tune Q and R until it diverges, then linearise a range-bearing model and watch the EKF's error.
Particles over a map, systematic resampling, effective sample size, particle deprivation and the kidnapped-robot problem.
A sequential-testing simulator: peek early and often, and watch the false-positive rate climb far past the level you chose.
An interactive DAG: fork, chain and collider. Condition on a collider and watch a spurious correlation appear from nothing.
Potential outcomes on one simulated dataset: matching, inverse-probability weighting and difference-in-differences, and the estimate moving as you adjust.