Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

Every subject on this site leans on statistics it never teaches. The multi-view geometry guide assumes noise models and RANSAC inlier probabilities; the nonlinear optimization guide assumes that least squares is a maximum-likelihood estimate; the LLM guides talk about sampling, temperature, cross-entropy, calibration and latency percentiles; and every robot that localises itself is running a Bayes filter. This series builds that machinery from pictures: a running sample you can redraw, an estimate you can watch wobble, a null distribution you can reshuffle by hand, and a covariance ellipse you can drag.

Parts build on each other but each one stands alone; the glossary collects every term, the distribution table and the formula cards in one place. Probability is a prerequisite, not part of this series. It assumes you already know sample spaces and conditional probability, Bayes' rule as an identity, random variables and their PMF/PDF/CDF, expectation, variance and covariance, the named distributions (Bernoulli through Beta), joint and conditional distributions, and the statements of the law of large numbers and the central limit theorem. Part 1 recaps only the load-bearing pieces; if the symbols are new, start with the probability course and come back.

The parts

Part 1
What statistics actually asks

A hidden population, one sample, and an estimate that moves every time you redraw it while the truth stays put.

Part 2
Describing one sample

Histogram, KDE and ECDF of the same data, and how a single outlier pulls the mean away from the median.

Part 3
Covariance, correlation, and the ellipse

Two variables, one ellipse: drag the correlation and the points, and meet the four datasets with identical summaries.

Part 4
The sampling distribution

Repeatedly draw a sample, drop its mean into a growing histogram, and watch the spread shrink as the sample grows.

Part 5
The CLT you can feel

Convolve a draggable population with itself, and read the standard error off a curve instead of a formula.

Part 6
Bias, variance, and mean squared error

Race three estimators of the same quantity across thousands of simulated samples, and split the error into the part that shrinks and the part that does not.

Part 7
Maximum likelihood

Data as ticks on an axis; drag the parameters over a live log-likelihood surface and watch the density slide onto the data.

Part 8
Fisher information and the Cramér–Rao bound

The curvature of the log-likelihood is the error bar: zoom the peak, read off the width, and compare it to the best any estimator can do.

Part 9
The bootstrap and the plug-in principle

Resample the data in front of you, build the bootstrap distribution, and get an interval without a formula.

Part 10
Linear regression as estimation

Drag points and watch the fit, its standard error, the residuals and the confidence band move together.

Part 11
What the 95% refers to

A hundred repeated experiments, a hundred intervals, and a count of how many miss. The interval is random; the parameter is not.

Part 12
Null distributions and the p-value

Shuffle the labels, rebuild the null distribution, and shade the tail - a permutation test you can watch, with no formula required.

Part 13
Errors, power, and sample size

Two overlapping densities: drag the effect size, the sample size and the threshold, and watch the four outcomes trade against each other.

Part 14
p-hacking and the garden of forking paths

A p-hacking sandbox: add covariates and subgroups until something is significant, then correct with Bonferroni or Benjamini-Hochberg.

Part 15
t, z, chi-square, F and the likelihood-ratio

One interactive reference: pick a question, and see which test, which null distribution and which assumption it needs.

Part 16
Prior × likelihood → posterior

Beta-Binomial inference: drag the prior, feed coin flips one at a time, and watch the data overwhelm it.

Part 17
Conjugacy and Gaussian fusion

Two Gaussians merging into a precision-weighted average - deliberately the seed of the Kalman filter.

Part 18
Credible intervals and decision theory

Credible vs confidence side by side, and a loss function that decides whether the mean, the median or the mode is the number to report.

Part 19
When conjugacy fails: MCMC

Live Metropolis-Hastings on a two-dimensional posterior: proposal width, trace plot, acceptance rate, burn-in and R-hat.

Part 20
Laplace and variational approximations

Fit a Gaussian to a skewed posterior, and toggle KL(q‖p) against KL(p‖q) to see mode-seeking and mass-covering.

Part 21
The bias–variance decomposition

A polynomial-degree slider and an ensemble of fits over resampled data, drawing the bias and variance terms directly.

Part 22
Regularization is a prior

Ridge is a Gaussian prior and lasso is a Laplace prior: watch the constraint region and the loss contours meet, and see MAP become penalised maximum likelihood.

Part 23
Cross-validation, AIC and BIC

A k-fold animation, and the optimism of training error made visible as a gap that grows with capacity.

Part 24
Evaluation, ROC/PR, and calibration

One threshold slider driving the confusion matrix, the ROC point, the PR point and the reliability diagram together.

Part 25
Predictive uncertainty

Ensembles, MC dropout and split conformal prediction: the coverage guarantee demonstrated, then broken by distribution shift.

Part 26
Gaussians in n dimensions

A draggable covariance, marginals on the margins and a conditioning slice: the multivariate Gaussian made geometric.

Part 27
The Bayes filter

A robot in a corridor: a histogram belief that blurs when it moves and multiplies when it sees a door.

Part 28
The Kalman filter and the EKF

Track in 2D with a live covariance ellipse, tune Q and R until it diverges, then linearise a range-bearing model and watch the EKF's error.

Part 29
Particle filters and MCL

Particles over a map, systematic resampling, effective sample size, particle deprivation and the kidnapped-robot problem.

Part 30
Randomization, A/B tests, and peeking

A sequential-testing simulator: peek early and often, and watch the false-positive rate climb far past the level you chose.

Part 31
Confounding, Simpson's paradox, and DAGs

An interactive DAG: fork, chain and collider. Condition on a collider and watch a spurious correlation appear from nothing.

Part 32
Estimating causal effects

Potential outcomes on one simulated dataset: matching, inverse-probability weighting and difference-in-differences, and the estimate moving as you adjust.

Reference

Start at Part 1 →