Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

Volume I built the rules. This volume does the work. Given data and a model, how do you estimate the parameters — and how do you know whether your estimator is any good? How do you draw samples when the distribution has no closed form, and how do you average over a space too large to enumerate? How much information does an observation carry, and what does a model that assigns probabilities to sentences actually believe? Each answer is an algorithm, and each algorithm has a canvas.

The last act is where the site cashes in: the language-model guides consume cross-entropy and the softmax, multi-view geometry consumes RANSAC and robust losses, and robot navigation consumes the Bayes filter in all three of its guises — histograms, Gaussians and particles. The first volume is the prerequisite; nothing here needs more than it taught. A routing table maps every idea to the page that uses it.

The parts

Part 1
Likelihood and maximum likelihood

Drag the parameter and watch the log-likelihood surface move with the fitted curve — and see why a likelihood is not a probability over θ.

Part 2
Prior, posterior, conjugacy

Shape a Beta prior by hand, stream in data a point at a time, and watch the posterior concentrate.

Part 3
Is this estimator any good?

Sampling distributions, the bias–variance trade-off on a dartboard, Fisher information, and the Cramér–Rao floor.

Part 4
Confidence vs credible

A hundred simulated intervals with about ninety-five covering the truth, beside the posterior interval on the same data.

Part 5
Hypothesis testing, and how it goes wrong

Null against alternative with a draggable effect size: α, β and power, the p-value under H₀, and p-hacking made visible.

Part 6
The bootstrap

Resample with replacement and build a sampling distribution from one dataset — and see where the method breaks.

Part 7
How to draw a sample

Inverse-CDF, rejection sampling with a draggable envelope, and importance sampling with its weight histogram.

Part 8
Monte Carlo integration

π by darts with its 1/√N error band, then antithetic variates, control variates and stratification raced on one plot.

Part 9
Random walks, Poisson processes, Brownian motion

Walks in one and two dimensions, √n scaling, exponential gaps, the Brownian limit, and a martingale with optional stopping.

Part 10
Markov chains

Edit a transition matrix, watch a token hop between states, and see the stationary distribution appear — then rank the web.

Part 11
MCMC

Metropolis–Hastings on a 2D posterior, the step size that helps and then hurts, Gibbs' axis-aligned moves, and HMC for contrast.

Part 12
Latent variables: EM and the ELBO

Fit a Gaussian mixture step by step, with responsibilities as colour and the ELBO as a lower bound you can watch tighten.

Part 13
Entropy, cross-entropy, KL, mutual information

Shape a distribution and watch entropy fall, read KL as a codelength excess, and see mode-seeking against mode-covering fits.

Part 14
Bayes nets and factor graphs

A clickable DAG with d-separation highlighted, explaining-away done numerically, and the same model drawn as a factor graph.

Part 15
Hidden Markov models

Forward–backward fills a trellis cell by cell and Viterbi's backtrace lights up, on a robot lost in a corridor.

Part 16
The Bayes filter

A histogram filter on a looped corridor: predict is convolve and blur, update is multiply and sharpen — the loop everything else specialises.

Part 17
The Kalman filter and EKF

The same recursion with Gaussians: a tunable tracker, then a 2D EKF whose linearisation error you can watch grow.

Part 18
Particle filters

Monte Carlo localization in a 2D map — sample, weight, resample — with particle depletion shown and then fixed.

Part 19
Outliers, RANSAC and robust estimation

Heavy-tailed noise breaking least squares, RANSAC's inlier search, and the Huber and Cauchy loss curves that repair it.

Part 20
Probability in machine learning

Softmax as a distribution with a temperature dial, cross-entropy as negative log-likelihood, calibration, a GP posterior, VAEs and diffusion.

Reference

Start at Part 1 →