Probability in Action, Interactively
The second volume - likelihood and Bayesian inference, estimators and intervals, sampling and Monte Carlo, Markov chains and MCMC, information theory, then the filters and probabilistic models the AI, vision and robotics guides actually run.
Volume I built the rules. This volume does the work. Given data and a model, how do you estimate the parameters — and how do you know whether your estimator is any good? How do you draw samples when the distribution has no closed form, and how do you average over a space too large to enumerate? How much information does an observation carry, and what does a model that assigns probabilities to sentences actually believe? Each answer is an algorithm, and each algorithm has a canvas.
The last act is where the site cashes in: the language-model guides consume cross-entropy and the softmax, multi-view geometry consumes RANSAC and robust losses, and robot navigation consumes the Bayes filter in all three of its guises — histograms, Gaussians and particles. The first volume is the prerequisite; nothing here needs more than it taught. A routing table maps every idea to the page that uses it.
The parts
Drag the parameter and watch the log-likelihood surface move with the fitted curve — and see why a likelihood is not a probability over θ.
Shape a Beta prior by hand, stream in data a point at a time, and watch the posterior concentrate.
Sampling distributions, the bias–variance trade-off on a dartboard, Fisher information, and the Cramér–Rao floor.
A hundred simulated intervals with about ninety-five covering the truth, beside the posterior interval on the same data.
Null against alternative with a draggable effect size: α, β and power, the p-value under H₀, and p-hacking made visible.
Resample with replacement and build a sampling distribution from one dataset — and see where the method breaks.
Inverse-CDF, rejection sampling with a draggable envelope, and importance sampling with its weight histogram.
π by darts with its 1/√N error band, then antithetic variates, control variates and stratification raced on one plot.
Walks in one and two dimensions, √n scaling, exponential gaps, the Brownian limit, and a martingale with optional stopping.
Edit a transition matrix, watch a token hop between states, and see the stationary distribution appear — then rank the web.
Metropolis–Hastings on a 2D posterior, the step size that helps and then hurts, Gibbs' axis-aligned moves, and HMC for contrast.
Fit a Gaussian mixture step by step, with responsibilities as colour and the ELBO as a lower bound you can watch tighten.
Shape a distribution and watch entropy fall, read KL as a codelength excess, and see mode-seeking against mode-covering fits.
A clickable DAG with d-separation highlighted, explaining-away done numerically, and the same model drawn as a factor graph.
Forward–backward fills a trellis cell by cell and Viterbi's backtrace lights up, on a robot lost in a corridor.
A histogram filter on a looped corridor: predict is convolve and blur, update is multiply and sharpen — the loop everything else specialises.
The same recursion with Gaussians: a tunable tracker, then a 2D EKF whose linearisation error you can watch grow.
Monte Carlo localization in a 2D map — sample, weight, resample — with particle depletion shown and then fixed.
Heavy-tailed noise breaking least squares, RANSAC's inlier search, and the Huber and Cauchy loss curves that repair it.
Softmax as a distribution with a temperature dial, cross-entropy as negative log-likelihood, calibration, a GP posterior, VAEs and diffusion.