Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

How to use this card

Filter, then follow the link

The first table routes an idea from the part that derives it to the page elsewhere that consumes it. The second names each algorithm and what it produces. Type in the box to filter both tables case-insensitively; clear it to see everything again. The foundational definitions these parts assume live in Probability, Interactively.

No row matches that — try a looser word, or clear the box.

💡 An appendix, not another part. Where the first volume collected notation, this one collects destinations: every row answers “I met this method — where is it actually run?”
2

Where each idea is used

From the part that derives it to the page that consumes it

IdeaVolume II partWhere it is used
Likelihood & maximum likelihood Part 1 · Likelihood and maximum likelihood Every loss in pretraining is a negative log-likelihood; the scaling arguments in scaling assume the MLE framing.
Prior, posterior, conjugacy Part 2 · Prior, posterior, conjugacy Posterior reasoning about model behaviour in evaluation, and the KL-controlled updates of RLHF.
Bias–variance, Fisher information, Cramér–Rao Part 3 · Is this estimator any good? Why a fit is biased or high-variance in Nonlinear Optimization, and the error floor of any estimator there.
Confidence vs credible intervals Part 4 · Confidence vs credible Reporting uncertainty for a benchmark in evaluation, and landmarks in SLAM.
Hypothesis testing, power, p-hacking Part 5 · Hypothesis testing, and how it goes wrong Comparing models or ablations honestly in benchmarks and evaluation.
The bootstrap Part 6 · The bootstrap Error bars on a fitted model without a closed-form variance, e.g. robust fits in RANSAC.
Sampling: inverse-CDF, rejection, importance Part 7 · How to draw a sample Sampling a token distribution in the decode loop, and the reparameterised draws diffusion and RLHF need.
Monte Carlo integration Part 8 · Monte Carlo integration Estimating expectations too hard to integrate, from rendering to the policy gradients behind RLHF.
Random walks, Brownian motion Part 9 · Random walks & Brownian motion Diffusion models as time-reversed noise, and the integrated pose error of odometry.
Markov chains & stationarity Part 10 · Markov chains The autoregressive state machine of the decode loop, and PageRank-style ranking of research in papers.
MCMC: Metropolis, Gibbs, HMC Part 11 · MCMC Drawing from an intractable posterior — calibration and uncertainty in evaluation, and post-processing in Nonlinear Optimization.
Latent variables: EM & the ELBO Part 12 · Latent variables: EM and the ELBO Clustering features before a vision pipeline, and the variational objective behind VAEs and the reparameterised sampling in diffusion.
Entropy, cross-entropy, KL, MI Part 13 · Entropy and information The training loss of every language model (language models), perplexity in evaluation, and the KL penalty in RLHF.
Bayes nets & factor graphs Part 14 · Bayes nets and factor graphs The factor graph a pose graph or SLAM backend solves — pose-graph optimization and SLAM.
Hidden Markov models Part 15 · Hidden Markov models Sequence decoding and mode-finding — the same recursion the Bayesian smoother runs in odometry.
The Bayes filter Part 16 · The Bayes filter The predict/update recursion every robot runs in odometry and SLAM.
Kalman filter & EKF Part 17 · The Kalman filter and EKF Sensor fusion and state estimation in odometry, kinematics and SLAM.
Particle filters Part 18 · Particle filters Monte Carlo localization in a map, the non-Gaussian half of robot navigation and SLAM.
Outliers, RANSAC & robust estimation Part 19 · Outliers, RANSAC and robust estimation Fitting with gross outliers: RANSAC and the robust losses of gradient descent.
Probability in machine learning Part 20 · Probability in machine learning Softmax and temperature in the decode loop and speculative decoding, calibration in evaluation, and the next-token distribution of language models.
3

The algorithms, collected

One method, one output, one part

AlgorithmWhat it producesBuilt in
$\hat\theta = argmax \ell(\theta)$The parameter whose model makes the data most likely.Part 1
$\hat\theta = E[\theta\mid D]$ or $argmax p(\theta\mid D)$A posterior summary — mean, median or mode.Part 2
I(θ) = -E[ℓ''(θ)]The curvature of the log-likelihood, and the variance floor of any unbiased estimator.Part 3
Resample with replacement, refit, repeatAn empirical sampling distribution from one dataset.Part 6
Proposals, accept/reject, weightsSamples from a target you can only evaluate, or a lower-variance expectation.Part 7
$\frac{1}{N}\sum f(x_i)$An expectation estimate with a 1/√N error band.Part 8
πP = πThe long-run distribution of a Markov chain.Part 10
A chain whose stationary distribution is the targetPosterior samples without an analytic form.Part 11
Responsibilities, then a maximised lower boundParameters of a latent-variable model.Part 12
Marginals over hidden states, or the optimal pathA decoded sequence.Part 15
A belief bel(x) over the stateThe recursive state estimate every filter specialises.Part 16
x, P and the gain KAn optimal linear-Gaussian state estimate.Part 17
A weighted particle setA non-Gaussian, multi-modal state estimate.Part 18
A consensus set and a robust fitParameters immune to gross outliers.Part 19
Normalised probabilities, calibrated scores, learned transformationsThe probabilistic outputs of modern ML systems.Part 20
4

Where to start, given where you came from

Three common entry points

From SLAM or odometry, start at the Bayes filter and work forward through the Kalman filter and particle filters; the graphical-model part is the formal backing. From LLM Training, start at information theory: cross-entropy is the loss, KL is the constraint, and perplexity is its exponential. From multi-view geometry, start at robust estimation, which is where the sampling arguments in that guide are actually proved.