Where each idea is used on this site
This volume is the bridge between the axioms and the algorithms the rest of the site runs. Each row below names one idea, links to the part that builds it, and points at the concrete page whose argument leans on it. The second table collects the algorithms the volume constructs, so a method you meet in a solver or a serving stack resolves to the part that derived it. Nothing here is proved — it is a map.
How to use this card
Filter, then follow the link
The first table routes an idea from the part that derives it to the page elsewhere that consumes it. The second names each algorithm and what it produces. Type in the box to filter both tables case-insensitively; clear it to see everything again. The foundational definitions these parts assume live in Probability, Interactively.
No row matches that — try a looser word, or clear the box.
Where each idea is used
From the part that derives it to the page that consumes it
| Idea | Volume II part | Where it is used |
|---|---|---|
| Likelihood & maximum likelihood | Part 1 · Likelihood and maximum likelihood | Every loss in pretraining is a negative log-likelihood; the scaling arguments in scaling assume the MLE framing. |
| Prior, posterior, conjugacy | Part 2 · Prior, posterior, conjugacy | Posterior reasoning about model behaviour in evaluation, and the KL-controlled updates of RLHF. |
| Bias–variance, Fisher information, Cramér–Rao | Part 3 · Is this estimator any good? | Why a fit is biased or high-variance in Nonlinear Optimization, and the error floor of any estimator there. |
| Confidence vs credible intervals | Part 4 · Confidence vs credible | Reporting uncertainty for a benchmark in evaluation, and landmarks in SLAM. |
| Hypothesis testing, power, p-hacking | Part 5 · Hypothesis testing, and how it goes wrong | Comparing models or ablations honestly in benchmarks and evaluation. |
| The bootstrap | Part 6 · The bootstrap | Error bars on a fitted model without a closed-form variance, e.g. robust fits in RANSAC. |
| Sampling: inverse-CDF, rejection, importance | Part 7 · How to draw a sample | Sampling a token distribution in the decode loop, and the reparameterised draws diffusion and RLHF need. |
| Monte Carlo integration | Part 8 · Monte Carlo integration | Estimating expectations too hard to integrate, from rendering to the policy gradients behind RLHF. |
| Random walks, Brownian motion | Part 9 · Random walks & Brownian motion | Diffusion models as time-reversed noise, and the integrated pose error of odometry. |
| Markov chains & stationarity | Part 10 · Markov chains | The autoregressive state machine of the decode loop, and PageRank-style ranking of research in papers. |
| MCMC: Metropolis, Gibbs, HMC | Part 11 · MCMC | Drawing from an intractable posterior — calibration and uncertainty in evaluation, and post-processing in Nonlinear Optimization. |
| Latent variables: EM & the ELBO | Part 12 · Latent variables: EM and the ELBO | Clustering features before a vision pipeline, and the variational objective behind VAEs and the reparameterised sampling in diffusion. |
| Entropy, cross-entropy, KL, MI | Part 13 · Entropy and information | The training loss of every language model (language models), perplexity in evaluation, and the KL penalty in RLHF. |
| Bayes nets & factor graphs | Part 14 · Bayes nets and factor graphs | The factor graph a pose graph or SLAM backend solves — pose-graph optimization and SLAM. |
| Hidden Markov models | Part 15 · Hidden Markov models | Sequence decoding and mode-finding — the same recursion the Bayesian smoother runs in odometry. |
| The Bayes filter | Part 16 · The Bayes filter | The predict/update recursion every robot runs in odometry and SLAM. |
| Kalman filter & EKF | Part 17 · The Kalman filter and EKF | Sensor fusion and state estimation in odometry, kinematics and SLAM. |
| Particle filters | Part 18 · Particle filters | Monte Carlo localization in a map, the non-Gaussian half of robot navigation and SLAM. |
| Outliers, RANSAC & robust estimation | Part 19 · Outliers, RANSAC and robust estimation | Fitting with gross outliers: RANSAC and the robust losses of gradient descent. |
| Probability in machine learning | Part 20 · Probability in machine learning | Softmax and temperature in the decode loop and speculative decoding, calibration in evaluation, and the next-token distribution of language models. |
The algorithms, collected
One method, one output, one part
| Algorithm | What it produces | Built in |
|---|---|---|
| $\hat\theta = argmax \ell(\theta)$ | The parameter whose model makes the data most likely. | Part 1 |
| $\hat\theta = E[\theta\mid D]$ or $argmax p(\theta\mid D)$ | A posterior summary — mean, median or mode. | Part 2 |
I(θ) = -E[ℓ''(θ)] | The curvature of the log-likelihood, and the variance floor of any unbiased estimator. | Part 3 |
| Resample with replacement, refit, repeat | An empirical sampling distribution from one dataset. | Part 6 |
| Proposals, accept/reject, weights | Samples from a target you can only evaluate, or a lower-variance expectation. | Part 7 |
| $\frac{1}{N}\sum f(x_i)$ | An expectation estimate with a 1/√N error band. | Part 8 |
πP = π | The long-run distribution of a Markov chain. | Part 10 |
| A chain whose stationary distribution is the target | Posterior samples without an analytic form. | Part 11 |
| Responsibilities, then a maximised lower bound | Parameters of a latent-variable model. | Part 12 |
| Marginals over hidden states, or the optimal path | A decoded sequence. | Part 15 |
A belief bel(x) over the state | The recursive state estimate every filter specialises. | Part 16 |
x, P and the gain K | An optimal linear-Gaussian state estimate. | Part 17 |
| A weighted particle set | A non-Gaussian, multi-modal state estimate. | Part 18 |
| A consensus set and a robust fit | Parameters immune to gross outliers. | Part 19 |
| Normalised probabilities, calibrated scores, learned transformations | The probabilistic outputs of modern ML systems. | Part 20 |
Where to start, given where you came from
Three common entry points
From SLAM or odometry, start at the Bayes filter and work forward through the Kalman filter and particle filters; the graphical-model part is the formal backing. From LLM Training, start at information theory: cross-entropy is the loss, KL is the constraint, and perplexity is its exponential. From multi-view geometry, start at robust estimation, which is where the sampling arguments in that guide are actually proved.