Probability, Interactively
The foundations every other subject assumes and none of them teaches - what a probability means, how to count, condition and update, what a random variable is, and where the law of large numbers and the central limit theorem come from. Pictures first, every idea attached to something you can drag.
Probability is the one language every other subject on this site already speaks: the SLAM and RANSAC parts of the vision guides sample and score hypotheses, a pose filter fuses noisy measurements, and the language-model guides are cross-entropy all the way down. This volume builds that language from the ground up. It starts at what a probability actually is and ends at the limit theorems that let a sample stand in for a distribution — the two results every estimator rests on.
The companion volume, Probability in Action, takes these foundations and runs them: maximum likelihood, Bayesian inference, estimation theory, Monte Carlo, Markov chains, MCMC, and the filters and models the rest of the site is built on. The density change-of-variables you may have met in Calculus in Motion is the calculus of this subject; here we do the probability of it, and the two are written to be read side by side. Keep the glossary open for notation.
The parts
Relative frequency settling into a limit, a belief you can bet on, and the three axioms both pictures have to obey.
Permutations against combinations, the birthday problem, and the hash collisions that make counting a practical tool.
Conditioning literally crops the sample space and renormalizes it — the unit square squeezed into a corner.
The medical test, natural frequencies, and the odds form that lets evidence accumulate one log at a time.
Conditional independence, explaining-away in a three-node network, and Simpson's paradox in two draggable clouds.
One sample, three views: outcomes to PMF bars to a CDF staircase, with the highlighting linked across all three.
The balance point of a distribution, why linearity survives dependence, LOTUS, and the indicator trick.
Variance as a moment of inertia, and Markov, Chebyshev and Hoeffding bounds drawn over the true tail.
Bernoulli, binomial, geometric, negative binomial and Poisson in one switchable panel, with the Poisson limit animated.
Narrow the bins until a histogram becomes a pdf, and meet the trap that a density is not a probability.
Uniform, exponential, gamma, beta and Student-t, and the Gaussian as the distribution entropy and the CLT both select.
Push a density through a monotone map, watch it squash by the Jacobian, and sample by inverting the CDF.
A joint table with its marginals on the edges: slicing a row is conditioning, and independence is an outer product.
Drag a point cloud and read off the covariance matrix, the correlation, and the 1σ and 2σ ellipses it defines.
Eigen-decomposition as the ellipse axes, Mahalanobis distance, live conditioning, whitening, and a 3D surface.
Slide one density across another to convolve them, and see why the Gaussian is stable while the Cauchy is not.
Many running averages at once, weak against strong, and the four modes of convergence as a map of counterexamples.
Pick a base distribution, slide n, and watch the sample mean become Gaussian — with the Berry–Esseen error readout.
Where the CLT fails: a Cauchy mean that never settles, Chernoff bounds, and concentration of measure in 100 dimensions.
Why not every set can have a probability, σ-algebras as the questions you may ask, filtrations, and Borel–Cantelli.