Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

Probability is the one language every other subject on this site already speaks: the SLAM and RANSAC parts of the vision guides sample and score hypotheses, a pose filter fuses noisy measurements, and the language-model guides are cross-entropy all the way down. This volume builds that language from the ground up. It starts at what a probability actually is and ends at the limit theorems that let a sample stand in for a distribution — the two results every estimator rests on.

The companion volume, Probability in Action, takes these foundations and runs them: maximum likelihood, Bayesian inference, estimation theory, Monte Carlo, Markov chains, MCMC, and the filters and models the rest of the site is built on. The density change-of-variables you may have met in Calculus in Motion is the calculus of this subject; here we do the probability of it, and the two are written to be read side by side. Keep the glossary open for notation.

The parts

Part 1
What a probability is

Relative frequency settling into a limit, a belief you can bet on, and the three axioms both pictures have to obey.

Part 2
Counting, and why it is the hard part

Permutations against combinations, the birthday problem, and the hash collisions that make counting a practical tool.

Part 3
Conditional probability

Conditioning literally crops the sample space and renormalizes it — the unit square squeezed into a corner.

Part 4
Bayes' rule

The medical test, natural frequencies, and the odds form that lets evidence accumulate one log at a time.

Part 5
Independence, and the paradoxes

Conditional independence, explaining-away in a three-node network, and Simpson's paradox in two draggable clouds.

Part 6
Random variables, PMFs and CDFs

One sample, three views: outcomes to PMF bars to a CDF staircase, with the highlighting linked across all three.

Part 7
Expectation

The balance point of a distribution, why linearity survives dependence, LOTUS, and the indicator trick.

Part 8
Spread, moments and concentration

Variance as a moment of inertia, and Markov, Chebyshev and Hoeffding bounds drawn over the true tail.

Part 9
The discrete families

Bernoulli, binomial, geometric, negative binomial and Poisson in one switchable panel, with the Poisson limit animated.

Part 10
From mass to density

Narrow the bins until a histogram becomes a pdf, and meet the trap that a density is not a probability.

Part 11
The continuous families and the Gaussian

Uniform, exponential, gamma, beta and Student-t, and the Gaussian as the distribution entropy and the CLT both select.

Part 12
Transforming a random variable

Push a density through a monotone map, watch it squash by the Jacobian, and sample by inverting the CDF.

Part 13
Joint, marginal, conditional

A joint table with its marginals on the edges: slicing a row is conditioning, and independence is an outer product.

Part 14
Covariance, correlation and Σ

Drag a point cloud and read off the covariance matrix, the correlation, and the 1σ and 2σ ellipses it defines.

Part 15
The multivariate Gaussian

Eigen-decomposition as the ellipse axes, Mahalanobis distance, live conditioning, whitening, and a 3D surface.

Part 16
Sums, convolution and generating functions

Slide one density across another to convolve them, and see why the Gaussian is stable while the Cauchy is not.

Part 17
The Law of Large Numbers

Many running averages at once, weak against strong, and the four modes of convergence as a map of counterexamples.

Part 18
The Central Limit Theorem

Pick a base distribution, slide n, and watch the sample mean become Gaussian — with the Berry–Esseen error readout.

Part 19
Tails, heavy tails, and high dimensions

Where the CLT fails: a Cauchy mean that never settles, Chernoff bounds, and concentration of measure in 100 dimensions.

Part 20
A measure-theoretic aside

Why not every set can have a probability, σ-algebras as the questions you may ask, filtrations, and Borel–Cantelli.

Reference

Start at Part 1 →