What a probability is
Before there is a formula there is a question: what is the thing the formula computes? Two answers have survived. The first is frequency — flip the coin again and again and watch the proportion of heads settle. The second is belief — a number you would buy and sell a bet at, a degree of confidence that has to behave consistently or someone can take your money. They look like different subjects, and for a century they were treated as rivals. But they obey exactly the same three rules, and once you see that, the rest of probability is just bookkeeping on top of them. This part builds both pictures, side by side, and extracts the axioms they share.
The question
One number, two jobs
Start with a coin that lands heads with some fixed but unknown tendency. What does it mean to say the probability of heads is one half? One answer points at the future: run the experiment many times, count the fraction of heads, and watch that fraction. If there is a stable value it is converging to, call it the probability. This is the frequentist answer, and it is why probability and long-run averages are so hard to separate.
The other answer points at you. You are about to bet on a single flip, one that will happen only once — no long run exists to invoke. A probability is then the price at which you would be indifferent between taking either side of a fair bet. If you say the chance is one half, you should be equally happy to pay fifty cents for a dollar paid on heads or on tails. This is the subjective answer, sometimes called Bayesian or simply the belief interpretation, and it is the one that lets you assign probabilities to a presidential election, a scientific hypothesis, or a cryptographic key being compromised — events with no repeated trials at all.
The puzzle is that both pictures produce the same arithmetic. Frequencies add when events are disjoint, because counts add. Bets satisfy the same rule, because otherwise a bookmaker could construct a combination of wagers that loses you money no matter what happens — a Dutch book. The two roads lead to the same three axioms, and the axioms are what the rest of this guide uses. Everything else follows by pure algebra: conditional probability, Bayes' rule, expectation, the law of large numbers.
So the plan is to build the frequentist picture first, with a running proportion that wobbles and then settles, because it is the one you can watch. Then we build the belief picture, where a probability becomes an odds ratio, and see that consistency forces the same rules. Finally we write the rules down and check that both pictures satisfy them.
Frequencies that settle
The running proportion and its wobble
Flip a coin n times and let Sn be the number of heads. The relative frequency is the ratio Sn/n. For small n it is wild: the first flip is either 0 or 1, so after one toss the frequency is as far from one half as it can be. As n grows the ratio has more opportunities to average out, and the swings shrink roughly like 1/√n. The curve on the canvas below is exactly this — a random walk of proportions that starts at one extreme and tightens around the true value.
Notice how it behaves. The frequency does not sneak up on the limit and stop; it keeps moving. What shrinks is the size of the excursions, not the motion itself. This distinction matters and will be made precise in the law of large numbers: for any tolerance you name, the fraction of runs that ever stray further than that tolerance from the truth goes to zero. The probability is the value the wobble is centred on.
Drag the true probability and the sample size. Two things are worth watching. First, the curve centres on the dashed line at p, not on one half — a biased coin converges to its own bias. Second, at a fixed n the distance between the curve and the line is a random quantity; run the simulation again with a different seed and it is different, while the spread of those distances is governed by p(1-p)/n. The next section turns those observations into rules.
The heavy line is the running proportion of heads. The dashed line is the true p you set; the shaded band is $\pm 2\sqrt{p(1-p)/n}$.
There is a subtlety the animation hides, and it is the reason "the frequency converges to the probability" is a theorem rather than a definition. A sequence can wander forever without settling, yet the measurable event on which it fails to settle can have probability zero. Convergence here is not pointwise but almost sure, and Part 17 makes the modes of convergence precise. For now, the picture is honest: the band narrows like $1/\sqrt{n}$, and inside it the curve has no memory of where it has been.
The three axioms
Everything else is a consequence
Write $\Omega$ for the sample space, the set of all possible outcomes of an experiment. An event is a subset $A\subseteq\Omega$; it happens when the outcome lands inside it. A probability is a function P that assigns a number to each event. The entire subject rests on three requirements.
From these three lines, by algebra alone, come the familiar rules. The complement has probability P(A^c)=1-P(A), because A and its complement are disjoint and together fill $\Omega$. Probabilities lie in the unit interval, because non-negativity applied to the complement gives $1-P(A)\ge 0$. Monotonicity $A\subseteq B\Rightarrow P(A)\le P(B)$ follows by splitting B into A and $B\setminus A$. Even the inclusion–exclusion rule for two events is just additivity applied to three disjoint pieces.
The frequentist picture satisfies all three almost by construction: ratios of counts are non-negative, the count of everything is n so the normalisation is automatic, and counts of disjoint sets add. The belief picture satisfies them for a different reason, which is the content of the next section: a violation is an arbitrage, and an arbitrage is money left on the table. Same axioms, opposite justifications.
One technical caveat belongs here even though its proof is Part 20. Axiom (iii) only makes sense if we are allowed to add up infinitely many probabilities, and that requires the collection of events to be closed under countable unions and complements — a σ-algebra. On the real line one can manufacture sets so pathological that no consistent probability can be assigned to them at all, so the domain of P is deliberately restricted. For every finite sample space, and for every set you will meet in practice, the full power set works and the caveat can be forgotten.
Probability as a price
Odds, and why inconsistency is expensive
Forget long runs and consider a single event A you must bet on. If you would pay p dollars for a contract that pays one dollar when A occurs, then p is your probability. The bet is fair when the price equals the chance, because then your expected profit is zero: you gain 1-p with probability p and lose p with probability 1-p, and p(1-p)-(1-p)p=0.
Bookmakers prefer odds. If the probability is p, the fair odds against are (1-p)/p to one: stake one dollar to win that many. The ratio p/(1-p) is the odds in favour, and its logarithm is the log-odds or logit, a quantity that runs over the whole real line and turns multiplication of odds into addition. Bayes' rule, in the next act, becomes almost trivial in log-odds: each independent observation shifts your log-odds by a fixed amount. Drag the belief slider and watch the three representations move together.
The marker is your stated probability on a scale from impossible (0) to certain (1). The tick below it is the fair price of a one-dollar bet.
Why should beliefs obey the axioms? Because if they do not, a bookmaker can offer you a set of bets that leaves you poorer whichever outcome occurs — a Dutch book — and no consistent agent would accept one. If your probabilities of two disjoint events do not add to your probability of their union, that gap is a free lunch for someone else. Additivity is therefore not an arbitrary convention imposed on beliefs; it is the price of coherence. The evaluation part of the language-model guide uses exactly this idea in reverse: a model whose stated probabilities are not coherent is scored badly by a proper scoring rule such as log loss.
Where this shows up
The axioms under everything
Sensor models as probabilities
A robot's odometry and range measurements are never exact, so every fusion step treats them as random variables with a probability distribution. Non-negativity and normalisation are what make a belief a distribution over poses rather than a score; the axioms are the sanity conditions every filter enforces on every update.
Calibrated predictions
A classifier that outputs p is claiming a long-run frequency: among all inputs it labels with 0.9, about ninety percent should be correct. That is the frequentist reading applied to a model's outputs, and it is what calibration and evaluation measure. Cross-entropy training pushes the outputs toward coherence with observed frequencies.
RANSAC and hypothesis counts
Robust fitting in multi-view geometry reasons about the chance that a random sample of correspondences is all-inlier. That is a counting calculation built directly on the axioms, and it decides how many iterations are enough.
The calculus of densities
When a continuous variable is pushed through a map, probabilities are preserved but densities rescale by a Jacobian. That story is told as calculus in Calculus of probability; the probability of it — what the density means and when it exists — is Part 10 of this volume.
Cheat sheet
Every formula in one place
| Idea | Formula | Reading |
|---|---|---|
| Sample space | $\Omega$, the set of outcomes | Every possible result of the experiment. |
| Event | $A\subseteq\Omega$ | A yes/no question about the outcome. |
| Axioms | $P(A)\ge0,\;P(\Omega)=1,\;P(\bigcup A_i)=\sum P(A_i)$ for disjoint A_i | Three rules; everything else is algebra. |
| Complement | P(A^c)=1-P(A) | Not-A has the leftover probability. |
| Range | $0\le P(A)\le 1$ | Follows from non-negativity and normalisation. |
| Odds | p/(1-p) | Fair stake-to-win ratio a bookmaker quotes. |
| Log-odds | $\log\frac{p}{1-p}$ | Adds across independent evidence; ranges over $\mathbb{R}$. |
| Relative frequency | S_n/n, spread $\approx\sqrt{p(1-p)/n}$ | The frequentist estimate and its $1/\sqrt n$ error. |
Further reading
Where to go deeper
- Andrey Kolmogorov, Foundations of the Theory of Probability, 1933 — the axiomatisation this part describes, in the original.
- Bruno de Finetti, Theory of Probability, 1970 — the betting argument for the axioms, and the case for subjective probability.
- David MacKay, Information Theory, Inference, and Learning Algorithms, chapter 2 — probability as a consistent degree of belief, written for exactly this audience.
- E. T. Jaynes, Probability Theory: The Logic of Science, chapters 1–2 — the Cox derivation of the rules from consistency requirements.
- Grant Sanderson, "What does probability mean?", 3Blue1Brown — the same two pictures, drawn.