Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

One number, two jobs

Start with a coin that lands heads with some fixed but unknown tendency. What does it mean to say the probability of heads is one half? One answer points at the future: run the experiment many times, count the fraction of heads, and watch that fraction. If there is a stable value it is converging to, call it the probability. This is the frequentist answer, and it is why probability and long-run averages are so hard to separate.

The other answer points at you. You are about to bet on a single flip, one that will happen only once — no long run exists to invoke. A probability is then the price at which you would be indifferent between taking either side of a fair bet. If you say the chance is one half, you should be equally happy to pay fifty cents for a dollar paid on heads or on tails. This is the subjective answer, sometimes called Bayesian or simply the belief interpretation, and it is the one that lets you assign probabilities to a presidential election, a scientific hypothesis, or a cryptographic key being compromised — events with no repeated trials at all.

The puzzle is that both pictures produce the same arithmetic. Frequencies add when events are disjoint, because counts add. Bets satisfy the same rule, because otherwise a bookmaker could construct a combination of wagers that loses you money no matter what happens — a Dutch book. The two roads lead to the same three axioms, and the axioms are what the rest of this guide uses. Everything else follows by pure algebra: conditional probability, Bayes' rule, expectation, the law of large numbers.

So the plan is to build the frequentist picture first, with a running proportion that wobbles and then settles, because it is the one you can watch. Then we build the belief picture, where a probability becomes an odds ratio, and see that consistency forces the same rules. Finally we write the rules down and check that both pictures satisfy them.

💡 By the end of this part you'll see why a probability can be read as a limiting frequency or as a consistent betting price, why both readings obey the same three axioms, and how the whole edifice of the subject is just those axioms applied over and over.
2

Frequencies that settle

The running proportion and its wobble

Flip a coin n times and let Sn be the number of heads. The relative frequency is the ratio Sn/n. For small n it is wild: the first flip is either 0 or 1, so after one toss the frequency is as far from one half as it can be. As n grows the ratio has more opportunities to average out, and the swings shrink roughly like 1/√n. The curve on the canvas below is exactly this — a random walk of proportions that starts at one extreme and tightens around the true value.

Notice how it behaves. The frequency does not sneak up on the limit and stop; it keeps moving. What shrinks is the size of the excursions, not the motion itself. This distinction matters and will be made precise in the law of large numbers: for any tolerance you name, the fraction of runs that ever stray further than that tolerance from the truth goes to zero. The probability is the value the wobble is centred on.

$$\text{relative frequency after }n\text{ trials}=\frac{S_n}{n}=\frac{X_1+X_2+\cdots+X_n}{n},\qquad X_i\in\{0,1\},\qquad \mathbb{E}\!\left[\frac{S_n}{n}\right]=p.$$

Drag the true probability and the sample size. Two things are worth watching. First, the curve centres on the dashed line at p, not on one half — a biased coin converges to its own bias. Second, at a fixed n the distance between the curve and the line is a random quantity; run the simulation again with a different seed and it is different, while the spread of those distances is governed by p(1-p)/n. The next section turns those observations into rules.

The heavy line is the running proportion of heads. The dashed line is the true p you set; the shaded band is $\pm 2\sqrt{p(1-p)/n}$.

There is a subtlety the animation hides, and it is the reason "the frequency converges to the probability" is a theorem rather than a definition. A sequence can wander forever without settling, yet the measurable event on which it fails to settle can have probability zero. Convergence here is not pointwise but almost sure, and Part 17 makes the modes of convergence precise. For now, the picture is honest: the band narrows like $1/\sqrt{n}$, and inside it the curve has no memory of where it has been.

3

The three axioms

Everything else is a consequence

Write $\Omega$ for the sample space, the set of all possible outcomes of an experiment. An event is a subset $A\subseteq\Omega$; it happens when the outcome lands inside it. A probability is a function P that assigns a number to each event. The entire subject rests on three requirements.

(i)Non-negativity. $P(A)\ge 0$ for every event A. A probability is never negative.
(ii)Normalisation. $P(\Omega)=1$. Something happens; the whole sample space has probability one.
(iii)Countable additivity. If $A_1,A_2,\dots$ are pairwise disjoint, then $P(\bigcup_i A_i)=\sum_i P(A_i)$. Disjoint events add.
$$P(A)\ge 0,\qquad P(\Omega)=1,\qquad A_i\cap A_j=\varnothing \;(i\ne j)\;\Longrightarrow\; P\!\left(\bigcup_{i} A_i\right)=\sum_i P(A_i).$$

From these three lines, by algebra alone, come the familiar rules. The complement has probability P(A^c)=1-P(A), because A and its complement are disjoint and together fill $\Omega$. Probabilities lie in the unit interval, because non-negativity applied to the complement gives $1-P(A)\ge 0$. Monotonicity $A\subseteq B\Rightarrow P(A)\le P(B)$ follows by splitting B into A and $B\setminus A$. Even the inclusion–exclusion rule for two events is just additivity applied to three disjoint pieces.

The frequentist picture satisfies all three almost by construction: ratios of counts are non-negative, the count of everything is n so the normalisation is automatic, and counts of disjoint sets add. The belief picture satisfies them for a different reason, which is the content of the next section: a violation is an arbitrage, and an arbitrage is money left on the table. Same axioms, opposite justifications.

One technical caveat belongs here even though its proof is Part 20. Axiom (iii) only makes sense if we are allowed to add up infinitely many probabilities, and that requires the collection of events to be closed under countable unions and complements — a σ-algebra. On the real line one can manufacture sets so pathological that no consistent probability can be assigned to them at all, so the domain of P is deliberately restricted. For every finite sample space, and for every set you will meet in practice, the full power set works and the caveat can be forgotten.

4

Probability as a price

Odds, and why inconsistency is expensive

Forget long runs and consider a single event A you must bet on. If you would pay p dollars for a contract that pays one dollar when A occurs, then p is your probability. The bet is fair when the price equals the chance, because then your expected profit is zero: you gain 1-p with probability p and lose p with probability 1-p, and p(1-p)-(1-p)p=0.

Bookmakers prefer odds. If the probability is p, the fair odds against are (1-p)/p to one: stake one dollar to win that many. The ratio p/(1-p) is the odds in favour, and its logarithm is the log-odds or logit, a quantity that runs over the whole real line and turns multiplication of odds into addition. Bayes' rule, in the next act, becomes almost trivial in log-odds: each independent observation shifts your log-odds by a fixed amount. Drag the belief slider and watch the three representations move together.

$$\text{odds}(A)=\frac{p}{1-p},\qquad \text{log-odds}(A)=\log\frac{p}{1-p},\qquad p=\frac{\text{odds}}{1+\text{odds}}.$$

The marker is your stated probability on a scale from impossible (0) to certain (1). The tick below it is the fair price of a one-dollar bet.

Why should beliefs obey the axioms? Because if they do not, a bookmaker can offer you a set of bets that leaves you poorer whichever outcome occurs — a Dutch book — and no consistent agent would accept one. If your probabilities of two disjoint events do not add to your probability of their union, that gap is a free lunch for someone else. Additivity is therefore not an arbitrary convention imposed on beliefs; it is the price of coherence. The evaluation part of the language-model guide uses exactly this idea in reverse: a model whose stated probabilities are not coherent is scored badly by a proper scoring rule such as log loss.

5

Where this shows up

The axioms under everything

Robotics

Sensor models as probabilities

A robot's odometry and range measurements are never exact, so every fusion step treats them as random variables with a probability distribution. Non-negativity and normalisation are what make a belief a distribution over poses rather than a score; the axioms are the sanity conditions every filter enforces on every update.

AI / ML

Calibrated predictions

A classifier that outputs p is claiming a long-run frequency: among all inputs it labels with 0.9, about ninety percent should be correct. That is the frequentist reading applied to a model's outputs, and it is what calibration and evaluation measure. Cross-entropy training pushes the outputs toward coherence with observed frequencies.

Vision

RANSAC and hypothesis counts

Robust fitting in multi-view geometry reasons about the chance that a random sample of correspondences is all-inlier. That is a counting calculation built directly on the axioms, and it decides how many iterations are enough.

Math

The calculus of densities

When a continuous variable is pushed through a map, probabilities are preserved but densities rescale by a Jacobian. That story is told as calculus in Calculus of probability; the probability of it — what the density means and when it exists — is Part 10 of this volume.

6

Cheat sheet

Every formula in one place

IdeaFormulaReading
Sample space$\Omega$, the set of outcomesEvery possible result of the experiment.
Event$A\subseteq\Omega$A yes/no question about the outcome.
Axioms$P(A)\ge0,\;P(\Omega)=1,\;P(\bigcup A_i)=\sum P(A_i)$ for disjoint A_iThree rules; everything else is algebra.
ComplementP(A^c)=1-P(A)Not-A has the leftover probability.
Range$0\le P(A)\le 1$Follows from non-negativity and normalisation.
Oddsp/(1-p)Fair stake-to-win ratio a bookmaker quotes.
Log-odds$\log\frac{p}{1-p}$Adds across independent evidence; ranges over $\mathbb{R}$.
Relative frequencyS_n/n, spread $\approx\sqrt{p(1-p)/n}$The frequentist estimate and its $1/\sqrt n$ error.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered