Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Why a positive test is not a positive diagnosis

Here is a sentence that sounds like a riddle and is not one. A disease affects one person in a thousand. A screening test catches it ninety percent of the time when it is present, and falsely alarms ten percent of the time when it is absent. You test positive. What is the chance you have the disease? Most people, including most doctors asked this question in a famous study, answer something near ninety percent. The correct answer is just under one percent, close to the false-alarm rate rather than the detection rate. The gap between the intuition and the arithmetic is the entire subject of this part.

The mistake is not carelessness. It is a confusion of conditioning direction. The test's quality is a statement about people who already have the disease: P(+|D), the chance of a positive given the disease, which is ninety percent. The question asks the reverse: P(D|+), the chance of the disease given the positive. Those are different numbers because the populations they condition on are different sizes. There are a thousand times more healthy people than sick ones, so a ten-percent false-alarm rate manufactures far more positives than the ninety-percent detection rate finds true cases.

Conditional probability is the tool for this, but the definition of conditioning computes the forward direction naturally. $P(A|B)=P(A\cap B)/P(B)$ tells you how to update a belief about A when you learn B; it does not hand you P(D|+) on its own. What we have instead is P(+|D), $P(+|\neg D)$ and a prior P(D). Bayes' rule is the single algebraic step that rearranges the definition to swap the two roles. It is not a new axiom or a different philosophy of probability — it is $P(A\cap B)=P(A|B)P(B)=P(B|A)P(A)$, written twice and divided.

Once the rule is written down, two forms matter. The first is the probability form, which multiplies prior by likelihood and normalises over the alternatives — the form you would compute by hand. The second is the odds form, which says that the ratio of posterior beliefs equals the ratio of priors times a likelihood ratio. The odds form is what makes evidence accumulate gracefully: repeated independent tests add a constant to log-odds, so beliefs can be updated one small integer at a time instead of re-deriving the whole table. The demos below make both forms tangible.

💡 By the end of this part you'll see why a positive test on a rare disease is usually a false alarm, how to derive P(D|+) from the definition of conditioning, and why the odds form turns a stream of evidence into simple addition in log-odds.
2

Bayes' rule from conditioning

One identity, written twice

Conditional probability is defined as the probability of the intersection renormalised by the condition. For events A and B with P(B)>0,

$$P(A\mid B)=\frac{P(A\cap B)}{P(B)},\qquad\text{and symmetrically}\qquad P(B\mid A)=\frac{P(A\cap B)}{P(A)}.$$

Two expressions, one shared numerator. Solve the second for the intersection, $P(A\cap B)=P(B|A)P(A)$, and substitute it into the first. The result is Bayes' rule in its cleanest form: P(A|B)=P(B|A)P(A)/P(B). There is no subtlety here at all; it is a rearranged definition. Everything that follows is a choice about how to name A and B and how to expand the denominator.

Name A the hypothesis D (disease) and B the evidence + (a positive test). The denominator P(+) is the total probability of a positive result across both worlds: the people who are sick and test positive, plus the people who are healthy and test positive. Splitting the sample space into D and its complement and applying additivity gives the law of total probability, $P(+)=P(+|D)P(D)+P(+|\neg D)P(\neg D)$. Substitute and the full formula appears:

$$P(D\mid +)=\frac{P(+\mid D)\,P(D)}{P(+\mid D)\,P(D)+P(+\mid\neg D)\,P(\neg D)}.$$

Read the four pieces. P(D) is the prior, what you believed before the test. P(+|D) is the likelihood of the evidence under the hypothesis, called the sensitivity or true-positive rate. $P(+|\neg D)$ is the likelihood under the alternative, the false-positive rate, equal to one minus the specificity. The prior appears twice — once in the numerator, once in the denominator — and that double appearance is what keeps the answer in the unit interval and what makes a small prior dominate a large sensitivity.

The pattern generalises in two directions that matter later. If the hypotheses form a partition $H_1,\dots,H_k$, the denominator becomes a sum over all of them, and the rule returns a full posterior distribution rather than a single number. And if new evidence arrives in pieces, you are free to use yesterday's posterior as today's prior; the rule composed with itself is the same as conditioning on everything at once, which is the mathematical statement that beliefs are consistent. That self-composition is what Section 4 exploits when it adds one test at a time.

One warning about the direction of the algebra. Bayes' rule is exact and uncontroversial; it is not a philosophical position about the nature of probability. Even a committed frequentist, who refuses to call P(D) a belief, uses the formula whenever the quantities have relative-frequency meaning — for example in a medical trial where P(D) is the disease prevalence in a well-defined population. The subjectivity, if any, lives in the prior, not in the rule. This part treats the prior as an input you declare, and shows what happens to the answer as you vary it.

3

The medical test and the base-rate fallacy

A thousand people, four colours

The cleanest way to compute P(D|+) without algebra is to imagine a concrete population and count. Take one thousand people. Colour each square by two facts at once: whether they have the disease, and whether the test fired. Prevalence fills the diseased block; sensitivity decides how much of that block turns positive; specificity decides how much of the healthy block stays negative. The four counts — true positives, false negatives, false positives, true negatives — are the entire content of Bayes' rule, laid out spatially.

Watch the denominator. The positives are the squares that turned the positive colour, whether or not the person is actually diseased. In the default setup below, one percent prevalence with ninety-percent sensitivity and ninety-percent specificity gives about ten diseased people, nine of whom test positive, and about nine hundred ninety healthy people, roughly ninety-nine of whom also test positive. So a positive result comes from a pool of about one hundred eight people of whom nine are sick. The posterior is nine over one hundred eight, close to eight percent — not ninety.

Each square is ten people out of a thousand. Move the sliders and watch which colour grows.

true positive false negative false positive true negative

Now raise the prevalence and watch the posterior climb. At ten percent prevalence the same test has a positive predictive value above fifty percent; at forty percent it is nearly ninety. The test never changed. What changed is the ratio of sick to healthy people walking through the door, and that ratio is exactly the prior. This dependence on prevalence is the base-rate fallacy: the tendency to judge P(D|+) as if it were P(+|D), ignoring the size of the base the likelihood is measured against.

The same geometry explains why screening programmes are controversial. A test for a rare condition, even a good one, produces mostly false positives, and those false positives are not free: they trigger follow-up procedures with their own risks and costs. The threshold for calling a screening test worthwhile is a statement about the trade between true positives and false positives, which is a statement about prevalence and about the relative harm of the two errors. Bayes' rule does not make that decision, but it tells you precisely which numbers the decision depends on. It also tells you what would change the verdict: an enriched population with a higher prior, a cheaper confirmatory test, or a better specificity that shrinks the false-positive block.

Notice also what a second, independent test does. A positive result on an unrelated second test is new evidence, and its likelihood ratio multiplies the first. In the probability form you would plug the first posterior back in as the prior and recompute; in the odds form, which is next, you simply add. This is why confirmatory testing works: one positive is weak, two positives on independent tests can be decisive, and the arithmetic of that improvement is additive on the log-odds scale.

4

Odds and log-odds

Evidence shifts a belief by a fixed amount

Probabilities are awkward for multiplication because they are bounded and because the denominator has to be renormalised after every update. Odds do not have that problem. Define the odds of a hypothesis as p/(1-p), the ratio of its probability to the probability of its negation. Then Bayes' rule collapses to a statement about ratios: the posterior odds equal the prior odds times the likelihood ratio,

$$\underbrace{\frac{P(D\mid +)}{P(\neg D\mid +)}}_{\text{posterior odds}}=\underbrace{\frac{P(+\mid D)}{P(+\mid\neg D)}}_{\text{likelihood ratio}}\times\underbrace{\frac{P(D)}{P(\neg D)}}_{\text{prior odds}}.$$

The normalising denominator — the whole messy sum over the partition — has cancelled. That is the appeal of the form: there is no term to compute, only two ratios to multiply. The likelihood ratio is the number that describes the test itself, independent of prevalence. A test with sensitivity ninety and specificity ninety has a likelihood ratio of 0.9/0.1=9, meaning a positive result makes the disease nine times more likely than it was, in odds terms, no matter where you started.

Take logarithms and multiplication becomes addition. With $\text{log-odds}(p)=\log\frac{p}{1-p}$, the odds form reads $\text{log-odds}$ posterior equals $\text{log-odds}$ prior plus $\log\text{LR}$. For independent tests, each contributing the same evidence, the shift is the same number every time:

$$\text{log-odds}(D\mid +_1,\dots,+_k)=\text{log-odds}(D)+\sum_{i=1}^{k}\log\text{LR}_i=\text{log-odds}(D)+k\log\text{LR}.$$

The demo below starts from the same prevalence you set above and adds one positive test per click. The marker on the log-odds axis moves by the same distance each time, because the likelihood ratio is fixed; only the starting point moves when you change the prior. This is the picture to keep: evidence is a translation along the log-odds line, and the prior is the coordinate you translate from. In the probability picture the same journey looks like a curve that flattens as it approaches certainty, which is precisely why the log-odds scale is the friendlier one.

The hollow marker is the prior; the filled marker is the posterior after the tests added so far. Each click shifts it by $\log\text{LR}$.

Two practical consequences fall out of the additive picture. First, independent weak evidence can add up to a strong conclusion, and you can budget for it: if each test contributes log nine, about 2.2 nats, then four independent positives move you roughly 8.8 nats, taking a one-in-a-thousand prior to a confident belief. Second, dependent evidence does not add, and pretending it does is a common error. Two tests that share the same failure mode give $\log\text{LR}$ once, not twice; the odds form only multiplies when the likelihoods are conditionally independent given the hypothesis.

The log-odds scale also explains why logistic regression and neural classifiers work the way they do. A model that outputs a logit is producing exactly this coordinate, and each feature is a term added to it. Training fits the coefficients that translate features into log-odds shifts, and the sigmoid that converts back to a probability is just the inverse of the log-odds map. Bayes' rule, in the odds form, is the ancestor of that whole architecture — the cleanest example of turning evidence into a linear score.

Finally, keep the two scales straight when you communicate. A log-odds shift of two nats sounds modest but multiplies the odds by $e^2\approx 7.4$; a probability that moves from 0.01 to 0.07 looks like a rounding error but is a sevenfold change in belief. The same update wears two faces, and choosing the face that matches the audience — odds for the mechanism, probability for the decision — is most of the art of applying Bayes' rule in practice.

5

Where this shows up

The same update, under many names

Robotics

Belief updates in navigation

A robot's odometry says where it thinks it is; a landmark measurement says where it sees. Fusing the two is a Bayes update, and the log-odds form is why sensor fusion can run as a running sum of evidence rather than a full recomputation each step. When odometry drifts, the prior becomes wide and the measurement carries more weight.

Vision

Priors in SLAM

Loop closure in SLAM is a decision about whether two observations are the same place. The likelihood ratio of a match against a false match, combined with a prior over revisits, decides acceptance. Get the base rate wrong and the system either attaches to phantom loops or misses real ones.

AI / ML

Calibration and log loss

A language model's probability for a token is a posterior over continuations, and training pushes it toward the observed frequencies. Evaluation measures whether those posteriors are calibrated, and the log-odds scale is where the signal lives: log loss is average negative log-likelihood, the same logarithm that turns evidence into addition here.

Math

Posteriors in continuous models

Replace the two hypotheses by a continuum and Bayes' rule becomes a statement about densities: prior density times likelihood, normalised. The calculus of that renormalisation, including the Jacobian when variables are transformed, is worked out in Calculus of probability.

6

Cheat sheet

Every formula in one place

IdeaFormulaReading
Conditioning$P(A|B)=P(A\cap B)/P(B)$Renormalise the intersection by the condition.
Bayes' ruleP(A|B)=P(B|A)P(A)/P(B)Swap which event is conditioned on.
Total probability$P(+)=P(+|D)P(D)+P(+|\neg D)P(\neg D)$Positive results come from both worlds.
Medical form$P(D|+)=\dfrac{P(+|D)P(D)}{P(+|D)P(D)+P(+|\neg D)P(\neg D)}$The posterior predictive value.
SensitivityP(+|D)True-positive rate; high is good.
Specificity$P(\neg|\neg D)=1-P(+|\neg D)$True-negative rate; high suppresses false alarms.
Odds$\text{odds}=\dfrac{p}{1-p},\quad p=\dfrac{\text{odds}}{1+\text{odds}}$Ratio of belief to disbelief.
Odds form$\text{posterior odds}=\text{LR}\times\text{prior odds}$The denominator cancels; only ratios remain.
Likelihood ratio$\text{LR}=\dfrac{P(+|D)}{P(+|\neg D)}$How much a positive result favours the hypothesis.
Log-odds$\text{log-odds}(p)=\log\dfrac{p}{1-p}$Unbounded; multiplication becomes addition.
Accumulation$\text{log-odds}\mathrel{+}=\log\text{LR}$ per independent testAdd a fixed shift for each piece of independent evidence.
7

Further reading

Where to go deeper

8

Check your understanding

0/6 answered