Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

What does a score actually promise?

A binary classifier produces a real-valued score for every input, and usually a threshold turns that score into a label. The same score can be thresholded low, so almost everything is called positive, or high, so almost everything is called negative. Each choice produces a different 2×2 tally of outcomes, and therefore different headline numbers. Accuracy, precision, recall, the false-positive rate — all of them are functions of where the threshold sits, not of the model alone.

There are really two questions buried in one score, and they are independent. The first is a question of ranking: does the score put true positives above true negatives? The second is a question of calibration: does a score of 0.8 mean an 80% chance? A model can answer the first question beautifully and the second question terribly. It can order examples correctly while its numbers are nonsense as probabilities, like a thermometer that ranks days from hot to cold perfectly but prints every reading ten degrees too high.

This part separates those two questions with pictures. The ROC curve, the precision-recall curve and AUC are about ranking. The reliability diagram, the Brier score, log-loss and the expected calibration error are about meaning. A single toggle in the demo below performs a rank-preserving transform and leaves the ranking world untouched while the meaning world straightens out.

💡 By the end of this part you'll see why one threshold turns a score into a confusion matrix, an ROC point and a PR point at the same time; why precision and recall trade against each other under class imbalance; why AUC measures ranking rather than correctness; and why a model can score near-perfect AUC yet be badly calibrated.
2

One slider, four views

The signature interaction

The demo below builds a small scored dataset of 420 examples. Each example has a hidden latent value $x$, drawn from a standard normal; the true probability of the positive label is a logistic function of $x$; the label is then drawn from that probability. The score is built from the same latent value but through a deliberately over-confident transform: its logit is the true logit stretched by a factor of $1.8$, plus a little noise. So the score separates the classes well — it ranks them in almost the right order — but it shouts much louder than the evidence warrants. It is a good ranker and a bad probability.

One threshold slider sits on the right. As you drag it, four pictures update together: the histogram of scores with the threshold drawn through it, the confusion matrix that the threshold produces, the operating point it marks on a pre-drawn ROC curve, and the operating point it marks on a pre-drawn precision-recall curve. The reliability diagram re-colours with the threshold as well, shading the bins that the threshold calls positive. A second control, Recalibrate, applies a strictly increasing map fitted to the data and lets you watch the calibration metrics improve while every ranking quantity stays fixed.

Score distribution and threshold

ROC — operating point

Precision–recall — operating point

Reliability diagram (calibration)

The four panels share one threshold. The ROC, PR and reliability curves themselves are fixed; only the marked operating point and the shading move with the slider.

⚠ A threshold is a decision, not a property of the model. Two teams can share one model and report opposite precision and recall because they chose different operating points. Before comparing numbers, ask where the threshold sat — and whether accuracy is even the right summary when one class is rare.
3

The confusion matrix and the threshold

Four counts that say everything

Fix a threshold and every example falls into one of four cells. A true positive is a positive example the classifier called positive. A false positive is a negative example called positive. A false negative is a positive example called negative, and a true negative is a negative example called negative. Those four counts — TP, FP, FN, TN — are the confusion matrix, and every evaluation number on this page is just arithmetic on them.

$$\mathrm{TPR}=\mathrm{recall}=\frac{TP}{TP+FN},\qquad \mathrm{FPR}=\frac{FP}{FP+TN}$$

The true-positive rate is the fraction of real positives that were caught; the false-positive rate is the fraction of real negatives that were wrongly flagged. Both move as the threshold moves, and always in the same direction: lowering the threshold catches more positives (TPR up) but also flags more negatives (FPR up). You cannot improve one without worsening the other, because one slider controls both. That coupling is the whole reason the trade-off needs a curve rather than a single number.

Precision and recall describe the same matrix from the other side:

$$\mathrm{precision}=\frac{TP}{TP+FP},\qquad \mathrm{recall}=\frac{TP}{TP+FN}$$

Recall is identical to TPR — it asks how much of the positive class you found. Precision asks, of everything you flagged, how much was real. The difference is the denominator: recall divides by the true positives, precision divides by the predicted positives. A classifier can have perfect recall by calling everything positive, at which point its precision collapses to the base rate. It can have perfect precision by calling only one, almost certainly correct, example positive, at which point its recall collapses. Neither number alone describes the model; together with the threshold they pin down the operating point.

Accuracy, the fraction of all examples classified correctly, is the number people reach for first and the one that misleads most often. If 95% of examples are negative, a classifier that always says "negative" is 95% accurate and useless. The confusion matrix exposes this immediately, because it shows the positive class with zero TP — which is why every careful evaluation reports the class balance alongside the score.

4

TPR, FPR, ROC and AUC

Every threshold at once

Instead of choosing one threshold, sweep it across every value the score takes. Each threshold lands on a point $(\mathrm{FPR},\mathrm{TPR})$ in the unit square, and the points trace out the receiver operating characteristic, or ROC curve. The curve begins at the origin when the threshold is so high that nothing is positive, and ends at the top-right corner when it is so low that everything is. In between it bows toward the upper-left corner — the region of high recall and low false alarms.

The diagonal from $(0,0)$ to $(1,1)$ is the performance of random guessing, because a coin flip has $\mathrm{TPR}=\mathrm{FPR}$ at every threshold. A useful model lives above the diagonal. The area under the ROC curve, AUC, has a clean probabilistic reading: it is the probability that a randomly chosen positive example receives a higher score than a randomly chosen negative one. AUC is therefore a pure statement about ranking. It does not care what the actual scores are, only what order they fall in, which is why applying any strictly increasing transform to the scores — multiplying by two, taking a sigmoid, squaring — leaves AUC exactly where it was.

That invariance is a strength and a trap. It is a strength because AUC summarises discrimination without committing to a threshold, so it compares models fairly on the one thing they agree on. It is a trap because a single number hides the shape of the curve and the operating point you actually care about. Two models with identical AUC can behave completely differently in the high-recall region that your application lives in. And AUC shares the ROC curve's indifference to class balance: it weighs the minority class only through the count of positive examples, so a model can post a flattering AUC while its positive-class predictions are poor. That is when the precision-recall curve earns its place.

5

Precision, recall, and when PR wins

The curve that notices the rare class

Plot precision against recall as the threshold sweeps and you get the precision-recall curve. It begins at full recall, where everything is flagged, and precision equals the prevalence of the positive class. As the threshold rises, recall falls but precision rises, until the classifier is flagging only its most confident examples. The PR curve does not have a fixed baseline: the floor is not zero but the prevalence, so the curve lives in a smaller box when positives are rare.

This is exactly what makes PR the honest picture under class imbalance. A flooding classifier that flags everything gets recall one and precision equal to prevalence. If positives are one in a thousand, precision is 0.001 — visibly terrible — while on the ROC curve the same operating point sits at $\mathrm{FPR}=1$, a corner the eye tends to dismiss. ROC is symmetric in the two classes; PR is not, because precision's denominator contains the false positives, and when negatives vastly outnumber positives a small FPR is a large number of false positives. For rare positives — disease detection, fraud, a tiny set of unsafe outputs — the PR curve shows the cost that ROC smooths over.

When a single summary is needed, the harmonic mean of precision and recall, the $F_1$ score, is the usual choice:

$$F_1 = 2\cdot\frac{\mathrm{precision}\cdot\mathrm{recall}}{\mathrm{precision}+\mathrm{recall}}$$

The harmonic mean punishes imbalance between the two: it is high only when both are high, and it collapses if either is near zero. It is still threshold-dependent, so it is a summary of one chosen operating point rather than of the model. The general lesson from both curves is the same: there is no threshold-free verdict on a classifier. There is only a ranking, a set of trade-offs, and a choice about which errors you would rather make.

6

Calibration is not discrimination

Does the number mean the number?

Discrimination is ranking; calibration is meaning. A model is calibrated if the examples it scores 0.8 are positive about 80% of the time. To check this, group the predictions into bins of confidence and compare, in each bin, the average predicted probability against the observed frequency of positives. The result is the reliability diagram: confidence along the horizontal axis, empirical positive rate along the vertical. A perfectly calibrated model lies on the diagonal. A point below the diagonal means the model is over-confident — it claimed more probability than reality delivered; above the diagonal means it is under-confident.

The shaded, size-weighted distance between the reliability curve and the diagonal has a name, the expected calibration error, which averages the gap in each bin weighted by how many examples fall there:

$$\mathrm{ECE}=\sum_{b}\frac{n_b}{N}\,\left|\mathrm{acc}_b-\mathrm{conf}_b\right|$$

Calibration needs a score of its own, and the honest ones are the strictly proper scoring rules, which are minimised in expectation only by the true probabilities. The Brier score is the mean squared error of the probabilities against the labels, and log-loss is the average negative log-likelihood of the true labels:

$$\mathrm{Brier}=\frac{1}{N}\sum_i(\hat p_i-y_i)^2,\qquad \mathrm{logloss}=-\frac{1}{N}\sum_i\big[y_i\log\hat p_i+(1-y_i)\log(1-\hat p_i)\big]$$

Log-loss is the more punishing of the two: it grows without bound when the model is confidently wrong, so it notices over-confidence sharply, while Brier stays bounded and gentler. In the demo, toggling Recalibrate applies a strictly increasing map fitted on the logit of the score — a monotone transform, so the order of any two examples is preserved. The ROC curve, the PR curve, the operating points and AUC are therefore untouched to the last digit: the ranking is identical. But the reliability diagram straightens toward the diagonal, and ECE, Brier and log-loss all fall. Same decisions; better probabilities.

This is how a model can rank well and still be badly calibrated, and modern neural networks very often are. Training with cross-entropy on increasingly large models tends to produce over-confident predictions, and a preference model that scores two candidate answers for reinforcement learning from human feedback is a ranker first and a probability second — its scores order answers usefully while its absolute values drift. The RLHF chapter leans on those scores as rewards, which is exactly a setting where calibration matters beyond ordering. The fix is cheap: fit a monotone map on held-out data, as the toggle does, and convert a good ranker into a trustworthy probability without retraining it.

7

Where this shows up

Reading a score defensibly

ML / AI

Benchmarks and scores

A held-out benchmark score is a summary of a ranking, and the evaluation chapter treats it as one: accuracy is threshold-dependent, and a model that tops a leaderboard may simply be ranking well while its confidence values are inflated. Reporting ROC/PR curves and a calibration check turns a single headline number into a defensible claim about the operating region you actually use.

ML / AI

Thresholding a classifier

Guardrail classifiers for toxicity or jailbreak detection are scores thresholded into decisions, and their whole job is the trade-off in this part. The guardrails chapter sets the operating point deliberately: where to sit on the PR curve depends on whether a missed harm or a false alarm costs more. In production, the same questions show up as the monitoring metrics in the serving metrics chapter, where a shifting class balance quietly moves precision even when the model does not change.

The same four views repeat anywhere a score becomes a decision. A medical triage model is a ranking plus a chosen sensitivity, and its calibration decides whether a risk number can be shown to a patient. A spam filter lives at a precision that keeps the inbox clean, accepting the recall it loses. A retrieval system's similarity scores are rankers that must be calibrated before they can be multiplied into a probability. In each case the discipline is identical: look at the ranking with ROC or PR, look at the meaning with a reliability diagram, and compare numbers only at a stated threshold.

Further reading

The references below treat evaluation as three separate questions — where to threshold, how well the model ranks, and whether the numbers are true — rather than one score to maximise. If you take away one thing, take away the split between discrimination and calibration.

Fawcett's tutorial is the standard derivation of the ROC curve and the AUC's probabilistic meaning; Saito and Rehmsmeier make the case for precision-recall under imbalance with examples; Guo and colleagues document how modern neural networks become over-confident and how a single temperature or monotone map repairs it; and Gneiting and Raftery is the definitive account of why Brier and log-loss are the right scoring rules to trust.

Cheat sheet

TermMeaning here
ThresholdThe score cut that turns a ranking into a decision; a policy choice, not a model property
TP / FP / FN / TNThe confusion matrix counts behind every metric on this page
TPR (recall)$TP/(TP+FN)$: the fraction of real positives caught
FPR$FP/(FP+TN)$: the fraction of real negatives flagged
Precision$TP/(TP+FP)$: of those flagged, the fraction that were real
ROC curveTPR versus FPR over every threshold; the diagonal is chance
AUCProbability a random positive outranks a random negative; invariant to monotone transforms
PR curvePrecision versus recall; baseline is prevalence, and it is the honest curve under imbalance
CalibrationWhether a score of 0.8 really means 80%; separate from ranking
Reliability diagramObserved frequency versus predicted confidence, against the diagonal
ECECount-weighted average gap between confidence and accuracy across bins
Brier / log-lossProper scoring rules for probabilities; log-loss punishes confident errors hardest
RecalibrationA strictly increasing map fit on held-out data; fixes meaning, preserves ranking
8

Check your understanding

0/4 answered