Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A judge is a model, and models are biased

Watch a wrong answer win, then fix the judge

An LLM judge is asked a question so simple it seems to need no engineering: which of these two answers is better? That is exactly the kind of short, context-light comparison where systematic preferences do the most damage, because there is no reasoning depth to average them out. Three biases are documented well enough to have names. Position bias — the answer presented first wins more often than it should; swap the order and the verdict flips. Verbosity bias — a longer answer wins more often than it should, whether or not the extra words carry information. Self-preference bias — a judge scores text in its own family's style higher, including text it produced itself.

The playground below pairs a short correct answer (A) with a long, confident, wrong one (B) drawn from the judge's own family. The sliders are the three biases; the verdicts are computed with the same mock judge the rest of this volume uses. Start at the authored defaults: B wins, and B is wrong.

Three verdicts for one pair: as shown, with the order swapped, and order-averaged. The wrong answer is winning because it is long and self-authored, not because it is right.

💡 The fix is mechanical, not moral. Position bias dies when you grade both orders and average. Verbosity bias dies when the rubric constrains length. Self-preference dies when the judge is a different model family, or a human. None of them dies by asking the judge to be fair.
2

Pointwise and pairwise are different instruments

Absolute scores drift; relative comparisons do not

Pointwise asks for a score: "rate this answer 1–10 against the rubric." It is the natural shape for regression testing and for tracking a single output over time, and it is the only shape that works when there is no second candidate to compare against. Its weakness is calibration: absolute scores drift with the prompt, the day and the model version, so a threshold that passed last week fails this week even though the answers did not change. Pairwise asks for a choice: "which of these two is better?" It is far more stable, because the comparison is local, and it is the right instrument for ranking two candidate systems on the same inputs. Its weakness is scale: pairwise tells you A beats B, not by how much, and it needs a candidate to compare against.

The rule of thumb: use pairwise to compare two systems or two prompts on a fixed set, use pointwise (with a tight rubric) to detect regressions on a single system, and never mix the two when you report a number. The demo shows why. The pair never changes, but the pointwise scores wander across re-runs and cross the pass threshold; the pairwise verdict does not.

Sixteen re-runs of the same pair. Dashed line = the pass threshold. Pointwise scores drift across it; the pairwise winner is the same every run.

⚠️ The trap: a pointwise harness that reports "the average score went from 6.4 to 6.1" invites you to ship a fix for noise. Report the pairwise win rate on a fixed set, or report pointwise with a confidence interval and a stated threshold — never a bare mean.
3

Rubrics come from the failure taxonomy

Grade the failure modes you actually have

A rubric is not a list of virtues. It is a list of the specific ways your answers fail, written as checkable criteria and weighted by the cost of the failure. The source of that list is the error-analysis step: you read real outputs, name the failure modes, count them, and then turn the top ones into rubric lines. If your taxonomy says the dominant failure is answers that are fluent but not grounded in the retrieved context, then a rubric without a groundedness criterion is measuring something else.

Below, the rubric criteria are the five failure modes from the retrieval taxonomy. Each case in the human-labelled set is tagged with the mode it exhibits. Toggle the criteria on and watch agreement with the human labels climb — not because the judge got smarter, but because the rubric started asking the right question.

Left: the rubric, as criteria you can switch on. Right: judge–human agreement on a labelled set, against the 50% chance line.

💡 A rubric is a hypothesis about failure. If agreement with humans does not rise when you add the criterion, the criterion is either wrong or the labels disagree — and that second possibility is the next section.
4

Inter-annotator agreement comes first

If two humans cannot agree, no judge can be validated

Before you can ask whether a judge agrees with humans, you have to ask whether humans agree with each other. Label a sample twice — two annotators, or one annotator on two days — and measure the agreement. If it is low, the task is underspecified, the labels are ambiguous, or the rubric is unusable; a judge cannot beat the ceiling the humans set, so a low ceiling means the problem is upstream of the judge.

The right statistic is not raw agreement, because raw agreement is inflated by class imbalance: two annotators who both say "pass" almost every time agree often by accident. Cohen's κ corrects for that chance agreement. It runs up to 1 for perfect agreement and is 0 when the annotators are no better than chance; negative values are worse than chance. A widely used reading is that below 0.4 is weak, 0.4–0.6 moderate, 0.6–0.8 substantial and above 0.8 strong — a convention, not a law. Move the disagreement slider and watch κ fall, dragging the judge's achievable ceiling down with it.

Left: the confusion matrix between the two annotators. Right: κ as a bar on a 0–1 scale, with the 0.6 "substantial" marker.

⚠️ Never trust a judge you have not calibrated. A judge that looks reasonable is not evidence. Report the confusion matrix against human labels, the κ, and the human κ it is bounded by — and re-run the calibration whenever the judge model or the rubric changes.

Cheat sheet

QuestionThe answer that shapes the harness
Pointwise or pairwise?Pairwise to compare two systems on fixed inputs; pointwise to detect regressions on one system.
Why not a bare mean score?Absolute scores drift with prompt, day and model version; a trend may be calibration, not quality.
Where does the rubric come from?The error-analysis failure taxonomy, weighted by the cost of each failure — not a list of virtues.
Position biasThe first answer wins. Mitigation: grade both orders and average.
Verbosity biasThe longer answer wins. Mitigation: a length-controlled rubric, or explicit conciseness criteria.
Self-preference biasThe judge favours its own family's style. Mitigation: a different judge model, or humans.
What must you measure first?Inter-annotator agreement. The humans' κ is the judge's ceiling.
What validates the judge?A confusion matrix against human labels and Cohen's κ — reported, not assumed.
What invalidates it?Raw agreement alone, a drifting pointwise score, or an uncalibrated judge on a new model version.

Further reading

5

Check your understanding

0/5 answered