LLM-as-judge and its biases
Once you have a task set, something has to grade it, and for open-ended answers a second model is the only affordable grader. That makes the judge a component of your system, with its own failure modes. This part builds the judge deliberately: pointwise versus pairwise, a rubric drawn from the failure taxonomy, the three biases that reliably bend verdicts, and the human-agreement measurement that has to happen before any of it counts for anything.
A judge is a model, and models are biased
Watch a wrong answer win, then fix the judge
An LLM judge is asked a question so simple it seems to need no engineering: which of these two answers is better? That is exactly the kind of short, context-light comparison where systematic preferences do the most damage, because there is no reasoning depth to average them out. Three biases are documented well enough to have names. Position bias — the answer presented first wins more often than it should; swap the order and the verdict flips. Verbosity bias — a longer answer wins more often than it should, whether or not the extra words carry information. Self-preference bias — a judge scores text in its own family's style higher, including text it produced itself.
The playground below pairs a short correct answer (A) with a long, confident, wrong one (B) drawn from the judge's own family. The sliders are the three biases; the verdicts are computed with the same mock judge the rest of this volume uses. Start at the authored defaults: B wins, and B is wrong.
Three verdicts for one pair: as shown, with the order swapped, and order-averaged. The wrong answer is winning because it is long and self-authored, not because it is right.
Pointwise and pairwise are different instruments
Absolute scores drift; relative comparisons do not
Pointwise asks for a score: "rate this answer 1–10 against the rubric." It is the natural shape for regression testing and for tracking a single output over time, and it is the only shape that works when there is no second candidate to compare against. Its weakness is calibration: absolute scores drift with the prompt, the day and the model version, so a threshold that passed last week fails this week even though the answers did not change. Pairwise asks for a choice: "which of these two is better?" It is far more stable, because the comparison is local, and it is the right instrument for ranking two candidate systems on the same inputs. Its weakness is scale: pairwise tells you A beats B, not by how much, and it needs a candidate to compare against.
The rule of thumb: use pairwise to compare two systems or two prompts on a fixed set, use pointwise (with a tight rubric) to detect regressions on a single system, and never mix the two when you report a number. The demo shows why. The pair never changes, but the pointwise scores wander across re-runs and cross the pass threshold; the pairwise verdict does not.
Sixteen re-runs of the same pair. Dashed line = the pass threshold. Pointwise scores drift across it; the pairwise winner is the same every run.
Rubrics come from the failure taxonomy
Grade the failure modes you actually have
A rubric is not a list of virtues. It is a list of the specific ways your answers fail, written as checkable criteria and weighted by the cost of the failure. The source of that list is the error-analysis step: you read real outputs, name the failure modes, count them, and then turn the top ones into rubric lines. If your taxonomy says the dominant failure is answers that are fluent but not grounded in the retrieved context, then a rubric without a groundedness criterion is measuring something else.
Below, the rubric criteria are the five failure modes from the retrieval taxonomy. Each case in the human-labelled set is tagged with the mode it exhibits. Toggle the criteria on and watch agreement with the human labels climb — not because the judge got smarter, but because the rubric started asking the right question.
Left: the rubric, as criteria you can switch on. Right: judge–human agreement on a labelled set, against the 50% chance line.
Inter-annotator agreement comes first
If two humans cannot agree, no judge can be validated
Before you can ask whether a judge agrees with humans, you have to ask whether humans agree with each other. Label a sample twice — two annotators, or one annotator on two days — and measure the agreement. If it is low, the task is underspecified, the labels are ambiguous, or the rubric is unusable; a judge cannot beat the ceiling the humans set, so a low ceiling means the problem is upstream of the judge.
The right statistic is not raw agreement, because raw agreement is inflated by class imbalance: two annotators who both say "pass" almost every time agree often by accident. Cohen's κ corrects for that chance agreement. It runs up to 1 for perfect agreement and is 0 when the annotators are no better than chance; negative values are worse than chance. A widely used reading is that below 0.4 is weak, 0.4–0.6 moderate, 0.6–0.8 substantial and above 0.8 strong — a convention, not a law. Move the disagreement slider and watch κ fall, dragging the judge's achievable ceiling down with it.
Left: the confusion matrix between the two annotators. Right: κ as a bar on a 0–1 scale, with the 0.6 "substantial" marker.
Cheat sheet
| Question | The answer that shapes the harness |
|---|---|
| Pointwise or pairwise? | Pairwise to compare two systems on fixed inputs; pointwise to detect regressions on one system. |
| Why not a bare mean score? | Absolute scores drift with prompt, day and model version; a trend may be calibration, not quality. |
| Where does the rubric come from? | The error-analysis failure taxonomy, weighted by the cost of each failure — not a list of virtues. |
| Position bias | The first answer wins. Mitigation: grade both orders and average. |
| Verbosity bias | The longer answer wins. Mitigation: a length-controlled rubric, or explicit conciseness criteria. |
| Self-preference bias | The judge favours its own family's style. Mitigation: a different judge model, or humans. |
| What must you measure first? | Inter-annotator agreement. The humans' κ is the judge's ceiling. |
| What validates the judge? | A confusion matrix against human labels and Cohen's κ — reported, not assumed. |
| What invalidates it? | Raw agreement alone, a drifting pointwise score, or an uncalibrated judge on a new model version. |
Further reading
- Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", NeurIPS 2023 — the paper that named position, verbosity and self-enhancement bias and measured judge–human agreement.
- Wang et al., "Large Language Models are not Fair Evaluators", 2023 — the order-swap calibration that removes position bias, and evidence of how large the effect is.
- Panickssery et al., "LLM Evaluators Recognize and Favor Their Own Generations", 2024 — self-preference measured directly, and its link to self-recognition.
- Liu et al., "G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment", EMNLP 2023 — rubric-driven grading with chain of thought, and the agreement gains it buys.
- Cohen, "A Coefficient of Agreement for Nominal Scales", Educational and Psychological Measurement, 1960 — the κ statistic the calibration readout uses.
- Es et al., "RAGAS: Automated Evaluation of Retrieval Augmented Generation", EACL 2024 — faithfulness and context metrics as judge-scored rubric lines.