Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

How a multiple-choice benchmark is actually scored

MMLU-style scoring

A base model isn't given "A/B/C/D" as buttons to press — it's a next-token predictor. The standard trick: append each answer choice's text to the question and score the whole continuation's log-likelihood under the model; the highest-scoring choice wins. Two protocol details change results: whether you divide by the answer's token length (length normalization — without it, longer correct answers are unfairly penalized just for having more tokens to get exactly right), and whether you show the choices as text (cloze) or ask the model to output a letter (multiple-choice formulation) — models are often measurably better at one than the other.

Q: The capital of Australia is:

⚠️ Toggle length normalization above and watch the ranked order change — "Canberra" (correct, 3 tokens) loses to "Sydney" (wrong, 1 token) under raw summed log-likelihood, simply because a shorter answer has fewer negative log-probabilities to add up; per-token averaging removes that length bias but isn't automatically "more correct" either — it's a different, also-defensible scoring convention, which is exactly why the same model gets different reported numbers from different eval harnesses.
2

Few-shot prompting as a protocol

In-context learning

Benchmarks are usually run "k-shot": k solved examples are prepended to the prompt before the real question, purely as context — no gradient update happens. More shots generally help (the model infers the expected answer format and task from the examples) but cost more prompt tokens per question and eventually plateau or even hurt if examples are unrepresentative.

3

LLM-as-judge, and its position bias

Judging open-ended output

For tasks without a single correct string (chat quality, helpfulness), it's common to have a strong model judge a pair of responses and pick the better one. A well-documented failure mode: judges are measurably more likely to prefer whichever response is shown first, independent of content — position bias. Rigorous evals run every pairwise comparison in both orders and check whether the verdict flips.

4

Arena ratings: Elo / Bradley-Terry

Reusing Part 8's math

Human-preference arenas (like Chatbot Arena) collect pairwise "which response was better" votes across many models and convert them into a single ranking with the same Bradley-Terry model from Part 8's reward modeling section: P(A beats B) = σ(sA − sB). Feed in a batch of pairwise match results below and watch the fitted scores settle via simple gradient ascent on the same likelihood.

5

pass@k for code

Unbiased estimation

For code, "correct" is checkable by running unit tests, so accuracy is usually reported as pass@k: does at least one of k sampled completions pass? Naively generating exactly k samples and checking is a biased, high-variance estimator; the standard unbiased estimator instead samples n ≥ k completions, counts c correct ones, and computes:

$$\text{pass@}k = 1 - \frac{\binom{n-c}{k}}{\binom{n}{k}}$$
6

Other ways benchmark numbers mislead

Named, not demoed

Contamination — if a benchmark's questions (or close paraphrases) leaked into pretraining data, the model may have memorized answers rather than reasoned to them. N-gram overlap checks (Part 4) against known benchmark text are a standard mitigation.
Error bars and variance — a benchmark score from a few hundred questions has real sampling noise; a 1-point difference between two models is often not statistically significant. Ai2 and others increasingly report confidence intervals, not just a point estimate.
Benchmark saturation — once most frontier models score 90%+, a benchmark stops discriminating between them, and remaining errors may just be label noise in the benchmark itself.
OLMES — Ai2's own standardized evaluation suite, built specifically to fix apples-to-oranges comparisons: fixed prompt formats, fixed few-shot counts, and documented scoring conventions per task, so numbers from different papers are actually comparable.
7

Inside OLMo 3's evaluation suite

Tasks, clusters, and signal

OLMo 3 grew Ai2's base evaluation suite from 11 tasks in OLMo 2 to 43 tasks. That expansion is not just "more benchmarks" — it is a deliberate design with three ideas behind it.

Task clustering. Thirty or forty benchmarks over-weight whatever they all share. To see the structure, the team took per-task correctness from 70 models across ~23,000 benchmark results and ran agglomerative (Ward) clustering. Tasks that models pass and fail together land in the same cluster, revealing the underlying capabilities.

Schematic dendrogram of benchmark tasks (Ward-style). Slide the cut height to choose how coarse the capability clusters are.

Signal-to-noise by averaging a cluster. A single benchmark is a noisy estimate. Averaging the tasks inside a cluster cancels independent noise, so the cluster mean is a more reliable readout than any single task — provided the tasks aren't so correlated that they all move together.

SNR of a cluster average versus the number of tasks, for different inter-task correlations ρ.

Base Easy validates the pipeline; Base Main measures the model. The full 43-task suite is hard and noisy for a base model mid-run. So OLMo 3 also defines a small Base Easy suite (bits-per-byte on easier tasks) that runs quickly and correlates strongly with the harder Base Main average. Developers use it to confirm a training recipe is on track before paying for the full evaluation.

Base Easy (bits-per-byte) versus Base Main average across checkpoints. Lower BPB should mean higher Main accuracy.

💡 Held-out tasks: a handful of benchmarks — MMLU Pro, DeepMind Math, LBPP, and BigBench Hard — are deliberately excluded from model-development decisions and kept as a clean held-out check on the final model.
8

Noise, variance, and contamination

Reading a number honestly

Sampling temperature and top-p are protocol choices. For pass@k on code, OLMo 3 swept temperature and nucleus sampling across five models and settled on temperature 0.6. The number you report is only comparable to another number reported under the same sampling settings.

Match k to how the model will be trained. OLMo 3 reports code pass@k with k = 16, the same group size its GRPO uses. That makes pass@16 an empirical upper bound on what RL on that prompt set could achieve, and it is large enough to capture variance without exploding cost.

Know which benchmarks are noisy. Repeating runs lets you bucket tasks by stability. The very stable tasks (MATH, MMLU, PopQA) are safe to compare within a point; the high-variance tasks (GPQA, IFEval) need wider error bars before you claim a win.

StabilityBenchmarks (std. dev. across runs)
High varianceGPQA (1.48), AlpacaEval 3 (1.24), IFEval (0.88)
ModerateZebraLogic (0.56), Omega (0.56), AIME 24 Avg@32 (0.54), HumanEvalPlus (0.46), AgiEval (0.43), BigBenchHard (0.39)
Very stableLiveCodeBench Avg@10 (0.29), MBPPPlus (0.27), MATH (0.25), MMLU (0.22), PopQA (0.16)

Contamination and overestimation diverge. As Part 4 covered, a highly contaminated benchmark is not automatically an inflated one. The table below is repeated here because it is the single best argument for measuring overestimation directly instead of assuming high overlap means a fake score.

Benchmark% contaminatedPerf. overestimation
DROP (generative)66.87%+13.99
GSM-Symbolic0.00%+5.54
Codex HumanEval @167.07%+3.54
SQuAD83.24%+1.70
GSM8K84.99%+1.61 (decontaminated perf. higher)
Minerva45.17%−1.95
📌 Cross-links: the decontamination method and the full table are in Part 4; the spurious-reward contamination test for RL is in Part 15; and the pass@k estimator on this page is the same one used to measure code RL.
✓

Cheat sheet

Recap

ConceptWhat it isWatch out for
Log-likelihood scoringScore each answer choice as a continuation; highest winsLength normalization can reorder the choices
k-shot promptingPrepend k solved examples as context, with no gradient updateMore shots help, then plateau; costs prompt tokens
LLM-as-judgeA strong model picks the better of two responsesPosition bias: run both orders and check for flips
Elo / Bradley-TerryFit per-model scores from pairwise votes: P(A beats B) = σ(sA − sB)Same model as Part 8’s reward modelling
pass@kP(at least one of k samples passes); unbiased estimator 1 − C(n−c,k)/C(n,k)pass@1 = c/n; report the sampling temperature
Task clustersAverage related tasks for a less noisy readoutCorrelated tasks cancel less noise
Variance & contaminationKnow which benchmarks are noisy; measure overestimation directlyHigh overlap ≠ inflated score
📚

Further reading

References

?

Check your understanding

0/5 answered
With training and evaluation both covered, the last piece is making sure the deployed model behaves — and fails — safely. Continue: guardrails, safety & deployment →