Evaluation
"OLMo 2 scores 63.7 on MMLU" sounds like a simple fact. It hides several protocol choices — how each answer choice is scored, how many examples the model sees first, who or what is judging open-ended output — each of which can move the number by several points without changing the model at all. This part makes those choices visible.
How a multiple-choice benchmark is actually scored
MMLU-style scoring
A base model isn't given "A/B/C/D" as buttons to press — it's a next-token predictor. The standard trick: append each answer choice's text to the question and score the whole continuation's log-likelihood under the model; the highest-scoring choice wins. Two protocol details change results: whether you divide by the answer's token length (length normalization — without it, longer correct answers are unfairly penalized just for having more tokens to get exactly right), and whether you show the choices as text (cloze) or ask the model to output a letter (multiple-choice formulation) — models are often measurably better at one than the other.
Q: The capital of Australia is:
Few-shot prompting as a protocol
In-context learning
Benchmarks are usually run "k-shot": k solved examples are prepended to the prompt before the real question, purely as context — no gradient update happens. More shots generally help (the model infers the expected answer format and task from the examples) but cost more prompt tokens per question and eventually plateau or even hurt if examples are unrepresentative.
LLM-as-judge, and its position bias
Judging open-ended output
For tasks without a single correct string (chat quality, helpfulness), it's common to have a strong model judge a pair of responses and pick the better one. A well-documented failure mode: judges are measurably more likely to prefer whichever response is shown first, independent of content — position bias. Rigorous evals run every pairwise comparison in both orders and check whether the verdict flips.
Arena ratings: Elo / Bradley-Terry
Reusing Part 8's math
Human-preference arenas (like Chatbot Arena) collect pairwise "which response was better" votes across many models and convert them into a single ranking with the same Bradley-Terry model from Part 8's reward modeling section: P(A beats B) = σ(sA − sB). Feed in a batch of pairwise match results below and watch the fitted scores settle via simple gradient ascent on the same likelihood.
pass@k for code
Unbiased estimation
For code, "correct" is checkable by running unit tests, so accuracy is usually reported as pass@k: does at least one of k sampled completions pass? Naively generating exactly k samples and checking is a biased, high-variance estimator; the standard unbiased estimator instead samples n ≥ k completions, counts c correct ones, and computes:
Other ways benchmark numbers mislead
Named, not demoed
Inside OLMo 3's evaluation suite
Tasks, clusters, and signal
OLMo 3 grew Ai2's base evaluation suite from 11 tasks in OLMo 2 to 43 tasks. That expansion is not just "more benchmarks" — it is a deliberate design with three ideas behind it.
Task clustering. Thirty or forty benchmarks over-weight whatever they all share. To see the structure, the team took per-task correctness from 70 models across ~23,000 benchmark results and ran agglomerative (Ward) clustering. Tasks that models pass and fail together land in the same cluster, revealing the underlying capabilities.
Schematic dendrogram of benchmark tasks (Ward-style). Slide the cut height to choose how coarse the capability clusters are.
Signal-to-noise by averaging a cluster. A single benchmark is a noisy estimate. Averaging the tasks inside a cluster cancels independent noise, so the cluster mean is a more reliable readout than any single task — provided the tasks aren't so correlated that they all move together.
SNR of a cluster average versus the number of tasks, for different inter-task correlations ρ.
Base Easy validates the pipeline; Base Main measures the model. The full 43-task suite is hard and noisy for a base model mid-run. So OLMo 3 also defines a small Base Easy suite (bits-per-byte on easier tasks) that runs quickly and correlates strongly with the harder Base Main average. Developers use it to confirm a training recipe is on track before paying for the full evaluation.
Base Easy (bits-per-byte) versus Base Main average across checkpoints. Lower BPB should mean higher Main accuracy.
Noise, variance, and contamination
Reading a number honestly
Sampling temperature and top-p are protocol choices. For pass@k on code, OLMo 3 swept temperature and nucleus sampling across five models and settled on temperature 0.6. The number you report is only comparable to another number reported under the same sampling settings.
Match k to how the model will be trained. OLMo 3 reports code pass@k with k = 16, the same group size its GRPO uses. That makes pass@16 an empirical upper bound on what RL on that prompt set could achieve, and it is large enough to capture variance without exploding cost.
Know which benchmarks are noisy. Repeating runs lets you bucket tasks by stability. The very stable tasks (MATH, MMLU, PopQA) are safe to compare within a point; the high-variance tasks (GPQA, IFEval) need wider error bars before you claim a win.
| Stability | Benchmarks (std. dev. across runs) |
|---|---|
| High variance | GPQA (1.48), AlpacaEval 3 (1.24), IFEval (0.88) |
| Moderate | ZebraLogic (0.56), Omega (0.56), AIME 24 Avg@32 (0.54), HumanEvalPlus (0.46), AgiEval (0.43), BigBenchHard (0.39) |
| Very stable | LiveCodeBench Avg@10 (0.29), MBPPPlus (0.27), MATH (0.25), MMLU (0.22), PopQA (0.16) |
Contamination and overestimation diverge. As Part 4 covered, a highly contaminated benchmark is not automatically an inflated one. The table below is repeated here because it is the single best argument for measuring overestimation directly instead of assuming high overlap means a fake score.
| Benchmark | % contaminated | Perf. overestimation |
|---|---|---|
| DROP (generative) | 66.87% | +13.99 |
| GSM-Symbolic | 0.00% | +5.54 |
| Codex HumanEval @16 | 7.07% | +3.54 |
| SQuAD | 83.24% | +1.70 |
| GSM8K | 84.99% | +1.61 (decontaminated perf. higher) |
| Minerva | 45.17% | −1.95 |
Cheat sheet
Recap
| Concept | What it is | Watch out for |
|---|---|---|
| Log-likelihood scoring | Score each answer choice as a continuation; highest wins | Length normalization can reorder the choices |
| k-shot prompting | Prepend k solved examples as context, with no gradient update | More shots help, then plateau; costs prompt tokens |
| LLM-as-judge | A strong model picks the better of two responses | Position bias: run both orders and check for flips |
| Elo / Bradley-Terry | Fit per-model scores from pairwise votes: P(A beats B) = σ(sA − sB) | Same model as Part 8’s reward modelling |
| pass@k | P(at least one of k samples passes); unbiased estimator 1 − C(n−c,k)/C(n,k) | pass@1 = c/n; report the sampling temperature |
| Task clusters | Average related tasks for a less noisy readout | Correlated tasks cancel less noise |
| Variance & contamination | Know which benchmarks are noisy; measure overestimation directly | High overlap ≠ inflated score |
Further reading
References
- Hendrycks et al., "Measuring Massive Multitask Language Understanding" (2020) — MMLU.
- Chen et al. (OpenAI), "Evaluating Large Language Models Trained on Code" (2021) — pass@k's unbiased estimator, HumanEval.
- Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (2023) — position bias.
- Gu et al. (Ai2), "OLMES: A Standard for Language Model Evaluations" (2024).