Your first eval, in twenty lines
Five labelled inputs, one binary check, one number. That is an evaluation, and it is the smallest piece of real infrastructure in this volume. It is deliberately early — before retrieval, before reasoning, before any agent — because every later part is only measurable against it. A prompt is disposable; a labelled set and the check that grades it survive every model upgrade. This part builds the harness you will use for the rest of the series, and makes the case that evaluation, not prompting, is the durable moat.
An eval is three things
Labelled inputs, a deterministic check, a pass rate
An evaluation is not a leaderboard, a vibes check, or an afternoon of reading outputs. It is three pieces of machinery: a labelled set of inputs where each input carries an expected result, a deterministic check that grades one output as pass or fail, and the pass rate over the set. That is the entire object. Twenty lines of code will build it, and you can build it today.
The reason to build it this early — before retrieval, before agents, before a real product — is that everything downstream becomes measurable the moment it exists. A prompt change, a new model version, a reranker, a compaction policy: each is only better in relation to a number. Without the harness, "better" is an opinion, and opinions do not survive a model upgrade.
Below is a five-case harness. Each row of the heatmap is a labelled input; each column is a prompt version; the colour is the probability that case passes under that version, drawn from a fixed seed. The bars on the right are the realised pass rate. Nothing here calls a model — the probabilities are the mock model, and the seeded coin flip per case is the flakiness you would see on a real one.
Left: per-case pass probability for each prompt version (fixed seed). Right: the realised pass rate of the selected version. Re-run to draw a new seed and watch a flaky case flip.
true/false, the first job is to reshape the task until you can.Five examples is already enough
Imprecise, but not useless
Five cases will not give you a precise pass rate, and it is worth being honest about how imprecise. With five trials, a 95% confidence interval around an 80% pass rate spans roughly 38 to 96 per cent. That number is nearly useless as a measurement and completely useful as an alarm: a gross regression — a prompt that breaks formatting, a model that stops following the output contract — turns five passes into one or two, and that is visible through the noise.
The mistake is to treat precision as a prerequisite. Early on you are not estimating a rate, you are detecting a catastrophe, and a small golden set does that on day one. Precision is what you buy later, by growing the set. The demo below makes the trade concrete: it draws a seeded sample as the eval set grows and reports the 95% interval around the running pass rate.
Running pass rate as the set grows (seeded sample). The band is the 95% Wilson interval — wide at n=5, workable at n=50. Error bars mark both.
Three sets, three jobs
Golden, held-out, production sample
Once the harness exists, the labelled data splits into three roles, and conflating them is how teams fool themselves.
A golden set is small, curated and fixed. You look at it constantly, you tune against it, and its job is to catch regressions — to shout when a case that used to pass stops passing. A held-out set is labelled but not looked at while you iterate; its job is to give an honest estimate, because a prompt tuned against the golden set will eventually memorise it. (This is overfitting a prompt to its examples, and it gets the full treatment in Agents in Action, Part 9.) A production sample is drawn from live traffic, labelled later, and its job is to catch drift — the world moving out from under a set you froze six months ago.
Below is the golden set doing its one job: comparing two prompt versions case by case and pointing at the single case that flipped. That highlighted row is the entire value of a golden set.
Per-case outcome for two versions. A red row is a regression (passed before, fails now); green is a fix; grey is unchanged.
What it costs to know
Calls × tokens × price × how often you run it
An eval is inference, and inference is billed. The cost of one run is cases × calls-per-case × tokens-per-call × price-per-token, and the cost of having an eval is that figure multiplied by how often CI runs it. The arithmetic is friendly: even a five-hundred-case set with one call per case is a few dollars a run at list prices, which is why the honest framing is not "can we afford evals" but "can we afford to ship blind".
Latency matters too, because a suite that takes an hour will not run on every pull request. Below, a small suite (five cases), a working suite (fifty) and a release suite (five hundred) are priced daily against the same token and price assumptions. Move the sliders and watch the release suite become the only one worth optimising.
Estimated daily cost of running each suite at the chosen cadence, using the selected token counts and price list. The release suite is highlighted.
Evaluation is the durable moat
The thesis of both volumes
Prompting is a real skill and it transfers — but a prompt is a fragile artifact. It is tuned to one model at one moment, and the next checkpoint can break it silently. What compounds instead is the labelled set, the harness that grades it, and the habits around them: the golden set that catches regressions, the held-out set that keeps you honest, the production sample that catches drift, and the CI wiring that makes all three run without being asked.
Put the argument at its strongest and it sounds like this: as models improve, most prompt engineering is automated away, but the ability to say what "good" means — as labelled examples and a check — only becomes more valuable, because it is the thing that directs the improvement. No model can be told to "make the product better"; it can be told to raise a pass rate on a set you defined. The set is the specification.
The counter-argument is worth stating too, because it is not weak. An eval is a measurement, and measurements get gamed: an engineer under a pass-rate target will find the shortest path to the number rather than the fix, and a set frozen too early ossifies into a target that no longer matches the product. The answer is not to abandon the harness but to treat it as living machinery — refresh it from production, hold out the honest estimate, and never let one number be the whole picture. The point stands: the moat is not the model, and not the prompt. It is the loop that measures them.
Cheat sheet
| Question | The answer that shapes everything |
|---|---|
| What is an eval? | Labelled inputs + a deterministic check + a pass rate. Nothing else. |
| When do I build it? | Before you need it — before the prompt has variations worth comparing. |
| Why is a binary check better than a score? | A pass/fail you can defend beats a 1–10 you cannot reproduce. |
| How many cases do I start with? | Five. It cannot measure, but it can detect a catastrophe. |
| Golden vs held-out vs production sample? | Regressions · honest estimate · drift. Keep the roles separate. |
| Why is n=5 noisy? | The 95% interval at n=5 spans tens of percentage points. Width falls as 1/√n. |
| What does an eval cost? | cases × calls × tokens × price × cadence. Cheaper than shipping blind. |
| Where is the moat? | In evaluation, not prompting: the labelled set and the harness around it. |
Further reading
- Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?", ICLR 2024 — the arXiv version of the benchmark behind the SWE-bench Verified subset.
- OpenAI, "Introducing SWE-bench Verified", 13 August 2024 — a 500-instance human-validated subset, and a worked example of curating a golden set.
- Yao et al., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024 — the paper that introduced pass^k as the reliability counterpart to pass@k.
- Mialon et al., "GAIA: A Benchmark for General AI Assistants", 2023 — multi-step assistant tasks graded against exact answers, and a lesson in how much a grader can hide.
- Liang et al., "Holistic Evaluation of Language Models", 2022 — HELM's argument for reporting many metrics rather than one number.
- Sculley et al., "Hidden Technical Debt in Machine Learning Systems", NeurIPS 2015 — why the code is a small part of the system, and why the data and the loop dominate.