Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

An eval is three things

Labelled inputs, a deterministic check, a pass rate

An evaluation is not a leaderboard, a vibes check, or an afternoon of reading outputs. It is three pieces of machinery: a labelled set of inputs where each input carries an expected result, a deterministic check that grades one output as pass or fail, and the pass rate over the set. That is the entire object. Twenty lines of code will build it, and you can build it today.

The reason to build it this early — before retrieval, before agents, before a real product — is that everything downstream becomes measurable the moment it exists. A prompt change, a new model version, a reranker, a compaction policy: each is only better in relation to a number. Without the harness, "better" is an opinion, and opinions do not survive a model upgrade.

💡 The durable idea: a prompt is a hypothesis, and an eval is the only thing that can tell you whether the hypothesis is still true. Write the check before you write the twenty-second prompt variation.

Below is a five-case harness. Each row of the heatmap is a labelled input; each column is a prompt version; the colour is the probability that case passes under that version, drawn from a fixed seed. The bars on the right are the realised pass rate. Nothing here calls a model — the probabilities are the mock model, and the seeded coin flip per case is the flakiness you would see on a real one.

Left: per-case pass probability for each prompt version (fixed seed). Right: the realised pass rate of the selected version. Re-run to draw a new seed and watch a flaky case flip.

⚠️ The trap: a check that needs a human to read the output is not a check — it is a review. If you cannot express "correct" as a function from output to true/false, the first job is to reshape the task until you can.
2

Five examples is already enough

Imprecise, but not useless

Five cases will not give you a precise pass rate, and it is worth being honest about how imprecise. With five trials, a 95% confidence interval around an 80% pass rate spans roughly 38 to 96 per cent. That number is nearly useless as a measurement and completely useful as an alarm: a gross regression — a prompt that breaks formatting, a model that stops following the output contract — turns five passes into one or two, and that is visible through the noise.

The mistake is to treat precision as a prerequisite. Early on you are not estimating a rate, you are detecting a catastrophe, and a small golden set does that on day one. Precision is what you buy later, by growing the set. The demo below makes the trade concrete: it draws a seeded sample as the eval set grows and reports the 95% interval around the running pass rate.

Running pass rate as the set grows (seeded sample). The band is the 95% Wilson interval — wide at n=5, workable at n=50. Error bars mark both.

💡 Rule of thumb: interval width falls as 1/√n. Going from 5 to 20 cases halves it; going from 20 to 80 halves it again. Buy precision in powers of four, and only when the decision needs it.
3

Three sets, three jobs

Golden, held-out, production sample

Once the harness exists, the labelled data splits into three roles, and conflating them is how teams fool themselves.

A golden set is small, curated and fixed. You look at it constantly, you tune against it, and its job is to catch regressions — to shout when a case that used to pass stops passing. A held-out set is labelled but not looked at while you iterate; its job is to give an honest estimate, because a prompt tuned against the golden set will eventually memorise it. (This is overfitting a prompt to its examples, and it gets the full treatment in Agents in Action, Part 9.) A production sample is drawn from live traffic, labelled later, and its job is to catch drift — the world moving out from under a set you froze six months ago.

Below is the golden set doing its one job: comparing two prompt versions case by case and pointing at the single case that flipped. That highlighted row is the entire value of a golden set.

Per-case outcome for two versions. A red row is a regression (passed before, fails now); green is a fix; grey is unchanged.

⚠️ The trap: tuning against the set you also report on. The moment a golden case is edited to make a prompt pass, it is no longer evidence. Freeze the golden set, and keep the tuning edits in the prompt.
4

What it costs to know

Calls × tokens × price × how often you run it

An eval is inference, and inference is billed. The cost of one run is cases × calls-per-case × tokens-per-call × price-per-token, and the cost of having an eval is that figure multiplied by how often CI runs it. The arithmetic is friendly: even a five-hundred-case set with one call per case is a few dollars a run at list prices, which is why the honest framing is not "can we afford evals" but "can we afford to ship blind".

Latency matters too, because a suite that takes an hour will not run on every pull request. Below, a small suite (five cases), a working suite (fifty) and a release suite (five hundred) are priced daily against the same token and price assumptions. Move the sliders and watch the release suite become the only one worth optimising.

Estimated daily cost of running each suite at the chosen cadence, using the selected token counts and price list. The release suite is highlighted.

5

Evaluation is the durable moat

The thesis of both volumes

Prompting is a real skill and it transfers — but a prompt is a fragile artifact. It is tuned to one model at one moment, and the next checkpoint can break it silently. What compounds instead is the labelled set, the harness that grades it, and the habits around them: the golden set that catches regressions, the held-out set that keeps you honest, the production sample that catches drift, and the CI wiring that makes all three run without being asked.

Put the argument at its strongest and it sounds like this: as models improve, most prompt engineering is automated away, but the ability to say what "good" means — as labelled examples and a check — only becomes more valuable, because it is the thing that directs the improvement. No model can be told to "make the product better"; it can be told to raise a pass rate on a set you defined. The set is the specification.

The counter-argument is worth stating too, because it is not weak. An eval is a measurement, and measurements get gamed: an engineer under a pass-rate target will find the shortest path to the number rather than the fix, and a set frozen too early ossifies into a target that no longer matches the product. The answer is not to abandon the harness but to treat it as living machinery — refresh it from production, hold out the honest estimate, and never let one number be the whole picture. The point stands: the moat is not the model, and not the prompt. It is the loop that measures them.

💡 Carry this forward: every remaining part of this volume adds a component — tokens, sampling, cost, instructions, examples, reasoning, retrieval — and each one is only an improvement if it moves a pass rate you built in this part.

Cheat sheet

QuestionThe answer that shapes everything
What is an eval?Labelled inputs + a deterministic check + a pass rate. Nothing else.
When do I build it?Before you need it — before the prompt has variations worth comparing.
Why is a binary check better than a score?A pass/fail you can defend beats a 1–10 you cannot reproduce.
How many cases do I start with?Five. It cannot measure, but it can detect a catastrophe.
Golden vs held-out vs production sample?Regressions · honest estimate · drift. Keep the roles separate.
Why is n=5 noisy?The 95% interval at n=5 spans tens of percentage points. Width falls as 1/√n.
What does an eval cost?cases × calls × tokens × price × cadence. Cheaper than shipping blind.
Where is the moat?In evaluation, not prompting: the labelled set and the harness around it.

Further reading

6

Check your understanding

0/4 answered