Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Golden sets, held-out sets and coverage

A small set you trust, a set you must not tune against, and a map of the gaps

A golden set is small, high-signal and human-labelled: a few dozen cases you have personally read and would defend. It is the thing you iterate against, and its value is that you trust every row. A held-out test set is separate and larger, and the one rule is that you never tune against it. You will overfit prompts exactly the way models overfit weights — a phrase that fixes case 7 and breaks case 19 — and the held-out set is the only instrument that tells you the difference between learning the task and memorising your examples.

Neither is enough without a coverage matrix: input dimensions down one axis, cases along the other, and a count in each cell. The empty cells are the answer. A suite can be large, green and blind at the same time — covering short factual lookups exhaustively while never testing a multi-hop question, a typo, a no-answer case, or an adversarial input. Below, the heatmap is that matrix; the red cells are blind spots. Add cases and watch the gaps move rather than disappear.

Rows: input dimensions. Columns: cases. Dark = covered, white = a blind spot, outlined red where the suite has never looked.

💡 The durable idea: a dataset is a statement about coverage, not a pile of rows. Its size is the least interesting thing about it; the dimensions it exercises, and the ones it does not, are the whole point.
2

Overfitting the golden set

Train rises, held-out peaks and falls

Iterating on a prompt against a small golden set feels like progress, because the golden-set number goes up every time. The problem is that it goes up for two different reasons and the metric cannot distinguish them: you are fixing real failures, and you are also fitting the specific phrasing of the cases you happen to have. The first generalises. The second is memorisation, and it shows up as a widening gap between the set you tune on and the set you do not.

The curves below are the shape you should expect. Train climbs monotonically with each iteration. Held-out climbs too — until the prompt starts encoding details of the golden set rather than the task — and then it rolls over and declines. The peak is the honest end of useful iteration; everything to the right of it makes the golden set look better and the product worse.

The dashed held-out curve peaks and falls while the train curve keeps rising. The gap is the overfitting.

⚠️ You overfit prompts too. Tuning against the held-out set converts it into a second golden set and you lose the only clean instrument you had. Freeze it, look at it rarely, and when you do look, accept that a peak on it means the search is over.
3

Code checks before judges

Deterministic where you can, a judge only where you must

Given a case, the question is not "how do I grade this" but "what is the cheapest thing that distinguishes pass from fail?". For a large share of agent failures the answer is a code check: does the output parse against the schema, does every citation resolve to a document that contains the claim, is the required tool argument present, did the same call repeat, is the expected field present with the right value. A code check is instant, free, perfectly repeatable, and it fails loudly. A model judge is none of those things — it costs money per case, its verdict drifts run to run, and it carries the position and verbosity biases you would spend Part 10 fighting.

So the default ordering is code first, judge last. Reach for the judge only where the criterion is genuinely irreducibly subjective — "is this retrieved document relevant to the question" — and even then, prefer to decompose it into checkable sub-questions. The panels below compare the two on the same cases: accuracy, cost, and run-to-run spread.

Three panels, two graders: a deterministic check and a model judge. Note which one is flat across reruns.

4

Regression gating, significance and the CI bill

A prompt change is a deploy

The last step is the gate. A prompt change, a model version bump, a retriever tweak — each is a deploy, and each should have to pass the suite before it ships. But a gate is only useful if it can tell a regression from noise, and that is a sample-size question, not a taste question. A two-point drop measured on twenty cases is indistinguishable from a two-point drop measured on zero cases; the confidence interval around the delta is wider than the delta. Compute the interval, and let the gate block only when the interval excludes zero.

The demo makes the arithmetic visible. The pale band is where a second measurement of the unchanged system would land; the point is the candidate, with its own interval. Move the case count and the band narrows. Move the delta and it takes a real change to escape. The readout also prices the eval run itself, because the honest answer to "why not 10,000 cases" is usually the bill.

The band is noise around the baseline. The candidate is significant only when its interval clears the band entirely.

Cheat sheet

QuestionThe answer that shapes the build
What is a golden set?A small, high-signal, human-labelled set you trust and iterate against.
What is a held-out set for?The instrument you do not tune against, so you can tell learning from memorising.
What does a coverage matrix show?Input dimensions against cases. The empty cells are the finding.
What check should a case use?Code first: parse, citation resolves, argument present, call did not repeat. A judge only for irreducibly subjective criteria.
Why prefer code over a judge?It is instant, free, perfectly repeatable, and has no position or verbosity bias.
What is a prompt change?A deploy. It goes through the suite and the gate like any other change.
When is a delta a regression?When its confidence interval excludes zero. Otherwise it is noise.
Is a 2-point drop on 20 cases real?No. The interval around it is far wider than 2 points.
What limits the case count?Usually the bill for running the suite, not the availability of examples.
What does the gate protect?The held-out number, and therefore the product — from a fix that only improves the golden set.

Further reading

5

Check your understanding

0/5 answered