Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Look at your data

You cannot fix what you cannot see

The first move is not to write an eval. It is to open ten real traces and read them end to end. Almost every team that "needs better evals" is actually a team that has never looked at its own failures, and the failure is usually obvious the moment it is on screen: the retriever returned the glossary page, the tool call used the wrong argument name, the agent re-issued the same search four times. None of those is visible in an aggregate accuracy number, and none of them is fixable from the prompt alone.

Which is why the single most valuable piece of infrastructure you can build is a trace viewer: one screen that shows a run's steps, the context composition at each step, and the tool calls in order. This is the page's most important widget, so it is the first one. Click any node on the trajectory to inspect that step; the panel below it is the context that was actually sent.

Top: the step track. Middle: the selected step's context, stacked by segment. Bottom: the tool calls around it.

💡 The durable idea: eval is a reading discipline before it is a measuring one. The metric is downstream of the taxonomy, and the taxonomy comes from a human reading failures. Skip the reading and you will build a precise measurement of the wrong thing.
2

Open coding

One line per failure, then group

Open coding is the first pass: for each failed trace, write a single line describing what went wrong. Not a category — a description. "It cited the KV-cache document for a question about reranking" is open coding. "Retrieval failure" is not. The discipline of writing one concrete line per trace is what stops you from collapsing everything into "the model was wrong", which is true and useless.

Axial coding is the second pass: read your own lines back and group them. Nine of them will say some version of "the wrong document was retrieved"; seven will say "the tool call used the wrong argument". That is how you discover the taxonomy — from the data, bottom-up, rather than from a list of failure kinds you imagined in advance. Below are thirty seeded failed traces. Click a row to select it, then assign the preset label that fits. Label them in any order; the groups build as you go. Resist the urge to pre-read all thirty: label as you read, the way the work is actually done.

Thirty failed traces. A coloured dot marks a labelled row; the right column is the code you gave it.

3

Axial coding and the Pareto

A handful of modes cause most of the failures

Grouped, the thirty traces collapse into six modes, and the distribution is badly uneven. That unevenness is the practical gift of error analysis: you do not have to fix everything, you have to fix the top two or three modes and re-measure. This is the Pareto shape, and it shows up in agent traces over and over because failure modes are not independent — a bad tool schema causes many downstream "reasons", and a bad chunking policy causes many "wrong answer" symptoms.

The chart below is the expert coding of the same thirty traces: the bars are mode frequency, the dashed line is the cumulative share. Read it as a work queue, not a report. Three modes account for the majority of the failures, so three fixes — plus three eval cases to prove they worked — buy most of the available improvement.

Bars: failures per mode. Dashed line: cumulative share, read against the right-hand percentage labels.

4

From taxonomy to eval cases

Each mode becomes a case with a check — code first, judge only where you must

The taxonomy is only worth the reading if it becomes cases. For each failure mode, write one example that reproduces it and one check that decides pass or fail. Prefer a deterministic check — does the output parse, does the citation resolve, did the same call repeat, is the label correct — because a code check is cheap, instant and identical every run. Reach for a model judge only where the criterion is genuinely irreducibly subjective, like whether the retrieved document is relevant.

Below, click a mode to add its eval case. The suite starts with the two easiest modes — structured output and refusal — which is exactly the trap: they are the cheapest to measure, they are barely present in your failures, and a suite built only from them reports high coverage of nothing. Watch the coverage bar climb as you add the modes that actually break.

Click a mode to toggle its eval case. Bar length is the mode's share of the failures; the tag is code or judge.

⚠️ The measurement trap: a suite built from what is easy to check will report a high pass rate because it is testing the easy things. Coverage of the failure volume, not the number of cases, is the honest denominator — and the taxonomy is what tells you which volume you are missing.

Cheat sheet

QuestionThe answer that shapes the build
What is the hard part of an agent product?Evaluation, not prompting — and it is the moat a model upgrade does not erase.
What is the first move?Read ten real traces end to end. Do not write an eval first.
What is open coding?One concrete line per failure: "it cited the KV-cache doc for a reranking question".
What is axial coding?Grouping those lines into a small taxonomy of failure modes, bottom-up.
Why does the Pareto shape appear?Failure modes are not independent: one bad schema or chunking policy produces many symptoms.
What should you build before agents?A trace viewer. You cannot fix what you cannot see.
What is the right check for a case?A deterministic one — parse, citation resolves, call did not repeat, label matches. Use a judge only for irreducibly subjective criteria.
What does the taxonomy feed?Eval cases now, and the datasets, judges and CI of the next three parts.
What is the trap?Measuring only what is easy to measure, then reporting a pass rate that covers none of your actual failures.

Further reading

5

Check your understanding

0/5 answered