Why eval is the hard part, and error analysis
Everyone can write a prompt. Almost nobody can say whether the system got better. That asymmetry is the whole opportunity: evaluation — not prompting — is the hard part, and it is the only moat that survives a model upgrade. The work is unglamorous and mostly manual: read real traces, write one line about what went wrong, group those lines until a small taxonomy falls out, and turn that taxonomy into cases. This part builds the tool you cannot do it without, and shows why the obvious shortcut makes it worse.
Look at your data
You cannot fix what you cannot see
The first move is not to write an eval. It is to open ten real traces and read them end to end. Almost every team that "needs better evals" is actually a team that has never looked at its own failures, and the failure is usually obvious the moment it is on screen: the retriever returned the glossary page, the tool call used the wrong argument name, the agent re-issued the same search four times. None of those is visible in an aggregate accuracy number, and none of them is fixable from the prompt alone.
Which is why the single most valuable piece of infrastructure you can build is a trace viewer: one screen that shows a run's steps, the context composition at each step, and the tool calls in order. This is the page's most important widget, so it is the first one. Click any node on the trajectory to inspect that step; the panel below it is the context that was actually sent.
Top: the step track. Middle: the selected step's context, stacked by segment. Bottom: the tool calls around it.
Open coding
One line per failure, then group
Open coding is the first pass: for each failed trace, write a single line describing what went wrong. Not a category — a description. "It cited the KV-cache document for a question about reranking" is open coding. "Retrieval failure" is not. The discipline of writing one concrete line per trace is what stops you from collapsing everything into "the model was wrong", which is true and useless.
Axial coding is the second pass: read your own lines back and group them. Nine of them will say some version of "the wrong document was retrieved"; seven will say "the tool call used the wrong argument". That is how you discover the taxonomy — from the data, bottom-up, rather than from a list of failure kinds you imagined in advance. Below are thirty seeded failed traces. Click a row to select it, then assign the preset label that fits. Label them in any order; the groups build as you go. Resist the urge to pre-read all thirty: label as you read, the way the work is actually done.
Thirty failed traces. A coloured dot marks a labelled row; the right column is the code you gave it.
Axial coding and the Pareto
A handful of modes cause most of the failures
Grouped, the thirty traces collapse into six modes, and the distribution is badly uneven. That unevenness is the practical gift of error analysis: you do not have to fix everything, you have to fix the top two or three modes and re-measure. This is the Pareto shape, and it shows up in agent traces over and over because failure modes are not independent — a bad tool schema causes many downstream "reasons", and a bad chunking policy causes many "wrong answer" symptoms.
The chart below is the expert coding of the same thirty traces: the bars are mode frequency, the dashed line is the cumulative share. Read it as a work queue, not a report. Three modes account for the majority of the failures, so three fixes — plus three eval cases to prove they worked — buy most of the available improvement.
Bars: failures per mode. Dashed line: cumulative share, read against the right-hand percentage labels.
From taxonomy to eval cases
Each mode becomes a case with a check — code first, judge only where you must
The taxonomy is only worth the reading if it becomes cases. For each failure mode, write one example that reproduces it and one check that decides pass or fail. Prefer a deterministic check — does the output parse, does the citation resolve, did the same call repeat, is the label correct — because a code check is cheap, instant and identical every run. Reach for a model judge only where the criterion is genuinely irreducibly subjective, like whether the retrieved document is relevant.
Below, click a mode to add its eval case. The suite starts with the two easiest modes — structured output and refusal — which is exactly the trap: they are the cheapest to measure, they are barely present in your failures, and a suite built only from them reports high coverage of nothing. Watch the coverage bar climb as you add the modes that actually break.
Click a mode to toggle its eval case. Bar length is the mode's share of the failures; the tag is code or judge.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| What is the hard part of an agent product? | Evaluation, not prompting — and it is the moat a model upgrade does not erase. |
| What is the first move? | Read ten real traces end to end. Do not write an eval first. |
| What is open coding? | One concrete line per failure: "it cited the KV-cache doc for a reranking question". |
| What is axial coding? | Grouping those lines into a small taxonomy of failure modes, bottom-up. |
| Why does the Pareto shape appear? | Failure modes are not independent: one bad schema or chunking policy produces many symptoms. |
| What should you build before agents? | A trace viewer. You cannot fix what you cannot see. |
| What is the right check for a case? | A deterministic one — parse, citation resolves, call did not repeat, label matches. Use a judge only for irreducibly subjective criteria. |
| What does the taxonomy feed? | Eval cases now, and the datasets, judges and CI of the next three parts. |
| What is the trap? | Measuring only what is easy to measure, then reporting a pass rate that covers none of your actual failures. |
Further reading
- Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?", ICLR 2024 — an eval built from real failures, with the check being the repository's own test suite.
- Yao et al., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024 — end-state grading, and why reading the trajectory beats reading the score.
- Cemri et al., "Why Do Multi-Agent LLM Systems Fail?", 2025 — an open-coded taxonomy of coordination failures, built by reading traces.
- Liang et al., "Holistic Evaluation of Language Models (HELM)", 2022 — multi-metric reporting, and the argument against a single headline number.
- Anthropic, "Effective context engineering for AI agents", 29 September 2025 — context as the object under inspection, which is what a context-composition panel shows.
- OpenAI, "Getting started with OpenAI Evals" — a worked example of turning observed failures into a dataset and a grader.