Evaluating RAG and agents
A single end-to-end score tells you that something is wrong and nothing about what. This part breaks the pipeline into stages so a number localises a failure, then takes on the part that makes agents genuinely hard to grade: two runs can reach the same correct end state while one of them loops, overspends and violates a constraint on the way. We separate those questions, add reliability and cost to the scoreboard, and finish on why a public benchmark number is a weak predictor of your task.
Grade the pipeline, not just the answer
A number that points at a stage
A retrieval-augmented pipeline has stages, and each stage fails in its own way. Retrieval can miss the evidence the answer needed — measured by context recall. Reranking can keep the wrong chunk in a wide candidate set — measured by context precision. Generation can produce an answer that is fluent but not grounded in the context it was given — measured by faithfulness. And a fourth, quieter failure sits between them: the right evidence was retrieved and then not used, because it landed in the middle of a long context. Only end-to-end task accuracy catches that one.
Measure every stage, and a low end-to-end score stops being a verdict and becomes a diagnosis. The demo below starts with a weak pipeline and lets you switch on the fixes. Watch the end-to-end number stay stubbornly low until you fix the stage the label names as the bottleneck — fixing the others does almost nothing.
Three stages with their own metrics. The red outline is the weakest stage; the wide bar is the end-to-end task accuracy it caps.
End state versus trajectory
Same destination, very different runs
For a question-answering system the answer is the artifact. For an agent, the artifact is a run: a sequence of tool calls that leaves the world in some state. There are two independent questions, and conflating them hides the failures that matter in production. Did the world end up right — is the database row correct, the file written, the ticket closed? That is the end-state check, and it is objective and cheap. Was the path sane — bounded steps, bounded cost, no forbidden action, no repeated identical call? That is the trajectory check, and it is what a safety incident or a runaway bill actually is.
The two runs below reach the same correct end state. Run A does it in six steps at $0.12 with no violation. Run B does it in thirty-four steps at $0.98 and writes a file outside its sandbox once on the way. An end-state-only eval passes both. Move the trajectory weight and decide for yourself where a run like B should land.
Both runs end correct. Steps, cost and the safety flag are the only things that separate them — and only a trajectory score sees them.
pass@k says capability; pass^k says reliability
The number that decays as you raise k
pass@k asks: did the agent solve the task in at least one of k attempts? It rises with k, and it is how capability benchmarks report. pass^k asks the opposite: did it solve the task in all k attempts? It falls with k, and it is the number a product needs, because a user who runs the same task tomorrow does not get to pick the best of ten tries. If a task succeeds 90% of the time, an agent's pass@10 is almost certain but its pass^10 — ten clean runs in a row — is 0.35. Move the per-attempt success rate and watch the two curves pull apart.
Both curves come from the same per-attempt rate. pass@k climbs toward 1; pass^k falls away from it.
Public benchmarks do not transfer
A leaderboard is a shortlist, not a measurement
The named benchmarks are worth knowing — SWE-bench Verified for repo issue resolution, τ²-bench for tool-agent-user dialogue with a verifiable database end state, GAIA for multi-step assistant work, OSWorld for computer use. Each is graded differently, and each has a specific reason its score is a weak predictor of yours: training-data leakage inflates any benchmark that has been public long enough; the scaffold (the prompt, the tools, the loop) often contributes more variance than the model; the environment version changes the score; the grader can be gamed or can hide partial credit. A high public score means "this family is worth trying", not "this will work on your task".
The scatter below plots a public benchmark score against a hypothetical held-out score on your own task, for twelve model-plus-scaffold combinations. At low shared structure the correlation is near zero. Raise the overlap slider only if your task genuinely resembles the benchmark — and note that even moderate overlap leaves most of the variance unexplained.
Each dot is one system. x = public benchmark score; y = your held-out task score. The line is the least-squares fit.
Cheat sheet
| Question | The answer that shapes the harness |
|---|---|
| Retrieval missed the evidence | Low context recall. Fix with hybrid retrieval, query transforms, a wider candidate set. |
| The wrong chunk was kept | Low context precision. Fix with reranking, better chunking, metadata filters. |
| Fluid but ungrounded answer | Low faithfulness. Fix with grounding instructions, citations, retrieval-gated answers. |
| Right evidence, unused | Retrieved-but-not-used (the middle-of-context effect). Only end-to-end accuracy sees it. |
| Did the agent succeed? | End state: is the world left correct? Objective and cheap — build it first. |
| Was the run acceptable? | Trajectory: steps, cost, permission violations, repeated calls. A separate question. |
| What does pass@k hide? | Reliability. pass^k — solved on every one of k attempts — is the product number. |
| What is the cost metric? | Cost per completed task, not cost per call. Failures are paid for too. |
| Can I use a public score? | As a shortlist only. Leakage, scaffold, environment and grader all break the transfer. |
| What do I measure on? | A held-out task set with verifiable end states of your own. |
Further reading
- Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?", ICLR 2024 — the benchmark, its harness, and the scaffold-dependence its own authors document.
- Yao et al., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024 — verifiable database end states and the pass^k reliability measure.
- Mialon et al., "GAIA: a benchmark for General AI Assistants", 2023 — multi-step assistant tasks and the exact-answer grading that hides partial credit.
- Xie et al., "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments", NeurIPS 2024 — computer-use grading against environment state, and the version sensitivity.
- Es et al., "RAGAS: Automated Evaluation of Retrieval Augmented Generation", EACL 2024 — faithfulness, answer relevance, context precision and context recall as the component metrics used here.
- Anthropic, "Introducing Contextual Retrieval", 19 September 2024 — component-level failure rates as hybrid retrieval, BM25 and reranking are added one at a time.