Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Grade the pipeline, not just the answer

A number that points at a stage

A retrieval-augmented pipeline has stages, and each stage fails in its own way. Retrieval can miss the evidence the answer needed — measured by context recall. Reranking can keep the wrong chunk in a wide candidate set — measured by context precision. Generation can produce an answer that is fluent but not grounded in the context it was given — measured by faithfulness. And a fourth, quieter failure sits between them: the right evidence was retrieved and then not used, because it landed in the middle of a long context. Only end-to-end task accuracy catches that one.

Measure every stage, and a low end-to-end score stops being a verdict and becomes a diagnosis. The demo below starts with a weak pipeline and lets you switch on the fixes. Watch the end-to-end number stay stubbornly low until you fix the stage the label names as the bottleneck — fixing the others does almost nothing.

Three stages with their own metrics. The red outline is the weakest stage; the wide bar is the end-to-end task accuracy it caps.

💡 Component metrics are diagnostic, not decorative. A retrieval failure and a grounding failure need opposite fixes; an end-to-end score that merges them sends you to the wrong one half the time.
2

End state versus trajectory

Same destination, very different runs

For a question-answering system the answer is the artifact. For an agent, the artifact is a run: a sequence of tool calls that leaves the world in some state. There are two independent questions, and conflating them hides the failures that matter in production. Did the world end up right — is the database row correct, the file written, the ticket closed? That is the end-state check, and it is objective and cheap. Was the path sane — bounded steps, bounded cost, no forbidden action, no repeated identical call? That is the trajectory check, and it is what a safety incident or a runaway bill actually is.

The two runs below reach the same correct end state. Run A does it in six steps at $0.12 with no violation. Run B does it in thirty-four steps at $0.98 and writes a file outside its sandbox once on the way. An end-state-only eval passes both. Move the trajectory weight and decide for yourself where a run like B should land.

Both runs end correct. Steps, cost and the safety flag are the only things that separate them — and only a trajectory score sees them.

⚠️ The trap: an agent that reaches the right answer through a forbidden path is not a success you can ship. Grade the trajectory on the constraints you actually care about — steps, spend, permissions, loop detection — or those failures reach production unmeasured.
3

pass@k says capability; pass^k says reliability

The number that decays as you raise k

pass@k asks: did the agent solve the task in at least one of k attempts? It rises with k, and it is how capability benchmarks report. pass^k asks the opposite: did it solve the task in all k attempts? It falls with k, and it is the number a product needs, because a user who runs the same task tomorrow does not get to pick the best of ten tries. If a task succeeds 90% of the time, an agent's pass@10 is almost certain but its pass^10 — ten clean runs in a row — is 0.35. Move the per-attempt success rate and watch the two curves pull apart.

Both curves come from the same per-attempt rate. pass@k climbs toward 1; pass^k falls away from it.

💡 Report both, or you are hiding the reliability gap. A capability number that rises with k and a reliability number that falls with k are computed from the same runs; publishing only the first is how a demo becomes an outage.
4

Public benchmarks do not transfer

A leaderboard is a shortlist, not a measurement

The named benchmarks are worth knowing — SWE-bench Verified for repo issue resolution, τ²-bench for tool-agent-user dialogue with a verifiable database end state, GAIA for multi-step assistant work, OSWorld for computer use. Each is graded differently, and each has a specific reason its score is a weak predictor of yours: training-data leakage inflates any benchmark that has been public long enough; the scaffold (the prompt, the tools, the loop) often contributes more variance than the model; the environment version changes the score; the grader can be gamed or can hide partial credit. A high public score means "this family is worth trying", not "this will work on your task".

The scatter below plots a public benchmark score against a hypothetical held-out score on your own task, for twelve model-plus-scaffold combinations. At low shared structure the correlation is near zero. Raise the overlap slider only if your task genuinely resembles the benchmark — and note that even moderate overlap leaves most of the variance unexplained.

Each dot is one system. x = public benchmark score; y = your held-out task score. The line is the least-squares fit.

⚠️ Build your own held-out set with verifiable end states. A task set you can grade automatically — a test that passes, a database row that matches, a file whose checksum is right — is worth more than any leaderboard position, and it is the only thing your change can be measured against.

Cheat sheet

QuestionThe answer that shapes the harness
Retrieval missed the evidenceLow context recall. Fix with hybrid retrieval, query transforms, a wider candidate set.
The wrong chunk was keptLow context precision. Fix with reranking, better chunking, metadata filters.
Fluid but ungrounded answerLow faithfulness. Fix with grounding instructions, citations, retrieval-gated answers.
Right evidence, unusedRetrieved-but-not-used (the middle-of-context effect). Only end-to-end accuracy sees it.
Did the agent succeed?End state: is the world left correct? Objective and cheap — build it first.
Was the run acceptable?Trajectory: steps, cost, permission violations, repeated calls. A separate question.
What does pass@k hide?Reliability. pass^k — solved on every one of k attempts — is the product number.
What is the cost metric?Cost per completed task, not cost per call. Failures are paid for too.
Can I use a public score?As a shortlist only. Leakage, scaffold, environment and grader all break the transfer.
What do I measure on?A held-out task set with verifiable end states of your own.

Further reading

5

Check your understanding

0/5 answered