Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

You cannot analyse a trace you did not record

The eval is a query; the trace is the table

The OpenTelemetry GenAI semantic conventions give an agent request a standard shape: a span for the top-level request, child spans for each model call and each tool call, and attributes that carry the model identity, the token usage and the parameters. The conventions name the attributes rather than inventing your own, so the same trace can be read by a generic backend: gen_ai.system, gen_ai.request.model, gen_ai.request.temperature, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. A chat span, a tool span and a retrieval span are the three shapes you will see most; an agent request is a tree of them, and the tree is what tells you where the time and the tokens actually went.

What the conventions buy you is that the eval harness reads runs, not reconstructions. A pass rate stops being a number you derived from a log line and becomes a filter over spans: "requests where the tool returned a 500" or "sessions where the model call used prompt v14". The demo below is one such request, drawn as a span waterfall. Click any span to see the attributes the conventions would attach to it.

One agent request as nested spans on a shared time axis. Indentation is nesting; the bar is duration; the detail panel shows the attributes for the selected span.

💡 Instrument once, evaluate forever. The same span tree answers "where did the latency go", "what did this run cost" and "why did this eval case regress". Retrofitting telemetry after a quality incident means the incident is unanswerable.
2

Log the context and the prompt version — or the score is unattributable

Four fields that turn a number into a diagnosis

Token counts and a latency are not enough. When an eval moves, you need to know what changed, and a run that did not record its inputs cannot tell you. Four fields do most of that work. The prompt version — a content hash or a semantic version, not "the current prompt" — pins the instructions and the examples that produced the output. The retrieved document ids (with their scores and the index revision) pin the evidence the model saw, which is the difference between "the model got worse" and "the index returned the wrong chunks". The tool calls, with arguments and results, pin the world the agent acted on. And the cost and latency ride along on the same span so a quality change can be judged against what it cost to buy. Join all of this with a request id and an eval result becomes a pointer back into the exact run.

// the attributes that make an eval result attributable { "gen_ai.request.model": "gpt-4o", "app.prompt.version": "v14@3f9c1a", // content hash of the prompt "app.retrieved.ids": ["kv-cache", "paged-memory", "batching"], "app.retrieved.index": "corpus-2026-09-01", // which index revision "app.tool.calls": ["fetch_doc", "search"], "gen_ai.usage.input_tokens": 13100, "gen_ai.usage.output_tokens": 520, "app.cost.usd": 0.0173, "app.latency.ms": 2600, "app.user.id_hash": "sha256:9c1f…" // never the raw id }

Redact before you write, not after you export. Logs are copied, sampled, shipped to a vendor and kept for months; a raw prompt is the highest-risk artifact in the system because it is where users paste secrets and personal data. Redaction belongs at the boundary, before the span is emitted: hash or drop direct identifiers, redact detected PII classes, and keep the raw text only under a separate, short retention with explicit access — or not at all. It is a lossy tradeoff and it is the right one: a redacted trace you can share across the team beats a raw one nobody is allowed to look at. Part 16 returns to this as a production control, but the cheap version of it belongs here, in the logging boundary itself.

⚠️ A trace with no prompt version is a number with no address. If you cannot say which instructions and which retrieved documents produced a run, you cannot reproduce it, cannot bisect a regression, and cannot tell a model upgrade from an index change.
3

Online eval and A/B testing with real power

How many samples before the difference means anything

Offline evals run on a frozen set; online evals run on live traffic, which is the only place the distribution of real requests lives. The tool for online change is the A/B test, and the tool is only as good as its power: the probability that it detects a real effect of a given size. Power is not a property of the test — it is a property of the test at a sample size, and the arithmetic is unforgiving. For two proportions, the samples needed per arm are roughly n = (z1−α/2 + z1−β)² · (p₁(1−p₁) + p₂(1−p₂)) / (p₂ − p₁)². At a 10% baseline, a 2-point lift needs a few thousand sessions per arm; a half-point lift needs sixteen times as many, because the required n scales with the inverse square of the effect.

Move the sliders below. The left panel is required sample size as a function of the effect you want to detect; the right panel is the reason a dashboard that updates every hour is dangerous — if you check the result and stop when it looks significant, the false-positive rate is not 5%, it is the rate of a test you ran many times.

Left: required sessions per arm against the effect size, with the current setting marked. Right: family-wise false-positive rate as interim looks accumulate, against a nominal 5% line.

⚠️ Peeking inflates false positives. The formula assumes one look after the samples are in. Checking repeatedly and stopping on the first significant result is a different test with a much higher error rate; if you must monitor continuously, use a sequential or always-valid method, not a fixed-horizon p-value read early.
4

Implicit signals are the label you already have

Users grade every answer, whether or not you ask

Explicit feedback — a thumbs-up, a rating — is sparse, biased toward the extremes and easy to game. The dense signal is behavioural, and it is already in the traces. A user who accepts an answer and moves on is a weak positive. A user who edits the answer before accepting it is telling you exactly where it was wrong. A user who retries, or who rephrases the question as a follow-up, is a negative — the first answer failed to land. A copy event that is not followed by an edit is a strong positive. None of these is labelled, all of them are proxy, and together the edit rate is one of the best cheap quality meters an agent product has.

The funnel below turns 20,000 sessions into accepted, edited and retried outcomes, and the scatter checks the proxy against the thing you already trust: the offline eval. If the two agree, the implicit signal is a usable online meter; if they do not, one of the two is measuring something else, and that is worth knowing before you wire it to a dashboard.

Top: sessions through answers to accepted, edited and retried outcomes. Bottom: each dot is one day; x is the offline eval score, y is the observed edit rate, and the line is the fit.

💡 Correlate the proxy with the eval, then trust the cheaper one. An offline set is slow and stale; a behavioural signal is free and current. Their correlation is the licence to use the cheap signal online — and the check that you are not optimising a metric users do not care about.
5

Quality, cost and latency on the same dashboard

They trade off; a dashboard that shows one hides the trade

A quality number in isolation invites a decision that quietly spends money and time to buy it. Every operating point in an agent system sits in three dimensions at once: what it costs per task, how long it takes, and how good it is. Concretely, the dials are mundane — retrieve more candidates and rerank (cost, latency up, quality up), call a bigger model (cost up, latency up, quality up), cache more aggressively (cost down, latency down, correctness risk up), run fewer reasoning steps (all three down). The points that survive are the Pareto frontier: no other feasible point is at least as cheap, at least as fast and at least as good. Everything else is dominated, and a team that ships a dominated point is paying for nothing.

The dashboard below plots candidate operating points with cost on one axis, latency on the other, and quality as the size of the bubble. Set a quality floor and the frontier recomputes over the points that clear it — and note how raising the floor moves the frontier out toward cost and latency, which is the trade made visible.

Each bubble is an operating point; bigger is higher quality. Dark points are on the Pareto frontier among those meeting the floor; hollow points are dominated; grey points miss the floor.

💡 Evals are the moat because prompts are not. A prompt, a scaffold, a tool schema — each is copyable the afternoon it ships. The labelled dataset, the trace corpus and the eval harness that turns a change into a decided bet are not: they compound with every shipped run, and they are the only part of the system that gets harder to copy the longer you do it. The moat is not the model you call; it is the evidence you can bring to the next decision.

Cheat sheet

QuestionThe answer that shapes the system
Why instrument before evaluating?An eval is a query over recorded runs. No trace, no answer.
What does a GenAI span carry?Model, parameters, input/output token usage, and one span per model and tool call.
Which field pins the instructions?The prompt version — a content hash or semantic version, not "current".
Which field pins the evidence?Retrieved document ids with scores, plus the index revision.
When do you redact PII?Before the span is written. Hash identifiers, redact data classes, keep raw text only under tight retention.
Offline or online eval?Both. Offline is fast and frozen; online sees the real request distribution.
How many samples for an A/B? n per arm grows with the inverse square of the effect. A half-size effect needs four times the samples.
Why is peeking wrong?Stopping on the first significant interim look inflates false positives far past the nominal rate.
What are implicit signals?Accepted, edited, retried, copied, rephrased. Dense, free, and only proxies — correlate them with the eval.
Why one dashboard?Quality, cost and latency trade off. Showing one hides the decision.
What is the moat?Evals, datasets and traces compound. Prompts and scaffolds get copied.

Further reading

6

Check your understanding

0/5 answered