Observability, online eval, and evals as the moat
An evaluation is a question you ask of recorded runs. If the run was not traced, the question has no data to answer it — so observability is not a separate ops concern bolted on at the end, it is the substrate the eval sits on. This part covers the OpenTelemetry spans that carry an agent's shape, the fields without which a score cannot be attributed (prompt version, retrieved document ids, tool calls, cost), the discipline of redacting PII before anything is written down, online evaluation against live traffic with the implicit signals users hand you for free, the statistics of A/B testing with real power, and the dashboard where quality, cost and latency finally meet. It ends on the thesis the whole volume has been building toward: prompts get copied, evals compound.
You cannot analyse a trace you did not record
The eval is a query; the trace is the table
The OpenTelemetry GenAI semantic conventions give an agent request a standard shape: a span for the top-level request, child spans for each model call and each tool call, and attributes that carry the model identity, the token usage and the parameters. The conventions name the attributes rather than inventing your own, so the same trace can be read by a generic backend: gen_ai.system, gen_ai.request.model, gen_ai.request.temperature, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. A chat span, a tool span and a retrieval span are the three shapes you will see most; an agent request is a tree of them, and the tree is what tells you where the time and the tokens actually went.
What the conventions buy you is that the eval harness reads runs, not reconstructions. A pass rate stops being a number you derived from a log line and becomes a filter over spans: "requests where the tool returned a 500" or "sessions where the model call used prompt v14". The demo below is one such request, drawn as a span waterfall. Click any span to see the attributes the conventions would attach to it.
One agent request as nested spans on a shared time axis. Indentation is nesting; the bar is duration; the detail panel shows the attributes for the selected span.
Log the context and the prompt version — or the score is unattributable
Four fields that turn a number into a diagnosis
Token counts and a latency are not enough. When an eval moves, you need to know what changed, and a run that did not record its inputs cannot tell you. Four fields do most of that work. The prompt version — a content hash or a semantic version, not "the current prompt" — pins the instructions and the examples that produced the output. The retrieved document ids (with their scores and the index revision) pin the evidence the model saw, which is the difference between "the model got worse" and "the index returned the wrong chunks". The tool calls, with arguments and results, pin the world the agent acted on. And the cost and latency ride along on the same span so a quality change can be judged against what it cost to buy. Join all of this with a request id and an eval result becomes a pointer back into the exact run.
Redact before you write, not after you export. Logs are copied, sampled, shipped to a vendor and kept for months; a raw prompt is the highest-risk artifact in the system because it is where users paste secrets and personal data. Redaction belongs at the boundary, before the span is emitted: hash or drop direct identifiers, redact detected PII classes, and keep the raw text only under a separate, short retention with explicit access — or not at all. It is a lossy tradeoff and it is the right one: a redacted trace you can share across the team beats a raw one nobody is allowed to look at. Part 16 returns to this as a production control, but the cheap version of it belongs here, in the logging boundary itself.
Online eval and A/B testing with real power
How many samples before the difference means anything
Offline evals run on a frozen set; online evals run on live traffic, which is the only place the distribution of real requests lives. The tool for online change is the A/B test, and the tool is only as good as its power: the probability that it detects a real effect of a given size. Power is not a property of the test — it is a property of the test at a sample size, and the arithmetic is unforgiving. For two proportions, the samples needed per arm are roughly n = (z1−α/2 + z1−β)² · (p₁(1−p₁) + p₂(1−p₂)) / (p₂ − p₁)². At a 10% baseline, a 2-point lift needs a few thousand sessions per arm; a half-point lift needs sixteen times as many, because the required n scales with the inverse square of the effect.
Move the sliders below. The left panel is required sample size as a function of the effect you want to detect; the right panel is the reason a dashboard that updates every hour is dangerous — if you check the result and stop when it looks significant, the false-positive rate is not 5%, it is the rate of a test you ran many times.
Left: required sessions per arm against the effect size, with the current setting marked. Right: family-wise false-positive rate as interim looks accumulate, against a nominal 5% line.
Implicit signals are the label you already have
Users grade every answer, whether or not you ask
Explicit feedback — a thumbs-up, a rating — is sparse, biased toward the extremes and easy to game. The dense signal is behavioural, and it is already in the traces. A user who accepts an answer and moves on is a weak positive. A user who edits the answer before accepting it is telling you exactly where it was wrong. A user who retries, or who rephrases the question as a follow-up, is a negative — the first answer failed to land. A copy event that is not followed by an edit is a strong positive. None of these is labelled, all of them are proxy, and together the edit rate is one of the best cheap quality meters an agent product has.
The funnel below turns 20,000 sessions into accepted, edited and retried outcomes, and the scatter checks the proxy against the thing you already trust: the offline eval. If the two agree, the implicit signal is a usable online meter; if they do not, one of the two is measuring something else, and that is worth knowing before you wire it to a dashboard.
Top: sessions through answers to accepted, edited and retried outcomes. Bottom: each dot is one day; x is the offline eval score, y is the observed edit rate, and the line is the fit.
Quality, cost and latency on the same dashboard
They trade off; a dashboard that shows one hides the trade
A quality number in isolation invites a decision that quietly spends money and time to buy it. Every operating point in an agent system sits in three dimensions at once: what it costs per task, how long it takes, and how good it is. Concretely, the dials are mundane — retrieve more candidates and rerank (cost, latency up, quality up), call a bigger model (cost up, latency up, quality up), cache more aggressively (cost down, latency down, correctness risk up), run fewer reasoning steps (all three down). The points that survive are the Pareto frontier: no other feasible point is at least as cheap, at least as fast and at least as good. Everything else is dominated, and a team that ships a dominated point is paying for nothing.
The dashboard below plots candidate operating points with cost on one axis, latency on the other, and quality as the size of the bubble. Set a quality floor and the frontier recomputes over the points that clear it — and note how raising the floor moves the frontier out toward cost and latency, which is the trade made visible.
Each bubble is an operating point; bigger is higher quality. Dark points are on the Pareto frontier among those meeting the floor; hollow points are dominated; grey points miss the floor.
Cheat sheet
| Question | The answer that shapes the system |
|---|---|
| Why instrument before evaluating? | An eval is a query over recorded runs. No trace, no answer. |
| What does a GenAI span carry? | Model, parameters, input/output token usage, and one span per model and tool call. |
| Which field pins the instructions? | The prompt version — a content hash or semantic version, not "current". |
| Which field pins the evidence? | Retrieved document ids with scores, plus the index revision. |
| When do you redact PII? | Before the span is written. Hash identifiers, redact data classes, keep raw text only under tight retention. |
| Offline or online eval? | Both. Offline is fast and frozen; online sees the real request distribution. |
| How many samples for an A/B? | n per arm grows with the inverse square of the effect. A half-size effect needs four times the samples. |
| Why is peeking wrong? | Stopping on the first significant interim look inflates false positives far past the nominal rate. |
| What are implicit signals? | Accepted, edited, retried, copied, rephrased. Dense, free, and only proxies — correlate them with the eval. |
| Why one dashboard? | Quality, cost and latency trade off. Showing one hides the decision. |
| What is the moat? | Evals, datasets and traces compound. Prompts and scaffolds get copied. |
Further reading
- OpenTelemetry, "Semantic Conventions for GenAI", specification, revisions dated 2024–2025 — the span and attribute names (
gen_ai.*) used throughout this part. - Yao et al., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024 — pass^k and the argument that reliability, not best-of-k capability, is the product number.
- Kohavi, Tang & Xu, "Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing", Cambridge University Press, 2020 — power, sample-size arithmetic and the peeking problem this part reduces to two sliders.
- Sculley et al., "Hidden Technical Debt in Machine Learning Systems", NeurIPS 2015 — the original argument that monitoring, configuration and data dependencies are where ML systems decay.
- Es et al., "RAGAS: Automated Evaluation of Retrieval Augmented Generation", EACL 2024 — the component metrics an offline harness computes, and the reason retrieved ids must be logged to compute them.
- Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL 2024 — why the placement of retrieved context, not just its presence, is a quality variable worth logging.