Shipping probabilistic products
The last gap between a working agent and a product is not model quality. It is that everything downstream of the model assumed a deterministic dependency: deploys you can reason about, behaviour you can specify, interfaces that promise a right answer. A probabilistic system breaks each of those assumptions, and the teams that ship well replace them rather than pretend. A prompt becomes a versioned artifact with a canary and a rollback; autonomy becomes a decision matched to risk and measured confidence; uncertainty becomes something the interface shows rather than hides; and every user correction becomes a label that makes the next release better.
A prompt change is a deploy
Version it, review it, canary it, and be able to take it back
The single most common production incident in an LLM product is an unreviewed prompt edit shipped straight to everyone. It is not a config tweak, because the prompt is the program: a reworded instruction, a reordered example, or a swapped model changes behaviour across every request at once, and there is no compiler to catch it. Treat it exactly like a code deploy.
- Version it. A prompt is an artifact with a hash and a history. "Which prompt produced this answer?" must be answerable from a trace, months later.
- Review and gate it. Run the candidate against the eval suite from Parts 8 to 11 — the labelled set, the judge, the failure taxonomy — and block the merge on a quality regression, not on a feeling.
- Canary it. Send a small slice of live traffic to the new version and watch quality and cost per request on real traffic, which is always stranger than the eval set.
- Automate the rollback. Define the threshold in advance so the decision is not made under pressure. A canary without an automatic revert is a slow-motion full rollout.
The rollout below puts a candidate that looks faster and cheaper on a slice of traffic. It is both — and it is worse on one slice. Move the canary share and the rollback threshold and watch the release decide itself.
Top: quality per batch on the canary slice. Bottom: cost per request. The candidate wins on cost and loses on quality.
Risk-tiered autonomy
Match the level of autonomy to the risk and the measured confidence
"Should the agent do it, or ask?" is not one question. It has two inputs and four sensible answers. The inputs are the risk of the action — how expensive, how irreversible, how visible outside your boundary — and the system's measured confidence on this task, which should come from evidence (a labelled slice, a check that passed, agreement across attempts) rather than from the model's own assertion. The answers form a ladder.
- Suggest. High risk, low confidence. The agent proposes; a human decides and acts.
- Confirm. Low risk, low confidence. The agent can act, but it stops for a yes first.
- Act and report. High risk, high confidence. The agent acts and tells a human what it did, with a correction path.
- Fully autonomous. Low risk, high confidence. The agent acts silently; this is the only tier that earns back the human attention the others spend.
Place the sample tasks, then move the two thresholds. The point of the exercise is how few tasks belong in the bottom-right corner — the tier that actually saves time is the tier you can only justify with evidence.
Task risk on the horizontal axis, measured confidence on the vertical. The thresholds move the four autonomy levels.
An interface for uncertainty
Show provenance, make correction cheap, never fake confidence
A probabilistic system will be wrong sometimes. The interface decides what happens next. An answer presented with the same formatting regardless of how well-grounded it is teaches users to trust everything equally, which is precisely the wrong lesson — and it is expensive, because unearned trust is only discovered when it costs someone. Three moves change the failure mode without changing the model.
- Show provenance. Cite the source for each claim, link to it, and let the user check. A citation turns a blind trust decision into a cheap verification.
- Make the correction path cheap. An edit box, a thumbs-down with a reason, a one-click "this is wrong" — each is a label and a repair in the same gesture. If correcting costs more than ignoring, users ignore.
- Never fabricate a confident answer. If the retrieved evidence is thin or the checks disagree, say so and lower the claim. A hedged "I could not find that in the sources" is a better product than a confident guess.
The three answer cards below are the same model's output presented three ways. Move the accuracy slider: the bare card's trust score does not move with it, which is exactly the problem.
Trust, correction rate and misled users per 1,000 answers for each presentation. Trust that does not track accuracy is the hazard.
The edit-to-dataset flywheel
Every correction is a free label, and the flywheel is the moat
The reason a shipped LLM product can get better faster than a competitor's is not the prompt or the model — both are available to everyone. It is the labelled data the product generates as a by-product of use, and how quickly that data returns to the eval suite and the prompt. A user edit is the most valuable event in the system: it is a human telling you the right answer, for free, on real traffic, on a case your eval set probably did not contain.
Closing the loop is the discipline. Capture the edit with enough context to reproduce the request; turn it into a labelled example; add it to the eval set that gates releases; use the recurring failures to change the prompt, the retrieval or a fine-tune; then measure the improvement on the same set. This is the loop Parts 9 and 14 described from the other ends — error analysis upstream, tuning downstream — and it is the one asset a competitor cannot copy, because it is produced by your users on your product.
Step the loop below and watch the counts move. The capture rate is the lever that decides how much of the signal survives; the rest is destroyed at the moment the user corrects and nobody records it.
Edits → labels → eval → prompt change → better answers → more usage → more edits. The packet is one week of signal.
Cheat sheet
| Question | The answer that shapes the release |
|---|---|
| Is a prompt edit a config change? | No. It is a deploy: version it, review it, gate it on the eval suite, and be able to roll it back. |
| What must a canary have? | A pre-defined threshold and an automatic revert. Otherwise it is just a slow rollout. |
| What should a canary watch? | The metric that encodes the promise. A cost-only canary ships quality regressions. |
| What are the two inputs to autonomy? | Risk of the action and measured confidence — evidence, not the model's own assertion. |
| What are the four autonomy levels? | Suggest, confirm, act-and-report, fully autonomous. |
| Which tier saves the most time? | Fully autonomous — and it is the one that needs the most evidence, so it stays the smallest quadrant. |
| How should uncertainty be shown? | Provenance, a cheap correction path, and honest hedging. Never a confident answer the system cannot support. |
| What is a user edit worth? | A free label on real traffic: the right answer, on a case your eval set likely missed. |
| How do you close the loop? | Capture with context → label → add to the eval set → change the prompt/model → measure on the same set. |
| What is the durable moat? | The labelled dataset your users generate, and a loop fast enough to use it. |
Further reading
- Sculley et al., "Hidden Technical Debt in Machine Learning Systems", NeurIPS 2015 — why an ML product is a software system, and why the code around the model is the system.
- Breck et al., "The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction", 2017 — the production-readiness checklist behind canaries, rollback and monitoring.
- Yao et al., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024 — pass^k as the reliability metric that a canary on the eval set is trying to protect.
- Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?", ICLR 2024 — a benchmark whose score depends on the harness, and therefore on the deploy.
- Amershi et al., "Guidelines for Human-AI Interaction", CHI 2019 — provenance, correction and honest confidence as design guidance rather than taste.
- Anthropic, "Building effective agents", December 2024 — when to automate and when to keep a human in the loop, which is the autonomy ladder restated.