Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A prompt change is a deploy

Version it, review it, canary it, and be able to take it back

The single most common production incident in an LLM product is an unreviewed prompt edit shipped straight to everyone. It is not a config tweak, because the prompt is the program: a reworded instruction, a reordered example, or a swapped model changes behaviour across every request at once, and there is no compiler to catch it. Treat it exactly like a code deploy.

The rollout below puts a candidate that looks faster and cheaper on a slice of traffic. It is both — and it is worse on one slice. Move the canary share and the rollback threshold and watch the release decide itself.

Top: quality per batch on the canary slice. Bottom: cost per request. The candidate wins on cost and loses on quality.

⚠️ A canary that only watches cost will ship a regression. Here the candidate is genuinely cheaper, so a cost dashboard would call it a win while a slice of users quietly get worse answers. Gate on the metric that encodes the promise, and treat cost as a constraint rather than the objective.
2

Risk-tiered autonomy

Match the level of autonomy to the risk and the measured confidence

"Should the agent do it, or ask?" is not one question. It has two inputs and four sensible answers. The inputs are the risk of the action — how expensive, how irreversible, how visible outside your boundary — and the system's measured confidence on this task, which should come from evidence (a labelled slice, a check that passed, agreement across attempts) rather than from the model's own assertion. The answers form a ladder.

Place the sample tasks, then move the two thresholds. The point of the exercise is how few tasks belong in the bottom-right corner — the tier that actually saves time is the tier you can only justify with evidence.

Task risk on the horizontal axis, measured confidence on the vertical. The thresholds move the four autonomy levels.

💡 The durable idea: autonomy is not a property of the agent, it is a property of a task, a risk and a measured confidence. Change any of the three and the right answer changes — which is why a single global setting ("the agent may act") is always wrong somewhere in the product.
3

An interface for uncertainty

Show provenance, make correction cheap, never fake confidence

A probabilistic system will be wrong sometimes. The interface decides what happens next. An answer presented with the same formatting regardless of how well-grounded it is teaches users to trust everything equally, which is precisely the wrong lesson — and it is expensive, because unearned trust is only discovered when it costs someone. Three moves change the failure mode without changing the model.

The three answer cards below are the same model's output presented three ways. Move the accuracy slider: the bare card's trust score does not move with it, which is exactly the problem.

Trust, correction rate and misled users per 1,000 answers for each presentation. Trust that does not track accuracy is the hazard.

💡 The durable idea: the interface is where a probabilistic system earns or loses trust. Provenance makes verification cheap, a correction affordance turns a complaint into a label, and honest hedging keeps the label honest — because the user can tell the difference between "I know" and "I think".
4

The edit-to-dataset flywheel

Every correction is a free label, and the flywheel is the moat

The reason a shipped LLM product can get better faster than a competitor's is not the prompt or the model — both are available to everyone. It is the labelled data the product generates as a by-product of use, and how quickly that data returns to the eval suite and the prompt. A user edit is the most valuable event in the system: it is a human telling you the right answer, for free, on real traffic, on a case your eval set probably did not contain.

Closing the loop is the discipline. Capture the edit with enough context to reproduce the request; turn it into a labelled example; add it to the eval set that gates releases; use the recurring failures to change the prompt, the retrieval or a fine-tune; then measure the improvement on the same set. This is the loop Parts 9 and 14 described from the other ends — error analysis upstream, tuning downstream — and it is the one asset a competitor cannot copy, because it is produced by your users on your product.

Step the loop below and watch the counts move. The capture rate is the lever that decides how much of the signal survives; the rest is destroyed at the moment the user corrects and nobody records it.

Edits → labels → eval → prompt change → better answers → more usage → more edits. The packet is one week of signal.

💡 The durable idea: the flywheel is the durable moat. Models and prompts are commoditised and substitutable; a labelled dataset produced by your own users, and a loop fast enough to act on it, is the thing that compounds — and pass^k on that set is how you prove it did.

Cheat sheet

QuestionThe answer that shapes the release
Is a prompt edit a config change?No. It is a deploy: version it, review it, gate it on the eval suite, and be able to roll it back.
What must a canary have?A pre-defined threshold and an automatic revert. Otherwise it is just a slow rollout.
What should a canary watch?The metric that encodes the promise. A cost-only canary ships quality regressions.
What are the two inputs to autonomy?Risk of the action and measured confidence — evidence, not the model's own assertion.
What are the four autonomy levels?Suggest, confirm, act-and-report, fully autonomous.
Which tier saves the most time?Fully autonomous — and it is the one that needs the most evidence, so it stays the smallest quadrant.
How should uncertainty be shown?Provenance, a cheap correction path, and honest hedging. Never a confident answer the system cannot support.
What is a user edit worth?A free label on real traffic: the right answer, on a case your eval set likely missed.
How do you close the loop?Capture with context → label → add to the eval set → change the prompt/model → measure on the same set.
What is the durable moat?The labelled dataset your users generate, and a loop fast enough to use it.

Further reading

5

Check your understanding

0/5 answered