Reliability, cost and latency engineering
An agent in production depends on something that is slow, occasionally wrong, and priced per token — and it depends on that thing many times per task. That combination is exactly the shape of a distributed-systems problem, and the standard tools apply once you accept the premise: the model is a fallible remote dependency, not a function you own. This part is about the three budgets that decide whether the product is viable — reliability under failure, dollars per task, and milliseconds to a useful answer — and about the order in which to spend effort on each.
Patterns for a probabilistic dependency
Timeouts, backoff, breakers, bulkheads, degraded modes
The moment a call can hang, fail, or return slowly, the client needs the same defenses it would need against any unreliable RPC. A timeout is a deadline with a purpose: without one, a slow dependency consumes the request's whole latency budget and hides its own failure. Exponential backoff with jitter makes a retry cheap when the dependency recovers and rare when it does not, and the jitter is not decoration — it is what stops every client retrying on the same tick. A circuit breaker counts consecutive failures and, past a threshold, fails fast for a cooldown instead of queueing retries against something that is already down. A bulkhead gives the dependency its own pool so a stall cannot consume every worker in the process. And an explicit degraded mode decides in advance what "unavailable" means: a smaller model, a cached answer, or an honest "try again in a minute".
The simulator below runs 140 seeded requests through timeout, retry and breaker policy. Raise the retry count and average latency improves until the timeout is the binding constraint; lower the breaker threshold and watch the run trip into degraded mode earlier. The important readout is the error rate after the fallback, not before it.
Top: the outcome of each seeded request. Bottom: its latency, with the timeout drawn as the dashed ceiling.
The ordered cost-lever list
Do the big things first, in this order
Cost optimisation usually fails because it is attempted in the wrong order: teams spend a week hand-tuning a prompt to shave 8% of input tokens while the output is four times the price of the input and twice as long as it needs to be. The levers are not equal, and they are not independent. Applied in order, from largest to smallest and cheapest to hardest:
- Shorten the output. Output is the expensive side at every vendor — three to five times the input price in the dated pricing table this page cites. A tighter answer format, a cap on the completion, or a template instead of prose moves the biggest number first.
- Cache the prefix. Stable content first, volatile content last, so the provider's prefix cache hits and the repeated input is billed at a fraction of the rate.
- Route easy traffic to a smaller model. Most requests do not need the frontier model; a cheap classifier or a confidence rule sends the easy majority down a cheaper path.
- Batch. Offline and bulk work trades latency for a discount, often a large one.
- Optimise the prompt. Finally — and only after the arithmetic above — trim the instructions and examples. It is real work and it is the smallest lever, which is why it goes last.
The waterfall applies the levers in that order to a baseline monthly bill. Move the request volume, move the share of easy traffic, and step through the levers; the final bar is the bill you would actually pay.
Baseline, then each lever's saving as a step down, then the final bill. Levers apply in order — the labels are the order.
Three cache layers, and the correctness cost of semantic caching
Prefix, result, semantic — one of them can be wrong
Caching an LLM application happens at three different layers, and they have different hit conditions and different failure modes.
- Provider prefix cache. The provider reuses the attention state of a byte-identical prompt prefix and bills the repeat at a fraction of the input price. The hit condition is exact at the token level, so it is always correct — the only cost is that your prompt layout must keep stable content first.
- Application result cache. Your own key-value store keyed on the request. It skips the model entirely on an exact repeat, which is correct by construction and is the best value when requests genuinely repeat.
- Semantic cache. Embed the request, look up the nearest cached request, and return its answer when the similarity passes a threshold. It buys the highest hit rate, because paraphrases count as hits — and it is the only layer that can be wrong. A near-miss returns a plausible answer to a question nobody asked, and the user cannot tell.
The demo counts exactly that. It runs the same request stream through all three layers, then adds a wrong-answer counter to the semantic layer, driven by a near-miss rate you can move. The hit-rate bar rises with the rate; so does the damage.
Hit rates for the three layers, then what the semantic layer did to cost, latency and correctness.
Routing, cascades, and streaming
Cheap-first is not free; perceived latency is a real latency
Routing sends each request to a model chosen for it: easy traffic cheap, hard traffic expensive. A cascade is the special case where the cheap model tries first and escalates on a signal — low confidence, a failed check, a retry. Cascades are popular because they sound like getting the expensive model's quality at the cheap model's price, and that is true only under a condition worth stating precisely: every escalated request pays for both calls. A cascade that escalates most of its traffic is not a cheap path with an expensive fallback; it is the expensive path with a tax.
Move the escalation share below and watch where the cascade's cost crosses the cost of simply calling the good model. The crossover sits where the cheap call stops being a saving and becomes a surcharge. Latency crosses over much earlier, because an escalated request waits for the cheap model before it waits for the good one.
Left: cost per request. Right: latency per request. Both compare a cascade with calling the good model directly.
Streaming is the last lever, and it is about perception rather than throughput. The end-to-end time to a complete answer does not change when you stream; what changes is the time to the first useful token, and users judge responsiveness by that. Speculative UX pushes the same idea further: show a skeleton or a draft that is likely to be right while the model works, then reconcile. Both are ways of spending engineering effort on the latency the user feels rather than the latency you can measure at the API boundary.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| What is a timeout for? | Bounding a stall, so a slow dependency cannot consume the request's latency budget or hide its own failure. |
| Why does backoff need jitter? | Without it, every client retries on the same tick and the dependency is hit by a synchronised wave. |
| What does a circuit breaker buy? | It fails fast once failures are consecutive, instead of queueing retries against something already down. |
| What is a degraded mode? | A pre-decided fallback — smaller model, cached answer, or an honest "try later" — so "unavailable" has a defined behaviour. |
| Which cost lever comes first? | Shorten the output. Output is the expensive side at every vendor in the dated table. |
| And last? | Optimise the prompt — real work, smallest saving, so it goes after caching, routing and batching. |
| Which cache is always correct? | Provider prefix and application result caches: both key on an exact match. Only staleness can hurt you. |
| What does a semantic cache cost? | Correctness. A near-miss returns a plausible answer to a question nobody asked, invisibly to the user. |
| When does a cascade cost more? | When the escalation share is high enough that the cheap call becomes a surcharge on top of the expensive one. |
| When does cascade latency cross over? | Much earlier than cost — every escalated request waits for the cheap model before the good one. |
| Does streaming make the model faster? | No. It shortens time-to-first-token, which is the part users actually perceive. |
Further reading
- Anthropic, Prompt caching — cache-read and cache-write multipliers and the 5-minute TTL, which is why prompt layout is a cost decision.
- OpenAI, Prompt caching guide, October 2024 — automatic prefix caching, and reporting cached input tokens back to you.
- Nygard, Release It!, 2nd edition, 2018 — circuit breakers, bulkheads and the stability patterns this part borrows wholesale.
- Dean & Barroso, "The Tail at Scale", Communications of the ACM, 2013 — why the 99th percentile dominates a fan-out, and why hedged requests and jitter exist.
- Brooker, "Timeouts, retries and backoff with jitter", Amazon Builders' Library — the practical version of the reliability half of this part.
- Khattab et al., "Don't Break the Cache", 2024 — prompt structure and prefix-cache hit rates measured end to end.