Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Patterns for a probabilistic dependency

Timeouts, backoff, breakers, bulkheads, degraded modes

The moment a call can hang, fail, or return slowly, the client needs the same defenses it would need against any unreliable RPC. A timeout is a deadline with a purpose: without one, a slow dependency consumes the request's whole latency budget and hides its own failure. Exponential backoff with jitter makes a retry cheap when the dependency recovers and rare when it does not, and the jitter is not decoration — it is what stops every client retrying on the same tick. A circuit breaker counts consecutive failures and, past a threshold, fails fast for a cooldown instead of queueing retries against something that is already down. A bulkhead gives the dependency its own pool so a stall cannot consume every worker in the process. And an explicit degraded mode decides in advance what "unavailable" means: a smaller model, a cached answer, or an honest "try again in a minute".

The simulator below runs 140 seeded requests through timeout, retry and breaker policy. Raise the retry count and average latency improves until the timeout is the binding constraint; lower the breaker threshold and watch the run trip into degraded mode earlier. The important readout is the error rate after the fallback, not before it.

Top: the outcome of each seeded request. Bottom: its latency, with the timeout drawn as the dashed ceiling.

💡 The durable idea: decide what "down" means before it happens. A timeout, a breaker and a stated degraded mode convert an unbounded stall into a bounded, observable failure with a known fallback — and the fallback is what the user actually experiences.
2

The ordered cost-lever list

Do the big things first, in this order

Cost optimisation usually fails because it is attempted in the wrong order: teams spend a week hand-tuning a prompt to shave 8% of input tokens while the output is four times the price of the input and twice as long as it needs to be. The levers are not equal, and they are not independent. Applied in order, from largest to smallest and cheapest to hardest:

  1. Shorten the output. Output is the expensive side at every vendor — three to five times the input price in the dated pricing table this page cites. A tighter answer format, a cap on the completion, or a template instead of prose moves the biggest number first.
  2. Cache the prefix. Stable content first, volatile content last, so the provider's prefix cache hits and the repeated input is billed at a fraction of the rate.
  3. Route easy traffic to a smaller model. Most requests do not need the frontier model; a cheap classifier or a confidence rule sends the easy majority down a cheaper path.
  4. Batch. Offline and bulk work trades latency for a discount, often a large one.
  5. Optimise the prompt. Finally — and only after the arithmetic above — trim the instructions and examples. It is real work and it is the smallest lever, which is why it goes last.

The waterfall applies the levers in that order to a baseline monthly bill. Move the request volume, move the share of easy traffic, and step through the levers; the final bar is the bill you would actually pay.

Baseline, then each lever's saving as a step down, then the final bill. Levers apply in order — the labels are the order.

⚠️ The numbers are dated and move. The prices below are historical list prices, not current quotes; the order of the levers is the durable part, not the dollar amounts. Verify against live vendor pages before you put any of these figures in a budget.
3

Three cache layers, and the correctness cost of semantic caching

Prefix, result, semantic — one of them can be wrong

Caching an LLM application happens at three different layers, and they have different hit conditions and different failure modes.

The demo counts exactly that. It runs the same request stream through all three layers, then adds a wrong-answer counter to the semantic layer, driven by a near-miss rate you can move. The hit-rate bar rises with the rate; so does the damage.

Hit rates for the three layers, then what the semantic layer did to cost, latency and correctness.

💡 The durable idea: a cache is a correctness decision disguised as a cost decision. Prefix and result caches are exact, so their only risk is staleness. A semantic cache trades a measurable hit rate for an unmeasurable-by-the-user wrong-answer rate, so it needs a similarity threshold you have tuned against labelled near-misses — not one you guessed.
4

Routing, cascades, and streaming

Cheap-first is not free; perceived latency is a real latency

Routing sends each request to a model chosen for it: easy traffic cheap, hard traffic expensive. A cascade is the special case where the cheap model tries first and escalates on a signal — low confidence, a failed check, a retry. Cascades are popular because they sound like getting the expensive model's quality at the cheap model's price, and that is true only under a condition worth stating precisely: every escalated request pays for both calls. A cascade that escalates most of its traffic is not a cheap path with an expensive fallback; it is the expensive path with a tax.

Move the escalation share below and watch where the cascade's cost crosses the cost of simply calling the good model. The crossover sits where the cheap call stops being a saving and becomes a surcharge. Latency crosses over much earlier, because an escalated request waits for the cheap model before it waits for the good one.

Left: cost per request. Right: latency per request. Both compare a cascade with calling the good model directly.

Streaming is the last lever, and it is about perception rather than throughput. The end-to-end time to a complete answer does not change when you stream; what changes is the time to the first useful token, and users judge responsiveness by that. Speculative UX pushes the same idea further: show a skeleton or a draft that is likely to be right while the model works, then reconcile. Both are ways of spending engineering effort on the latency the user feels rather than the latency you can measure at the API boundary.

💡 The durable idea: optimise the latency the user perceives. Time-to-first-token, streaming and a speculative draft all reduce the interval between asking and seeing something useful, without pretending to make the model faster.

Cheat sheet

QuestionThe answer that shapes the build
What is a timeout for?Bounding a stall, so a slow dependency cannot consume the request's latency budget or hide its own failure.
Why does backoff need jitter?Without it, every client retries on the same tick and the dependency is hit by a synchronised wave.
What does a circuit breaker buy?It fails fast once failures are consecutive, instead of queueing retries against something already down.
What is a degraded mode?A pre-decided fallback — smaller model, cached answer, or an honest "try later" — so "unavailable" has a defined behaviour.
Which cost lever comes first?Shorten the output. Output is the expensive side at every vendor in the dated table.
And last?Optimise the prompt — real work, smallest saving, so it goes after caching, routing and batching.
Which cache is always correct?Provider prefix and application result caches: both key on an exact match. Only staleness can hurt you.
What does a semantic cache cost?Correctness. A near-miss returns a plausible answer to a question nobody asked, invisibly to the user.
When does a cascade cost more?When the escalation share is high enough that the cheap call becomes a surcharge on top of the expensive one.
When does cascade latency cross over?Much earlier than cost — every escalated request waits for the cheap model before the good one.
Does streaming make the model faster?No. It shortens time-to-first-token, which is the part users actually perceive.

Further reading

5

Check your understanding

0/5 answered