Cost and latency from first principles
Two numbers decide whether a feature ships: what it costs and how long it takes. Both follow from arithmetic you can write on a napkin, and both are governed by the same asymmetry — output tokens are the expensive side, and the loop around the model is what makes them multiply. This part builds a cost model, a latency model, the quadratic curve of an agent loop, and the effort dial that trades thinking for both.
Output is the expensive side
One formula, and where the lever is
The bill for one call is input × pin + output × pout, plus a discounted rate for any cached input. The first thing to notice is that pout is larger than pin at every vendor, by a factor of three to five in the historical list prices below. The second is that the output side is the one you can usually shrink: a retrieved passage is a fixed cost of the task, but the model's own verbosity, its reasoning traces and its retries are choices. Shortening the output is almost always the highest-leverage cost edit.
The calculator uses the historical pricing table. Pick a model, set the input and output tokens, and decide how much of the input is a cache read. The stacked bars break the per-call cost into uncached input, cached input and output, and repeat it across calls and retries for the whole task.
Cost per call and per task, split into uncached input, cached input and output.
TTFT, TPOT, and end to end
Two latencies, not one
Users feel two different delays. Time to first token (TTFT) is how long the system waits before any text appears: the request queueing, then prefill processing the whole prompt to build the first token's state. Time per output token (TPOT) is the cadence of the stream after that — the decode loop, one token at a time, bounded by memory bandwidth rather than arithmetic. End to end is TTFT + output × TPOT, and the shape matters: a long prompt hurts TTFT, a long answer hurts the decode tail, and prefix caching attacks the prefill term without touching either.
The timeline below starts at the request and ends at the last token. Adjust the queue, the prompt length, the cached fraction, the prefill rate and the per-token decode time, and watch where the time actually goes.
Request to last token: queue, prefill, then decode. TTFT ends where prefill does; end to end is the full bar.
An agent loop is quadratic in turns
The prefix is re-sent every turn
An agent alternates a model call with a tool call, and each call sends the whole growing transcript. Turn five does not pay for turn five's new text; it pays for turns one through five, re-sent. Over N turns with a steady per-turn addition h, the cumulative input is roughly h·N(N+1)/2 — quadratic, not linear. That single fact explains why long-running agents are expensive out of proportion to the work they appear to do, and why every harness eventually invents compaction.
Prefix caching is the brake. If the opening of the prompt is byte-identical across turns, a provider can reuse the cached computation and bill the repeat at a fraction of the input price. It does not change the shape — the input is still triangular — but it cuts the constant dramatically, and it is the reason prompt layout (stable content first, volatile content last) is an architectural decision.
Cumulative billed input tokens over twenty turns. Solid = cold every turn; dashed = a warm prefix cache.
Effort is a dial
Diminishing returns, with a knee
Reasoning models expose an effort control: a budget for how many thinking tokens to spend before answering. More thinking helps on hard problems and wastes money on easy ones, and the benefit decays. The curve below models that: accuracy rises quickly at first and then flattens toward a ceiling, while cost grows linearly in reasoning tokens. The knee marks where the marginal return drops below roughly one accuracy point per thousand tokens — past it you are buying very little.
Move the effort slider to spend tokens and watch both curves; change the task difficulty to see the knee move. The cost side uses the output price of the model selected in the first demo, because thinking tokens are output tokens.
Accuracy (solid) and normalised cost (dashed) versus reasoning tokens spent, with the knee marked.
Cheat sheet
| Question | Short answer |
|---|---|
| What is the cost formula? | input × p_in + output × p_out, with cached input at a discounted rate. |
| Which side is expensive? | Output. It costs 3×–5× input at every vendor in the historical table. |
| What is the usual lever? | Shorten the output: verbosity, reasoning traces and retries are all output tokens. |
| TTFT versus TPOT? | TTFT is queue plus prefill to the first token; TPOT is the per-token decode cadence after it. |
| End to end? | TTFT + output × TPOT. Long prompts hurt TTFT; long answers hurt the tail. |
| Why is an agent loop expensive? | Each turn re-sends the whole transcript, so cumulative input is triangular — quadratic in turns. |
| What does prefix caching do? | Cuts the price of a byte-identical prefix. The input is still triangular; the constant shrinks. |
| Are reasoning tokens special? | No. They are output tokens: billed as output and added to end-to-end latency. |
| How much effort should I spend? | Enough to clear the knee. Beyond it, marginal accuracy per dollar collapses. |
Further reading
- Anthropic, Prompt caching, 2024 — cache-write and cache-read multipliers, and the TTL that makes prompt layout a cost decision.
- OpenAI, Prompt caching guide, October 2024 — automatic prefix caching and the
cached_tokensaccounting used in the loop chart. - DeepSeek, API pricing, 2025 — the cache-hit / cache-miss / output split that makes the asymmetry concrete.
- Anthropic, Context management, 2025 — compaction and memory as first-class tools for long-running agents.
- Wei et al., "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models", NeurIPS 2022 (arXiv:2201.11903) — the source of the "thinking helps, and helps less as tasks get easier" intuition behind the effort curve.
- Snell et al., "Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters", 2024 (arXiv:2408.03314) — the empirical case that test-time compute has strongly diminishing returns and is best allocated per difficulty.