Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Output is the expensive side

One formula, and where the lever is

The bill for one call is input × pin + output × pout, plus a discounted rate for any cached input. The first thing to notice is that pout is larger than pin at every vendor, by a factor of three to five in the historical list prices below. The second is that the output side is the one you can usually shrink: a retrieved passage is a fixed cost of the task, but the model's own verbosity, its reasoning traces and its retries are choices. Shortening the output is almost always the highest-leverage cost edit.

The calculator uses the historical pricing table. Pick a model, set the input and output tokens, and decide how much of the input is a cache read. The stacked bars break the per-call cost into uncached input, cached input and output, and repeat it across calls and retries for the whole task.

Cost per call and per task, split into uncached input, cached input and output.

💡 Durable idea: reasoning and thinking tokens are output tokens. Every "let the model think harder" switch is a direct multiplier on the most expensive line item, which is why the effort dial at the end of this page is a cost decision as much as a quality one.
2

TTFT, TPOT, and end to end

Two latencies, not one

Users feel two different delays. Time to first token (TTFT) is how long the system waits before any text appears: the request queueing, then prefill processing the whole prompt to build the first token's state. Time per output token (TPOT) is the cadence of the stream after that — the decode loop, one token at a time, bounded by memory bandwidth rather than arithmetic. End to end is TTFT + output × TPOT, and the shape matters: a long prompt hurts TTFT, a long answer hurts the decode tail, and prefix caching attacks the prefill term without touching either.

The timeline below starts at the request and ends at the last token. Adjust the queue, the prompt length, the cached fraction, the prefill rate and the per-token decode time, and watch where the time actually goes.

Request to last token: queue, prefill, then decode. TTFT ends where prefill does; end to end is the full bar.

⚠️ Trap: streaming hides latency, it does not remove it. The user still waits for TTFT, and a slow decode makes the last token arrive long after the first — which is exactly the delay that decides whether a realtime feature feels realtime.
3

An agent loop is quadratic in turns

The prefix is re-sent every turn

An agent alternates a model call with a tool call, and each call sends the whole growing transcript. Turn five does not pay for turn five's new text; it pays for turns one through five, re-sent. Over N turns with a steady per-turn addition h, the cumulative input is roughly h·N(N+1)/2 — quadratic, not linear. That single fact explains why long-running agents are expensive out of proportion to the work they appear to do, and why every harness eventually invents compaction.

Prefix caching is the brake. If the opening of the prompt is byte-identical across turns, a provider can reuse the cached computation and bill the repeat at a fraction of the input price. It does not change the shape — the input is still triangular — but it cuts the constant dramatically, and it is the reason prompt layout (stable content first, volatile content last) is an architectural decision.

Cumulative billed input tokens over twenty turns. Solid = cold every turn; dashed = a warm prefix cache.

💡 Carry this forward: the loop length is a design parameter. Fewer, larger steps and aggressive compaction attack the quadratic directly; caching attacks its price, not its shape.
4

Effort is a dial

Diminishing returns, with a knee

Reasoning models expose an effort control: a budget for how many thinking tokens to spend before answering. More thinking helps on hard problems and wastes money on easy ones, and the benefit decays. The curve below models that: accuracy rises quickly at first and then flattens toward a ceiling, while cost grows linearly in reasoning tokens. The knee marks where the marginal return drops below roughly one accuracy point per thousand tokens — past it you are buying very little.

Move the effort slider to spend tokens and watch both curves; change the task difficulty to see the knee move. The cost side uses the output price of the model selected in the first demo, because thinking tokens are output tokens.

Accuracy (solid) and normalised cost (dashed) versus reasoning tokens spent, with the knee marked.

⚠️ Trap: "reasoning tokens" are not free scratch paper. They are billed as output, they lengthen end-to-end latency, and on an easy task they can reduce accuracy by inviting the model to second-guess a correct answer. Default to low effort and raise it per task.

Cheat sheet

QuestionShort answer
What is the cost formula?input × p_in + output × p_out, with cached input at a discounted rate.
Which side is expensive?Output. It costs 3×–5× input at every vendor in the historical table.
What is the usual lever?Shorten the output: verbosity, reasoning traces and retries are all output tokens.
TTFT versus TPOT?TTFT is queue plus prefill to the first token; TPOT is the per-token decode cadence after it.
End to end?TTFT + output × TPOT. Long prompts hurt TTFT; long answers hurt the tail.
Why is an agent loop expensive?Each turn re-sends the whole transcript, so cumulative input is triangular — quadratic in turns.
What does prefix caching do?Cuts the price of a byte-identical prefix. The input is still triangular; the constant shrinks.
Are reasoning tokens special?No. They are output tokens: billed as output and added to end-to-end latency.
How much effort should I spend?Enough to clear the knee. Beyond it, marginal accuracy per dollar collapses.

Further reading

5

Check your understanding

0/5 answered