Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Multimodal and voice: a budget with no slack

Roughly 500–1000 ms end-to-end, commonly cited as ~800 ms

A voice agent is a pipeline with a hard user-facing contract: the gap between the moment a person stops talking and the moment the agent starts talking back has to feel like a reply, not a lookup. That rules out most of what makes a text agent comfortable. There is no time to retrieve a large candidate set, no time to run a verification pass, and no time for a second model to read the first model's output. The budget is roughly 500–1000 ms end-to-end, commonly cited as about 800 ms, and it has to cover the whole loop: turn detection, ASR (speech to text), the language model's time to first token, TTS (text to speech), and the network in both directions.

The cited figures below are illustrative — a shape, not a spec — but the shape is the lesson. The four ordinary stages already consume the entire 800 ms: roughly 150 ms of ASR, 300 ms to first token, 250 ms of speech synthesis and 100 ms of network. That leaves nothing for the genuinely hard part. Turn detection — deciding that the user has finished speaking rather than merely paused — is where the variability lives. Be too eager and you interrupt a person mid-thought; be too slow and every reply feels sluggish. It is hard because human turn-taking is not a silence threshold: people pause to think, they backchannel ("mm-hm") while you are talking, and the agent's own voice must not be mistaken for input. Every millisecond a detector spends being careful is a millisecond added to the experience.

A waterfall of the realtime budget. Each bar starts where the previous stage ended; the shaded band is the 500–1000 ms conversational range; the dashed line is the ~800 ms target. Drag any stage past the target and the overrun turns red.

⚠️ The trap: treating the latency budget as a per-component target instead of a total one. Each stage can hit its number and the loop still overrun, because they add up — and turn detection is the term nobody budgeted for. The only honest way to design a voice agent is to measure the end-to-end gap in the conditions people actually use it: bad networks, interruptions, background noise.
💡 The durable idea: a realtime interface is a latency allocation problem. You decide how much of the gap each stage may spend before you choose any model, and then every design decision — how small the model, how short the answer, whether to stream — is made against that allocation.
2

Small models on the device

Distilled, quantised, and close to the user

The other half of the realtime story is not a faster network, it is not needing the network. A model that runs locally erases three costs at once: the round trip and the queue (latency), the per-token bill (cost), and the trip to someone else's servers (privacy). What it does not erase is the quality gap. Two techniques close most of it for a narrow job. Distillation trains a small student model to reproduce the outputs — or the reasoning traces — of a large teacher on your task, so the student inherits competence on that task rather than on everything. Quantisation stores weights at lower precision (4-bit or 8-bit) so the model fits in the memory of a phone or a laptop's NPU, at a small, measurable accuracy cost.

The productive framing is not "small versus large" but "which quality does this step need". Most agent traffic is easy: classify an intent, extract a field, decide whether to call a tool, reformat a result. A distilled model that clears the eval bar on that 80% of traffic is not a compromise — it is a product. The hard 20% escalates. That is a routing cascade: try the cheap local model first, check a confidence or verification signal, and hand the remainder to a large remote model. It is the same retrieve-wide-then-rerank-narrow shape from earlier parts, applied to model choice.

Quality against latency for the cloud-large, cloud-small and on-device options; bubble size is cost per million output tokens and the ring marks the selected row. The faint points are other points on the same curve — the tradeoff is a continuum, not three buckets.

💡 The durable idea: run the cheapest model that clears the eval bar for each step, and escalate on a measured signal. The eval set you built in Act II is what makes that decision safe; without it, "small enough" is a guess.
3

Protocols: MCP, A2A, AP2, x402

What they standardise, and what they cannot express

The protocol layer has been filling in fast. MCP (the Model Context Protocol) standardises how a model discovers and calls tools and resources, so a tool written once can be used by many clients; its spec revisions are dated, and the 2026 revision streamlines the core toward a stateless request/response model with explicit session handles. A2A standardises agent-to-agent communication: how one agent advertises its capabilities and delegates a task to another. AP2 and x402 standardise the money: agent-initiated payment mandates and authorisation (AP2), and an HTTP-native payment handshake (x402). Each one removes a class of bespoke integration work, and that is real.

None of them, however, standardises the things that actually break. A protocol can carry the bytes of a request; it cannot tell you whether the request expressed the user's intent, whether the counterparty deserves trust, who carries liability when the action was wrong, or what a promise means. Those are semantic and institutional problems wearing a technical costume. Two agents can exchange perfectly well-formed messages and still be in complete disagreement about what was agreed — because the message said "I will book the cheaper flight" and the verifier, the user and the airline each read a different commitment into it. Protocols move syntax and transport. The unsolved layer is meaning, and it will not be closed by a spec revision.

Each protocol standardises a different interface. The band underneath is what none of them standardises — and it is where the interesting failures live. Select a protocol to highlight it.

⚠️ The trap: assuming that because two systems speak the same protocol they share the same meaning. Interoperability of transport is not interoperability of intent, and an agent that pays a counterparty correctly is not thereby authorised, trusted or indemnified.
4

Reliability compounds

A good per-step number is a bad product

A text system fails one time. An agent fails at every step, and a task is only a success if every step succeeds. If each step is independent with success probability p, the chance the whole task succeeds is pn for an n-step task — and that exponent is brutal. A step that succeeds 95% of the time sounds excellent in isolation. Twenty of them in a row succeed 0.95²⁰ ≈ 0.36: the agent fails most of the tasks it attempts. This is arithmetic, not pessimism, and it explains why harness design — a step budget, loop detection, verification between steps, checkpoints you can resume from — is load-bearing rather than polish.

$$P_{\text{task}} \;=\; p^{\,n}$$

End-to-end success for an n-step task at per-step reliability p, assuming independence. It is the same compounding that makes a small per-step improvement worth more than it looks.

$$\frac{P_{p+\delta}}{P_p} \;=\; \left(1+\frac{\delta}{p}\right)^{n}$$

Raising the per-step rate by a small δ is amplified by the exponent. Moving 0.95 to 0.97 turns 0.36 into 0.54 — a per-step gain of two points, a task gain of eighteen.

Move the per-step reliability and watch the curve fall away, then note how much the dashed "improved by two points" line buys. The lesson is the same one the deployment parts made: the lever is rarely a bigger model — it is removing steps, verifying them, and retrying the ones you can.

End-to-end success pn against task length. The solid line is the current per-step rate; the dashed line is the same task with the per-step rate improved by two points.

💡 The durable idea: with an exponent in the path, step count is a first-class quality metric. A change that removes two steps or makes one of them verifiable can beat a model upgrade outright — and it is measurable on your own held-out set.
5

How to keep up without chasing

Fundamentals versus dated facts

The field resets its surface every few months and its foundations almost never. The practical skill is telling the two apart. Durable: the loop is still a loop, the context window is still a budget reassembled every turn, tools still need non-overlapping schemas and errors written for the model, retrieval still trades recall against precision, evaluation is still a labelled set plus a check, and cost is still arithmetic you can do on a napkin. Dated: prices, context-window sizes, benchmark scores, leaderboard positions and spec revisions. Every one of those changes, and every one of them is checkable — so memorising them is a way of being confidently wrong within a quarter.

A reading strategy beats a hype cycle. Prefer primary sources with dates — a spec revision, a model card, a paper — over a launch post or a thread interpreting one. Keep a dated reference card rather than a memorised leaderboard: record what each benchmark measures, how it is graded and its main weakness, and fill in live numbers with the date you read them. When a claim arrives — "X is dead", "Y solves prompt injection", "Z makes agents reliable" — restate it as a hypothesis and test it against your own held-out set, because that is the only thing your change can be measured on. And when you cite a volatile number in a design doc, write the date next to it; an undated figure is indistinguishable from a wrong one once the next release ships.

⚠️ The trap: following the surface. A team that rebuilds its stack around each new model or framework never accumulates the one asset that compounds — the eval set, the traces and the failure taxonomy that tell them what actually works for their task.
💡 The durable idea: invest in the parts that survive a model swap — the harness, the eval set, the traces, the failure taxonomy — and read the volatile parts with a date attached. The frontier moves; the fundamentals are the raft.

Cheat sheet

QuestionThe answer that keeps you oriented
What is the voice latency target?Roughly 500–1000 ms end-to-end, commonly cited as ~800 ms — a total budget, not a per-stage one.
What fills it?Turn detection, ASR, LLM time to first token, TTS and network. The four ordinary stages already spend the 800 ms.
What is the hard, variable part?Turn detection. Human turn-taking is not a silence threshold; interruptions and backchannels make it a moving target.
What buys on-device?Latency, privacy and cost. Distillation and quantisation close most of the quality gap for a narrow task.
Which model should I run?The smallest that clears the eval bar for that step; escalate the hard fraction in a routing cascade.
What does MCP standardise?How a model discovers and calls tools and resources. Spec revisions are dated.
What does A2A standardise?Agent-to-agent discovery and task handoff.
What do AP2 and x402 standardise?Agent payment mandates, authorisation and an HTTP-native payment handshake.
What can no protocol express?Intent, trust, liability and the semantics of a promise. Protocols move syntax; meaning is unsolved.
Why does reliability matter so much?Tasks require every step to succeed: 0.95²⁰ ≈ 0.36. Step count is a quality metric.
How do I keep up?Read primary sources with dates, keep a dated benchmark card, and measure claims on your own held-out set.
What compounds?The eval set, the traces and the failure taxonomy — the assets that survive a model swap.

Further reading

6

Check your understanding

0/5 answered