The frontier, and what's unsolved
Every part before this one was about a system you can measure: a loop, a tool, a context budget, an eval set. This part is about the edge, where the measurements get harder. Voice and multimodal agents live or die inside a latency budget that has no slack. Small models are moving onto the device, changing what "good enough" costs. Protocols are standardising tools, agent-to-agent handoff and payments, and yet cannot express the things that actually go wrong. And reliability compounds, so a per-step number that looks fine is a product that fails. We finish on the one durable skill in a field that resets every few months: telling the fundamentals apart from the dated facts, and reading rather than chasing.
Multimodal and voice: a budget with no slack
Roughly 500–1000 ms end-to-end, commonly cited as ~800 ms
A voice agent is a pipeline with a hard user-facing contract: the gap between the moment a person stops talking and the moment the agent starts talking back has to feel like a reply, not a lookup. That rules out most of what makes a text agent comfortable. There is no time to retrieve a large candidate set, no time to run a verification pass, and no time for a second model to read the first model's output. The budget is roughly 500–1000 ms end-to-end, commonly cited as about 800 ms, and it has to cover the whole loop: turn detection, ASR (speech to text), the language model's time to first token, TTS (text to speech), and the network in both directions.
The cited figures below are illustrative — a shape, not a spec — but the shape is the lesson. The four ordinary stages already consume the entire 800 ms: roughly 150 ms of ASR, 300 ms to first token, 250 ms of speech synthesis and 100 ms of network. That leaves nothing for the genuinely hard part. Turn detection — deciding that the user has finished speaking rather than merely paused — is where the variability lives. Be too eager and you interrupt a person mid-thought; be too slow and every reply feels sluggish. It is hard because human turn-taking is not a silence threshold: people pause to think, they backchannel ("mm-hm") while you are talking, and the agent's own voice must not be mistaken for input. Every millisecond a detector spends being careful is a millisecond added to the experience.
A waterfall of the realtime budget. Each bar starts where the previous stage ended; the shaded band is the 500–1000 ms conversational range; the dashed line is the ~800 ms target. Drag any stage past the target and the overrun turns red.
Small models on the device
Distilled, quantised, and close to the user
The other half of the realtime story is not a faster network, it is not needing the network. A model that runs locally erases three costs at once: the round trip and the queue (latency), the per-token bill (cost), and the trip to someone else's servers (privacy). What it does not erase is the quality gap. Two techniques close most of it for a narrow job. Distillation trains a small student model to reproduce the outputs — or the reasoning traces — of a large teacher on your task, so the student inherits competence on that task rather than on everything. Quantisation stores weights at lower precision (4-bit or 8-bit) so the model fits in the memory of a phone or a laptop's NPU, at a small, measurable accuracy cost.
The productive framing is not "small versus large" but "which quality does this step need". Most agent traffic is easy: classify an intent, extract a field, decide whether to call a tool, reformat a result. A distilled model that clears the eval bar on that 80% of traffic is not a compromise — it is a product. The hard 20% escalates. That is a routing cascade: try the cheap local model first, check a confidence or verification signal, and hand the remainder to a large remote model. It is the same retrieve-wide-then-rerank-narrow shape from earlier parts, applied to model choice.
Quality against latency for the cloud-large, cloud-small and on-device options; bubble size is cost per million output tokens and the ring marks the selected row. The faint points are other points on the same curve — the tradeoff is a continuum, not three buckets.
Protocols: MCP, A2A, AP2, x402
What they standardise, and what they cannot express
The protocol layer has been filling in fast. MCP (the Model Context Protocol) standardises how a model discovers and calls tools and resources, so a tool written once can be used by many clients; its spec revisions are dated, and the 2026 revision streamlines the core toward a stateless request/response model with explicit session handles. A2A standardises agent-to-agent communication: how one agent advertises its capabilities and delegates a task to another. AP2 and x402 standardise the money: agent-initiated payment mandates and authorisation (AP2), and an HTTP-native payment handshake (x402). Each one removes a class of bespoke integration work, and that is real.
None of them, however, standardises the things that actually break. A protocol can carry the bytes of a request; it cannot tell you whether the request expressed the user's intent, whether the counterparty deserves trust, who carries liability when the action was wrong, or what a promise means. Those are semantic and institutional problems wearing a technical costume. Two agents can exchange perfectly well-formed messages and still be in complete disagreement about what was agreed — because the message said "I will book the cheaper flight" and the verifier, the user and the airline each read a different commitment into it. Protocols move syntax and transport. The unsolved layer is meaning, and it will not be closed by a spec revision.
Each protocol standardises a different interface. The band underneath is what none of them standardises — and it is where the interesting failures live. Select a protocol to highlight it.
Reliability compounds
A good per-step number is a bad product
A text system fails one time. An agent fails at every step, and a task is only a success if every step succeeds. If each step is independent with success probability p, the chance the whole task succeeds is pn for an n-step task — and that exponent is brutal. A step that succeeds 95% of the time sounds excellent in isolation. Twenty of them in a row succeed 0.95²⁰ ≈ 0.36: the agent fails most of the tasks it attempts. This is arithmetic, not pessimism, and it explains why harness design — a step budget, loop detection, verification between steps, checkpoints you can resume from — is load-bearing rather than polish.
End-to-end success for an n-step task at per-step reliability p, assuming independence. It is the same compounding that makes a small per-step improvement worth more than it looks.
Raising the per-step rate by a small δ is amplified by the exponent. Moving 0.95 to 0.97 turns 0.36 into 0.54 — a per-step gain of two points, a task gain of eighteen.
Move the per-step reliability and watch the curve fall away, then note how much the dashed "improved by two points" line buys. The lesson is the same one the deployment parts made: the lever is rarely a bigger model — it is removing steps, verifying them, and retrying the ones you can.
End-to-end success pn against task length. The solid line is the current per-step rate; the dashed line is the same task with the per-step rate improved by two points.
How to keep up without chasing
Fundamentals versus dated facts
The field resets its surface every few months and its foundations almost never. The practical skill is telling the two apart. Durable: the loop is still a loop, the context window is still a budget reassembled every turn, tools still need non-overlapping schemas and errors written for the model, retrieval still trades recall against precision, evaluation is still a labelled set plus a check, and cost is still arithmetic you can do on a napkin. Dated: prices, context-window sizes, benchmark scores, leaderboard positions and spec revisions. Every one of those changes, and every one of them is checkable — so memorising them is a way of being confidently wrong within a quarter.
A reading strategy beats a hype cycle. Prefer primary sources with dates — a spec revision, a model card, a paper — over a launch post or a thread interpreting one. Keep a dated reference card rather than a memorised leaderboard: record what each benchmark measures, how it is graded and its main weakness, and fill in live numbers with the date you read them. When a claim arrives — "X is dead", "Y solves prompt injection", "Z makes agents reliable" — restate it as a hypothesis and test it against your own held-out set, because that is the only thing your change can be measured on. And when you cite a volatile number in a design doc, write the date next to it; an undated figure is indistinguishable from a wrong one once the next release ships.
Cheat sheet
| Question | The answer that keeps you oriented |
|---|---|
| What is the voice latency target? | Roughly 500–1000 ms end-to-end, commonly cited as ~800 ms — a total budget, not a per-stage one. |
| What fills it? | Turn detection, ASR, LLM time to first token, TTS and network. The four ordinary stages already spend the 800 ms. |
| What is the hard, variable part? | Turn detection. Human turn-taking is not a silence threshold; interruptions and backchannels make it a moving target. |
| What buys on-device? | Latency, privacy and cost. Distillation and quantisation close most of the quality gap for a narrow task. |
| Which model should I run? | The smallest that clears the eval bar for that step; escalate the hard fraction in a routing cascade. |
| What does MCP standardise? | How a model discovers and calls tools and resources. Spec revisions are dated. |
| What does A2A standardise? | Agent-to-agent discovery and task handoff. |
| What do AP2 and x402 standardise? | Agent payment mandates, authorisation and an HTTP-native payment handshake. |
| What can no protocol express? | Intent, trust, liability and the semantics of a promise. Protocols move syntax; meaning is unsolved. |
| Why does reliability matter so much? | Tasks require every step to succeed: 0.95²⁰ ≈ 0.36. Step count is a quality metric. |
| How do I keep up? | Read primary sources with dates, keep a dated benchmark card, and measure claims on your own held-out set. |
| What compounds? | The eval set, the traces and the failure taxonomy — the assets that survive a model swap. |
Further reading
- Anthropic, "Building effective agents", December 2024 — the workflow-versus-agent split that still frames which frontier capability you actually need.
- Anthropic, "Effective context engineering for AI agents", 29 September 2025 — the write/select/compress/isolate taxonomy, and why the loop outlives any single model.
- Anthropic, "How we built our multi-agent research system", 13 June 2025 — a worked account of delegation, cost and the harness around a long-horizon agent.
- Model Context Protocol, specification and revision history, 2024-11-05 through 2026 — the dated revisions behind the tool-interface section.
- Hinton, Vinyals & Dean, "Distilling the Knowledge in a Neural Network", 2015 — the original teacher/student distillation idea behind the small-model shift.
- Yao et al., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024 — pass^k, the reliability counterpart to pass@k, and the arithmetic this part leans on.