Multi-agent: the argument
Two engineering teams published opposite conclusions within weeks of each other in 2025, and both were right. One built a research system out of an orchestrator and parallel subagents and reported it beating a single agent on breadth-first work, at a large token cost. The other argued that splitting an agent into subagents quietly breaks the thing that makes it good. This part puts the two claims side by side, gives you a slider for the property that decides between them, and marks the point where the fan-out stops paying.
Two camps, two conditions
Both published in June 2025, both correct
Anthropic's case. In "How we built our multi-agent research system" (13 June 2025), the team describes a lead agent that plans, spawns parallel subagents, and synthesises their results. On breadth-first research — cast a wide net over many sources, then pull the findings together — that orchestrator-worker design beat a single agent. The cost was reported plainly: multi-agent systems used roughly an order of magnitude more tokens than a plain chat interaction, because every subagent carries its own context and every finding is summarised on the way back.
Cognition's case. In "Don't Build Multi-Agents" (June 2025), the counter-argument is not that parallelism is slow — it is that it is incoherent. The claim is that actions carry implicit decisions. A subagent handed only a task description makes its own assumptions about conventions, formats and prior context, and two subagents working in parallel can make different assumptions that only collide at merge time. Their principle: share full context, not just the task — and their recommended default is a single-threaded linear agent, or a deliberately serial hand-off.
The rest of this part makes that line measurable. A single property — how tightly the subtasks are coupled through shared state — runs underneath every demo, and the reader is the one who finds the crossover.
The crossover
One slider decides the architecture
Think of coupling as the fraction of the work that is not independent: 0 is a set of read-only subtasks whose results are combined once at the end, 1 is a set of subtasks that all read and write the same evolving state. At low coupling the fan-out is a pure win — the subagents never need to agree, so their independence costs nothing. As coupling rises, two things happen at once: the single thread gets better, because one context sees every change and never has to reconcile two sets of assumptions, while the orchestrator-worker gets worse, because its subagents increasingly make decisions the others cannot see.
Both curves are modelled, not measured — the seeded dots are the "observations" the model is drawn through. Move the coupling slider and find the point where the orchestrator stops winning on quality. It is marked, but the interesting exercise is knowing why it sits where it does.
Left: task quality against coupling. Right: wall-clock latency. The quality crossover and the latency crossover are not in the same place.
The token multiplier
Why the fan-out is expensive even when it wins
Multi-agent's cost is not a rounding error: it is the price of breadth, and it scales with the number of agents, not with the difficulty of the question. Each subagent carries its own system prompt, its own growing transcript of explorations, and its own output; the orchestrator then pays again to dispatch work and to read back summaries. Meanwhile the single thread answering the same question with the same number of explorations has one context, so it pays for the system prompt once.
This is the honest half of the Anthropic result: the design won on breadth-first research and used far more tokens doing it. Below, each column is one subagent's token bill; the dashed line is what a single thread would send for the same subtasks. Multiply the agents and watch the multiplier move roughly one-for-one.
One column per agent, stacked by system prompt, explorations and output. The dashed line is the single-thread baseline.
Coordination overhead, drawn
The same steps, run two ways
Here is the Cognition objection as a trace. Both tracks run the same twelve steps. The top track is a single thread: every step sees every earlier step, so its assumptions stay consistent no matter how coupled the work is. The bottom track is an orchestrator with parallel workers: as coupling and headcount rise, steps start appearing twice (two workers solved the same sub-problem) and start conflicting (two workers solved it under different assumptions).
Redundancy is the cheap failure — it burns tokens. Conflict is the expensive one: it produces a merge that is internally inconsistent, and it is invisible in an aggregate metric. Move the coupling slider from the previous section and add workers; the conflict rate climbs faster than the redundancy rate, which is exactly why the quality curve bends before the latency curve does.
Top: one thread, one set of assumptions. Bottom: doubled steps are redundancy, crossed steps are a conflict.
The evidence board
Each claim, and the condition it holds under
Both claims are conditional, and the conditions are stated in the posts themselves. Anthropic's design is a research system: subtasks are searches and reads, and the merge is a synthesis step. Cognition's objection is about agents that edit shared files, drive a single conversation, or write code in one repository — where the implicit decisions are the work and must not diverge. Put the two boards side by side, then let the coupling slider and the read-heavy toggle cast the verdict.
The verdict is deliberately mechanical: independence and coupling decide it, not taste.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| What did Anthropic report? | An orchestrator-worker research system beating a single agent on breadth-first tasks, at roughly an order of magnitude more tokens (2025-06-13). |
| What did Cognition argue? | Actions carry implicit decisions; subagents given only a task description make conflicting assumptions. Share full context, or stay single-threaded (2025-06). |
| When does multi-agent win? | When subtasks are read-heavy and independent: parallel search, fan-out over sources, merged once at the end. |
| When does it lose? | When subtasks share mutable state or must produce one coherent voice — the coupling is the work. |
| What is the crossover? | The coupling value at which coordination cost exceeds the parallelism win on quality. On this model it sits near 0.40. |
| Why is quality before latency? | Conflicting assumptions damage results long before the merge becomes the slow part; latency stays ahead across the whole range. |
| What scales the bill? | The number of agents: each carries its own system prompt, transcript and output, plus orchestrator dispatch and summarisation. |
| What must you build first? | Tracing. An untraceable multi-agent run cannot be evaluated, debugged or attributed. |
Further reading
- Anthropic, "How we built our multi-agent research system", 13 June 2025 — the orchestrator-worker research system, its breadth-first win, and its reported token multiplier.
- Cognition, "Don't Build Multi-Agents", June 2025 — the implicit-decision argument, "share full context, not just the task", and the case for a single thread.
- Anthropic, "Building effective agents", December 2024 — the workflow patterns, including orchestrator-worker, and when not to reach for any of them.
- Cemri et al., "Why Do Multi-Agent LLM Systems Fail?", 2025 — a taxonomy of multi-agent failure modes, most of them coordination rather than capability.
- Yao et al., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024 — pass^k and the reliability gap that parallel agents widen.
- Anthropic, "Effective context engineering for AI agents", 29 September 2025 — context isolation as an explicit operation, which is what a subagent boundary really is.