Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Two camps, two conditions

Both published in June 2025, both correct

Anthropic's case. In "How we built our multi-agent research system" (13 June 2025), the team describes a lead agent that plans, spawns parallel subagents, and synthesises their results. On breadth-first research — cast a wide net over many sources, then pull the findings together — that orchestrator-worker design beat a single agent. The cost was reported plainly: multi-agent systems used roughly an order of magnitude more tokens than a plain chat interaction, because every subagent carries its own context and every finding is summarised on the way back.

Cognition's case. In "Don't Build Multi-Agents" (June 2025), the counter-argument is not that parallelism is slow — it is that it is incoherent. The claim is that actions carry implicit decisions. A subagent handed only a task description makes its own assumptions about conventions, formats and prior context, and two subagents working in parallel can make different assumptions that only collide at merge time. Their principle: share full context, not just the task — and their recommended default is a single-threaded linear agent, or a deliberately serial hand-off.

💡 The reconciliation, in one line: multi-agent wins when the subtasks are read-heavy and independent — parallel search, fan-out over sources, results merged at the end. It loses when subtasks need shared mutable state or a single coherent voice, because then the "implicit decisions" are the work.

The rest of this part makes that line measurable. A single property — how tightly the subtasks are coupled through shared state — runs underneath every demo, and the reader is the one who finds the crossover.

2

The crossover

One slider decides the architecture

Think of coupling as the fraction of the work that is not independent: 0 is a set of read-only subtasks whose results are combined once at the end, 1 is a set of subtasks that all read and write the same evolving state. At low coupling the fan-out is a pure win — the subagents never need to agree, so their independence costs nothing. As coupling rises, two things happen at once: the single thread gets better, because one context sees every change and never has to reconcile two sets of assumptions, while the orchestrator-worker gets worse, because its subagents increasingly make decisions the others cannot see.

Both curves are modelled, not measured — the seeded dots are the "observations" the model is drawn through. Move the coupling slider and find the point where the orchestrator stops winning on quality. It is marked, but the interesting exercise is knowing why it sits where it does.

Left: task quality against coupling. Right: wall-clock latency. The quality crossover and the latency crossover are not in the same place.

3

The token multiplier

Why the fan-out is expensive even when it wins

Multi-agent's cost is not a rounding error: it is the price of breadth, and it scales with the number of agents, not with the difficulty of the question. Each subagent carries its own system prompt, its own growing transcript of explorations, and its own output; the orchestrator then pays again to dispatch work and to read back summaries. Meanwhile the single thread answering the same question with the same number of explorations has one context, so it pays for the system prompt once.

This is the honest half of the Anthropic result: the design won on breadth-first research and used far more tokens doing it. Below, each column is one subagent's token bill; the dashed line is what a single thread would send for the same subtasks. Multiply the agents and watch the multiplier move roughly one-for-one.

One column per agent, stacked by system prompt, explorations and output. The dashed line is the single-thread baseline.

⚠️ The multiplier is not the only cost: reliability compounds too. A twenty-step run at 95% per-step reliability succeeds about 36% of the time end to end (0.9520 ≈ 0.36). Ten parallel agents each at 95% do not average that away; they multiply the number of places a wrong assumption can enter.
4

Coordination overhead, drawn

The same steps, run two ways

Here is the Cognition objection as a trace. Both tracks run the same twelve steps. The top track is a single thread: every step sees every earlier step, so its assumptions stay consistent no matter how coupled the work is. The bottom track is an orchestrator with parallel workers: as coupling and headcount rise, steps start appearing twice (two workers solved the same sub-problem) and start conflicting (two workers solved it under different assumptions).

Redundancy is the cheap failure — it burns tokens. Conflict is the expensive one: it produces a merge that is internally inconsistent, and it is invisible in an aggregate metric. Move the coupling slider from the previous section and add workers; the conflict rate climbs faster than the redundancy rate, which is exactly why the quality curve bends before the latency curve does.

Top: one thread, one set of assumptions. Bottom: doubled steps are redundancy, crossed steps are a conflict.

5

The evidence board

Each claim, and the condition it holds under

Both claims are conditional, and the conditions are stated in the posts themselves. Anthropic's design is a research system: subtasks are searches and reads, and the merge is a synthesis step. Cognition's objection is about agents that edit shared files, drive a single conversation, or write code in one repository — where the implicit decisions are the work and must not diverge. Put the two boards side by side, then let the coupling slider and the read-heavy toggle cast the verdict.

The verdict is deliberately mechanical: independence and coupling decide it, not taste.

⚠️ The eval problem: you cannot evaluate a system you cannot trace. An orchestrator-worker run that produced a wrong answer is a tree of contexts, each contributing a fragment; if you only kept the final message, you have no way to tell a bad retrieval from a bad synthesis from a bad assumption in subagent three. Budget for tracing before you budget for agents.

Cheat sheet

QuestionThe answer that shapes the build
What did Anthropic report?An orchestrator-worker research system beating a single agent on breadth-first tasks, at roughly an order of magnitude more tokens (2025-06-13).
What did Cognition argue?Actions carry implicit decisions; subagents given only a task description make conflicting assumptions. Share full context, or stay single-threaded (2025-06).
When does multi-agent win?When subtasks are read-heavy and independent: parallel search, fan-out over sources, merged once at the end.
When does it lose?When subtasks share mutable state or must produce one coherent voice — the coupling is the work.
What is the crossover?The coupling value at which coordination cost exceeds the parallelism win on quality. On this model it sits near 0.40.
Why is quality before latency?Conflicting assumptions damage results long before the merge becomes the slow part; latency stays ahead across the whole range.
What scales the bill?The number of agents: each carries its own system prompt, transcript and output, plus orchestrator dispatch and summarisation.
What must you build first?Tracing. An untraceable multi-agent run cannot be evaluated, debugged or attributed.

Further reading

6

Check your understanding

0/5 answered