Workflows before agents
An agent is the most expensive, hardest-to-debug thing you can build out of an LLM call, so the first question is never “how do I build an agent?” but “does this task need one?”. Most products that call themselves agentic are workflows: a fixed code path through one or more model calls. This part names the five workflow patterns, gives you a gate for deciding whether the path really is unknowable, and shows the arithmetic behind the claim that a bounded task is better served by two deterministic steps than by a twenty-step loop.
Two things people both call an agent
A code path you can read, versus a loop the model directs
A workflow is a predefined code path through LLM calls and tools: you decide the sequence, the branches and the exit conditions, and the model fills in the content of each step. An agent is a loop: the model chooses which tool to call next, decides when it has enough, and stops when it says so. The distinction is not about how clever the model is or how many tools are attached — it is about who owns the control flow. In a workflow, you do. In an agent, the model does.
That ownership is why the trade is so lopsided at the small end. A workflow’s behaviour can be read off the source, tested step by step, and reasoned about by someone who has never seen the prompt. An agent’s behaviour is a trajectory: it emerges at runtime from the model’s choices, so the unit of debugging is a trace, and the failure modes are the ones this volume spends three acts on.
The five workflow patterns
Pick the task shape, and the pattern follows
Anthropic’s December 2024 essay “Building effective agents” names five ways to wire LLM calls together without giving up control flow. They are not a ladder of sophistication; they are answers to different structural questions about a task. Most real systems are one or two of them glued along a boundary.
Click a task below. The track shows the pattern that fits it, and pressing play animates the calls in order. Watch how few of these need a model deciding what to do next — and note the last one, which genuinely does.
One node per model or tool call. The highlighted node is the call currently being made.
Two of the five deserve a second look because they are easy to confuse. Parallelisation has two flavours: sectioning breaks one task into independent subtasks run at once, while voting runs the same task several times and aggregates. Orchestrator-worker looks like sectioning but the subtasks are not known when you write the code — a lead call reads the task and decides how many workers to spawn and what each should do. That is still a workflow: a human designed the shape, and the orchestration step has a fixed contract, even though its output is dynamic.
The gate: complexity, value, viability, cost of error
Four questions before you write a loop
Four properties decide whether a bounded workflow or a self-directed loop is the better engineering tool. Path predictability — can you enumerate the steps in advance? Required flexibility — does the task change shape part-way through in ways you cannot foresee? Task value — is the outcome worth the extra calls, latency and surface area an agent brings? Cost of error — what does a wrong answer cost, and can it be caught mechanically?
The gate is deliberately biased toward the boring answer. If the path is predictable and an error is cheap, a workflow wins: there is nothing for a loop to figure out, and a bounded pipeline fails in bounded, countable ways. An agent earns its complexity only when the path cannot be specified and the loop’s self-correction is what produces the value — and even then, expensive errors demand an explicit check step rather than an unattended loop.
Each factor pushes the verdict left (toward a workflow) or right (toward an agent). The marker is the combined verdict.
Reliability is arithmetic, and it compounds
Why a twenty-step loop is not five times worse than a four-step one
If a step succeeds 95% of the time, a run of ten steps succeeds 0.9510 ≈ 0.60 of the time — and a run of twenty, ≈ 0.36. That is not a model quality problem; it is multiplication. The same arithmetic applies to a workflow, but a workflow’s step count is a design choice you make and can keep small, while an agent’s step count is decided at runtime by the model and is usually larger than you hoped.
For a bounded task, the comparison is therefore stark: a two-step workflow against a loop that happens to take a dozen. Drag the per-step reliability and the loop length and watch the two lines separate. The pale dashed line is the two-step workflow.
End-to-end success against run length. The marker is the loop length you chose; the dashed line is a two-step workflow at the same per-step reliability.
The same task, two ways
Orchestrator-worker against one big prompt
Take one task — “read a set of documents and write a report” — and run it two ways. The single-prompt version puts every document in one context and asks for the report: one call, one big prefill, one long generation. The orchestrator-worker version spends one call planning the sections, runs the section workers in parallel, then spends one call assembling their output. It makes far more calls and moves more tokens, but the workers run concurrently and each worker sees a much smaller context than the single prompt did.
Which wins depends on what you are paying for. The bars compare calls, tokens and wall-clock latency for the same task; the readout totals each side and names the winner per metric.
Three metrics, two approaches. Latency assumes the section workers run in parallel, so the loop’s cost is the slowest worker, not their sum.
Cheat sheet
| Question | The answer that keeps the system boring |
|---|---|
| Workflow or agent? | Workflow unless the path genuinely cannot be written down in advance. |
| Who owns control flow in a workflow? | Your code. The model fills in the content of each step. |
| Who owns it in an agent? | The model, subject to the tools and the harness you gave it. |
| The five workflow patterns? | Prompt chaining, routing, parallelisation, orchestrator-worker, evaluator-optimiser. |
| Routing versus an agent? | The router picks from a finite branch set you wrote; the agent picks the next action at runtime. |
| When is an agent worth it? | When the path is unknowable and self-correction is where the value lives. |
| Why is a long loop unreliable? | Success multiplies: 0.9520 ≈ 0.36, no matter how good each step looks. |
| What does an error cost buy you? | Demand for verification. Expensive errors argue for a check step, not an unattended loop. |
Further reading
- Anthropic, “Building effective agents”, December 2024 — the workflow-versus-agent split and the five patterns this part is built on.
- Anthropic, “Effective context engineering for AI agents”, 29 September 2025 — the write/select/compress/isolate taxonomy and the reliability arithmetic that follows from it.
- Cognition, “Don’t Build Multi-Agents”, June 2025 — the case for owning context and control flow rather than delegating them to a fleet.
- Anthropic, “How we built our multi-agent research system”, 13 June 2025 — an orchestrator-worker deployment with the token and latency bill measured, not estimated.
- Wu, Terry & Cai, “AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation”, 2022 — the framework view, useful for reading claims about “agents” critically.