Long-horizon agents and harness design
A one-shot prompt is a function call. A sixty-step agent run is a control system, and the thing doing the controlling is not the model. Once a task spans many steps and many tool calls, the context grows, small mistakes stop being independent, and the run can drift or fall into a loop without ever raising an error. This part is about the harness around the model — context folding, step budgets, loop detection, sandboxing, permission gates, checkpointing — and about the uncomfortable finding that the harness moves benchmark scores about as much as swapping the model does.
What makes a task long-horizon
Many steps, many calls, a growing context, a real risk of drift
A task is long-horizon when it cannot be finished in one model call and the intermediate work has to be carried forward. Three properties define it. It takes many steps, so an error at step nine is inherited by every later step. It makes many tool calls, so the transcript accumulates raw results that are individually useful and collectively enormous. And its context grows monotonically, so the same prompt that was cheap at step one is expensive and diluted by step forty.
That last point is why long-horizon runs fail differently from short ones. A short task fails because the model answered badly. A long run fails because a fact was folded away, a plan silently changed, or the agent re-issued the same call for the tenth time — none of which is an exception you can catch. Reliability compounds arithmetically: at 95% per step, a twenty-step run finishes about 36% of the time (0.9520 ≈ 0.36). Raising per-step reliability is the model's job; keeping the run from needing perfection is the harness's.
Below is a sixty-step research run, drawn one stacked column per step, with a step budget you can set and a loop detector you can switch on. From step 18 the scripted agent re-issues the same search; switch detection on and the harness halts the run instead of paying for the rest of the budget.
One stacked column per step, by context segment. The grey block is budget the loop detector refused to spend.
The harness's jobs
Six responsibilities the model does not have
The harness is everything around the model call: the loop, the context policy, the tool layer, the sandbox, and the point at which a human is consulted. Six responsibilities are load-bearing once a run gets long.
- Context folding (compaction). Keep the live window bounded as history grows. Folding is lossy and adversarial: an evicted fact is indistinguishable from one that never existed, so the agent silently redoes work.
- Step budgets. A hard cap on steps, tool calls, tokens and wall-clock. A budget does not make the agent smarter; it converts an unbounded failure into a bounded one you can retry, escalate or report.
- Loop and repetition detection. The single most common long-run failure is not a wrong action, it is the same action. Detect repeated tool calls, repeated state hashes or repeated low-information turns and break the cycle.
- Sandboxing. Give write and execution capability only inside a disposable domain — a container, a scratch branch, a filesystem jail — so a bad step is recoverable rather than permanent.
- Permission gates. Classify actions by risk and require approval for the irreversible ones. Read is free; spend and send are not.
- Checkpointing and resumption. Persist state at step boundaries so a crashed, halted or escalated run can resume instead of restarting — and so a halted run is still an inspectable artifact.
The detector is the one most often missing. The demo below is a stuck agent: thirty identical tool calls in a row. An n-gram detector over the recent action signatures — or a state hash that is unchanged for k turns — catches it after a handful of repeats. The readout is the point: the wasted steps and the dollars they would have cost.
Thirty identical calls. The trace turns red where the detector fires; the bar underneath is the part of the run it saved.
Permission gates and the human in the loop
Classify by risk, then decide who approves
A long-horizon agent with tools has four kinds of capability, and they do not deserve the same trust. A read touches nothing the world can see. A write changes state, but inside a sandbox it is reversible. A spend moves money or quota. A send leaves your boundary — an email, a webhook, a posted message — and cannot be recalled. The gate policy is a threshold over that ordering: auto-approve everything below it, require explicit human approval above it.
The trap is that the safe-looking policy is not the cheap one. Gating reads adds human reviews without adding safety, and each review has a real cost in attention and latency. The defensible design gates the irreversible actions, keeps the reversible ones in a sandbox, and accepts that the human-in-the-loop cost is the price of the actions that need it. Move the threshold and watch the mix.
Each action is classified by risk, then routed: auto-approved, or blocked on a human. The bottom bar is the price of that decision per 1,000 actions.
Computer use, deep research, and why the harness is half the variance
The composite cases, and the empirical claim
Computer use is the purest long-horizon case: the tool is a screen. The loop is screenshot → decide → act, where the action is a click, a keystroke or a scroll, and the result is another screenshot. The token cost is dominated by images, the latency is dominated by round trips rather than by generation, and the error profile is nasty — a misclick is not a wrong answer, it is a wrong world state, and every later step reasons from it. OSWorld was built precisely to grade the end state rather than the transcript, and its scores move when the environment version changes, which is why an OSWorld number without a pinned environment is not a number.
Deep research is the composite case that exercises every harness responsibility at once: search → read → synthesise → cite, over dozens of sources. It is read-heavy, so it tolerates parallel fan-out; it is context-hungry, so folding is mandatory; it is citation-dependent, so an evicted fact is not just a redo, it is an uncited claim. Anthropic's multi-agent research system is the reference design, and it reports the honest cost: breadth was bought with roughly an order of magnitude more tokens.
Both cases point at the same conclusion, which is the part's sharpest claim. On SWE-bench Verified the scaffold — the harness, the tooling, the retry policy — contributes at least as much variance as the model does; on τ²-bench the score moves with the user simulator and the harness around it. So a public benchmark number is a statement about a model plus a harness, and reporting the model alone is misleading. The demo below takes one model and runs it under two harnesses. The gap between them is 15 points, decomposed into the three harness decisions that caused it — and it is the same size as the gap you get by changing the model.
Left: the same model under two harnesses, always 15 points apart. Right: where those 15 points come from.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| When is a task long-horizon? | Many steps, many tool calls, a growing context, and a real risk of drift — not just a long prompt. |
| Why does reliability matter so much? | It compounds: 0.95 per step over twenty steps is about 36% end to end. |
| What does context folding buy? | A bounded window. It is lossy and adversarial: an evicted fact looks like one that never existed, so work is silently redone. |
| What is a step budget for? | Converting an unbounded failure into a bounded one you can retry, escalate or report. |
| What is the commonest long-run failure? | The same action, repeated. Detect repeated calls, n-grams or unchanged state hashes and halt the cycle. |
| What does sandboxing change? | It makes write and execute steps reversible, so a bad step is recoverable. |
| Which actions should a gate block? | The irreversible ones: spend and send. Reads are free; writes belong in a sandbox. |
| Does a gate stop injection? | No. It gates a capability; it does not verify that the instruction authorising it was legitimate. |
| What does computer use add? | A screenshot action space with a wrong-world-state failure mode, and image-dominated token cost. |
| Why is deep research the composite case? | Read-heavy fan-out, mandatory folding, and citations that fail when a fact is evicted. |
| Why is a benchmark score not the model's? | It is a model-plus-harness score; the scaffold contributes at least as much variance as the weights. |
Further reading
- Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?", ICLR 2024 — the scaffold-dependence that makes a bare model score misleading.
- Yao et al., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024 — pass^k, and a benchmark whose score moves with the harness.
- Xie et al., "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments", 2024 — computer use graded on the end state, with a pinned environment.
- Mialon et al., "GAIA: a benchmark for General AI Assistants", 2023 — multi-step assistant work, and how answer-string grading hides partial trajectories.
- Anthropic, "How we built our multi-agent research system", 13 June 2025 — deep research as the reference composite case, and its token multiplier.
- Anthropic, "Effective context engineering for AI agents", 29 September 2025 — write/select/compress/isolate, the taxonomy behind context folding.
- Debenedetti et al., "Defeating Prompt Injections by Design", 2025 — capability-based information flow, and why a permission gate is not an injection defense.