Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

What makes a task long-horizon

Many steps, many calls, a growing context, a real risk of drift

A task is long-horizon when it cannot be finished in one model call and the intermediate work has to be carried forward. Three properties define it. It takes many steps, so an error at step nine is inherited by every later step. It makes many tool calls, so the transcript accumulates raw results that are individually useful and collectively enormous. And its context grows monotonically, so the same prompt that was cheap at step one is expensive and diluted by step forty.

That last point is why long-horizon runs fail differently from short ones. A short task fails because the model answered badly. A long run fails because a fact was folded away, a plan silently changed, or the agent re-issued the same call for the tenth time — none of which is an exception you can catch. Reliability compounds arithmetically: at 95% per step, a twenty-step run finishes about 36% of the time (0.9520 ≈ 0.36). Raising per-step reliability is the model's job; keeping the run from needing perfection is the harness's.

Below is a sixty-step research run, drawn one stacked column per step, with a step budget you can set and a loop detector you can switch on. From step 18 the scripted agent re-issues the same search; switch detection on and the harness halts the run instead of paying for the rest of the budget.

One stacked column per step, by context segment. The grey block is budget the loop detector refused to spend.

💡 The durable idea: the unit of reliability is the step, not the answer. A harness earns its keep by making the run survivable — bounded context, bounded steps, bounded blast radius — so per-step reliability does not have to be near-perfect.
2

The harness's jobs

Six responsibilities the model does not have

The harness is everything around the model call: the loop, the context policy, the tool layer, the sandbox, and the point at which a human is consulted. Six responsibilities are load-bearing once a run gets long.

// the loop's contract, not the model's run step: assert step < budget && tokens < ceiling && wall < limit if loop_detector(history) -> halt(reason="repeat"), checkpoint() if context > ceiling -> fold(oldest_turns) // lossy: record what folded action = model(context) if risk(action) > threshold -> approval_gate(action) // blocks the run result = sandbox.run(action) checkpoint(step, notes, action, result)

The detector is the one most often missing. The demo below is a stuck agent: thirty identical tool calls in a row. An n-gram detector over the recent action signatures — or a state hash that is unchanged for k turns — catches it after a handful of repeats. The readout is the point: the wasted steps and the dollars they would have cost.

Thirty identical calls. The trace turns red where the detector fires; the bar underneath is the part of the run it saved.

3

Permission gates and the human in the loop

Classify by risk, then decide who approves

A long-horizon agent with tools has four kinds of capability, and they do not deserve the same trust. A read touches nothing the world can see. A write changes state, but inside a sandbox it is reversible. A spend moves money or quota. A send leaves your boundary — an email, a webhook, a posted message — and cannot be recalled. The gate policy is a threshold over that ordering: auto-approve everything below it, require explicit human approval above it.

The trap is that the safe-looking policy is not the cheap one. Gating reads adds human reviews without adding safety, and each review has a real cost in attention and latency. The defensible design gates the irreversible actions, keeps the reversible ones in a sandbox, and accepts that the human-in-the-loop cost is the price of the actions that need it. Move the threshold and watch the mix.

Each action is classified by risk, then routed: auto-approved, or blocked on a human. The bottom bar is the price of that decision per 1,000 actions.

⚠️ The gate is not a defense against injection. A gate decides whether a capability may fire; it does not decide whether the instruction that asked for it was authorised. An agent with a send tool and untrusted content in its context is one prompt away from using a legitimately granted capability for someone else's purpose.
4

Computer use, deep research, and why the harness is half the variance

The composite cases, and the empirical claim

Computer use is the purest long-horizon case: the tool is a screen. The loop is screenshot → decide → act, where the action is a click, a keystroke or a scroll, and the result is another screenshot. The token cost is dominated by images, the latency is dominated by round trips rather than by generation, and the error profile is nasty — a misclick is not a wrong answer, it is a wrong world state, and every later step reasons from it. OSWorld was built precisely to grade the end state rather than the transcript, and its scores move when the environment version changes, which is why an OSWorld number without a pinned environment is not a number.

Deep research is the composite case that exercises every harness responsibility at once: search → read → synthesise → cite, over dozens of sources. It is read-heavy, so it tolerates parallel fan-out; it is context-hungry, so folding is mandatory; it is citation-dependent, so an evicted fact is not just a redo, it is an uncited claim. Anthropic's multi-agent research system is the reference design, and it reports the honest cost: breadth was bought with roughly an order of magnitude more tokens.

Both cases point at the same conclusion, which is the part's sharpest claim. On SWE-bench Verified the scaffold — the harness, the tooling, the retry policy — contributes at least as much variance as the model does; on τ²-bench the score moves with the user simulator and the harness around it. So a public benchmark number is a statement about a model plus a harness, and reporting the model alone is misleading. The demo below takes one model and runs it under two harnesses. The gap between them is 15 points, decomposed into the three harness decisions that caused it — and it is the same size as the gap you get by changing the model.

Left: the same model under two harnesses, always 15 points apart. Right: where those 15 points come from.

Cheat sheet

QuestionThe answer that shapes the build
When is a task long-horizon?Many steps, many tool calls, a growing context, and a real risk of drift — not just a long prompt.
Why does reliability matter so much?It compounds: 0.95 per step over twenty steps is about 36% end to end.
What does context folding buy?A bounded window. It is lossy and adversarial: an evicted fact looks like one that never existed, so work is silently redone.
What is a step budget for?Converting an unbounded failure into a bounded one you can retry, escalate or report.
What is the commonest long-run failure?The same action, repeated. Detect repeated calls, n-grams or unchanged state hashes and halt the cycle.
What does sandboxing change?It makes write and execute steps reversible, so a bad step is recoverable.
Which actions should a gate block?The irreversible ones: spend and send. Reads are free; writes belong in a sandbox.
Does a gate stop injection?No. It gates a capability; it does not verify that the instruction authorising it was legitimate.
What does computer use add?A screenshot action space with a wrong-world-state failure mode, and image-dominated token cost.
Why is deep research the composite case?Read-heavy fan-out, mandatory folding, and citations that fail when a fact is evicted.
Why is a benchmark score not the model's?It is a model-plus-harness score; the scaffold contributes at least as much variance as the weights.

Further reading

5

Check your understanding

0/5 answered