Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The loop is two message types

Reason, tool_use, tool_result, repeat

An agent turn has exactly two new message shapes. The model emits a tool_use block: a typed request naming a tool and an argument object. Your code reads it, executes the corresponding function — or refuses to — and appends a tool_result block keyed to the same call id. The model then runs again with that result in context. That is the whole loop. There is no hidden channel: the tool's output reaches the model only because your code put it back into the messages.

Two things follow, and both are easy to get wrong. First, the model never runs anything. It proposes a call; the decision to execute it is a policy your code owns, which is where permissions, idempotency and rate limits live. Second, a failing tool is not an exception that ends the loop — it is a tool_result like any other, and the loop continues. What the model does next depends entirely on how you phrased that result.

// the loop, in the shape your code actually writes it messages = [system, user] while True: reply = llm(messages, tools=TOOLS) if reply.stop_reason == "tool_use": results = [run(c) for c in reply.tool_calls] # your policy, your side effects messages += [reply, tool_results(results)] else: return reply.text

Step through the loop below. Each node is one message; the JSON panel shows the actual block that would be appended at that point.

The track is one task. The highlighted node is the message currently being produced; the panel shows its block.

💡 The durable idea: the loop is not magic, it is a while that appends two message types. Everything you can control about an agent — what it can do, what it sees, when it stops — is a decision in that while, not a property of the model.
2

Schemas are the interface, and descriptions are prompts

The model chooses from what you describe, not from what you built

A tool definition has three parts: a name, a description, and a typed input schema. The model never sees your implementation — it sees only those three strings. A description is therefore a prompt in the strictest sense: it is the text the model conditions on when it decides which tool to call and how to fill the arguments. “Search” is a prompt. “Search the documentation corpus by keyword and return the top passages; use this for facts that appear verbatim in the docs, and use ask_user when the request is ambiguous” is a better one.

The schema does two jobs at once. It documents the argument types for the model, and it is a contract your code validates before executing anything. A well-typed schema with an enum where the choices are finite removes a whole class of malformed calls before they reach your function.

Then there is the design principle that matters most and is least intuitive: keep the tool set few and non-overlapping. Two tools that plausibly answer the same request split the probability mass between them, and the model's error rate rises with the overlap, not with the raw count. Drag the controls below — with no overlap, a larger set is survivable; with high overlap, even a handful of tools becomes a coin flip.

Selection accuracy against the overlap coefficient. The marker is your current setting; the dashed line is the analytic curve for the chosen tool count.

⚠️ The trap: adding a tool is free to write and expensive to own. Every definition is tokens in every request, and every near-duplicate is a permanent chance to pick wrong. If two tools overlap, merge them or delete one — that is a cheaper fix than any amount of prompt tuning.
3

Parallel calls, partial failure, and errors as data

The model may batch; your code must survive it

When the calls are independent, the model will often emit several tool_use blocks in a single turn — a search and two reads, say, rather than a search, then a read, then a read. That is a latency win: the calls can run concurrently and the results return together. It is also a failure-handling problem, because a batch can fail partially. One tool succeeds, one times out, one is rejected by a permission check. The loop must still return a result for every call id in the turn, or the next model request is malformed.

How you describe those failures is the highest-leverage text you write all day. A raw stack trace tells the model nothing it can act on: it does not know whether the failure was transient, whether the argument was invalid, or whether retrying would help. A structured error — a type, the offending field, the constraint, and a suggested next action — is content the model can reason about, and it turns a dead end into a branch.

The same failure, returned two ways. Bars are seeded recovery rates over 800 trials with the chosen number of retries.

Two habits make the difference. Make side-effecting tools idempotent where you can — pass an idempotency key so a retried write is a no-op rather than a duplicate — and separate reads from writes so the model can explore freely and commit deliberately. And return errors in the same structured shape every time, so the model learns the contract rather than guessing at prose.

💡 Write the error message for its reader. The reader is a model that will decide the next action from your text alone. Type, field, constraint, fix — four lines of JSON beat forty lines of traceback, every time.
4

Tool returns are spent forever

The context window is not a data bus

A tool result appended to the transcript does not just cost tokens once. Because the model is stateless, that result is re-sent on every subsequent turn, so a fat return is paid for again at every step of the loop. Ten steps that each stuff four thousand tokens of raw rows into context are not a ten-thousand-token problem — they are a growing-transcript problem, and the bill compounds exactly the way the whole volume has warned.

The fix is the same as everywhere else in context engineering: return less, on purpose. Paginate and let the model ask for the next page. Trim to the fields the model needs, not the fields the database has. Summarise a large result behind a smaller handle the model can expand. Crucially, the model's own note about a result survives better than the result itself: “searched eval harness, found five passages, the relevant one says the harness runs five cases” costs a fraction of the five passages.

Below, each tool return is counted for real, and each step's context is checked against the window. Watch the overflow appear — it is a property of the return size, not the task.

Context composition per step. The dashed line is the usable budget (window minus the output reserve); a marked column has overflowed.

⚠️ Foreshadowing Part 4: “the context window is not a data bus” is the rule that motivates code execution and skills. If a step would move ten thousand tokens through the window just to compute one number, let code compute the number and put only the number in context.

The training-side counterpart — why a model is taught to emit structured calls at all, and how function-calling is evaluated — is owned by Part 14 of the training guide. This page assumes the capability and concerns itself with the interface you hand it.

Cheat sheet

QuestionThe answer that keeps the loop alive
What does the model actually do for a tool call?It emits a typed tool_use block. Your code decides whether and how to execute it.
How does the result get back to the model?Your code appends a tool_result keyed to the call id and calls the model again.
What is a tool description, really?A prompt. It is the only thing the model sees when choosing between tools.
How many tools should I expose?As few as the task needs, and none that overlap. Overlap splits the probability mass.
What does the model get when a tool fails?Whatever you put in the tool_result. Make it a structured error, not a stack trace.
What is a good error shape?Type, field, constraint, suggested fix — written for the model to branch on.
How do I handle a batch of calls?Return a result for every call id, even the failed ones. Run independent calls concurrently.
What does a big return cost?Every subsequent turn of the loop, because the transcript is re-sent. Paginate, trim, summarise.

Further reading

5

Check your understanding

0/5 answered