Tool use: the loop, and tool design as the craft
Last part ended with a decision: the task genuinely has no writable path, so the model owns the control flow. For the model to act on anything, it needs tools, and tools are not a prompt trick — they are an interface you design. This part walks the loop turn by turn, then treats tool design as the lever it is: names, descriptions and typed parameters are what the model chooses between; overlapping tools split the probability mass; errors are data you write for the model, not for a stack trace; and everything a tool returns is spent from the context budget, forever.
The loop is two message types
Reason, tool_use, tool_result, repeat
An agent turn has exactly two new message shapes. The model emits a tool_use block: a typed request naming a tool and an argument object. Your code reads it, executes the corresponding function — or refuses to — and appends a tool_result block keyed to the same call id. The model then runs again with that result in context. That is the whole loop. There is no hidden channel: the tool's output reaches the model only because your code put it back into the messages.
Two things follow, and both are easy to get wrong. First, the model never runs anything. It proposes a call; the decision to execute it is a policy your code owns, which is where permissions, idempotency and rate limits live. Second, a failing tool is not an exception that ends the loop — it is a tool_result like any other, and the loop continues. What the model does next depends entirely on how you phrased that result.
Step through the loop below. Each node is one message; the JSON panel shows the actual block that would be appended at that point.
The track is one task. The highlighted node is the message currently being produced; the panel shows its block.
while that appends two message types. Everything you can control about an agent — what it can do, what it sees, when it stops — is a decision in that while, not a property of the model.Schemas are the interface, and descriptions are prompts
The model chooses from what you describe, not from what you built
A tool definition has three parts: a name, a description, and a typed input schema. The model never sees your implementation — it sees only those three strings. A description is therefore a prompt in the strictest sense: it is the text the model conditions on when it decides which tool to call and how to fill the arguments. “Search” is a prompt. “Search the documentation corpus by keyword and return the top passages; use this for facts that appear verbatim in the docs, and use ask_user when the request is ambiguous” is a better one.
The schema does two jobs at once. It documents the argument types for the model, and it is a contract your code validates before executing anything. A well-typed schema with an enum where the choices are finite removes a whole class of malformed calls before they reach your function.
Then there is the design principle that matters most and is least intuitive: keep the tool set few and non-overlapping. Two tools that plausibly answer the same request split the probability mass between them, and the model's error rate rises with the overlap, not with the raw count. Drag the controls below — with no overlap, a larger set is survivable; with high overlap, even a handful of tools becomes a coin flip.
Selection accuracy against the overlap coefficient. The marker is your current setting; the dashed line is the analytic curve for the chosen tool count.
Parallel calls, partial failure, and errors as data
The model may batch; your code must survive it
When the calls are independent, the model will often emit several tool_use blocks in a single turn — a search and two reads, say, rather than a search, then a read, then a read. That is a latency win: the calls can run concurrently and the results return together. It is also a failure-handling problem, because a batch can fail partially. One tool succeeds, one times out, one is rejected by a permission check. The loop must still return a result for every call id in the turn, or the next model request is malformed.
How you describe those failures is the highest-leverage text you write all day. A raw stack trace tells the model nothing it can act on: it does not know whether the failure was transient, whether the argument was invalid, or whether retrying would help. A structured error — a type, the offending field, the constraint, and a suggested next action — is content the model can reason about, and it turns a dead end into a branch.
The same failure, returned two ways. Bars are seeded recovery rates over 800 trials with the chosen number of retries.
Two habits make the difference. Make side-effecting tools idempotent where you can — pass an idempotency key so a retried write is a no-op rather than a duplicate — and separate reads from writes so the model can explore freely and commit deliberately. And return errors in the same structured shape every time, so the model learns the contract rather than guessing at prose.
Tool returns are spent forever
The context window is not a data bus
A tool result appended to the transcript does not just cost tokens once. Because the model is stateless, that result is re-sent on every subsequent turn, so a fat return is paid for again at every step of the loop. Ten steps that each stuff four thousand tokens of raw rows into context are not a ten-thousand-token problem — they are a growing-transcript problem, and the bill compounds exactly the way the whole volume has warned.
The fix is the same as everywhere else in context engineering: return less, on purpose. Paginate and let the model ask for the next page. Trim to the fields the model needs, not the fields the database has. Summarise a large result behind a smaller handle the model can expand. Crucially, the model's own note about a result survives better than the result itself: “searched eval harness, found five passages, the relevant one says the harness runs five cases” costs a fraction of the five passages.
Below, each tool return is counted for real, and each step's context is checked against the window. Watch the overflow appear — it is a property of the return size, not the task.
Context composition per step. The dashed line is the usable budget (window minus the output reserve); a marked column has overflowed.
The training-side counterpart — why a model is taught to emit structured calls at all, and how function-calling is evaluated — is owned by Part 14 of the training guide. This page assumes the capability and concerns itself with the interface you hand it.
Cheat sheet
| Question | The answer that keeps the loop alive |
|---|---|
| What does the model actually do for a tool call? | It emits a typed tool_use block. Your code decides whether and how to execute it. |
| How does the result get back to the model? | Your code appends a tool_result keyed to the call id and calls the model again. |
| What is a tool description, really? | A prompt. It is the only thing the model sees when choosing between tools. |
| How many tools should I expose? | As few as the task needs, and none that overlap. Overlap splits the probability mass. |
| What does the model get when a tool fails? | Whatever you put in the tool_result. Make it a structured error, not a stack trace. |
| What is a good error shape? | Type, field, constraint, suggested fix — written for the model to branch on. |
| How do I handle a batch of calls? | Return a result for every call id, even the failed ones. Run independent calls concurrently. |
| What does a big return cost? | Every subsequent turn of the loop, because the transcript is re-sent. Paginate, trim, summarise. |
Further reading
- Anthropic, “Writing effective tools for AI agents — using AI agents”, September 2025 — tool descriptions as prompts, consolidation, and returning token-efficient results.
- Anthropic, Tool use overview, docs — the
tool_use/tool_resultmessage contract and parallel tool use. - Schick et al., “Toolformer: Language Models Can Teach Themselves to Use Tools”, arXiv:2302.04761, 2023 — the origin story of learned tool invocation.
- Patil et al., “Gorilla: Large Language Model Connected with Massive APIs”, arXiv:2305.15334, 2023 — why API documentation quality is a first-class variable.
- Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models”, arXiv:2210.03629, 2022 — the interleaved reason/act loop this part implements.
- Berkeley Function Calling Leaderboard, gorilla.cs.berkeley.edu — function-selection and argument-correctness scores, and the failure modes behind them.