Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Why tools

Parametric knowledge vs retrieval

Ask a model a factual question and it answers from its weights — “parametric” knowledge. That knowledge is broad but stale, unverifiable, and unreliable for rare facts. A tool converts an open-book problem into a closed-book one: the model still has to know to search, formulate a query, and read the result, but the fact itself comes from a database. The effect is dramatic on factuality benchmarks. On SimpleQA (short, hard factual questions), OLMo 3 Instruct 7B jumps from 3.3% correct with no tools to 79.2% with a Google-search tool — a 76-point swing. Every Qwen baseline shows the same pattern.

On LitQA2 (literature QA) the story is more nuanced and worth internalizing. OLMo 3 Instruct 7B improves from 24.4% to 38.2% using the Asta Scientific Corpus tools, but Qwen 3 VL 8B and Qwen 2.5 7B actually lose 4–6 points when given the same tools, because they lean on parametric knowledge and get pulled off course by retrieval. Tools are a capability, not a free upgrade.

No-tools vs with-tools accuracy by model. Toggle the benchmark.

⚠️ Tools can hurt. Giving a model a tool it hasn't been trained to use well can lower accuracy. The format, the training trajectories, and the model's willingness to trust retrieval all have to line up.
2

The function-calling format

One format to rule them all

Tool use is a syntax problem before it is a reasoning problem. The model needs to know what functions exist, emit a call that the runtime can parse, and recognize the reply. OLMo 3 unifies all of this:

  • Tools are declared with the OpenAPI specification and injected into the system prompt.
  • Function calls are written as pythonic code blocks, wrapped in XML tags inside the assistant message.
  • Environment outputs come back in a dedicated environment role.
  • The tokenizer's vocabulary is extended with special tokens for those tags. Interestingly, the Think model went the other way and avoided special tokens during midtraining; for Instruct tool use, dedicated tokens measurably helped.

The team's strongest finding here: unifying the format across every tool-use dataset was crucial for stable behavior. Mixing incompatible conventions is a common cause of unreliable tool use.

Build a call. Pick a function and fill its arguments to see the full round trip: system tool spec → assistant call → environment result.

1 · tool spec (system)
2 · assistant call
3 · environment result
3

Real vs simulated environments

Executable MCP servers vs LLM-simulated tools

Tool-use training data comes in two flavors, and they teach different things. Real trajectories execute against actual MCP (Model Context Protocol) servers — OLMo 3 uses the Asta Scientific Corpus server and the Serper search API. These are small but naturally complex, with many environment interactions per request, and they teach the model to cope with real, messy, sometimes-erroring outputs. Simulated trajectories (the SimFC dataset) are generated end-to-end by an LLM against a large pool of tool definitions, so they scale to hundreds of thousands of examples and tens of thousands of unique functions.

DatasetEnvironmentTrajectoriesUnique functionsMulti-turnMulti-step

Multi-turn = several user turns per trajectory. Multi-step = several assistant–environment interactions per user request. Real data is smaller but richer in multi-step structure; simulated data wins on scale and function diversity.

4

Evaluating tool use

Intrinsic vs extrinsic

There are two separate questions. Intrinsic function calling asks: did the model choose the right function and fill the arguments correctly? That is what BFCLv3 (Berkeley Function Calling Leaderboard) measures, including multi-turn and multi-step scenarios. Extrinsic task completion asks: did the model actually solve the task with tools? OLMo 3 measures that on LitQA2 via the Asta MCP server and SimpleQA via Serper, allowing the agent at most 10 turns, sampling at temperature 0, and averaging three runs.

Mini BFCL scorer. Read the request, pick the function, fill the key argument, and check.

available functions

Illustrative: task accuracy if each tool call succeeds independently with probability p, versus the turn budget.

5

When tools help or hurt

A balanced view

  • Retrieval-friendly tasks benefit hugely. Short factual questions and literature lookup are where tools shine, because the bottleneck is knowledge, not reasoning.
  • Reasoning-heavy tasks can regress. A model that already knows the answer may be derailed by an imperfect retrieval; two of three Qwen baselines lose points on LitQA2 with tools.
  • Format is the hidden variable. The same functions, declared or encoded differently, produce very different reliability. Unify the format before tuning anything else.
  • MCP is the interface standard. The Model Context Protocol gives tools a common transport, which is what makes a single agent harness able to drive many different servers.
📌 Cross-links: function-calling SFT is a specialization of Part 7's instruction tuning; the multi-turn preference data that keeps agents coherent across turns is Part 15's territory; and the context window that holds a long trajectory is Part 13's extension stage.
✓

Cheat sheet

Recap

ConceptWhat it is
Function callingModel emits a structured call; runtime executes it; result returns to the context.
OpenAPI tool specStandard schema for declaring tool names, descriptions, and arguments in the system prompt.
Environment roleA dedicated message role that carries tool outputs back to the model, distinct from user/assistant.
Tool special tokensTokenizer tokens delimiting calls; unify the format or tool use becomes unstable.
MCPModel Context Protocol — the transport standard for connecting agents to tools/servers.
Real trajectoriesExecuted against live MCP servers (Asta, Serper); smaller but richly multi-step.
SimFCLLM-simulated trajectories: 200K examples, 42.6K unique functions, scalable.
BFCLv3Intrinsic function-calling benchmark (function choice + argument correctness).
Extrinsic completionEnd-task accuracy with tools: LitQA2 via Asta, SimpleQA via Serper, ≤10 turns.
Multi-turn vs multi-stepMultiple user turns vs multiple tool calls within one user request.
📚

Further reading

References

?

Check your understanding

0/5 answered
We've now seen the three post-training surfaces — reasoning, tools, and instruction — as if they were separate. The final part steps back and looks at the single recipe that produces all of them at scale. Continue: post-training at scale →