Tool Use, Function Calling & Agents
A language model's knowledge is frozen at training time and cannot, by itself, look anything up, run code, or change the world. Tool use fixes that: the model emits a structured function call, a runtime executes it, and the result is fed back into the conversation. Done well, this turns a model into an agent that can search, calculate, and act. OLMo 3's Instruct models are explicitly trained for it, with a single unified format and newly released datasets. This part covers the format, the data, and how it is measured.
Why tools
Parametric knowledge vs retrieval
Ask a model a factual question and it answers from its weights — “parametric” knowledge. That knowledge is broad but stale, unverifiable, and unreliable for rare facts. A tool converts an open-book problem into a closed-book one: the model still has to know to search, formulate a query, and read the result, but the fact itself comes from a database. The effect is dramatic on factuality benchmarks. On SimpleQA (short, hard factual questions), OLMo 3 Instruct 7B jumps from 3.3% correct with no tools to 79.2% with a Google-search tool — a 76-point swing. Every Qwen baseline shows the same pattern.
On LitQA2 (literature QA) the story is more nuanced and worth internalizing. OLMo 3 Instruct 7B improves from 24.4% to 38.2% using the Asta Scientific Corpus tools, but Qwen 3 VL 8B and Qwen 2.5 7B actually lose 4–6 points when given the same tools, because they lean on parametric knowledge and get pulled off course by retrieval. Tools are a capability, not a free upgrade.
No-tools vs with-tools accuracy by model. Toggle the benchmark.
The function-calling format
One format to rule them all
Tool use is a syntax problem before it is a reasoning problem. The model needs to know what functions exist, emit a call that the runtime can parse, and recognize the reply. OLMo 3 unifies all of this:
- Tools are declared with the OpenAPI specification and injected into the system prompt.
- Function calls are written as pythonic code blocks, wrapped in XML tags inside the assistant message.
- Environment outputs come back in a dedicated environment role.
- The tokenizer's vocabulary is extended with special tokens for those tags. Interestingly, the Think model went the other way and avoided special tokens during midtraining; for Instruct tool use, dedicated tokens measurably helped.
The team's strongest finding here: unifying the format across every tool-use dataset was crucial for stable behavior. Mixing incompatible conventions is a common cause of unreliable tool use.
Build a call. Pick a function and fill its arguments to see the full round trip: system tool spec → assistant call → environment result.
Real vs simulated environments
Executable MCP servers vs LLM-simulated tools
Tool-use training data comes in two flavors, and they teach different things. Real trajectories execute against actual MCP (Model Context Protocol) servers — OLMo 3 uses the Asta Scientific Corpus server and the Serper search API. These are small but naturally complex, with many environment interactions per request, and they teach the model to cope with real, messy, sometimes-erroring outputs. Simulated trajectories (the SimFC dataset) are generated end-to-end by an LLM against a large pool of tool definitions, so they scale to hundreds of thousands of examples and tens of thousands of unique functions.
| Dataset | Environment | Trajectories | Unique functions | Multi-turn | Multi-step |
|---|
Multi-turn = several user turns per trajectory. Multi-step = several assistant–environment interactions per user request. Real data is smaller but richer in multi-step structure; simulated data wins on scale and function diversity.
Evaluating tool use
Intrinsic vs extrinsic
There are two separate questions. Intrinsic function calling asks: did the model choose the right function and fill the arguments correctly? That is what BFCLv3 (Berkeley Function Calling Leaderboard) measures, including multi-turn and multi-step scenarios. Extrinsic task completion asks: did the model actually solve the task with tools? OLMo 3 measures that on LitQA2 via the Asta MCP server and SimpleQA via Serper, allowing the agent at most 10 turns, sampling at temperature 0, and averaging three runs.
Mini BFCL scorer. Read the request, pick the function, fill the key argument, and check.
Illustrative: task accuracy if each tool call succeeds independently with probability p, versus the turn budget.
When tools help or hurt
A balanced view
- Retrieval-friendly tasks benefit hugely. Short factual questions and literature lookup are where tools shine, because the bottleneck is knowledge, not reasoning.
- Reasoning-heavy tasks can regress. A model that already knows the answer may be derailed by an imperfect retrieval; two of three Qwen baselines lose points on LitQA2 with tools.
- Format is the hidden variable. The same functions, declared or encoded differently, produce very different reliability. Unify the format before tuning anything else.
- MCP is the interface standard. The Model Context Protocol gives tools a common transport, which is what makes a single agent harness able to drive many different servers.
Cheat sheet
Recap
| Concept | What it is |
|---|---|
| Function calling | Model emits a structured call; runtime executes it; result returns to the context. |
| OpenAPI tool spec | Standard schema for declaring tool names, descriptions, and arguments in the system prompt. |
| Environment role | A dedicated message role that carries tool outputs back to the model, distinct from user/assistant. |
| Tool special tokens | Tokenizer tokens delimiting calls; unify the format or tool use becomes unstable. |
| MCP | Model Context Protocol — the transport standard for connecting agents to tools/servers. |
| Real trajectories | Executed against live MCP servers (Asta, Serper); smaller but richly multi-step. |
| SimFC | LLM-simulated trajectories: 200K examples, 42.6K unique functions, scalable. |
| BFCLv3 | Intrinsic function-calling benchmark (function choice + argument correctness). |
| Extrinsic completion | End-task accuracy with tools: LitQA2 via Asta, SimpleQA via Serper, ≤10 turns. |
| Multi-turn vs multi-step | Multiple user turns vs multiple tool calls within one user request. |
Further reading
References
- OLMo 3 Team (Ai2), "OLMo 3" (2025) — function-calling data and tool-use evaluation (§5.2.1, Tables 9–10).
- Patil et al., "BFCL: Berkeley Function Calling Leaderboard" (2024); BFCLv3.
- Liu et al., "xLAM: A Family of Large Action Models" (2024).
- Liu et al., "ToolACE: Winning the Points of LLM Function Calling" (2024).
- Xu et al., "DR Tulu" (2025) — search/browse agent data.
- Skarlinski et al., "Language agents achieve superhuman synthesis of scientific knowledge" (2024) — LitQA2.
- Wei et al., "Measuring short-form factuality in large language models" (2024) — SimpleQA.
- Asta Scientific Corpus MCP server: allenai.org/asta/resources/mcp.
- Model Context Protocol: modelcontextprotocol.io.
- OpenAI Agent SDK: openai.github.io/openai-agents-python.
- Serving consequence: LLM Serving, Part 18 covers what agent traces do to KV residency and pool sizing.