MCP, skills, and code execution
An agent is only as useful as the things you connect to it, and two connector stories now dominate: the Model Context Protocol for tools and data, and Agent Skills for know-how. Both arrive with the same hidden bill — everything you expose has to be described somewhere, and the description is paid for in the model's context. This part follows a protocol exchange lane by lane, then measures what tool definitions actually cost, then shows the loading trick that keeps skills small, and ends on the rule that saves the most tokens of all: stop ferrying data through the window and let code do it.
The protocol, lane by lane
Host, client, server — and a stateless core
The Model Context Protocol (MCP) is a wire format for letting a model talk to tools and data sources through a uniform interface instead of bespoke glue per integration. Anthropic introduced it in November 2024 and donated it to the Linux Foundation; the spec is versioned by date. The revisions you will see named are 2024-11-05, 2025-03-26 and 2025-06-18, with a 2026 revision that streamlines the core toward a stateless request/response model — explicit session handles rather than per-connection state — so a server can be scaled horizontally behind a load balancer. Treat any specific revision as checkable against modelcontextprotocol.io; clients and servers lag the spec by months.
There are three roles, and keeping them apart is the whole mental model. The host is your application, the thing holding the model and its context. The client is the protocol endpoint inside the host: it speaks JSON-RPC, matches request ids to responses, and owns the session. The server is what exposes capabilities — tools, resources, prompts — over that protocol. The host never speaks MCP; the client does.
Step through the exchange below. The three lanes are the three roles; each arrow is one JSON-RPC message, and the animation is only a walk through the handshake an agent performs before it can do anything at all.
One arrow per message. Without session state, every request could be served by a different replica.
Tool definitions are a system-prompt tax
N servers × M tools × schema tokens, before the user speaks
Here is the part of MCP that surprises people. The tools/list result is not a catalogue the model browses on demand; the client hands it to the model as a tool-definition block that sits in the prompt on every call. A tool is not just a name. It is a name, a description written for the model, and a JSON Schema for its arguments — and a faithful schema for one non-trivial tool can run to several hundred tokens once you describe every parameter, enum and nested object.
So the cost is multiplicative. Five servers exposing a dozen tools each, at two or three hundred schema tokens per tool, is fifteen thousand tokens of standing instruction before the user types a word. Those tokens are billed as input on every turn. They compete with the retrieved documents and the transcript for the same budget. And they are the reason a "small" agent that connects to ten services starts to feel slow and expensive without any of its actual work getting harder.
Drag the sliders below. The stacked column is the prompt against a fixed window; the line is where the output reserve begins. The readout reports the point at which the user's own query no longer fits — the number of servers at which you have spent the window on describing what you could do.
Tool schemas are a line item like any other. The interesting number is the one at which the query gets squeezed out.
Agent Skills and SKILL.md's three stages
Metadata always, instructions on invocation, resources on demand
Tools tell an agent what it can do; skills tell it how your team does things. An Agent Skill is a folder with a SKILL.md file — a name, a description, and a body of instruction that may reference scripts and other files. The design trick is not the file format but the progressive disclosure: the skill is loaded in three stages, and only the first is unconditional.
The metadata — name and description, a few dozen tokens — is the only part resident in context all the time, so the model knows the skill exists and when it might apply. The instructions, the body of SKILL.md, are pulled in only when the skill is invoked. The resources it points at — scripts, templates, reference tables — are loaded only when a step actually needs them, and often they are never loaded into the model at all because they are executed rather than read.
That is what makes a library of skills affordable: twenty skills can be resident for the price of the metadata, because nineteen of them are sitting on disk unread. The comparison below puts progressive disclosure next to the naive alternative — loading every skill in full — against a small agent's window. Toggle how many skills are invoked and whether their resources are loaded.
Two columns, same skill library: progressive disclosure versus loading everything. The dashed line is the window.
The context window is not a data bus
Let code move the data; bring back only the answer
Once an agent has tools, the tempting design is to use the context window as the channel between them: call a tool, paste the result into the transcript, call the next tool, paste that result, and so on. Each hop is a full round trip, each result is billed as input on every subsequent turn, and the intermediate values you will never look at again accumulate forever. Ten steps of a data-shuffling analysis can put more tokens through the window than the final answer is worth by two orders of magnitude.
The alternative is to treat code as the tool. Give the agent a sandbox with your data reachable and let it write one program that does the whole job — filter, join, aggregate, compare — and return only the result. The intermediate rows, the fifty megabytes of JSON, the loop over a thousand files: none of it enters the context, because none of it needs to. This is also exactly how a skill's resources work: a reference table is not read into the prompt, it is queried.
The comparison below runs the same multi-step analysis both ways. Increase the number of steps or the size of each intermediate result and watch the sequential path's token count bend quadratic while the code path barely moves.
Three measures, two designs. Each bar is normalised within its own group; the labels carry the real values.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| What are MCP's three roles? | Host (your app), client (the JSON-RPC endpoint inside it), server (what exposes tools, resources, prompts). |
| What changed in the 2026 spec revision? | The core moves toward stateless request/response with explicit session handles, so servers scale horizontally. Verify the revision against the spec. |
| Where do tool definitions live? | In the system prompt of every call — name, description and JSON Schema, billed as input every time. |
| How do you size the bill? | Servers × tools per server × average schema tokens, against the window you actually have left after the output reserve. |
| How is a SKILL.md loaded? | Three stages: metadata always in context, instructions on invocation, resources on demand. |
| Why does that matter? | A library of skills costs the metadata of all of them plus the body of the one in use, not the sum of every body. |
| Why is the window not a data bus? | Every intermediate result is re-sent on every later turn. Code moves data locally and returns only the answer. |
| When is code the wrong answer? | When there is no sandbox, no runtime, or no safe permission model — then sequential calls are the honest fallback. |
Further reading
- Anthropic, "Introducing the Model Context Protocol", 25 November 2024 — the announcement, and the host/client/server framing.
- Model Context Protocol, specification — dated revisions; the place to check what the core actually says before quoting a revision number.
- Linux Foundation, Agentic AI Foundation announcement — MCP's donation and the vendor-neutral home it now sits in.
- Anthropic, "Code execution with MCP", 2025 — the argument this part's fourth section makes: present tools as code APIs and keep intermediate results out of the context.
- Anthropic, "Introducing Agent Skills", 2025 —
SKILL.mdand the three-stage progressive disclosure it is built around. - Anthropic, "Effective context engineering for AI agents", 29 September 2025 — the write/select/compress/isolate taxonomy that tool definitions and skills are both spending against.