Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The protocol, lane by lane

Host, client, server — and a stateless core

The Model Context Protocol (MCP) is a wire format for letting a model talk to tools and data sources through a uniform interface instead of bespoke glue per integration. Anthropic introduced it in November 2024 and donated it to the Linux Foundation; the spec is versioned by date. The revisions you will see named are 2024-11-05, 2025-03-26 and 2025-06-18, with a 2026 revision that streamlines the core toward a stateless request/response model — explicit session handles rather than per-connection state — so a server can be scaled horizontally behind a load balancer. Treat any specific revision as checkable against modelcontextprotocol.io; clients and servers lag the spec by months.

There are three roles, and keeping them apart is the whole mental model. The host is your application, the thing holding the model and its context. The client is the protocol endpoint inside the host: it speaks JSON-RPC, matches request ids to responses, and owns the session. The server is what exposes capabilities — tools, resources, prompts — over that protocol. The host never speaks MCP; the client does.

// the shape of every exchange: a JSON-RPC request, id-matched to a response { "jsonrpc": "2.0", "id": 2, "method": "tools/list", "params": {} } { "jsonrpc": "2.0", "id": 2, "result": { "tools": [ { "name": "search_docs", "description": "Search the docs corpus", "inputSchema": { "type": "object", "properties": { "query": { "type": "string" } } } } ] } }

Step through the exchange below. The three lanes are the three roles; each arrow is one JSON-RPC message, and the animation is only a walk through the handshake an agent performs before it can do anything at all.

One arrow per message. Without session state, every request could be served by a different replica.

💡 The durable idea: MCP standardises how capabilities are described and invoked. It does not decide what goes in the model's context — that is still your call, and it is where the next three sections look.
2

Tool definitions are a system-prompt tax

N servers × M tools × schema tokens, before the user speaks

Here is the part of MCP that surprises people. The tools/list result is not a catalogue the model browses on demand; the client hands it to the model as a tool-definition block that sits in the prompt on every call. A tool is not just a name. It is a name, a description written for the model, and a JSON Schema for its arguments — and a faithful schema for one non-trivial tool can run to several hundred tokens once you describe every parameter, enum and nested object.

So the cost is multiplicative. Five servers exposing a dozen tools each, at two or three hundred schema tokens per tool, is fifteen thousand tokens of standing instruction before the user types a word. Those tokens are billed as input on every turn. They compete with the retrieved documents and the transcript for the same budget. And they are the reason a "small" agent that connects to ten services starts to feel slow and expensive without any of its actual work getting harder.

Drag the sliders below. The stacked column is the prompt against a fixed window; the line is where the output reserve begins. The readout reports the point at which the user's own query no longer fits — the number of servers at which you have spent the window on describing what you could do.

Tool schemas are a line item like any other. The interesting number is the one at which the query gets squeezed out.

⚠️ The trap: "just connect every server" is a context decision, not an integration decision. A tool nobody calls still costs tokens on every request, and it also gives the model more near-miss choices to be wrong about.
3

Agent Skills and SKILL.md's three stages

Metadata always, instructions on invocation, resources on demand

Tools tell an agent what it can do; skills tell it how your team does things. An Agent Skill is a folder with a SKILL.md file — a name, a description, and a body of instruction that may reference scripts and other files. The design trick is not the file format but the progressive disclosure: the skill is loaded in three stages, and only the first is unconditional.

The metadata — name and description, a few dozen tokens — is the only part resident in context all the time, so the model knows the skill exists and when it might apply. The instructions, the body of SKILL.md, are pulled in only when the skill is invoked. The resources it points at — scripts, templates, reference tables — are loaded only when a step actually needs them, and often they are never loaded into the model at all because they are executed rather than read.

That is what makes a library of skills affordable: twenty skills can be resident for the price of the metadata, because nineteen of them are sitting on disk unread. The comparison below puts progressive disclosure next to the naive alternative — loading every skill in full — against a small agent's window. Toggle how many skills are invoked and whether their resources are loaded.

Two columns, same skill library: progressive disclosure versus loading everything. The dashed line is the window.

4

The context window is not a data bus

Let code move the data; bring back only the answer

Once an agent has tools, the tempting design is to use the context window as the channel between them: call a tool, paste the result into the transcript, call the next tool, paste that result, and so on. Each hop is a full round trip, each result is billed as input on every subsequent turn, and the intermediate values you will never look at again accumulate forever. Ten steps of a data-shuffling analysis can put more tokens through the window than the final answer is worth by two orders of magnitude.

The alternative is to treat code as the tool. Give the agent a sandbox with your data reachable and let it write one program that does the whole job — filter, join, aggregate, compare — and return only the result. The intermediate rows, the fifty megabytes of JSON, the loop over a thousand files: none of it enters the context, because none of it needs to. This is also exactly how a skill's resources work: a reference table is not read into the prompt, it is queried.

The comparison below runs the same multi-step analysis both ways. Increase the number of steps or the size of each intermediate result and watch the sequential path's token count bend quadratic while the code path barely moves.

Three measures, two designs. Each bar is normalised within its own group; the labels carry the real values.

💡 Carry this forward: the window is a working set for reasoning, not a channel for bytes. Push data movement into code, keep decisions in the context, and the same agent gets cheaper, faster and more accurate at once.
⚠️ Not free: code execution needs a sandbox, a language runtime, permissions on the data, and a story for what happens when the script fails. Where the environment cannot provide that — a pure API agent with no shell — sequential tool calls are the honest choice, and the bill above is what they cost.

Cheat sheet

QuestionThe answer that shapes the build
What are MCP's three roles?Host (your app), client (the JSON-RPC endpoint inside it), server (what exposes tools, resources, prompts).
What changed in the 2026 spec revision?The core moves toward stateless request/response with explicit session handles, so servers scale horizontally. Verify the revision against the spec.
Where do tool definitions live?In the system prompt of every call — name, description and JSON Schema, billed as input every time.
How do you size the bill?Servers × tools per server × average schema tokens, against the window you actually have left after the output reserve.
How is a SKILL.md loaded?Three stages: metadata always in context, instructions on invocation, resources on demand.
Why does that matter?A library of skills costs the metadata of all of them plus the body of the one in use, not the sum of every body.
Why is the window not a data bus?Every intermediate result is re-sent on every later turn. Code moves data locally and returns only the answer.
When is code the wrong answer?When there is no sandbox, no runtime, or no safe permission model — then sequential calls are the honest fallback.

Further reading

5

Check your understanding

0/5 answered