Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Two doors into the same room

The user tries to override the system prompt; or the content does

Direct injection is the familiar one. The user types "ignore your previous instructions and …" into their own turn. The model sees a conflict between two instructions — the system prompt that defines its job and the user text that contradicts it — and the failure is a matter of which one wins. This is a jailbreak, and it is the easy case, because the untrusted text is not pretending to be anything. It is in the user role, and every guardrail you own is pointed at the user turn.

Indirect injection is the one that matters for agents. The instruction never comes from the user. It arrives inside a retrieved document, a web page the agent fetched, an email in the mailbox it can read, a tool result it asked for, a code comment in a file it opened, or a markdown block in someone else's pull request. To the model, that text is data — the very thing it was told to summarise, answer questions about, or cite. But data and instructions are both just tokens in the same window, so an instruction hidden in a document is read with exactly the same authority as the system prompt. In the injection sandbox of this series, the poisoned document is the variant that runs away with the agent; the direct override is the variant that most defenses stop.

Move the hardening slider below. It spends more of the system prompt telling the model to ignore embedded instructions — and you can watch what that buys. The direct family bends. The indirect family barely notices, because an instruction that arrives as data never announces itself as an instruction.

Bypass rate for each family, before and after hardening. The faint bar is no hardening; the solid bar is the slider. Dots are the individual attack variants in the family.

💡 The durable idea: direct injection is a conflict of roles — user text versus system text. Indirect injection is a conflict of kinds — instructions versus data, which are the same kind of thing inside the model. That is why the second one is not a jailbreak you can talk your way out of.
2

A benign task, hijacked by a document

Watch the tool call change, not the tone

Here is the smallest version of the failure, with no cleverness in it. An agent is asked to summarise the release notes for a documentation set. It retrieves a page, reads it, and writes a summary. That is the benign path, and it is what the agent does ninety-nine times out of a hundred.

Now the retrieved page carries a line of text that looks like a vendor comment or an editorial note. It is not addressed to a human; it is addressed to the model, and it reads as an instruction because that is what it is. The summary is no longer the goal. The agent still does its job — it still produces a fluent answer — while quietly taking an action nobody asked for: reading a private configuration file and sending it out over the network tool it already has. The tell is not the prose. The tell is the tool call.

Toggle the poison on and off and step through the pipeline. In the clean run the model's action is a summarise call and the answer leaves with nothing attached. In the poisoned run the same task produces a network call, and the private data goes with it.

Task → retrieved document → model's tool call → what leaves the system. The document is the only thing that changes between the two runs.

Two details are worth holding onto. The first is that the injected line did not have to look like a command; it can be phrased as a system notice, a policy reminder, or a maintenance note, and the model reads all of them as instructions. The second is that the agent is not "broken" when this happens — it followed the most recent and most specific instruction it could see, which is exactly the behaviour that makes an agent useful when the instructions are yours.

⚠️ The trap: looking for a suspicious-looking string in the retrieved document. The payload is ordinary text with an unusually strong effect. Nothing about its shape distinguishes it from content you legitimately wanted to retrieve, which is why a filter on the text cannot be the boundary.
3

The lethal trifecta

Three conditions, and the attack exists

Simon Willison's framing makes the failure operational instead of hand-waving (June 2025). An agent is exposed to a real, exploitable prompt injection when three things are true at once:

When all three are present, an attack exists — not "may exist one day", but is constructible today by anyone who can get text in front of the agent. Remove any one of them and the trifecta breaks. That single sentence is the whole of least privilege for agents, and it is why the architectural defenses of Part 16 are framed as which condition does this pattern remove, rather than as a scoreboard of how clever the prompt is.

Turn the three conditions on and off. The verdict stays honest: with two present you have a risk to watch, with all three present you have a working attack path, and the readout always names the condition to remove first.

Three conditions. When all three are lit, the agent is one retrieved document away from acting for someone else.

💡 The durable idea: injection is not a property of the model or of the prompt. It is a property of the system — the coincidence of three capabilities in one loop. You cannot make the model immune, but you can make the coincidence impossible, and that is a design decision.
4

The way out: exfiltration channels

A leak needs somewhere to go

The third condition is the one teams forget, because it does not look like a security control. It looks like a feature. A browsing tool, an email-sending tool, an image renderer, a form that writes a field, a webhook, a "send to Slack" integration — every one of them is an exfiltration channel, and every one of them is there because it is useful. The private data does not need to be spoken aloud to leave; it only needs a route, and the routes are exactly the tools you gave the agent to do its job.

Three channels recur. A tool with network access is the broadest: the agent can put a secret in a URL, a request body, or a header, and it will arrive at an endpoint the attacker names. A rendered markdown image is the quietest: the model emits an image whose URL contains the secret, and the leak happens when the response is displayed and the browser fetches it — no network tool required. A writable output field is the narrowest but the most common: an agent that can write to a shared document, a ticket, or a public profile can simply leave the value somewhere the attacker can read later.

Remove the channels in the demo and watch the leak count fall. Nothing about the agent's reasoning changes — the poisoned instruction is still in the context, still read with full authority — but a leak needs a channel, and a channel you deleted has no secret for it.

One context holding a secret, three ways for the secret to leave. Each channel removed is one exfiltration path closed.

💡 The durable idea: you rarely have to remove the agent's reach to break the trifecta. Removing the channel — dropping the network tool, disabling remote image loading, making the output field readable only by the user — removes the same condition at a fraction of the product cost. This is the move Part 16 expands into a full set of architectural patterns.

Cheat sheet

QuestionThe answer that shapes the build
What is direct prompt injection?The user's own turn tries to override the system prompt. A jailbreak, and the easier case.
What is indirect prompt injection?Instructions arrive inside retrieved documents, web pages, emails, tool results or code comments — as data the agent was told to use.
Why can the model not tell data from instructions?Both are text in one context. There is no enforced boundary, only a request not to be fooled.
What is the lethal trifecta?Private data access + exposure to untrusted content + an exfiltration channel. All three present means an attack exists.
Does a better system prompt fix it?No. Hardening is a request to a probabilistic function; it bends direct override and barely moves indirect injection.
What are the exfiltration channels?A tool with network access, a rendered markdown image URL, a writable output field, a webhook or any outbound integration.
What is the cheapest way to break the trifecta?Remove one condition — usually the channel — rather than trying to make the model resistant.
Is this a model problem?No. It is a systems problem in which untrusted text and privileged capability share one context.
Where is the fix?Architecture: quarantine the untrusted read, gate actions on provenance, and never let private data and untrusted content share a live channel.

Further reading

5

Check your understanding

0/5 answered