Prompt injection and the lethal trifecta
An agent that retrieves, browses or reads your inbox is doing something no chat window does: it lets text it did not write influence what it does next. Most of that text is helpful. Some of it is written by someone who wants the agent to act on their behalf, and the agent cannot tell the two apart, because both arrive as language. This part names the failure — direct injection, indirect injection, and the three conditions that turn a curious agent into a leaking one — and ends with the honest conclusion: the fix is architecture, not wording. Part 16 is where the architecture lives.
Two doors into the same room
The user tries to override the system prompt; or the content does
Direct injection is the familiar one. The user types "ignore your previous instructions and …" into their own turn. The model sees a conflict between two instructions — the system prompt that defines its job and the user text that contradicts it — and the failure is a matter of which one wins. This is a jailbreak, and it is the easy case, because the untrusted text is not pretending to be anything. It is in the user role, and every guardrail you own is pointed at the user turn.
Indirect injection is the one that matters for agents. The instruction never comes from the user. It arrives inside a retrieved document, a web page the agent fetched, an email in the mailbox it can read, a tool result it asked for, a code comment in a file it opened, or a markdown block in someone else's pull request. To the model, that text is data — the very thing it was told to summarise, answer questions about, or cite. But data and instructions are both just tokens in the same window, so an instruction hidden in a document is read with exactly the same authority as the system prompt. In the injection sandbox of this series, the poisoned document is the variant that runs away with the agent; the direct override is the variant that most defenses stop.
Move the hardening slider below. It spends more of the system prompt telling the model to ignore embedded instructions — and you can watch what that buys. The direct family bends. The indirect family barely notices, because an instruction that arrives as data never announces itself as an instruction.
Bypass rate for each family, before and after hardening. The faint bar is no hardening; the solid bar is the slider. Dots are the individual attack variants in the family.
A benign task, hijacked by a document
Watch the tool call change, not the tone
Here is the smallest version of the failure, with no cleverness in it. An agent is asked to summarise the release notes for a documentation set. It retrieves a page, reads it, and writes a summary. That is the benign path, and it is what the agent does ninety-nine times out of a hundred.
Now the retrieved page carries a line of text that looks like a vendor comment or an editorial note. It is not addressed to a human; it is addressed to the model, and it reads as an instruction because that is what it is. The summary is no longer the goal. The agent still does its job — it still produces a fluent answer — while quietly taking an action nobody asked for: reading a private configuration file and sending it out over the network tool it already has. The tell is not the prose. The tell is the tool call.
Toggle the poison on and off and step through the pipeline. In the clean run the model's action is a summarise call and the answer leaves with nothing attached. In the poisoned run the same task produces a network call, and the private data goes with it.
Task → retrieved document → model's tool call → what leaves the system. The document is the only thing that changes between the two runs.
Two details are worth holding onto. The first is that the injected line did not have to look like a command; it can be phrased as a system notice, a policy reminder, or a maintenance note, and the model reads all of them as instructions. The second is that the agent is not "broken" when this happens — it followed the most recent and most specific instruction it could see, which is exactly the behaviour that makes an agent useful when the instructions are yours.
The lethal trifecta
Three conditions, and the attack exists
Simon Willison's framing makes the failure operational instead of hand-waving (June 2025). An agent is exposed to a real, exploitable prompt injection when three things are true at once:
- Access to private data. The agent can read something an attacker wants: a customer record, an internal document, a set of credentials, the contents of a mailbox.
- Exposure to untrusted content. The agent ingests text that someone outside your trust boundary could have written — retrieved documents, web pages, emails, tool results, code comments.
- An exfiltration channel. There is a way for the agent's output to leave your system: a tool with network access, a rendered markdown image whose URL the agent controls, a writable output field a third party can read.
When all three are present, an attack exists — not "may exist one day", but is constructible today by anyone who can get text in front of the agent. Remove any one of them and the trifecta breaks. That single sentence is the whole of least privilege for agents, and it is why the architectural defenses of Part 16 are framed as which condition does this pattern remove, rather than as a scoreboard of how clever the prompt is.
Turn the three conditions on and off. The verdict stays honest: with two present you have a risk to watch, with all three present you have a working attack path, and the readout always names the condition to remove first.
Three conditions. When all three are lit, the agent is one retrieved document away from acting for someone else.
The way out: exfiltration channels
A leak needs somewhere to go
The third condition is the one teams forget, because it does not look like a security control. It looks like a feature. A browsing tool, an email-sending tool, an image renderer, a form that writes a field, a webhook, a "send to Slack" integration — every one of them is an exfiltration channel, and every one of them is there because it is useful. The private data does not need to be spoken aloud to leave; it only needs a route, and the routes are exactly the tools you gave the agent to do its job.
Three channels recur. A tool with network access is the broadest: the agent can put a secret in a URL, a request body, or a header, and it will arrive at an endpoint the attacker names. A rendered markdown image is the quietest: the model emits an image whose URL contains the secret, and the leak happens when the response is displayed and the browser fetches it — no network tool required. A writable output field is the narrowest but the most common: an agent that can write to a shared document, a ticket, or a public profile can simply leave the value somewhere the attacker can read later.
Remove the channels in the demo and watch the leak count fall. Nothing about the agent's reasoning changes — the poisoned instruction is still in the context, still read with full authority — but a leak needs a channel, and a channel you deleted has no secret for it.
One context holding a secret, three ways for the secret to leave. Each channel removed is one exfiltration path closed.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| What is direct prompt injection? | The user's own turn tries to override the system prompt. A jailbreak, and the easier case. |
| What is indirect prompt injection? | Instructions arrive inside retrieved documents, web pages, emails, tool results or code comments — as data the agent was told to use. |
| Why can the model not tell data from instructions? | Both are text in one context. There is no enforced boundary, only a request not to be fooled. |
| What is the lethal trifecta? | Private data access + exposure to untrusted content + an exfiltration channel. All three present means an attack exists. |
| Does a better system prompt fix it? | No. Hardening is a request to a probabilistic function; it bends direct override and barely moves indirect injection. |
| What are the exfiltration channels? | A tool with network access, a rendered markdown image URL, a writable output field, a webhook or any outbound integration. |
| What is the cheapest way to break the trifecta? | Remove one condition — usually the channel — rather than trying to make the model resistant. |
| Is this a model problem? | No. It is a systems problem in which untrusted text and privileged capability share one context. |
| Where is the fix? | Architecture: quarantine the untrusted read, gate actions on provenance, and never let private data and untrusted content share a live channel. |
Further reading
- Greshake et al., "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection", 2023 — the paper that named indirect injection and demonstrated it against real integrated applications.
- Simon Willison, "The lethal trifecta for AI agents", 16 June 2025 — private data, untrusted content and an exfiltration channel; the framing this part is built on.
- OWASP, "LLM01: Prompt Injection", OWASP Top 10 for LLM Applications — injection as the first entry, with the direct/indirect split and the systems framing.
- Yi et al., "Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models", 2023 — measurement of indirect attacks and why simple defences leak.
- Debenedetti et al., "Defeating Prompt Injections by Design" (CaMeL), 2025 — capability-based information-flow control; the architecture Part 16 unpacks.
- Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL 2024 — background on why instructions placed inside long retrieved context are read so unevenly.