Architectural defenses, guardrails and red-teaming
The previous part showed that an agent with private data, untrusted content and an exfiltration channel is one retrieved document away from being someone else's tool. The temptation is to look for the sentence that stops it — a better system prompt, a cleverer filter. There is no such sentence, because prompt injection is not a text problem. It is a systems problem in which untrusted text and privileged capability share one context, and the fix is to change who is allowed to influence what. This part lays out the six design patterns, shows honestly which attacks each one misses, and ends with the practice that keeps a release honest: red-teaming in CI.
The six design patterns
An architectural answer to a systems problem
Because the failure is a systems failure, the defenses that survive contact with reality are the ones that change the system around the model rather than the words given to it. Six patterns recur, roughly in order of how much they cost to build and how much they buy.
- No defense. Untrusted text and privileged tools share one context. This is the default, and it is a pattern only in the sense that doing nothing is a choice.
- Prompt hardening. The system prompt instructs the model to ignore embedded instructions and to treat retrieved text as data. It raises the bar and it is nearly free, but it is a request to a probabilistic function, not a boundary.
- Rails and classifiers. Input and output filters flag known payload shapes and known exfiltration channels. They are cheap, they run outside the model, and they are defeated by a novel encoding they have not seen.
- Dual-LLM quarantine. A second, unprivileged model reads the untrusted content and may only emit a narrow structured result — a summary, a field extraction — to the privileged model, which never sees the raw text.
- Least privilege. Break the Lethal Trifecta: never put private data, untrusted content and a live exfiltration path in the same context. Remove the network tool, scope the credentials, or keep the untrusted content out of the privileged turn entirely.
- Capability-based information flow. Tag values with a capability — where they came from and what they may influence — and let the runtime decide whether a value may flow into an action, independent of the text that carries it. This is the strongest single control, and it is the one non-obvious idea on the list.
The sandbox below is the whole lesson. It puts all six patterns against the same bank of eight attack variants and shows, cell by cell, which ones each stops — and, more importantly, which ones it does not. Step through the defenses and watch the red cells move; no row is clean.
Rows are defenses, columns are attack variants. Blue means stopped, red means it gets through — the misses are the point.
Capability information flow, and why text filters cannot do it
Let provenance decide permission, not the words
Every content filter is asking the same question: does this text look like an instruction? That is the wrong question, because the payload can be made to not look like one — encoded, split across turns, hidden in a homoglyph, or wearing the uniform of a trusted tool result. The right question is structural: where did this value come from, and is its provenance allowed to influence this action?
Capability-based information flow, as implemented in CaMeL, attaches that answer to the value rather than the sentence. Content read from a document is tagged untrusted. Any value derived from it inherits that tag. A privileged tool — send an email, make a payment, call an endpoint — declares which capabilities its arguments must satisfy, and the runtime refuses the call when a tagged value reaches it, no matter how convincing the surrounding prose is. The payload can say anything at all; it still cannot authorise the send.
The diagram below shows the flow two ways. With capability control off, the hidden instruction rides through the privileged model and reaches the action. With it on, the same value is stopped at the gate because of where it came from — the text never gets a vote. Press play to animate the packet.
Untrusted content → quarantined model → capability gate → privileged action. The gate reads the tag, not the prose.
Rails, classifiers, and never logging the secret
Defence in depth, and the copy you forgot you made
Rails and classifiers are the layer you keep even though it is bypassable, for the same reason a lock is worth having even though a determined thief can pick it: it stops the cheap attack, it buys detection signal, and it costs almost nothing per call. An input classifier scores the payload shape; an output classifier scores the response for an exfiltration channel — a markdown image with a private value in the URL, an outbound HTTP call with a customer record in the body. The failure mode is precise and must be stated in the threat model: a novel encoding that the classifier has never seen passes. So rails are defence in depth, never the defence.
The companion habit is smaller and easier to get wrong. The moment you add "do not reveal the customer's email" to a system prompt, that email is now in a prompt, and prompts get logged. An agent that refuses to say a secret out loud while writing it into a trace, a debug log, or a vendor's dashboard has not protected anything — it has just moved the leak somewhere less watched. Redact before the write, not on the way out of the log viewer.
The pipeline below logs a seeded stream of events twice: raw, and with the redactor in front. Turn the redactor off, or turn its coverage down, and watch the leak counter — the number of secrets that reached the log store — climb.
Left: what the application held. Right: what the log store kept. Anything that is not redacted before the write is a leak.
Red-teaming in CI
Measure attack-success rate, then gate the release on it
You cannot know whether a defense works until you attack it, and a one-off audit tells you about the day it ran. The durable practice is a bank of attack variants that runs on every change, scoring the agent's attack-success rate — the fraction of attacks that achieve their goal — and gating the release on it, exactly like a test suite gates a merge. Three families of attack make a useful bank.
- GCG (Greedy Coordinate Gradient) search finds adversarial suffixes — token sequences optimised to make a model comply. They transfer between models more often than intuition suggests, so a suffix found against an open model is worth trying against yours.
- PAIR uses a second model as an attacker, iteratively refining a prompt against the target in a handful of queries. It is cheap, black-box, and it finds the semantic attacks a static list misses.
- AgentDojo supplies the missing piece for agents: a dynamic environment where the attack rides in on real tool results, and where the score is whether the agent's end state was compromised, not whether a string was emitted.
Run the bank against every candidate prompt and model. If the attack-success rate rises — a new model regresses on a defense the old one passed, a prompt change opens a channel — block the release. The chart below compares a team that does not red-team in CI with one that does, across ten releases, against a gate you can move.
Attack-success rate per release. The dashed line is the release gate; releases above it do not ship.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| Is prompt injection a text or a systems problem? | Systems. Untrusted text and privileged capability sharing one context is the defect; no sentence fixes it. |
| What are the six patterns? | None, prompt hardening, rails/classifiers, dual-LLM quarantine, least privilege, capability information flow. |
| Which single pattern is strongest? | Capability information flow — it stops seven of the eight variants, gating actions on a value's provenance rather than its words. |
| What does capability control not stop? | An authorised direct instruction. The user is the authority; capability tokens govern untrusted influence. |
| Why keep rails if they are bypassable? | They stop the cheap attack and buy detection signal; a novel encoding defeats them, so they are depth, not the defense. |
| What is the dual-LLM pattern? | A quarantined model reads untrusted content and may only emit structured data to the privileged model. |
| What is least privilege here? | Break the Lethal Trifecta: private data, untrusted content and a live exfiltration path never share a context. |
| Where should PII be redacted? | Before it is written, not when it is read. A prompt is a future log line. |
| What do GCG, PAIR and AgentDojo cover? | Adversarial suffixes, iterative black-box attacks, and end-state compromise inside a tool-using environment. |
| How do you gate a release? | Run the attack bank in CI and block when attack-success rate exceeds the threshold. |
Further reading
- Greshake et al., "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection", 2023 — the paper that named indirect injection and demonstrated it in real systems.
- Debenedetti et al., "Defeating Prompt Injections by Design" (CaMeL), 2025 — capability-based information-flow control, the pattern this part calls the strongest single control.
- Zou et al., "Universal and Transferable Adversarial Attacks on Aligned Language Models", 2023 — GCG adversarial suffixes and their transferability.
- Chao et al., "Jailbreaking Black Box Large Language Models in Twenty Queries" (PAIR), 2023 — the attacker-model loop that makes a black-box red-team cheap.
- Debenedetti et al., "AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents", 2024 — end-state attack success inside a tool-using agent.
- Simon Willison, "The lethal trifecta for AI agents", 16 June 2025 — the framing that makes least privilege operational.
- OWASP, "LLM01: Prompt Injection", OWASP Top 10 for LLM Applications — injection as the first entry, with the systems framing.