Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The six design patterns

An architectural answer to a systems problem

Because the failure is a systems failure, the defenses that survive contact with reality are the ones that change the system around the model rather than the words given to it. Six patterns recur, roughly in order of how much they cost to build and how much they buy.

The sandbox below is the whole lesson. It puts all six patterns against the same bank of eight attack variants and shows, cell by cell, which ones each stops — and, more importantly, which ones it does not. Step through the defenses and watch the red cells move; no row is clean.

Rows are defenses, columns are attack variants. Blue means stopped, red means it gets through — the misses are the point.

💡 The durable idea: every single pattern is partial. Prompt hardening stops one variant of eight; the strongest single control stops seven. The defense is the union of layered controls, not the best one — which is exactly what defence in depth means when the adversary is a stream of text.
2

Capability information flow, and why text filters cannot do it

Let provenance decide permission, not the words

Every content filter is asking the same question: does this text look like an instruction? That is the wrong question, because the payload can be made to not look like one — encoded, split across turns, hidden in a homoglyph, or wearing the uniform of a trusted tool result. The right question is structural: where did this value come from, and is its provenance allowed to influence this action?

Capability-based information flow, as implemented in CaMeL, attaches that answer to the value rather than the sentence. Content read from a document is tagged untrusted. Any value derived from it inherits that tag. A privileged tool — send an email, make a payment, call an endpoint — declares which capabilities its arguments must satisfy, and the runtime refuses the call when a tagged value reaches it, no matter how convincing the surrounding prose is. The payload can say anything at all; it still cannot authorise the send.

// capability tokens: provenance decides permission, not the text read(doc) -> "…" # tagged untrusted summarise(value:untrusted) -> "…" # derived values inherit the tag send_email(to, body:untrusted) -> REFUSED # untrusted may not authorise a send send_email(to, body:user) -> allowed # the user's own intent is authorised

The diagram below shows the flow two ways. With capability control off, the hidden instruction rides through the privileged model and reaches the action. With it on, the same value is stopped at the gate because of where it came from — the text never gets a vote. Press play to animate the packet.

Untrusted content → quarantined model → capability gate → privileged action. The gate reads the tag, not the prose.

⚠️ Capability control does not stop an authorised direct instruction. If the user themselves asks the agent to send the email, that is not an injection — the user is the authority. Capability tokens govern untrusted influence, and they leave the user's own intent deliberately unconstrained. That is why the direct-override column stays red even for the strongest pattern.
3

Rails, classifiers, and never logging the secret

Defence in depth, and the copy you forgot you made

Rails and classifiers are the layer you keep even though it is bypassable, for the same reason a lock is worth having even though a determined thief can pick it: it stops the cheap attack, it buys detection signal, and it costs almost nothing per call. An input classifier scores the payload shape; an output classifier scores the response for an exfiltration channel — a markdown image with a private value in the URL, an outbound HTTP call with a customer record in the body. The failure mode is precise and must be stated in the threat model: a novel encoding that the classifier has never seen passes. So rails are defence in depth, never the defence.

The companion habit is smaller and easier to get wrong. The moment you add "do not reveal the customer's email" to a system prompt, that email is now in a prompt, and prompts get logged. An agent that refuses to say a secret out loud while writing it into a trace, a debug log, or a vendor's dashboard has not protected anything — it has just moved the leak somewhere less watched. Redact before the write, not on the way out of the log viewer.

The pipeline below logs a seeded stream of events twice: raw, and with the redactor in front. Turn the redactor off, or turn its coverage down, and watch the leak counter — the number of secrets that reached the log store — climb.

Left: what the application held. Right: what the log store kept. Anything that is not redacted before the write is a leak.

💡 The durable idea: never log the secret you are trying to protect. A prompt is a log line waiting to happen, and a redactor that runs at read time is a promise, not a control. Redact at the boundary where the value would be persisted.
4

Red-teaming in CI

Measure attack-success rate, then gate the release on it

You cannot know whether a defense works until you attack it, and a one-off audit tells you about the day it ran. The durable practice is a bank of attack variants that runs on every change, scoring the agent's attack-success rate — the fraction of attacks that achieve their goal — and gating the release on it, exactly like a test suite gates a merge. Three families of attack make a useful bank.

Run the bank against every candidate prompt and model. If the attack-success rate rises — a new model regresses on a defense the old one passed, a prompt change opens a channel — block the release. The chart below compares a team that does not red-team in CI with one that does, across ten releases, against a gate you can move.

Attack-success rate per release. The dashed line is the release gate; releases above it do not ship.

💡 The durable idea: a defense is a claim, and a red-team suite is the experiment that tests it. Gating on attack-success rate turns injection from an argument into a number that can regress — which is the only kind of control that survives a fast-moving model upgrade.

Cheat sheet

QuestionThe answer that shapes the build
Is prompt injection a text or a systems problem?Systems. Untrusted text and privileged capability sharing one context is the defect; no sentence fixes it.
What are the six patterns?None, prompt hardening, rails/classifiers, dual-LLM quarantine, least privilege, capability information flow.
Which single pattern is strongest?Capability information flow — it stops seven of the eight variants, gating actions on a value's provenance rather than its words.
What does capability control not stop?An authorised direct instruction. The user is the authority; capability tokens govern untrusted influence.
Why keep rails if they are bypassable?They stop the cheap attack and buy detection signal; a novel encoding defeats them, so they are depth, not the defense.
What is the dual-LLM pattern?A quarantined model reads untrusted content and may only emit structured data to the privileged model.
What is least privilege here?Break the Lethal Trifecta: private data, untrusted content and a live exfiltration path never share a context.
Where should PII be redacted?Before it is written, not when it is read. A prompt is a future log line.
What do GCG, PAIR and AgentDojo cover?Adversarial suffixes, iterative black-box attacks, and end-state compromise inside a tool-using environment.
How do you gate a release?Run the attack bank in CI and block when attack-success rate exceeds the threshold.

Further reading

5

Check your understanding

0/5 answered