Instructions, roles and the right altitude
Every call begins with a policy: who the model is, what the roles mean, and where the instructions stop and the data begins. Get the altitude wrong and the model is either too unconstrained to be reliable or too brittle to survive a paraphrased input. Get the boundaries wrong and untrusted text is read as a command — which is the first security lesson in this series, not an advanced one. This part fixes the roles, the altitude, the delimiters, and the separation that everything about injection later depends on.
Roles: a strong prior, not a hard boundary
System, user, assistant
A chat request carries a list of messages, each tagged with a role. The system message carries the standing policy: identity, constraints, tone, what to do when unsure. The user messages carry the current request. The assistant messages carry the model's own prior turns, replayed back so the transcript is coherent. The order matters less than the tags: the role is metadata the model was trained to respect.
Respect, not enforce. On most chat-tuned models the system message dominates the opening distribution — it is the strongest prior in the call — but it is a prior, not a security boundary. A sufficiently determined user turn, a conflicting later instruction, or a document in the context can pull behaviour away from it. The same instruction placed in the data position is weaker still, because the model has no reason to treat retrieved text as policy rather than content.
The chart below is a mock compliance rate for one instruction placed in each of the three positions, under a configurable amount of conflicting pressure. It is a simulation, but the ordering is the lesson: put policy in the system message, and never assume the system message is a wall.
Mock compliance for the same instruction in system, user and data position, at the chosen conflicting pressure. Seeded and stable across reloads.
The right altitude
Specific enough to constrain, general enough not to overfit
Instructions live at an altitude, and both ends fail. Too vague — "be helpful and accurate" — and the model has no constraint to satisfy, so behaviour drifts from case to case. Too brittle — forty lines of MUSTs, exact phrasings and rigid output templates — and the prompt overfits to the examples you had in mind: it passes those and breaks on any input that was phrased differently, or any model version that reads the constraints literally.
The right altitude is specific about the things that must not vary (the job, the boundaries, the output contract) and general about everything else, so the model can do its own reasoning on the long tail. The trade is not hypothetical — it shows up as a curve on an eval set, rising with coverage and falling with brittleness.
Mock pass rate over a seeded 20-case set against instruction altitude. The point estimate and its interval are drawn at the slider's altitude.
Delimiters and structured framing
Make the boundary visible
Instructions and data are both just tokens, so the only way to mark the boundary is to make it visible in the text. The standard tools are cheap and effective: a system message that states the contract, delimiters that wrap the payload (<context>, <document>, triple backticks, a fenced block), and an explicit output contract that says what shape the answer takes. A tag tells the model where a region starts and stops; a label tells it what the region is.
Structure also buys tokens. The composer below assembles a prompt from four parts and reports its exact token cost, its share of the window, and whether the instruction/data separation is explicit or merely positional. Toggle the parts and watch the verdict change without the text changing meaning.
The assembled prompt, section by section. Colours mark the system policy, the delimiters, the retrieved data and the output contract.
The first security lesson
Untrusted content is data, never instructions
If retrieved text is placed in the context without a boundary, the model has no principled way to know that a sentence inside it is content to read rather than policy to follow. That is indirect prompt injection: an attacker does not need to talk to your user, only to get a sentence into a document your retriever will fetch. Instructions inside data are the whole attack.
The demo stages it with one poisoned chunk. In the naive layout the instruction runs and the answer is diverted. Delimiters plus a "this is data" contract raise the bar and make the boundary explicit — but the defence matrix is honest about the limit: prompt hardening alone does not stop this variant. What does is removing the ability to act, either by quarantining the untrusted read (a second model that returns only structured fields) or by breaking the lethal trifecta — private data, untrusted content and an exfiltration channel in the same context.
A retrieved document containing an instruction, and the model's answer under three layouts: naive, delimited, and quarantined.
Positive instructions and output contracts
Why a wall of MUSTs degrades
Two habits separate instructions that hold up from instructions that decay. The first is writing positive instructions — say what to do rather than what to avoid. "Answer only from the provided context" gives the model a target; "don't hallucinate" gives it a prohibition it cannot act on, and prohibitions are weak constraints in a next-token model. The second is stating an output contract: the shape, the length, the citation format, and what to do when the context does not contain the answer.
A wall of MUSTs degrades for three compounding reasons. Instructions compete with each other and with the data for attention, and the loudest, latest or most specific one tends to win. Absolute rules collide on the long tail — "always cite a source" and "never make one up" conflict on an unanswerable question unless you say which wins. And every rule that does not change an outcome on your eval set spends tokens and attention to add a failure mode. The discipline is the same as the altitude: state the few constraints that matter, define the escape hatch, and let the eval decide which rules stay.
Cheat sheet
| Question | The answer that shapes everything |
|---|---|
| What does the system role actually do? | Sets the strongest prior in the call — not an enforceable boundary. |
| Where should policy live? | In the system message; enforcement lives in code, tools and permissions. |
| What is the right altitude? | Specific about the job, the boundary and the contract; general about the rest. |
| What fails at high altitude? | Brittle prompts overfit to their own wording and break on paraphrase or upgrade. |
| Why delimiters? | They make the instruction/data boundary visible in the text itself. |
| What is the first security lesson? | Untrusted content is data, never instructions. |
| Do delimiters stop indirect injection? | They raise the bar; they do not close the hole. Remove the capability instead. |
| How do you write constraints? | Positively, with an output contract and a stated fallback. |
Further reading
- Greshake et al., "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection", 2023 — the paper that named indirect prompt injection and showed it through retrieved content.
- Debenedetti et al., "Defeating Prompt Injections by Design", 2025 — CaMeL and capability-based information-flow control: the idea that permissions, not phrasing, decide what an agent may do.
- Debenedetti et al., "AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents", 2024 — a benchmark where defences are measured against adaptive attacks rather than a fixed payload list.
- Zou et al., "Universal and Transferable Adversarial Attacks on Aligned Language Models", 2023 — the GCG suffix attack, the reason "the model will refuse" is not a defence.
- OWASP, "LLM01: Prompt Injection", OWASP Top 10 for LLM Applications — the industry framing of prompt injection as the first risk, not an edge case.