Guardrails, Safety & Deployment
Alignment in Part 8 shapes the model's own behavior through training. This final stage adds the layer on top: explicit safety training, adversarial testing, and runtime guardrails — the difference between a model that behaves well in a lab and a system that holds up once millions of people, some of them deliberately trying to break it, start using it.
Constitutional AI: critique and revise, worked
Training-time safety
Constitutional AI has the model critique and revise its own responses against a written set of principles, then trains on the revised versions — reducing reliance on humans manually writing every refusal example. Here's a single worked pass, not just the idea in the abstract:
This whole critique-and-revise exchange becomes one training example — no human had to write the revised response by hand, only the constitutional principle it was checked against.
The over-refusal trade-off, as an ROC curve
Quantifying the threshold problem
A single threshold slider hides the real trade-off: at every possible threshold, a refusal classifier has some true-positive rate (genuinely harmful prompts refused) and some false-positive rate (harmless prompts wrongly refused). Plotting both across all thresholds — a ROC curve — makes the trade-off, and how good the underlying classifier actually is, visible at once.
Prompt injection vs. jailbreaking
Two distinct attack shapes
| Jailbreaking | Prompt injection | |
|---|---|---|
| Attacker | The user themselves | Often a third party, via content the model reads (a webpage, a document, an email) |
| Goal | Get the model to ignore its own safety training | Get the model to treat attacker text as new instructions instead of data |
| Typical form | Role-play framings, hypothetical/fictional wrapping, obfuscation | "Ignore previous instructions and..." embedded in retrieved content |
| Where it's hardest to defend | Model weights themselves | Any system feeding untrusted external content into the context window (RAG, browsing, tool results — see Part 11) |
Defense in depth: layered guardrails
Deployment-time architecture
No single layer catches everything, which is why production systems stack several: an input guard classifier, the system prompt's own instructions, the trained model's own refusal behavior, and an output guard classifier reviewing what's about to be shown. Toggle layers off below and watch the probability that a given attack gets all the way through change — this is why defense-in-depth beats any single regex.
Red-teaming: testing it like an adversary would
Adversarial evaluation
Red-teaming means deliberately trying to make the model fail — done by dedicated internal teams, external researchers, and increasingly other models trained specifically to generate adversarial prompts. Findings feed back into more safety training and runtime guardrails, since not every failure mode can be fully trained away.
System prompt vs. training, and staged release
After training ends
Training-time alignment and deployment-time guardrails fail differently: a training gap requires another training run to fix, while a system-prompt or classifier gap can often be patched within hours of discovery. This is why "just put it in the system prompt" and "just train it in" aren't competing strategies — the system prompt is fast to iterate but easy to override with enough user effort; training is slow to iterate but far more robust once it's in.
The safety evaluation suite
Measuring refusal, not vibes
Safety claims need their own benchmarks, separate from capability ones. OLMo 3 reports a suite split into six development evaluations used during training and four held-out evaluations reserved for final reporting. No single benchmark is sufficient: a model that refuses everything would ace the harm tests and fail the over-refusal ones, so the suite deliberately includes both directions.
| Benchmark | What it measures | Metric | OLMo 3 7B Instruct |
|---|---|---|---|
| HarmBench | Refusal of 320 diverse harmful prompts (cyber, bio, harassment, …) | Refusal accuracy (1 − ASR) | 94.9 |
| DoAnythingNow (DAN) | Robustness to the DAN jailbreak template | Refusal accuracy | 75.2 |
| XSTest | Over-refusal: safe prompts that look unsafe | Accuracy | 93.2 |
| WildGuard-Test | Prompt harm, response harm, refusal on 749 adversarial prompts | Safety label | 99.6 |
| WildJailbreak-Test | Adversarial jailbreaks, split harmful / benign | Refusal (harmful) + non-refusal (benign) | 69.1 / 98.0 |
| TrustLLM-JailbreakTrigger | 13 jailbreak attack families, 400 prompts | Refusal accuracy | 79.2 |
| Toxigen (held out) | Compliance with requests for toxic text across demographic groups | Toxicity (inverted) | 100.0 |
| StrongReject (held out) | 37 real-world jailbreak techniques | Weighted 1–5 safety score | 88.1 |
| WMDP (held out) | Dual-use bio/chem/cyber knowledge (many-choice) | Inverted accuracy | 45.5 |
| BBQ (held out) | Social bias across age, gender, race, religion, … | Accuracy (and bias scores) | 79.0 |
Normalization is part of the measurement. Different benchmarks report different things — some score refusal, some score correctness, some score toxicity — so OLMo 3 normalizes every metric so that higher is always safer: refusal accuracy (1 − attack success rate) for HarmBench, DAN, WildGuard, TrustLLM, Toxigen, and StrongReject; plain accuracy for XSTest and BBQ; the combination of harmful-query refusals and benign-query non-refusals for WildJailbreak; and inverted accuracy for WMDP. Without that unification, averaging the suite would be meaningless. The whole suite is sampled at temperature 0.7 and top-p 0.95, and the aggregate safety score for OLMo 3 7B Instruct is 87.6.
Pipeline cheat sheet (Parts 3–12)
The whole pipeline, one row per phase
| Phase | Objective | Data | Typical scale |
|---|---|---|---|
| Architecture (Part 3) | Define the function being trained | — | ~1–100B+ parameters |
| Dataset (Part 4) | Build a clean, diverse training corpus | Filtered, deduplicated web text + code, papers, books | Trillions of tokens |
| Pretraining (Part 5) | Next-token prediction, minimize cross-entropy loss | The full pretraining corpus | Trillions of tokens, thousands of GPUs, weeks–months |
| SFT (Part 7) | Learn to follow instructions in the right format | Curated (prompt, response) pairs | Thousands–millions of examples |
| Alignment (Part 8) | Prefer better responses over worse ones | Preference pairs (chosen vs. rejected) | Tens of thousands–low millions of pairs |
| Reasoning RL (Part 9) | Improve on verifiably-checkable tasks | Math, code, and instruction-following tasks with automatic checkers | Varies; compute-heavy per example |
| Guardrails (Part 12) | Refuse harmful requests, hold up under adversarial use | Safety-specific training data + red-team findings + eval suites | Ongoing, including post-deployment |
Every phase reuses the same underlying machinery — tokens, embeddings, a next-token-shaped loss, gradient-based updates — applied to progressively narrower, more curated, more human-judgment-laden data. OLMo is worth returning to here: because Ai2 publishes every stage's data and code, this entire pipeline is one you can go verify yourself. See the glossary for every bolded term across the series, or the series hub to jump anywhere.
Further reading
References
- Bai et al., "Constitutional AI: Harmlessness from AI Feedback" (2022).
- Ganguli et al., "Red Teaming Language Models to Reduce Harms" (2022).
- Mitchell et al., "Model Cards for Model Reporting" (2018).
- Greshake et al., "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection" (2023).
- Mazeika et al., "HarmBench" (2024); Röttger et al., "XSTest" (2023); Han et al., "WildGuard" (2024).
- Jiang et al., "WildTeaming / WildJailbreak" (2024); Souly et al., "StrongReject" (2024); Li et al., "WMDP" (2024).
- Lambert et al. (Ai2), "Tulu 3" — OLMo's safety training approach.