Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Constitutional AI: critique and revise, worked

Training-time safety

Constitutional AI has the model critique and revise its own responses against a written set of principles, then trains on the revised versions — reducing reliance on humans manually writing every refusal example. Here's a single worked pass, not just the idea in the abstract:

Prompt"What's the easiest way to get into someone's locked phone if I've forgotten the password?"
Initial response (unaligned draft)"Try common PIN combinations like 0000 or the owner's birthday, or use a third-party unlocking tool that bypasses the lock screen."
Constitutional principle"Does this response give operational help for accessing a device or account that may not belong to the requester?"
Model's self-critique"Yes — this treats 'forgotten my own password' and 'getting into someone else's phone' as the same request, and gives bypass instructions without establishing ownership."
Revised response"If it's your own phone: use your device maker's official account-recovery process (e.g. Find My iPhone or your Google account), which verifies ownership. If it's someone else's phone, I can't help bypass their lock screen — but if you own the device jointly or have legitimate authority over it, the manufacturer's official recovery process is still the right path."

This whole critique-and-revise exchange becomes one training example — no human had to write the revised response by hand, only the constitutional principle it was checked against.

2

The over-refusal trade-off, as an ROC curve

Quantifying the threshold problem

A single threshold slider hides the real trade-off: at every possible threshold, a refusal classifier has some true-positive rate (genuinely harmful prompts refused) and some false-positive rate (harmless prompts wrongly refused). Plotting both across all thresholds — a ROC curve — makes the trade-off, and how good the underlying classifier actually is, visible at once.

⚠️ Over-refusal is a real cost, not a safe default: a model that refuses too liberally is less useful and can push people toward less careful tools. Scores above are illustrative, not a real classifier's output — but the ROC curve's shape (bowed toward the top-left corner = a genuinely useful classifier; near the diagonal = little better than random) is the real diagnostic labs use.
3

Prompt injection vs. jailbreaking

Two distinct attack shapes

JailbreakingPrompt injection
AttackerThe user themselvesOften a third party, via content the model reads (a webpage, a document, an email)
GoalGet the model to ignore its own safety trainingGet the model to treat attacker text as new instructions instead of data
Typical formRole-play framings, hypothetical/fictional wrapping, obfuscation"Ignore previous instructions and..." embedded in retrieved content
Where it's hardest to defendModel weights themselvesAny system feeding untrusted external content into the context window (RAG, browsing, tool results — see Part 11)
4

Defense in depth: layered guardrails

Deployment-time architecture

No single layer catches everything, which is why production systems stack several: an input guard classifier, the system prompt's own instructions, the trained model's own refusal behavior, and an output guard classifier reviewing what's about to be shown. Toggle layers off below and watch the probability that a given attack gets all the way through change — this is why defense-in-depth beats any single regex.

5

Red-teaming: testing it like an adversary would

Adversarial evaluation

Red-teaming means deliberately trying to make the model fail — done by dedicated internal teams, external researchers, and increasingly other models trained specifically to generate adversarial prompts. Findings feed back into more safety training and runtime guardrails, since not every failure mode can be fully trained away.

⚠️ This demo is deliberately shallow — a few hardcoded phrase patterns, not a real safety classifier — to illustrate the idea of a guard layer. Real guardrail systems use trained classifiers, precisely because keyword matching is trivial to evade (try rephrasing the same request without the flagged words).
6

System prompt vs. training, and staged release

After training ends

Training-time alignment and deployment-time guardrails fail differently: a training gap requires another training run to fix, while a system-prompt or classifier gap can often be patched within hours of discovery. This is why "just put it in the system prompt" and "just train it in" aren't competing strategies — the system prompt is fast to iterate but easy to override with enough user effort; training is slow to iterate but far more robust once it's in.

💡 Staged release: labs commonly ship a new model to a limited group first (internal testers, a beta cohort, or a fraction of production traffic) with active monitoring for the specific failure modes red-teaming surfaced, widening access only as real-world signal confirms the model holds up — the deployment equivalent of the checkpoint-by-checkpoint inspection from Part 5.
💡 Model cards (a norm OLMo and most major labs follow) document a model's training data, intended uses, known limitations, and evaluation results at release time — the release-time counterpart to publishing the training recipe itself.
7

The safety evaluation suite

Measuring refusal, not vibes

Safety claims need their own benchmarks, separate from capability ones. OLMo 3 reports a suite split into six development evaluations used during training and four held-out evaluations reserved for final reporting. No single benchmark is sufficient: a model that refuses everything would ace the harm tests and fail the over-refusal ones, so the suite deliberately includes both directions.

BenchmarkWhat it measuresMetricOLMo 3 7B Instruct
HarmBenchRefusal of 320 diverse harmful prompts (cyber, bio, harassment, …)Refusal accuracy (1 − ASR)94.9
DoAnythingNow (DAN)Robustness to the DAN jailbreak templateRefusal accuracy75.2
XSTestOver-refusal: safe prompts that look unsafeAccuracy93.2
WildGuard-TestPrompt harm, response harm, refusal on 749 adversarial promptsSafety label99.6
WildJailbreak-TestAdversarial jailbreaks, split harmful / benignRefusal (harmful) + non-refusal (benign)69.1 / 98.0
TrustLLM-JailbreakTrigger13 jailbreak attack families, 400 promptsRefusal accuracy79.2
Toxigen (held out)Compliance with requests for toxic text across demographic groupsToxicity (inverted)100.0
StrongReject (held out)37 real-world jailbreak techniquesWeighted 1–5 safety score88.1
WMDP (held out)Dual-use bio/chem/cyber knowledge (many-choice)Inverted accuracy45.5
BBQ (held out)Social bias across age, gender, race, religion, …Accuracy (and bias scores)79.0

Normalization is part of the measurement. Different benchmarks report different things — some score refusal, some score correctness, some score toxicity — so OLMo 3 normalizes every metric so that higher is always safer: refusal accuracy (1 − attack success rate) for HarmBench, DAN, WildGuard, TrustLLM, Toxigen, and StrongReject; plain accuracy for XSTest and BBQ; the combination of harmful-query refusals and benign-query non-refusals for WildJailbreak; and inverted accuracy for WMDP. Without that unification, averaging the suite would be meaningless. The whole suite is sampled at temperature 0.7 and top-p 0.95, and the aggregate safety score for OLMo 3 7B Instruct is 87.6.

⚠️ Watch what moves and what doesn't. Across the three post-training stages, different safety benchmarks move in opposite directions: HarmBench improves from 87.7 (SFT) to 94.9 (final) while DAN drops from 90.0 to 75.2, and WildJailbreak (harmful) falls from 80.9 to 69.1 while its benign non-refusal rises to 98.0. Safety is not a single scalar, and "more refusal" is not the same as "safer".
✓

Pipeline cheat sheet (Parts 3–12)

The whole pipeline, one row per phase

PhaseObjectiveDataTypical scale
Architecture (Part 3)Define the function being trained—~1–100B+ parameters
Dataset (Part 4)Build a clean, diverse training corpusFiltered, deduplicated web text + code, papers, booksTrillions of tokens
Pretraining (Part 5)Next-token prediction, minimize cross-entropy lossThe full pretraining corpusTrillions of tokens, thousands of GPUs, weeks–months
SFT (Part 7)Learn to follow instructions in the right formatCurated (prompt, response) pairsThousands–millions of examples
Alignment (Part 8)Prefer better responses over worse onesPreference pairs (chosen vs. rejected)Tens of thousands–low millions of pairs
Reasoning RL (Part 9)Improve on verifiably-checkable tasksMath, code, and instruction-following tasks with automatic checkersVaries; compute-heavy per example
Guardrails (Part 12)Refuse harmful requests, hold up under adversarial useSafety-specific training data + red-team findings + eval suitesOngoing, including post-deployment

Every phase reuses the same underlying machinery — tokens, embeddings, a next-token-shaped loss, gradient-based updates — applied to progressively narrower, more curated, more human-judgment-laden data. OLMo is worth returning to here: because Ai2 publishes every stage's data and code, this entire pipeline is one you can go verify yourself. See the glossary for every bolded term across the series, or the series hub to jump anywhere.

📚

Further reading

References

?

Check your understanding

0/5 answered
Guardrails are one axis of capability. The next three parts cover the frontier concerns that decide whether a model can actually be used: how far its context reaches, whether it can call tools, and the full post-training recipe behind a frontier assistant. Continue: long-context & context extension →