Post-training at Scale
Parts 7–9 introduced supervised fine-tuning, preference optimization, and verifiable-reward RL one at a time. A frontier model uses all three, in sequence, across four domains, on top of a base model that was itself carefully midtrained. This part is the integration: the measured three-stage recipe, the counter-intuitive Delta Learning trick, the upgrades that turned GRPO into OLMoRL, the verifiers behind each domain, and the systems work that makes RL fast enough to run at all.
The three-stage recipe
SFT → DPO → RLVR
OLMo 3 Think is trained in three stages on top of the midtrained base: supervised fine-tuning on curated thinking traces (Dolci-Think-SFT), preference tuning (Dolci-Think-DPO), then reinforcement learning with verifiable rewards (OLMoRL). Many open reasoning models stop after SFT — OpenThoughts3 and s1 used only SFT; SmolLM3 used SFT and DPO with no RL. OLMo 3's claim is that each stage adds something the previous one cannot, and that RL works better when it starts from a DPO checkpoint than directly from SFT. The stage table below shows the pattern: DPO tends to lift the reasoning and code benchmarks, and RL then lifts instruction-following and the hardest math.
OLMo 3 Think 7B across post-training stages. Bars are SFT / DPO / final RL. Hover-free: values are printed at the bar ends.
Delta Learning
The signal is the contrast, not the quality
Here is the counter-intuitive finding. By the time you have a strong SFT model, generating better completions with a bigger model and doing more SFT on them can hurt. On OLMo 3's dev checkpoint, continued SFT on Qwen3-32B thinking traces dropped the average from 70.3 to 64.5 — a 5.8-point regression. The model was already saturated on imitation. Delta Learning extracts signal from those same, now-useless completions by turning them into preference pairs: pick the strong completion as chosen and deliberately pair it with a much weaker one as rejected. The quality of the pairing depends on the difference between the two, not on either alone. Pairing Qwen3-32B with Qwen3-0.6B and running DPO lifted the average to 72.9 — better than the original SFT checkpoint.
Illustrative model of imitation vs contrastive signal. Lowering the rejected quality restores gradient even when the chosen model is saturated.
The Instruct model uses the same trick without thinking traces, and supplements it with delta-aware GPT-judged pairs: instead of trusting a judge to pick a good rejected answer from a pool of strong models, the pipeline forces weaker models into the pool and selects the worst response as rejected. Raising only the chosen model's quality had failed to improve over the OLMo 2 preference baseline; maximizing the delta fixed it.
Preference-tuning pitfalls
Dataset size, length bias, multi-turn
DPO has three failure modes that are easy to miss. First, more data is not monotonically better: AlpacaEval and ZebraLogic peak around 75–100K preference pairs and then decline, while AIME is still improving at that budget. The optimum depends on the downstream task, so dataset size is swept as a hyperparameter and early stopping is essential. Second, preference data has a length bias — chosen responses are systematically longer than rejected ones (80th percentile difference: 538 tokens for GPT-judged pairs, 564 for delta-heuristic pairs), and the model learns verbosity along with quality. OLMo 3 caps the difference at 100 tokens for chat and multi-turn subsets. Third, single-turn preferences don't teach multi-turn coherence; OLMo 3 synthesizes multi-turn conversations two ways (self-talk and synthetic-context) and changes only the final turn, so the ranking is unambiguous.
Preference dataset size (K pairs) vs benchmark score. Drag the budget to see which tasks are past their peak.
Chosen-minus-rejected token difference (illustrative distribution). Pairs above the cap are filtered from chat/multi-turn data.
GRPO, upgraded
OLMoRL
Part 9's GRPO used a group's own mean and standard deviation as the baseline. OLMoRL keeps the group-relative idea but changes seven things, most of them from DAPO and Dr.GRPO:
- Zero-gradient filtering — drop groups where every reward is identical (zero advantage, zero gradient).
- Active sampling — keep resampling prompts until the batch is full of non-zero-gradient completions, so filtering doesn't shrink batches over time.
- Token-level loss — normalize by total tokens in the batch, not per sample, removing a length bias.
- No KL loss — remove the reference-model penalty; updates are less restricted and training stays stable.
- Clip-higher — set the upper clipping bound above the lower one so probability can rise more than it falls.
- Truncated importance sampling — correct for the gap between inference-engine and trainer log-probabilities.
- No std normalization — divide by nothing: Ai = ri − mean(r) rather than (ri − mean)/std.
The last change fixes a difficulty bias: dividing by a small group standard deviation inflates the advantages of prompts that are almost always right or almost always wrong — exactly the prompts that carry the least learning signal. Removing the division stops those prompts from dominating the update.
Illustrative: how much std-normalization inflates a fixed reward gap as the group's reward spread shrinks.
The clipped surrogate objective at a fixed positive advantage. Clip-higher raises only the upper bound.
Per-token gradient weight by sequence length. Sample-level normalization makes long sequences count less per token; token-level makes every token equal.
Reward design across domains
One verifier per domain
RLVR replaces a learned reward model with a deterministic verifier, but “deterministic” looks different in each domain. OLMo 3 uses a separate verifier (or LM judge) per domain, and each has a characteristic failure mode. Pick a domain to see a candidate completion, its reward, and what typically goes wrong.
RL systems
Inference is the bottleneck
RL for long reasoning is an inference problem wearing a training costume. OLMo 3 generated rollouts up to 32K tokens (mean ≈ 14,628 for the reasoner models), and the learner spends 75% of its time waiting for data: the 32B run used 8 H100 nodes for training and 20 for inference, roughly 5× the compute for inference; the 7B run used 2 training nodes and 7 inference nodes, about 14×. The fix is a fully asynchronous actor–learner setup with three big wins.
Six variable-length rollouts (columns = generation steps). Grey cells are idle GPU slots. Continuous batching backfills as soon as a slot frees.
The other two wins: active sampling replaces DAPO's 3× oversampling with continuous resampling from a result queue, stabilizing batch size and loss; and inflight weight updates push new weights to the actors without pausing generation or invalidating the KV cache — up to 4× faster. Stacked together, the infrastructure ablation below takes throughput from 881 to 2,949 tokens/second and memory-bandwidth utilization from 12.9% to 43.2%.
OLMoRL infrastructure ablation: training throughput (tokens/second) as each component is added. MBU/MFU shown below each bar.
RL from base + contamination control
RL-Zero and spurious rewards
OLMo 3 also releases RL-Zero: RLVR applied directly to the midtrained base model with no SFT or DPO warm-up, over math, code, instruction-following, and chat. Two findings matter beyond the numbers. First, “simple” prompt templates beat standard post-trained templates when starting from a base model that never saw chat special tokens — the same special-token lesson from Part 5. Second, offline difficulty filtering: for each prompt, sample eight rollouts and remove prompts the model solves more than 62.5% of the time, keeping only prompts with learning signal. The 32B run then relies on active sampling at train time rather than offline filtering.
Pass-rate histogram from eight rollouts per prompt. Prompts right of the cutoff are filtered out as too easy.
Contamination is the silent failure of RLVR: if the evaluation data leaked into pretraining or midtraining, a model can improve on benchmarks by memorizing rather than reasoning. The check is elegant — train with spurious rewards, assigning random binary scores independent of correctness. Real learning cannot occur, so benchmark performance stays flat. If it rises, the eval was contaminated. OLMo 3's curve stays flat or degrades, evidence that decontamination worked.
Illustrative reward curves: a real verifier improves, a spurious (random) reward does not.
Cheat sheet
Recap
| Concept | What it is |
|---|---|
| Three-stage recipe | SFT → DPO → RLVR; DPO-before-RL beats SFT-before-RL. |
| Delta Learning | Preference signal comes from the chosen-minus-rejected contrast, not the chosen quality alone. |
| Delta-aware judging | Force weak models into the candidate pool and pick the worst response as rejected. |
| Dataset-size U-shape | Preference performance peaks (~75–100K pairs) then declines by task; sweep size and early-stop. |
| Length bias | Chosen responses skew long; cap the chosen-minus-rejected token difference. |
| Multi-turn preferences | Synthesize conversations and vary only the final turn so the ranking is unambiguous. |
| Zero-gradient filtering | Drop groups with identical rewards (no advantage). |
| Active sampling | Resample until the batch has enough non-zero-gradient completions. |
| Token-level loss | Normalize by total batch tokens, removing length bias. |
| Clip-higher | Upper clip bound above lower bound; let probabilities rise more than fall. |
| Truncated importance sampling | Correct inference-engine vs trainer log-prob mismatch. |
| No std normalization | Advantage = reward − group mean; avoids difficulty bias. |
| Verifier zoo | SymPy math, test-case code (cloud), IF constraints, LM-judge chat. |
| Offline difficulty filtering | Remove prompts solved >62.5% of 8 rollouts. |
| Spurious-reward control | Random rewards should not improve benchmarks; if they do, data is contaminated. |
| Inflight updates | Update actor weights without pausing generation — up to 4× throughput. |
Further reading
References
- OLMo 3 Team (Ai2), "OLMo 3" (2025) — the full post-training recipe (§4–6).
- Geng et al., "Delta Learning" (2025) — the contrast is the signal.
- Yu et al., "DAPO: An Open-Source LLM Reinforcement Learning System at Scale" (2025) — clip-higher, token-level loss, dynamic sampling.
- Liu et al., "Understanding R1-Zero-Like Training: A Critical Perspective" (2025) — Dr.GRPO, no std normalization, prompt templates.
- Shao et al., "Spurious Rewards: Rethinking Training Signals in RLVR" (2025) — contamination control.
- Hu et al., "Open-Reasoner-Zero" (2025) and DeepSeek-AI, "DeepSeek-R1" (2025).
- Noukhovitch et al., "Asynchronous RLHF" (2024) and the PipelineRL approach — async actor–learner updates.
- Lambert et al., "Tulu 3" (2024) — the recipe OLMo 3 builds on.