Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The three-stage recipe

SFT → DPO → RLVR

OLMo 3 Think is trained in three stages on top of the midtrained base: supervised fine-tuning on curated thinking traces (Dolci-Think-SFT), preference tuning (Dolci-Think-DPO), then reinforcement learning with verifiable rewards (OLMoRL). Many open reasoning models stop after SFT — OpenThoughts3 and s1 used only SFT; SmolLM3 used SFT and DPO with no RL. OLMo 3's claim is that each stage adds something the previous one cannot, and that RL works better when it starts from a DPO checkpoint than directly from SFT. The stage table below shows the pattern: DPO tends to lift the reasoning and code benchmarks, and RL then lifts instruction-following and the hardest math.

OLMo 3 Think 7B across post-training stages. Bars are SFT / DPO / final RL. Hover-free: values are printed at the bar ends.

📌 Why order matters: DPO first gives the policy a cleaner starting distribution for RL, and filtering the RL data with the DPO checkpoint (offline difficulty filtering) selects prompts that still have signal. The rest of this part is mostly about making each stage actually work.
2

Delta Learning

The signal is the contrast, not the quality

Here is the counter-intuitive finding. By the time you have a strong SFT model, generating better completions with a bigger model and doing more SFT on them can hurt. On OLMo 3's dev checkpoint, continued SFT on Qwen3-32B thinking traces dropped the average from 70.3 to 64.5 — a 5.8-point regression. The model was already saturated on imitation. Delta Learning extracts signal from those same, now-useless completions by turning them into preference pairs: pick the strong completion as chosen and deliberately pair it with a much weaker one as rejected. The quality of the pairing depends on the difference between the two, not on either alone. Pairing Qwen3-32B with Qwen3-0.6B and running DPO lifted the average to 72.9 — better than the original SFT checkpoint.

$$\text{gain} \;\propto\; \big(y_c - y_r\big), \qquad y_c \succ y_r$$

Illustrative model of imitation vs contrastive signal. Lowering the rejected quality restores gradient even when the chosen model is saturated.

The Instruct model uses the same trick without thinking traces, and supplements it with delta-aware GPT-judged pairs: instead of trusting a judge to pick a good rejected answer from a pool of strong models, the pipeline forces weaker models into the pool and selects the worst response as rejected. Raising only the chosen model's quality had failed to improve over the OLMo 2 preference baseline; maximizing the delta fixed it.

3

Preference-tuning pitfalls

Dataset size, length bias, multi-turn

DPO has three failure modes that are easy to miss. First, more data is not monotonically better: AlpacaEval and ZebraLogic peak around 75–100K preference pairs and then decline, while AIME is still improving at that budget. The optimum depends on the downstream task, so dataset size is swept as a hyperparameter and early stopping is essential. Second, preference data has a length bias — chosen responses are systematically longer than rejected ones (80th percentile difference: 538 tokens for GPT-judged pairs, 564 for delta-heuristic pairs), and the model learns verbosity along with quality. OLMo 3 caps the difference at 100 tokens for chat and multi-turn subsets. Third, single-turn preferences don't teach multi-turn coherence; OLMo 3 synthesizes multi-turn conversations two ways (self-talk and synthetic-context) and changes only the final turn, so the ranking is unambiguous.

Preference dataset size (K pairs) vs benchmark score. Drag the budget to see which tasks are past their peak.

Chosen-minus-rejected token difference (illustrative distribution). Pairs above the cap are filtered from chat/multi-turn data.

4

GRPO, upgraded

OLMoRL

Part 9's GRPO used a group's own mean and standard deviation as the baseline. OLMoRL keeps the group-relative idea but changes seven things, most of them from DAPO and Dr.GRPO:

  • Zero-gradient filtering — drop groups where every reward is identical (zero advantage, zero gradient).
  • Active sampling — keep resampling prompts until the batch is full of non-zero-gradient completions, so filtering doesn't shrink batches over time.
  • Token-level loss — normalize by total tokens in the batch, not per sample, removing a length bias.
  • No KL loss — remove the reference-model penalty; updates are less restricted and training stays stable.
  • Clip-higher — set the upper clipping bound above the lower one so probability can rise more than it falls.
  • Truncated importance sampling — correct for the gap between inference-engine and trainer log-probabilities.
  • No std normalization — divide by nothing: Ai = ri − mean(r) rather than (ri − mean)/std.
$$\hat{A}_i = r_i - \operatorname{mean}\big(\{r_1,\dots,r_G\}\big)$$

The last change fixes a difficulty bias: dividing by a small group standard deviation inflates the advantages of prompts that are almost always right or almost always wrong — exactly the prompts that carry the least learning signal. Removing the division stops those prompts from dominating the update.

Illustrative: how much std-normalization inflates a fixed reward gap as the group's reward spread shrinks.

The clipped surrogate objective at a fixed positive advantage. Clip-higher raises only the upper bound.

Per-token gradient weight by sequence length. Sample-level normalization makes long sequences count less per token; token-level makes every token equal.

5

Reward design across domains

One verifier per domain

RLVR replaces a learned reward model with a deterministic verifier, but “deterministic” looks different in each domain. OLMo 3 uses a separate verifier (or LM judge) per domain, and each has a characteristic failure mode. Pick a domain to see a candidate completion, its reward, and what typically goes wrong.

⚠️ Recall the format sensitivity from Part 9: an answer the verifier cannot parse gets zero reward no matter how correct it is. That is why OLMo 3 cleans evaluation prompts (stripping \boxed{}) to look like training prompts, and why format is a first-class part of reward design.
6

RL systems

Inference is the bottleneck

RL for long reasoning is an inference problem wearing a training costume. OLMo 3 generated rollouts up to 32K tokens (mean ≈ 14,628 for the reasoner models), and the learner spends 75% of its time waiting for data: the 32B run used 8 H100 nodes for training and 20 for inference, roughly 5× the compute for inference; the 7B run used 2 training nodes and 7 inference nodes, about 14×. The fix is a fully asynchronous actor–learner setup with three big wins.

Six variable-length rollouts (columns = generation steps). Grey cells are idle GPU slots. Continuous batching backfills as soon as a slot frees.

The other two wins: active sampling replaces DAPO's 3× oversampling with continuous resampling from a result queue, stabilizing batch size and loss; and inflight weight updates push new weights to the actors without pausing generation or invalidating the KV cache — up to 4× faster. Stacked together, the infrastructure ablation below takes throughput from 881 to 2,949 tokens/second and memory-bandwidth utilization from 12.9% to 43.2%.

OLMoRL infrastructure ablation: training throughput (tokens/second) as each component is added. MBU/MFU shown below each bar.

7

RL from base + contamination control

RL-Zero and spurious rewards

OLMo 3 also releases RL-Zero: RLVR applied directly to the midtrained base model with no SFT or DPO warm-up, over math, code, instruction-following, and chat. Two findings matter beyond the numbers. First, “simple” prompt templates beat standard post-trained templates when starting from a base model that never saw chat special tokens — the same special-token lesson from Part 5. Second, offline difficulty filtering: for each prompt, sample eight rollouts and remove prompts the model solves more than 62.5% of the time, keeping only prompts with learning signal. The 32B run then relies on active sampling at train time rather than offline filtering.

Pass-rate histogram from eight rollouts per prompt. Prompts right of the cutoff are filtered out as too easy.

Contamination is the silent failure of RLVR: if the evaluation data leaked into pretraining or midtraining, a model can improve on benchmarks by memorizing rather than reasoning. The check is elegant — train with spurious rewards, assigning random binary scores independent of correctness. Real learning cannot occur, so benchmark performance stays flat. If it rises, the eval was contaminated. OLMo 3's curve stays flat or degrades, evidence that decontamination worked.

Illustrative reward curves: a real verifier improves, a spurious (random) reward does not.

📌 The through-line: the recipe only works because the base model was midtrained with the right data (Part 5), the context window was extended to hold long chains of thought (Part 13), the preference stage was redesigned around contrast (Part 8), and RL ran on infrastructure fast enough to iterate (Part 11). Post-training at scale is the integration of everything before it.
✓

Cheat sheet

Recap

ConceptWhat it is
Three-stage recipeSFT → DPO → RLVR; DPO-before-RL beats SFT-before-RL.
Delta LearningPreference signal comes from the chosen-minus-rejected contrast, not the chosen quality alone.
Delta-aware judgingForce weak models into the candidate pool and pick the worst response as rejected.
Dataset-size U-shapePreference performance peaks (~75–100K pairs) then declines by task; sweep size and early-stop.
Length biasChosen responses skew long; cap the chosen-minus-rejected token difference.
Multi-turn preferencesSynthesize conversations and vary only the final turn so the ranking is unambiguous.
Zero-gradient filteringDrop groups with identical rewards (no advantage).
Active samplingResample until the batch has enough non-zero-gradient completions.
Token-level lossNormalize by total batch tokens, removing length bias.
Clip-higherUpper clip bound above lower bound; let probabilities rise more than fall.
Truncated importance samplingCorrect inference-engine vs trainer log-prob mismatch.
No std normalizationAdvantage = reward − group mean; avoids difficulty bias.
Verifier zooSymPy math, test-case code (cloud), IF constraints, LM-judge chat.
Offline difficulty filteringRemove prompts solved >62.5% of 8 rollouts.
Spurious-reward controlRandom rewards should not improve benchmarks; if they do, data is contaminated.
Inflight updatesUpdate actor weights without pausing generation — up to 4× throughput.
📚

Further reading

References

?

Check your understanding

0/5 answered
You've reached the end of the series. Every term introduced across all 15 parts — linked back to where it appears — is collected in the glossary. Continue: glossary & numbers →