Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Chain-of-thought as computation-in-tokens

Setup

A transformer does a fixed amount of computation per generated token (one forward pass through L layers). If a problem needs more computation than that, the model has no way to "think longer" on a single token — but it can spend more tokens: writing out intermediate reasoning steps effectively lets the model use its next-token machinery as a scratchpad, trading generation length for solved-problem difficulty. Models explicitly trained to do this at length (o1-style, DeepSeek-R1, OLMo 3 Think) can spend thousands of tokens "thinking" before answering.

2

RLVR: reward from a verifier, not a model

Tulu 3's third stage

RLHF's reward comes from a learned reward model trained on human preference judgments — useful for open-ended quality, but noisy and gameable (Part 8's reward-hacking demo). For tasks with a checkable correct answer, Tulu 3 introduced RLVR (Reinforcement Learning with Verifiable Rewards): skip the reward model, and instead score each completion with a small deterministic verifier function — does the final numeric answer match, does the code pass its unit tests, was the requested format followed exactly.

The verifier here just extracts the last number in the response and checks it against 408 — crude, but this is genuinely close to how math RLVR verifiers work in practice (answer extraction + exact or numeric match), which is also exactly why format matters so much: an otherwise-correct answer the verifier can't parse gets zero reward.

3

GRPO: dropping the value model

Group-relative advantages

PPO (Part 8) needs a learned value model to estimate how good a state is, so it can compute an advantage (reward minus baseline) — that's a second model roughly the size of the policy, trained alongside it, doubling memory and compute. GRPO (Group Relative Policy Optimization) removes it: sample a group of G completions for the same prompt, score each with the verifier, and use the group's own mean and standard deviation as the baseline.

$$\hat{A}_i = \frac{r_i - \text{mean}(r_1,\dots,r_G)}{\text{std}(r_1,\dots,r_G)}$$

Tulu 3.1 switched its RLVR stage from PPO to GRPO for exactly this reason — comparable results, less compute per step, no separate value model to train and keep synchronized.

⚠️ Notice: a completion with reward 1 gets a negative advantage if every other sample in its group also scored 1 (mean≈1, nothing to distinguish it) — advantage is entirely relative to what else was sampled for the same prompt, which is both GRPO's efficiency trick and its main quirk (very easy or very hard prompts, where every sample gets the same reward, contribute almost no learning signal).
4

Test-time compute: more samples, higher accuracy

A second axis besides training compute

If a model solves a given problem with probability p per independent attempt, sampling k attempts and checking if any is correct (pass@k, reused from Part 10) succeeds with probability $1-(1-p)^k$. Majority voting — sample k times, take the most common final answer — does worse than pass@k (it needs the correct answer to be the plurality, not just present) but doesn't require an external verifier at inference time, which is what makes it usable in production.

5

From GRPO to OLMoRL — and RL from base

The 2025 upgrades

GRPO as introduced above is now the starting point, not the final word. OLMoRL builds on it with ideas from DAPO and Dr.GRPO, and each change fixes a concrete failure of vanilla GRPO. In one line: filter the zero-signal groups, refill the batch, normalize the loss by tokens, drop the KL term, clip asymmetrically, correct for engine mismatch, and stop dividing by the group's standard deviation. Part 15 opens each one up with demos.

  • Zero-gradient filtering — drop groups whose rewards are all identical (no advantage, no gradient).
  • Active sampling — keep pulling prompt–completion pairs from the result queue so filtering doesn't shrink the batch over training.
  • Token-level loss — normalize by total batch tokens, not per sample, removing a length bias.
  • No KL loss — less-restricted updates without observed over-optimization.
  • Clip-higher — asymmetric clipping lets probability rise more than it falls.
  • Truncated importance sampling — correct the log-prob gap between the vLLM inference engine and the trainer.
  • No std normalization — advantage becomes r − mean(r), removing the difficulty bias that inflates near-saturated prompts.

Verifiers go multi-domain. Tulu 3 and early RLVR mostly rewarded math. OLMo 3 expands to four domains, each with its own checker: math (SymPy answer equivalence), code (test-case execution on AWS Lambda, so verification never blocks the trainer), instruction following (constraint checkers, all-or-nothing), and chat (an LM judge, with or without a reference answer). This is what turns RLVR from a math trick into a general post-training stage.

Difficulty filtering is where the sample efficiency comes from. For each prompt, sample eight rollouts from the starting checkpoint and discard prompts the model already solves more than 62.5% of the time. Those prompts produce near-zero advantage and waste compute; keeping only the hard-but-solvable middle makes each RL step count. (The 32B run then relies on active sampling instead of offline filtering.)

RL-Zero applies all of this directly to the midtrained base model, with no SFT or DPO warm start. Two findings generalize: simple prompt templates beat standard chat templates when the base model never saw special tokens, and the choice of midtraining data determines whether RL learns to reason longer — a base with too little reasoning data plateaus instead of growing its response length. Because RL-Zero starts from a fully open pipeline, it also enables a clean contamination test: train with spurious random rewards. If benchmark performance rises, the eval leaked into pretraining; OLMo 3's flat curve is evidence it didn't. Part 15 walks through both the difficulty filter and the spurious-reward control.

⚠️ Multi-objective RL is harder than single-domain RL. A model trained on math alone reaches a higher math reward than one trained on math + code + instruction + chat. But the single-domain model overfits and generalizes worse. The mixed run is the better benchmark for real assistants — and the reason OLMo 3's final RL mix is roughly balanced across domains.
📌 Cross-links: Part 15 has the OLMoRL objective, the domain-verifier zoo, the difficulty-filter histogram, and the async RL systems that make long-chain-of-thought training practical.
✓

Cheat sheet

Recap

ConceptWhat it is
Chain-of-thoughtSpending generated tokens as computation, not just output
RLVRRL reward from a deterministic verifier instead of a learned reward model
GRPOAdvantage from a sampled group's own mean/std — no value model
pass@kP(at least one of k samples correct) = 1−(1−p)ᵏ
Majority votingSample k, answer with the most common output — usable without a verifier at inference
Zero-gradient filteringDrop groups with identical rewards — no advantage means no gradient
Active samplingResample prompts until the batch is full of non-zero-gradient completions
Clip-higher / token-level lossAsymmetric clipping and total-token normalization (DAPO-style GRPO upgrades)
No std normalizationAdvantage = reward − group mean; avoids difficulty bias (Dr.GRPO)
Multi-domain verifiersSymPy math, test-case code, IF constraint checks, LM-judge chat
Difficulty filteringRemove prompts solved >62.5% of 8 rollouts to keep only learnable signal
RL-ZeroRLVR directly from the base model, no SFT/DPO warm start
Spurious-reward controlRandom rewards should not improve benchmarks; if they do, data is contaminated
📚

Further reading

References

?

Check your understanding

0/5 answered
All of these techniques need a way to measure whether they worked — which is a surprisingly deep question in its own right. Continue: evaluation →