Reasoning & RL with Verifiable Rewards
Part 8 covered alignment with human preferences. A parallel, more recent development trains models to be better at tasks with an automatically checkable answer — math, code, precise instruction-following — using reinforcement learning where the reward comes from a verifier function, not a learned reward model. This is the biggest change to post-training since RLHF itself, and it's what powers "thinking" models that write out long chains of reasoning before answering.
Chain-of-thought as computation-in-tokens
Setup
A transformer does a fixed amount of computation per generated token (one forward pass through L layers). If a problem needs more computation than that, the model has no way to "think longer" on a single token — but it can spend more tokens: writing out intermediate reasoning steps effectively lets the model use its next-token machinery as a scratchpad, trading generation length for solved-problem difficulty. Models explicitly trained to do this at length (o1-style, DeepSeek-R1, OLMo 3 Think) can spend thousands of tokens "thinking" before answering.
RLVR: reward from a verifier, not a model
Tulu 3's third stage
RLHF's reward comes from a learned reward model trained on human preference judgments — useful for open-ended quality, but noisy and gameable (Part 8's reward-hacking demo). For tasks with a checkable correct answer, Tulu 3 introduced RLVR (Reinforcement Learning with Verifiable Rewards): skip the reward model, and instead score each completion with a small deterministic verifier function — does the final numeric answer match, does the code pass its unit tests, was the requested format followed exactly.
The verifier here just extracts the last number in the response and checks it against 408 — crude, but this is genuinely close to how math RLVR verifiers work in practice (answer extraction + exact or numeric match), which is also exactly why format matters so much: an otherwise-correct answer the verifier can't parse gets zero reward.
GRPO: dropping the value model
Group-relative advantages
PPO (Part 8) needs a learned value model to estimate how good a state is, so it can compute an advantage (reward minus baseline) — that's a second model roughly the size of the policy, trained alongside it, doubling memory and compute. GRPO (Group Relative Policy Optimization) removes it: sample a group of G completions for the same prompt, score each with the verifier, and use the group's own mean and standard deviation as the baseline.
Tulu 3.1 switched its RLVR stage from PPO to GRPO for exactly this reason — comparable results, less compute per step, no separate value model to train and keep synchronized.
Test-time compute: more samples, higher accuracy
A second axis besides training compute
If a model solves a given problem with probability p per independent attempt, sampling k attempts and checking if any is correct (pass@k, reused from Part 10) succeeds with probability $1-(1-p)^k$. Majority voting — sample k times, take the most common final answer — does worse than pass@k (it needs the correct answer to be the plurality, not just present) but doesn't require an external verifier at inference time, which is what makes it usable in production.
From GRPO to OLMoRL — and RL from base
The 2025 upgrades
GRPO as introduced above is now the starting point, not the final word. OLMoRL builds on it with ideas from DAPO and Dr.GRPO, and each change fixes a concrete failure of vanilla GRPO. In one line: filter the zero-signal groups, refill the batch, normalize the loss by tokens, drop the KL term, clip asymmetrically, correct for engine mismatch, and stop dividing by the group's standard deviation. Part 15 opens each one up with demos.
- Zero-gradient filtering — drop groups whose rewards are all identical (no advantage, no gradient).
- Active sampling — keep pulling prompt–completion pairs from the result queue so filtering doesn't shrink the batch over training.
- Token-level loss — normalize by total batch tokens, not per sample, removing a length bias.
- No KL loss — less-restricted updates without observed over-optimization.
- Clip-higher — asymmetric clipping lets probability rise more than it falls.
- Truncated importance sampling — correct the log-prob gap between the vLLM inference engine and the trainer.
- No std normalization — advantage becomes r − mean(r), removing the difficulty bias that inflates near-saturated prompts.
Verifiers go multi-domain. Tulu 3 and early RLVR mostly rewarded math. OLMo 3 expands to four domains, each with its own checker: math (SymPy answer equivalence), code (test-case execution on AWS Lambda, so verification never blocks the trainer), instruction following (constraint checkers, all-or-nothing), and chat (an LM judge, with or without a reference answer). This is what turns RLVR from a math trick into a general post-training stage.
Difficulty filtering is where the sample efficiency comes from. For each prompt, sample eight rollouts from the starting checkpoint and discard prompts the model already solves more than 62.5% of the time. Those prompts produce near-zero advantage and waste compute; keeping only the hard-but-solvable middle makes each RL step count. (The 32B run then relies on active sampling instead of offline filtering.)
RL-Zero applies all of this directly to the midtrained base model, with no SFT or DPO warm start. Two findings generalize: simple prompt templates beat standard chat templates when the base model never saw special tokens, and the choice of midtraining data determines whether RL learns to reason longer — a base with too little reasoning data plateaus instead of growing its response length. Because RL-Zero starts from a fully open pipeline, it also enables a clean contamination test: train with spurious random rewards. If benchmark performance rises, the eval leaked into pretraining; OLMo 3's flat curve is evidence it didn't. Part 15 walks through both the difficulty filter and the spurious-reward control.
Cheat sheet
Recap
| Concept | What it is |
|---|---|
| Chain-of-thought | Spending generated tokens as computation, not just output |
| RLVR | RL reward from a deterministic verifier instead of a learned reward model |
| GRPO | Advantage from a sampled group's own mean/std — no value model |
| pass@k | P(at least one of k samples correct) = 1−(1−p)ᵏ |
| Majority voting | Sample k, answer with the most common output — usable without a verifier at inference |
| Zero-gradient filtering | Drop groups with identical rewards — no advantage means no gradient |
| Active sampling | Resample prompts until the batch is full of non-zero-gradient completions |
| Clip-higher / token-level loss | Asymmetric clipping and total-token normalization (DAPO-style GRPO upgrades) |
| No std normalization | Advantage = reward − group mean; avoids difficulty bias (Dr.GRPO) |
| Multi-domain verifiers | SymPy math, test-case code, IF constraint checks, LM-judge chat |
| Difficulty filtering | Remove prompts solved >62.5% of 8 rollouts to keep only learnable signal |
| RL-Zero | RLVR directly from the base model, no SFT/DPO warm start |
| Spurious-reward control | Random rewards should not improve benchmarks; if they do, data is contaminated |
Further reading
References
- Lambert et al. (Ai2), "Tulu 3: Pushing Frontiers in Open Language Model Post-Training" (2024) — RLVR.
- Shao et al. (DeepSeek), "DeepSeekMath: Pushing the Limits of Mathematical Reasoning" (2024) — GRPO.
- DeepSeek-AI, "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning" (2025).
- Wang et al., "Self-Consistency Improves Chain of Thought Reasoning" (2022) — majority voting.