Alignment: RLHF, Reward Models, and DPO
SFT teaches a model to produce a plausible response in the right format. It has no mechanism for the much more common real situation: two responses are both reasonable, but one is clearly better. That comparative judgment is what alignment is built to train on, using human (or model) preferences rather than single correct answers — and it comes with its own failure mode, reward hacking, that this part now makes concrete rather than just naming.
Preference data: comparisons, not corrections
Where alignment data comes from
Instead of writing "the correct answer," labelers (human annotators, or sometimes another model acting as a judge — Part 10 covers the pitfalls of that) are shown a prompt with two candidate responses and asked which is better. Each judgment becomes one training example: a chosen response and a rejected response for the same prompt.
Training a reward model
Method 1 of 2 — the RLHF route
The classic RLHF recipe first trains a separate reward model: a copy of the language model with its output head replaced by a single scalar score, trained so the chosen response scores higher than the rejected one — the Bradley-Terry loss, reused in Part 10's arena-rating math:
Reward models are themselves evaluated — Ai2's RewardBench is a standard benchmark specifically for how well a reward model's scores match held-out human preference judgments, since a flawed reward model quietly poisons everything trained against it downstream.
PPO: the clipped objective, and the KL–reward frontier
Method 1 of 2, continued
With a trained reward model, the SFT model (now the policy) is optimized with PPO to generate higher-scoring responses. PPO's defining trick is the clipped surrogate objective: it caps how much the policy is allowed to change in one update relative to the policy that generated the data, based on the probability ratio $r = \pi_\theta(y|x)/\pi_{old}(y|x)$:
With positive advantage (this action was better than expected), the objective flattens once the policy has already increased that action's probability past 1+ε — no incentive to keep pushing it further in one step. With negative advantage, clipping works the mirror way. This is what keeps a single batch from swinging the policy too far.
Separately, PPO's objective also includes a KL-divergence penalty against the frozen SFT reference, controlled by β, trading off reward against how far the policy drifts:
Best-of-n: alignment without any training
A simpler baseline
Before reaching for RL, a surprisingly strong baseline: sample n responses from the SFT model, score each with the reward model, and simply return the highest-scoring one at inference time (rejection sampling / best-of-n). No policy update at all — just spending more inference compute to pick better among what the model could already produce.
Best-of-n's achieved reward rises with n, but so does its own KL divergence from the reference policy — it's not a free lunch, and it costs n× the inference compute per response actually returned.
DPO: skipping the reward model entirely
Method 2 of 2 — the simpler route
Direct Preference Optimization (DPO) substitutes the KL-regularized objective's closed-form optimal policy back into the reward-model loss, producing an objective that trains the policy directly on preference pairs — no separate reward model, no RL loop:
The bracketed quantity is DPO's implicit reward — the model's own log-probability ratio versus the reference, standing in for a learned reward score. Edit the four values below (each is a toy log-probability) and watch the implicit reward margin and per-pair gradient weight move:
No reward model to train, no PPO rollouts — closer in engineering complexity to SFT than to a full RL pipeline. Since its introduction, a family of variants has emerged tweaking exactly how the implicit reward is shaped: IPO (a bounded loss less prone to overfitting confident preferences), KTO (trains on unpaired binary "good/bad" labels instead of pairs), SimPO and ORPO (drop the reference model entirely, trading a little robustness for simplicity). β itself controls how tightly the policy is tied to the reference in every variant — small β allows bigger behavior changes per preference pair, large β is more conservative.
Preference-tuning at scale: delta, size, length, and turns
OLMo 3's preference stage
OLMo 3 turns preference tuning from a nicety into a capability stage, and in doing so surfaces four lessons that apply to any DPO run.
1 · Delta Learning — the contrast is the signal. On a strong SFT checkpoint, continuing to imitate better responses from a larger model can hurt: OLMo 3's dev checkpoint dropped from 70.3 → 64.5 when supervised on Qwen3-32B thinking traces. Pairing those same strong completions (chosen) with deliberately weak ones from Qwen3-0.6B (rejected) and running DPO instead lifted the average to 72.9. The quality of a preference pair depends on the gap between the two responses, not on the chosen response alone.
2 · Dataset size is a U-shape. Preference performance does not improve monotonically with more pairs. AlpacaEval and ZebraLogic peak around 75–100K pairs and then decline, while AIME is still improving at that budget. Early stopping matters, and the best size depends on the downstream task — so size is swept as a hyperparameter rather than fixed.
3 · Watch the length bias. Chosen responses are systematically longer than rejected ones; at the 80th percentile the difference is 538 tokens for GPT-judged pairs and 564 for delta-heuristic pairs. Models learn verbosity along with quality. OLMo 3 filters chat and multi-turn pairs to a 100-token chosen-minus-rejected difference, accepting slightly lower length-sensitive benchmark scores for much more usable output.
4 · Teach multi-turn coherence explicitly. Single-turn pairs say nothing about staying on topic across turns. OLMo 3 synthesizes multi-turn conversations two ways — self-talk (the model generates follow-ups) and synthetic-context (related independent questions become earlier turns) — and varies only the final turn, so there is no ambiguity about which turn the preference applies to.
Finally, combine complementary signals: delta-learning heuristic pairs and delta-aware GPT-judged pairs produce different spreads of gains, and using both beats either alone. The GPT-judged pipeline had to be redesigned to maximize the delta (force weaker models into the candidate pool, pick the worst response as rejected) — improving the judge or the generator pool alone did nothing.
Cheat sheet
Recap
| RLHF (PPO) | DPO | |
|---|---|---|
| Needs a separate reward model? | Yes | No |
| Needs an RL loop? | Yes (PPO, clipped objective) | No — direct supervised-style loss |
| Models held in memory | Policy + reward model + frozen reference | Policy + frozen reference (or none, for SimPO/ORPO) |
| Main risk | Reward hacking against the reward model | Less flexible for reward signals beyond static preference pairs |
| Delta Learning | Pair a strong chosen response with a deliberately weak rejected one; the contrast, not the chosen quality, carries the signal. | |
| Dataset size | U-shaped: peaks around 75–100K pairs for some evals while others keep rising — sweep size and early-stop. | |
| Length bias | Chosen responses skew long; cap the chosen-minus-rejected difference (OLMo 3 uses 100 tokens). | |
| Multi-turn preferences | Synthesize conversations and vary only the final turn so the ranking is unambiguous. | |
Further reading
References
- Ouyang et al., "Training Language Models to Follow Instructions with Human Feedback" (2022).
- Schulman et al., "Proximal Policy Optimization Algorithms" (2017).
- Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023).
- Lambert et al. (Ai2), "RewardBench: Evaluating Reward Models for Language Modeling" (2024).
- Lambert et al. (Ai2), "Tulu 3" (2024) — SFT + DPO + RLVR recipe used for OLMo.