Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Preference data: comparisons, not corrections

Where alignment data comes from

Instead of writing "the correct answer," labelers (human annotators, or sometimes another model acting as a judge — Part 10 covers the pitfalls of that) are shown a prompt with two candidate responses and asked which is better. Each judgment becomes one training example: a chosen response and a rejected response for the same prompt.

Response A

Response B

Pick the response you think a labeler would prefer, then compare with the intended answer.
⚠️ Real preference data has real biases: annotators (human or model) measurably favor longer responses (length/verbosity bias) and responses that flatter the asker's stated view (sycophancy bias), independent of actual quality — both get baked into whatever is trained on this data unless specifically corrected for.
2

Training a reward model

Method 1 of 2 — the RLHF route

The classic RLHF recipe first trains a separate reward model: a copy of the language model with its output head replaced by a single scalar score, trained so the chosen response scores higher than the rejected one — the Bradley-Terry loss, reused in Part 10's arena-rating math:

$$L = -\log\sigma\big(r_\theta(x,y_w) - r_\theta(x,y_l)\big), \qquad y_w = \text{chosen},\ y_l = \text{rejected}$$

Reward models are themselves evaluated — Ai2's RewardBench is a standard benchmark specifically for how well a reward model's scores match held-out human preference judgments, since a flawed reward model quietly poisons everything trained against it downstream.

3

PPO: the clipped objective, and the KL–reward frontier

Method 1 of 2, continued

With a trained reward model, the SFT model (now the policy) is optimized with PPO to generate higher-scoring responses. PPO's defining trick is the clipped surrogate objective: it caps how much the policy is allowed to change in one update relative to the policy that generated the data, based on the probability ratio $r = \pi_\theta(y|x)/\pi_{old}(y|x)$:

$$L^{CLIP} = \min\big(r \cdot \hat{A},\ \text{clip}(r, 1-\epsilon, 1+\epsilon)\cdot \hat{A}\big)$$

With positive advantage (this action was better than expected), the objective flattens once the policy has already increased that action's probability past 1+ε — no incentive to keep pushing it further in one step. With negative advantage, clipping works the mirror way. This is what keeps a single batch from swinging the policy too far.

Separately, PPO's objective also includes a KL-divergence penalty against the frozen SFT reference, controlled by β, trading off reward against how far the policy drifts:

$$\text{maximize}\ \ E[r_\theta(x,y)] - \beta \cdot \mathrm{KL}(\pi_\theta \Vert \pi_{ref})$$
⚠️ Reward hacking, made concrete: as β shrinks, reward-model score keeps climbing but drifts past the point where it reflects real quality — the curve above is built so that past a KL "cliff," reward-model score and actual quality visibly diverge. This is the literal mechanism behind the classic failure of RLHF policies learning to pad responses, use confident-sounding hedge phrases, or otherwise exploit reward-model quirks rather than genuinely improve.
4

Best-of-n: alignment without any training

A simpler baseline

Before reaching for RL, a surprisingly strong baseline: sample n responses from the SFT model, score each with the reward model, and simply return the highest-scoring one at inference time (rejection sampling / best-of-n). No policy update at all — just spending more inference compute to pick better among what the model could already produce.

Best-of-n's achieved reward rises with n, but so does its own KL divergence from the reference policy — it's not a free lunch, and it costs n× the inference compute per response actually returned.

5

DPO: skipping the reward model entirely

Method 2 of 2 — the simpler route

Direct Preference Optimization (DPO) substitutes the KL-regularized objective's closed-form optimal policy back into the reward-model loss, producing an objective that trains the policy directly on preference pairs — no separate reward model, no RL loop:

$$L = -\log\sigma\Big(\beta\big[\log\tfrac{\pi_\theta(y_w|x)}{\pi_{ref}(y_w|x)} - \log\tfrac{\pi_\theta(y_l|x)}{\pi_{ref}(y_l|x)}\big]\Big)$$

The bracketed quantity is DPO's implicit reward — the model's own log-probability ratio versus the reference, standing in for a learned reward score. Edit the four values below (each is a toy log-probability) and watch the implicit reward margin and per-pair gradient weight move:

No reward model to train, no PPO rollouts — closer in engineering complexity to SFT than to a full RL pipeline. Since its introduction, a family of variants has emerged tweaking exactly how the implicit reward is shaped: IPO (a bounded loss less prone to overfitting confident preferences), KTO (trains on unpaired binary "good/bad" labels instead of pairs), SimPO and ORPO (drop the reference model entirely, trading a little robustness for simplicity). β itself controls how tightly the policy is tied to the reference in every variant — small β allows bigger behavior changes per preference pair, large β is more conservative.

💡 Online vs. offline preference data: DPO as described trains on a fixed, pre-collected preference dataset (offline). Newer setups instead sample fresh completions from the current policy and get them freshly judged during training (online DPO/RLHF) — closer to PPO's loop, at the cost of needing a live judge or reward model rather than a static dataset.
💡 OLMo's actual recipe: Ai2's Tulu 3 post-training pipeline combines SFT, then DPO on preference data, then RLVR (Part 9) — a concrete, published example of mixing these techniques rather than treating them as mutually exclusive.
6

Preference-tuning at scale: delta, size, length, and turns

OLMo 3's preference stage

OLMo 3 turns preference tuning from a nicety into a capability stage, and in doing so surfaces four lessons that apply to any DPO run.

1 · Delta Learning — the contrast is the signal. On a strong SFT checkpoint, continuing to imitate better responses from a larger model can hurt: OLMo 3's dev checkpoint dropped from 70.3 → 64.5 when supervised on Qwen3-32B thinking traces. Pairing those same strong completions (chosen) with deliberately weak ones from Qwen3-0.6B (rejected) and running DPO instead lifted the average to 72.9. The quality of a preference pair depends on the gap between the two responses, not on the chosen response alone.

2 · Dataset size is a U-shape. Preference performance does not improve monotonically with more pairs. AlpacaEval and ZebraLogic peak around 75–100K pairs and then decline, while AIME is still improving at that budget. Early stopping matters, and the best size depends on the downstream task — so size is swept as a hyperparameter rather than fixed.

3 · Watch the length bias. Chosen responses are systematically longer than rejected ones; at the 80th percentile the difference is 538 tokens for GPT-judged pairs and 564 for delta-heuristic pairs. Models learn verbosity along with quality. OLMo 3 filters chat and multi-turn pairs to a 100-token chosen-minus-rejected difference, accepting slightly lower length-sensitive benchmark scores for much more usable output.

4 · Teach multi-turn coherence explicitly. Single-turn pairs say nothing about staying on topic across turns. OLMo 3 synthesizes multi-turn conversations two ways — self-talk (the model generates follow-ups) and synthetic-context (related independent questions become earlier turns) — and varies only the final turn, so there is no ambiguity about which turn the preference applies to.

Finally, combine complementary signals: delta-learning heuristic pairs and delta-aware GPT-judged pairs produce different spreads of gains, and using both beats either alone. The GPT-judged pipeline had to be redesigned to maximize the delta (force weaker models into the candidate pool, pick the worst response as rejected) — improving the judge or the generator pool alone did nothing.

📌 Go deeper: Part 15 shows the Delta Learning numbers, the DPO size sweep, and how the length-controlled DPO checkpoint becomes a better starting point for RL.
✓

Cheat sheet

Recap

RLHF (PPO)DPO
Needs a separate reward model?YesNo
Needs an RL loop?Yes (PPO, clipped objective)No — direct supervised-style loss
Models held in memoryPolicy + reward model + frozen referencePolicy + frozen reference (or none, for SimPO/ORPO)
Main riskReward hacking against the reward modelLess flexible for reward signals beyond static preference pairs
Delta LearningPair a strong chosen response with a deliberately weak rejected one; the contrast, not the chosen quality, carries the signal.
Dataset sizeU-shaped: peaks around 75–100K pairs for some evals while others keep rising — sweep size and early-stop.
Length biasChosen responses skew long; cap the chosen-minus-rejected difference (OLMo 3 uses 100 tokens).
Multi-turn preferencesSynthesize conversations and vary only the final turn so the ranking is unambiguous.
📚

Further reading

References

?

Check your understanding

0/6 answered
For tasks with an automatically checkable answer, there's now a whole parallel track that skips human preferences entirely. Continue: reasoning & RL with verifiable rewards →