Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Base model vs. instruction-tuned model

The behavior shift

Same prompt, same underlying knowledge from pretraining — very different behavior after SFT. These example completions are illustrative (written for this page, not sampled from a live model), but they show the qualitative shift accurately.

Base model continuation
SFT (instruction-tuned) response
🎯 Learning goal: the base model isn't "worse" at language — it's answering a different implicit question ("what text plausibly follows this?") than the one you meant to ask. SFT doesn't inject new facts so much as it teaches the model which completion behavior you actually want.
2

Instruction-tuning data: curation, synthesis, and distillation

What SFT trains on

SFT data is a curated set of (prompt, response) pairs — ideally covering a wide range of task types and response styles. For OLMo, Ai2 uses and publishes the Tulu mixture: a combination of human-written examples, carefully filtered existing instruction datasets, and synthetically generated examples — a real mix of curation and distillation (generating training examples from a stronger existing model's outputs, then filtering the results), rather than either alone.

⚠️ Quality over quantity: unlike pretraining, SFT data is measured in thousands to low millions of examples, not trillions of tokens — and the field's own research (including Ai2's) has repeatedly found that a smaller, carefully curated instruction set often outperforms a much larger, noisier one. Bad SFT examples (wrong answers, inconsistent formatting, sycophantic tone) get imitated just as readily as good ones.

SFT hyperparameters reflect this: learning rates are typically ~100× smaller than peak pretraining rates (the model is already good at language; SFT should nudge, not overwrite), and training usually runs only 2–3 epochs over the data, since with so few examples relative to pretraining, more epochs risks the overfitting-to-surface-patterns problem below.

3

Chat templates and multi-turn masking

Mechanics

A multi-turn conversation has to be serialized into a single flat token sequence, using special tokens to mark role boundaries (exact tokens vary by model family). SFT reuses the exact cross-entropy loss from Part 5 — the only change is which tokens count toward it: loss is masked on the system prompt and every user turn, and computed only on assistant-response tokens, across however many turns the conversation has.

$$L = -\frac{1}{|R|}\sum_{t \in R} \log P_\theta(x_t \mid x_{

Build a small multi-turn conversation below and see it formatted, with the mask applied per turn:

highlighted = contributes to the SFT loss (response tokens, both turns). Faded = present in the sequence, masked out of the loss.

4

LoRA: fine-tuning without touching most of the weights

Parameter-efficient fine-tuning

Full fine-tuning updates every one of a model's weight matrices, which for a 7B+ model means storing optimizer state for billions of parameters — expensive, and every fine-tune produces a full-size copy of the model. LoRA (Low-Rank Adaptation) instead freezes the original weight matrix W entirely and learns a small additive update as a product of two thin matrices, $\Delta W = BA$, where B is d×r and A is r×d for a small rank r (often 8–64) — far fewer trainable parameters than d×d.

$$W' = W + BA, \qquad \text{trainable params} = r(d_{in}+d_{out}) \ll d_{in} d_{out}$$

Below, a toy 8×8 weight matrix has a "true" full-fine-tune update it would need to learn (a fixed, moderately complex target). Drag the rank slider and watch LoRA's B·A approximate it via real gradient descent — trainable-parameter count and reconstruction error both update live.

QLoRA layers one more idea on top: keep the frozen base weights W quantized to 4 bits (Part 11 covers quantization) while the small LoRA matrices train in higher precision — cutting base-model memory further, at the cost of a little extra quantization error on top of LoRA's own approximation error.

Full fine-tuningLoRAQLoRA
Trainable params (7B model, typical)~7B (100%)~0.1–1%~0.1–1%
Base weightsUpdated directlyFrozen, bf16Frozen, 4-bit
Optimizer memoryFull AdamW state over all paramsOnly over LoRA matricesOnly over LoRA matrices
Output artifactA new full-size model copyA small adapter file (MBs) on top of the baseA small adapter file on top of the quantized base
5

What SFT changes, and the risk of forgetting

Scope

SFT's effect is mostly on behavior and format, not the underlying knowledge set during pretraining on vastly more data. Two related risks: over-training on a narrow SFT set can cause the model to overfit to surface patterns (a fixed response length, a signature phrase); and fine-tuning too aggressively on a narrow domain can cause catastrophic forgetting, where capabilities the base model had — general knowledge, other task types — measurably degrade, sometimes called the alignment tax when the specific capability lost is general-purpose quality traded for narrower instruction-following. LoRA's frozen base weights are one structural mitigation: since most of the network is literally untouched, broad pretrained knowledge is mechanically harder to overwrite than with full fine-tuning.

6

Specialized SFT: tools and the reasoning-first warm start

OLMo 3 Instruct

SFT is not one homogeneous stage. OLMo 3 trains two different assistants from the same base, and each adds its own SFT flavor. Function-calling SFT teaches the model to emit structured tool calls: the training mix combines real trajectories executed against MCP servers (the Asta Scientific Corpus and Serper) with the much larger LLM-simulated SimFC dataset. The single most important lesson from that work is that the format has to be unified across every dataset — OpenAPI tool specs, pythonic code-block calls wrapped in XML tags, a dedicated environment role, and (for Instruct) tool-specific special tokens. Mixing conventions is a direct cause of unreliable tool use. Part 14 has the full format and the evaluation harness.

The reasoning-first warm start is the other notable choice. Rather than teaching instruction-following from the base model directly, OLMo 3 initializes the Instruct SFT stage from the already-trained Think SFT checkpoint. The instruction model then learns to suppress the thinking traces and answer concisely, but it inherits the reasoning skill underneath. The gain is broad and real, without making responses longer:

Instruct SFT starting pointAvg.GPQAMATHGSM8KOMEGAMBPPIFEval
No thinking SFT first44.529.760.387.68.654.181.0
With thinking SFT first (OLMo 3)47.834.465.991.112.257.184.7
Gain+3.3+4.7+5.6+3.5+3.6+3.0+3.7

Source: OLMo 3 report, arXiv:2512.13961v2, Table 13 (intermediate OLMo 3 Instruct 7B).

⚠️ The special-token caution, revisited. Chat special tokens like <|im_start|> must be introduced at SFT, not midtraining. Part 5 showed that including them in midtraining data makes the base model emit them at inference and collapses GSM8K from 49.43 to 0. OLMo 3's Think stage therefore avoided special tokens entirely, while Instruct deliberately adds them (including tool tokens) because the SFT data and inference harness are built around them. Timing is everything.
📌 Cross-links: the tool-use data and format live in Part 14; the Thinking-then-Instruct pipeline and its preference stages are Part 15; the special-token finding is Part 5.
✓

Cheat sheet

Recap

AspectPretrainingSFT
ObjectivePredict next token, everywherePredict next token, only on response spans
Data scaleTrillions of tokensThousands–millions of examples
Learning ratePeak ~1e-3 to 1e-4 scale~100× smaller
What it teachesLanguage, facts, world knowledgeFormat, directness, instruction-following behavior
LoRA—Freezes base weights, trains a low-rank BA update instead
📚

Further reading

References

?

Check your understanding

0/5 answered
SFT teaches format, but not which of two reasonable responses is actually better. That comparative judgment is what alignment — RLHF and DPO — is built to train on. Continue: alignment, RLHF & DPO →