Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A real training loop, small enough to run in this page

Foundations

Each training step: take a batch of token sequences, run the forward pass to get a predicted next-token distribution at every position, compare it against the actual next token using cross-entropy loss, then use backpropagation to compute how every parameter should change to reduce that loss, and apply an update via an optimizer — almost always AdamW for language models.

$$L = -\frac{1}{T}\sum_t \log P_\theta(x_t \mid x_{

Below is a genuine (if tiny) instance of this loop: a linear "neural bigram" model — one learned row of logits per previous word, no hand-coded counting table — trained by real gradient descent on Part 1's toy corpus. Click step repeatedly and watch the loss actually fall and the weight row for a chosen word actually shift.

This model has no attention, no depth, no context beyond one word — it's the smallest possible thing that "trains" in the sense Part 1 through Part 3 describe. Everything else in this series scales this exact loop up: bigger model, more context, vastly more data.

2

The learning-rate schedule

Training dynamics

The optimizer's step size — the learning rate — isn't fixed through training. The classic recipe is linear warmup followed by cosine decay: ramp up over the first ~1–2% of steps (starting at full learning rate on randomly-initialized weights is unstable), hold near a peak, then decay smoothly toward a small final value. A newer alternative, WSD (warmup–stable–decay), holds a constant peak rate for most of training and only decays sharply at the very end — its appeal is that you can extend a run (train on more tokens) without having pre-committed to a total step count the way cosine's smooth decay curve requires.

💡 Too little warmup and early training can blow up; too high a peak learning rate and loss oscillates or diverges later; decaying too early wastes the compute budget on preventable under-training. Published training reports like OLMo's include the exact schedule used, which is part of what makes a run reproducible rather than folklore.
3

Batch size, gradient accumulation, and packing

Making one "step" efficient

A training step's batch is rarely one GPU's worth of data — gradient accumulation runs several smaller forward/backward passes, sums their gradients, and applies one optimizer update, simulating a much larger batch than fits in memory at once. Separately, since real documents rarely divide evenly into fixed-length training sequences, sequence packing concatenates multiple documents (with a separator token) to fill each sequence completely rather than wasting compute on padding.

4

Gradient clipping and loss spikes

Stability at scale

Occasionally a batch produces an unusually large gradient — a rare, weirdly-formatted document, or accumulated numerical drift in mixed precision — and an uncontrolled update can knock the model's weights into a bad region, visible as a sudden loss spike. Gradient clipping caps the gradient's norm before the update is applied, the single cheapest stability measure available. OLMo's training reports are unusually candid about this: published loss curves show real spikes and recoveries, not a manicured monotone line.

5

Mid-training: OLMo 3's stage 2

A distinct stage between pretraining and fine-tuning

OLMo 2 introduced the anneal; OLMo 3 turns it into a full, measured stage. After roughly 5.9T tokens of pretraining, the model trains on Dolmino Mix 2 — 100B high-quality tokens blending top-quality web and PDF data with synthetic math and code, QA, instruction data, and thinking traces. The goal isn't only a lower loss: midtraining deliberately lays the groundwork for post-training, so the base model already contains the raw material that SFT and RL will later sharpen.

Microanneals are the fast loop: sample 5B tokens of a candidate source, mix with 5B web tokens, anneal, and compare against a web-only baseline. They let a team test dozens of datasets in parallel and decide what deserves a slot in the mix. Integration tests are the slow loop: a full 100B-token anneal of the whole candidate mixture, followed by SFT, to see whether a dataset still helps next to everything else. OLMo 3 ran five rounds; the average base score climbed 49.7 → 50.7 → 53.1 from the first candidate mix to the last, and the gain carried through to post-SFT evals.

Compute spent testing candidate datasets. A microanneal is ~10× cheaper per dataset; a full integration test costs far more but reveals interactions.

Midtraining also exposes hard domain tradeoffs. An exploratory math/code/thinking-heavy mix raised Math (57.3 → 60.8) and Code (31.2 → 35.6) but cost MCQA-STEM (66.4 → 62.5) and GenQA (73.1 → 65.9); a GenQA-heavy mix did the reverse. There is no free lunch, which is precisely why the mix is chosen by integration tests rather than a single benchmark. And there is a formatting trap: including chat special tokens (<|im_start|>) in midtraining data makes the base model emit them at inference. GSM8K collapsed from 49.43 to 0 with special tokens, versus 46.02 with a plain chat template — so OLMo 3 strips both from midtraining and leaves them for SFT.

Model souping closes the stage. Merging two independent 32B midtraining runs (different seeds) added roughly 1 point on MCQA-STEM, 0.4 on GenQA, 2.9/1.6 on Math versus the two runs, ~1 on MMLU, and 5/2 on GSM-Symbolic — enough that the merged checkpoint became the final 32B midtrained model. The 7B saw no such gain and stayed a single run.

💡 Where this feeds: the long-context extension of Part 13 starts from this midtrained checkpoint, and the whole post-training recipe of Part 15 consumes it. Midtraining is the hinge between the base model and everything that makes it an assistant.
6

Watching it happen: loss curves and checkpoints

Evidence, not folklore

Most labs report a final loss number, if that. Ai2 published OLMo's entire loss curve, along with model checkpoints saved roughly every 1,000 training steps — so instead of taking "capability emerges gradually" on faith, you can download a checkpoint from any point in training and test it yourself. Drag the slider below through a representative (illustrative, shape-matched, including the late anneal drop from Section 5) training run.

💡 OlmoTrace takes this further at inference time: given a specific generated response, it searches back through the training corpus for documents whose text plausibly shaped that output — turning "the model said X" into "the model said X, and here's training text that looks like why."
✓

Cheat sheet

Recap

PieceRole
Cross-entropy lossScores how well predicted next-token probabilities match the true next token
AdamWAdaptive optimizer used to turn loss gradients into parameter updates
Warmup + cosine / WSDLearning-rate schedules that stabilize early training and control late-training decay
Gradient accumulationSimulates a larger batch than fits in memory by summing several micro-batch gradients
Sequence packingConcatenates documents to fill training sequences without padding waste
Gradient clippingCaps gradient norm to prevent a single bad batch from destabilizing training
Mid-training annealLate-training shift to a smaller, higher-quality data mixture (OLMo 2's Dolmino, OLMo 3's Dolmino Mix 2)
Microanneal~10B-token test of one candidate dataset against a web-only baseline, run in parallel across many candidates
Integration testFull 100B-token midtraining run on a candidate mix plus SFT, revealing how sources interact
Model soupingAveraging weights of independent runs; OLMo 3 used it for the 32B midtrained and long-context checkpoints
Special tokens at midtrainingChat tokens make the base model emit them at inference (GSM8K 49.43 → 0); leave them for SFT
CheckpointsSnapshots of the model at intervals through training, enabling inspection of how capability develops
📚

Further reading

References

?

Check your understanding

0/5 answered
Before scaling any of this up, it's worth knowing the actual math for how much compute a given model-and-data size costs, and how to fit it on real hardware. Continue: scaling laws & training systems →