Pretraining: the Training Loop
With a cleaned, trillion-token dataset from Part 4 and the transformer block from Part 3, pretraining is the process that turns them into a base model: the network is shown token sequences over and over, scored on how well it predicts the next token, and nudged to do slightly better each time — repeated for hundreds of thousands of steps across thousands of GPUs. This part builds the actual training loop, including one you can step through yourself, and checks each piece against what Ai2 published for OLMo.
A real training loop, small enough to run in this page
Foundations
Each training step: take a batch of token sequences, run the forward pass to get a predicted next-token distribution at every position, compare it against the actual next token using cross-entropy loss, then use backpropagation to compute how every parameter should change to reduce that loss, and apply an update via an optimizer — almost always AdamW for language models.
Below is a genuine (if tiny) instance of this loop: a linear "neural bigram" model — one learned row of logits per previous word, no hand-coded counting table — trained by real gradient descent on Part 1's toy corpus. Click step repeatedly and watch the loss actually fall and the weight row for a chosen word actually shift.
This model has no attention, no depth, no context beyond one word — it's the smallest possible thing that "trains" in the sense Part 1 through Part 3 describe. Everything else in this series scales this exact loop up: bigger model, more context, vastly more data.
The learning-rate schedule
Training dynamics
The optimizer's step size — the learning rate — isn't fixed through training. The classic recipe is linear warmup followed by cosine decay: ramp up over the first ~1–2% of steps (starting at full learning rate on randomly-initialized weights is unstable), hold near a peak, then decay smoothly toward a small final value. A newer alternative, WSD (warmup–stable–decay), holds a constant peak rate for most of training and only decays sharply at the very end — its appeal is that you can extend a run (train on more tokens) without having pre-committed to a total step count the way cosine's smooth decay curve requires.
Batch size, gradient accumulation, and packing
Making one "step" efficient
A training step's batch is rarely one GPU's worth of data — gradient accumulation runs several smaller forward/backward passes, sums their gradients, and applies one optimizer update, simulating a much larger batch than fits in memory at once. Separately, since real documents rarely divide evenly into fixed-length training sequences, sequence packing concatenates multiple documents (with a separator token) to fill each sequence completely rather than wasting compute on padding.
Gradient clipping and loss spikes
Stability at scale
Occasionally a batch produces an unusually large gradient — a rare, weirdly-formatted document, or accumulated numerical drift in mixed precision — and an uncontrolled update can knock the model's weights into a bad region, visible as a sudden loss spike. Gradient clipping caps the gradient's norm before the update is applied, the single cheapest stability measure available. OLMo's training reports are unusually candid about this: published loss curves show real spikes and recoveries, not a manicured monotone line.
Mid-training: OLMo 3's stage 2
A distinct stage between pretraining and fine-tuning
OLMo 2 introduced the anneal; OLMo 3 turns it into a full, measured stage. After roughly 5.9T tokens of pretraining, the model trains on Dolmino Mix 2 — 100B high-quality tokens blending top-quality web and PDF data with synthetic math and code, QA, instruction data, and thinking traces. The goal isn't only a lower loss: midtraining deliberately lays the groundwork for post-training, so the base model already contains the raw material that SFT and RL will later sharpen.
Microanneals are the fast loop: sample 5B tokens of a candidate source, mix with 5B web tokens, anneal, and compare against a web-only baseline. They let a team test dozens of datasets in parallel and decide what deserves a slot in the mix. Integration tests are the slow loop: a full 100B-token anneal of the whole candidate mixture, followed by SFT, to see whether a dataset still helps next to everything else. OLMo 3 ran five rounds; the average base score climbed 49.7 → 50.7 → 53.1 from the first candidate mix to the last, and the gain carried through to post-SFT evals.
Compute spent testing candidate datasets. A microanneal is ~10× cheaper per dataset; a full integration test costs far more but reveals interactions.
Midtraining also exposes hard domain tradeoffs. An exploratory math/code/thinking-heavy mix raised Math (57.3 → 60.8) and Code (31.2 → 35.6) but cost MCQA-STEM (66.4 → 62.5) and GenQA (73.1 → 65.9); a GenQA-heavy mix did the reverse. There is no free lunch, which is precisely why the mix is chosen by integration tests rather than a single benchmark. And there is a formatting trap: including chat special tokens (<|im_start|>) in midtraining data makes the base model emit them at inference. GSM8K collapsed from 49.43 to 0 with special tokens, versus 46.02 with a plain chat template — so OLMo 3 strips both from midtraining and leaves them for SFT.
Model souping closes the stage. Merging two independent 32B midtraining runs (different seeds) added roughly 1 point on MCQA-STEM, 0.4 on GenQA, 2.9/1.6 on Math versus the two runs, ~1 on MMLU, and 5/2 on GSM-Symbolic — enough that the merged checkpoint became the final 32B midtrained model. The 7B saw no such gain and stayed a single run.
Watching it happen: loss curves and checkpoints
Evidence, not folklore
Most labs report a final loss number, if that. Ai2 published OLMo's entire loss curve, along with model checkpoints saved roughly every 1,000 training steps — so instead of taking "capability emerges gradually" on faith, you can download a checkpoint from any point in training and test it yourself. Drag the slider below through a representative (illustrative, shape-matched, including the late anneal drop from Section 5) training run.
Cheat sheet
Recap
| Piece | Role |
|---|---|
| Cross-entropy loss | Scores how well predicted next-token probabilities match the true next token |
| AdamW | Adaptive optimizer used to turn loss gradients into parameter updates |
| Warmup + cosine / WSD | Learning-rate schedules that stabilize early training and control late-training decay |
| Gradient accumulation | Simulates a larger batch than fits in memory by summing several micro-batch gradients |
| Sequence packing | Concatenates documents to fill training sequences without padding waste |
| Gradient clipping | Caps gradient norm to prevent a single bad batch from destabilizing training |
| Mid-training anneal | Late-training shift to a smaller, higher-quality data mixture (OLMo 2's Dolmino, OLMo 3's Dolmino Mix 2) |
| Microanneal | ~10B-token test of one candidate dataset against a web-only baseline, run in parallel across many candidates |
| Integration test | Full 100B-token midtraining run on a candidate mix plus SFT, revealing how sources interact |
| Model souping | Averaging weights of independent runs; OLMo 3 used it for the 32B midtrained and long-context checkpoints |
| Special tokens at midtraining | Chat tokens make the base model emit them at inference (GSM8K 49.43 → 0); leave them for SFT |
| Checkpoints | Snapshots of the model at intervals through training, enabling inspection of how capability develops |
Further reading
References
- Kingma & Ba, "Adam: A Method for Stochastic Optimization" (2014); Loshchilov & Hutter, "Decoupled Weight Decay Regularization" (AdamW, 2017).
- Hu et al., "MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies" (2024) — WSD schedule.
- OLMo 2 Team (Ai2), "2 OLMo 2 Furious" (2024) — Dolmino mid-training mix, loss spikes, checkpoint souping.
- Ai2, OlmoTrace announcement — tracing model outputs to training data.