LLM Training, Interactively
Fifteen parts, one pipeline: raw web text becomes a corpus, a corpus trains a base model, a base model becomes an assistant, an assistant is aligned and hardened, and the result is served to actual users. Every stage below is a real, interactive part of the series — not a static description.
The pipeline
Click any stage to jump to its part.
All 15 parts
What is a language model
Next-token prediction, autoregressive generation, and why OLMo is the running example for this series.
Part 2 of 15Tokenization & embeddings
Real BPE training, multilingual token cost, and learned (not hand-placed) word vectors.
Part 3 of 15The transformer, block by block
Attention, multi-head, RoPE, residual stream, norms, SwiGLU, parameter counting, and MoE.
Part 4 of 15Building the pretraining dataset
Sourcing, filtering, deduplication, contamination, and mixing at trillion-token scale.
Part 5 of 15Pretraining: the training loop
Loss, perplexity, learning-rate schedules, gradient clipping, and mid-training annealing.
Part 6 of 15Scaling laws & training systems
Chinchilla compute-optimal scaling, memory accounting, and distributed-training parallelism.
Part 7 of 15Supervised fine-tuning & PEFT
Chat templates, loss masking, and LoRA/QLoRA for parameter-efficient fine-tuning.
Part 8 of 15Alignment: RLHF & DPO
Reward models, PPO's clipped objective, DPO's implicit reward, and the KL-reward frontier.
Part 9 of 15Reasoning & RL with verifiable rewards
Chain-of-thought, RLVR, GRPO, and test-time compute scaling.
Part 10 of 15Evaluation
How benchmarks are actually scored, LLM-as-judge, arena ratings, and pass@k.
Part 11 of 15Inference, serving & efficiency
Sampling, the KV cache, prefill vs. decode, batching, and quantization.
Part 12 of 15Guardrails, safety & deployment
Refusal training, layered defenses, and staged release.
Part 13 of 15Long-context & context extension
Stretching RoPE, long-context data, best-fit packing, and why extension is its own late stage.
Part 14 of 15Tool use, function calling & agents
OpenAPI tool specs, environment roles, real vs simulated trajectories, and how tool use is evaluated.
Part 15 of 15Post-training at scale
The three-stage recipe, delta learning, upgraded GRPO, domain verifiers, and RL systems.
ReferenceGlossary & numbers to know
Every bolded term across the series, linked back to where it's introduced, plus a quick-reference numbers card.
Already know some of this? Start here
| If you... | Start at |
|---|---|
| Have never touched an ML paper | Part 1 — What is a language model |
| Know tokenization/embeddings, want the architecture | Part 3 — The transformer, block by block |
| Know transformers, want the data/training-systems side | Part 4 — Building the pretraining dataset |
| Have a base model, want to fine-tune it yourself | Part 7 — SFT & PEFT (LoRA) |
| Know SFT/RLHF, want reasoning-model RL specifically | Part 9 — Reasoning & RLVR |
| Just want to know how to read a benchmark number | Part 10 — Evaluation |
| Care about running models cheaply, not training them | Part 11 — Inference, serving & efficiency |
| Finished Part 11 and want serving as an engineering discipline | The sequel — LLM Serving, Interactively |
| Want the calculus behind backprop and gradient descent taught from scratch | Calculus in Motion, Interactively |
| Have a base model but it only knows a few thousand tokens | Part 13 — Long-context & context extension |
| Want to build tool-calling agents or MCP integrations | Part 14 — Tool use, function calling & agents |
| Want the full frontier post-training recipe end to end | Part 15 — Post-training at scale |