Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

The pipeline

Click any stage to jump to its part.

Raw web → corpus Part 4 ~200TB → 3T tokens Pretraining (base model) Parts 3, 5, 6 Weeks–months, 1000s of GPUs Mid-training (anneal) Part 5 Hours–days Long-context extension Part 13 8K → 65K, RoPE stretch SFT (instruction tuning) Part 7 Thousands–millions of examples Alignment + RL Parts 8–9 DPO / PPO / GRPO Post-training at scale Part 15 Delta learning, upgraded GRPO Tool use & agents Part 14 Function calling, MCP Evaluation + guardrails Parts 10, 12 Ongoing, incl. post-release Serving (inference) Part 11 Milliseconds/token

All 15 parts

Part 1 of 15

What is a language model

Next-token prediction, autoregressive generation, and why OLMo is the running example for this series.

Part 2 of 15

Tokenization & embeddings

Real BPE training, multilingual token cost, and learned (not hand-placed) word vectors.

Part 3 of 15

The transformer, block by block

Attention, multi-head, RoPE, residual stream, norms, SwiGLU, parameter counting, and MoE.

Part 4 of 15

Building the pretraining dataset

Sourcing, filtering, deduplication, contamination, and mixing at trillion-token scale.

Part 5 of 15

Pretraining: the training loop

Loss, perplexity, learning-rate schedules, gradient clipping, and mid-training annealing.

Part 6 of 15

Scaling laws & training systems

Chinchilla compute-optimal scaling, memory accounting, and distributed-training parallelism.

Part 7 of 15

Supervised fine-tuning & PEFT

Chat templates, loss masking, and LoRA/QLoRA for parameter-efficient fine-tuning.

Part 8 of 15

Alignment: RLHF & DPO

Reward models, PPO's clipped objective, DPO's implicit reward, and the KL-reward frontier.

Part 9 of 15

Reasoning & RL with verifiable rewards

Chain-of-thought, RLVR, GRPO, and test-time compute scaling.

Part 10 of 15

Evaluation

How benchmarks are actually scored, LLM-as-judge, arena ratings, and pass@k.

Part 11 of 15

Inference, serving & efficiency

Sampling, the KV cache, prefill vs. decode, batching, and quantization.

Part 12 of 15

Guardrails, safety & deployment

Refusal training, layered defenses, and staged release.

Part 13 of 15

Long-context & context extension

Stretching RoPE, long-context data, best-fit packing, and why extension is its own late stage.

Part 14 of 15

Tool use, function calling & agents

OpenAPI tool specs, environment roles, real vs simulated trajectories, and how tool use is evaluated.

Part 15 of 15

Post-training at scale

The three-stage recipe, delta learning, upgraded GRPO, domain verifiers, and RL systems.

Reference

Glossary & numbers to know

Every bolded term across the series, linked back to where it's introduced, plus a quick-reference numbers card.

Already know some of this? Start here

If you...Start at
Have never touched an ML paperPart 1 — What is a language model
Know tokenization/embeddings, want the architecturePart 3 — The transformer, block by block
Know transformers, want the data/training-systems sidePart 4 — Building the pretraining dataset
Have a base model, want to fine-tune it yourselfPart 7 — SFT & PEFT (LoRA)
Know SFT/RLHF, want reasoning-model RL specificallyPart 9 — Reasoning & RLVR
Just want to know how to read a benchmark numberPart 10 — Evaluation
Care about running models cheaply, not training themPart 11 — Inference, serving & efficiency
Finished Part 11 and want serving as an engineering disciplineThe sequel — LLM Serving, Interactively
Want the calculus behind backprop and gradient descent taught from scratchCalculus in Motion, Interactively
Have a base model but it only knows a few thousand tokensPart 13 — Long-context & context extension
Want to build tool-calling agents or MCP integrationsPart 14 — Tool use, function calling & agents
Want the full frontier post-training recipe end to endPart 15 — Post-training at scale
Start from the beginning, or jump straight into whichever part you need. Start with Part 1 →