LLM Training, Interactively
Every chat assistant you've used — Claude, GPT, Gemini, Llama — went through the same broad pipeline: a base model learns language by predicting text, then it's shaped into an assistant, then it's aligned to be helpful and safe. Most labs won't show you the middle of that pipeline. This guide will, using OLMo, the model family from the Allen Institute for AI (Ai2) that publishes its dataset, its code, and hundreds of checkpoints from an actual training run — so every claim here points at something you can go inspect yourself.
Why OLMo, and what this series covers
Setup
Most "open" models today are really only open-weight: you get the final numbers, not how they got there. The exact web-crawl recipe, the filtering thresholds, the preference data, the reward model — those stay inside the lab. That's a problem if you actually want to understand training, not just call an API.
OLMo (Ai2) is built differently. Ai2 releases the pretraining corpus (Dolma, a multi-trillion-token corpus, along with the exact curation code), checkpoints saved roughly every 1,000 training steps, full loss curves and hardware logs, and a tool called OlmoTrace that lets you trace a specific generated response back to the pretraining documents that most likely shaped it. We'll lean on OLMo throughout this series as the concrete, verifiable example behind each abstract idea.
Prerequisites: comfort with vectors, matrices, probability, and basic calculus — the level of a first-year graduate course. We'll define every model-specific term as it comes up. See the series hub for the full map, or the glossary for any term out of order.
A language model is a probability distribution over the next word
Foundations
Strip away the transformer architecture, the GPUs, the trillions of tokens, and a language model is answering one narrow question, over and over: given the text so far, what word (or token) is likely to come next? Formally, it defines a conditional probability distribution
over its vocabulary, where xt is the next token and x1…xt-1 is everything said before it. Chain that one-step prediction together — predict a token, append it, predict the next one — and you get a model that can generate whole passages. This is called autoregressive generation, and it's the objective every GPT-style model, including OLMo, is pretrained on.
Below is a genuine (if tiny) example: a bigram model — it predicts the next word using only the previous word — trained on a 110-word toy passage. Type a prompt ending in one of the passage's words (try the, robot, or window) and watch the actual learned distribution.
Bars show P(next word | last word), estimated by counting the toy passage below.
Toy passage:
Why not just use a bigger counting table?
Motivating neural language models
The obvious fix for a bigram model's short memory is to condition on more previous words — a trigram, a 4-gram, and so on. Try it below: as n grows, generated text gets locally more coherent, because the model has more context to work with. But watch the coverage number — as n grows, the number of n-word sequences that actually appeared in this tiny passage collapses toward zero, and the model has nothing to fall back on.
Perplexity: turning loss into a number you can feel
Foundations
Training minimizes the average negative log-probability the model assigns to the actual next token — the cross-entropy loss. That number in nats or bits isn't very intuitive on its own, so it's usually reported as perplexity, its exponential:
Perplexity has a clean reading: it's the effective number of equally-likely choices the model was uncertain between, on average, at each step. A perplexity of 1 means the model was completely certain every time (loss = 0); a perplexity equal to the vocabulary size means it was no better than guessing uniformly at random.
Try a sequence the bigram model has never seen the like of (e.g. the window read the robot) and watch perplexity spike — an ungrammatical or surprising continuation is, by definition, one the model assigned low probability to.
Before prediction: tokens and embeddings
Foundations
Models don't operate on raw characters or whole words — they operate on tokens from a fixed vocabulary, each looked up in a learned embedding table to become a vector. Both of those deserve real treatment rather than a paragraph, so Part 2 is entirely devoted to them: a byte-level BPE tokenizer we actually trained, and word embeddings we actually learned from co-occurrence statistics — not illustrative placeholders.
Putting it together: autoregressive generation
Foundations
Tokenize the prompt, look up embeddings, run them through the network to get a probability distribution over the next token, pick one, append it, and repeat. That loop — one forward pass per new token — is exactly how every response you've ever gotten from a chat model was produced (Part 11 covers the sampling choices — temperature, top-k, top-p — in full). Use the same bigram model from Step 1 to generate a few tokens yourself:
Cheat sheet
Recap
| Concept | What it is | Where it fits |
|---|---|---|
| Next-token distribution | P(xt | x<t), a probability over the whole vocabulary | What the model outputs at every position |
| n-gram sparsity | Exact-context counting tables need exponentially more data as context grows | The reason neural LMs replaced n-gram models |
| Cross-entropy loss | Average negative log-probability of the true next token | What training actually minimizes |
| Perplexity | e^loss — effective number of equally-likely choices | The human-readable form of the loss curve (Part 5) |
| Autoregressive generation | Predict → sample → append → repeat | How a distribution over one token becomes a whole response |
Further reading
References
- Vaswani et al., "Attention Is All You Need" (2017) — the transformer architecture.
- Jurafsky & Martin, "Speech and Language Processing", Ch. 3 — n-gram language models and perplexity.
- OLMo 2 Team (Ai2), "2 OLMo 2 Furious" (2024) — the OLMo 2 model report.
- Ai2's OLMo project page — models, data, and code.