Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

Why OLMo, and what this series covers

Setup

Most "open" models today are really only open-weight: you get the final numbers, not how they got there. The exact web-crawl recipe, the filtering thresholds, the preference data, the reward model — those stay inside the lab. That's a problem if you actually want to understand training, not just call an API.

OLMo (Ai2) is built differently. Ai2 releases the pretraining corpus (Dolma, a multi-trillion-token corpus, along with the exact curation code), checkpoints saved roughly every 1,000 training steps, full loss curves and hardware logs, and a tool called OlmoTrace that lets you trace a specific generated response back to the pretraining documents that most likely shaped it. We'll lean on OLMo throughout this series as the concrete, verifiable example behind each abstract idea.

💡 The shape of the series: Part 1 (this page) covers next-token prediction and perplexity. Part 2 covers tokenization and embeddings for real. Part 3 opens up the transformer block itself. Part 4 covers the pretraining dataset. Part 5 covers pretraining — the training loop. Part 6 covers scaling laws and training systems. Part 7 covers supervised fine-tuning and LoRA. Part 8 covers alignment — RLHF and DPO. Part 9 covers reasoning and RL with verifiable rewards. Part 10 covers evaluation. Part 11 covers inference and serving. Part 12 covers guardrails — safety and deployment. Parts 13–15 then cover the frontier stages: long-context extension, tool use and agents, and the full post-training recipe at scale.

Prerequisites: comfort with vectors, matrices, probability, and basic calculus — the level of a first-year graduate course. We'll define every model-specific term as it comes up. See the series hub for the full map, or the glossary for any term out of order.

1

A language model is a probability distribution over the next word

Foundations

Strip away the transformer architecture, the GPUs, the trillions of tokens, and a language model is answering one narrow question, over and over: given the text so far, what word (or token) is likely to come next? Formally, it defines a conditional probability distribution

$$P(x_t \mid x_1, x_2, \ldots, x_{t-1})$$

over its vocabulary, where xt is the next token and x1…xt-1 is everything said before it. Chain that one-step prediction together — predict a token, append it, predict the next one — and you get a model that can generate whole passages. This is called autoregressive generation, and it's the objective every GPT-style model, including OLMo, is pretrained on.

Below is a genuine (if tiny) example: a bigram model — it predicts the next word using only the previous word — trained on a 110-word toy passage. Type a prompt ending in one of the passage's words (try the, robot, or window) and watch the actual learned distribution.

Bars show P(next word | last word), estimated by counting the toy passage below.

Toy passage:

🎯 Learning goal: a language model doesn't "know" the next word — it estimates how likely each candidate is, given everything it's seen. A bigram model estimates this from one previous word; OLMo estimates it from up to thousands of previous tokens, using a transformer instead of a counting table — but the object being predicted is exactly the same shape.
2

Why not just use a bigger counting table?

Motivating neural language models

The obvious fix for a bigram model's short memory is to condition on more previous words — a trigram, a 4-gram, and so on. Try it below: as n grows, generated text gets locally more coherent, because the model has more context to work with. But watch the coverage number — as n grows, the number of n-word sequences that actually appeared in this tiny passage collapses toward zero, and the model has nothing to fall back on.

⚠️ This is exactly why n-gram language models lost out to neural ones: the number of possible n-word contexts grows exponentially with n (vocabulary size V raised to the n), so a counting table needs exponentially more data to have seen most of them even once — data sparsity. A neural network doesn't store a table entry per exact context; it learns a smooth function of the context's embedding, so it can generalize to word sequences it never saw verbatim during training. This — not just "more context" — is the real reason transformers replaced n-gram models rather than just scaling n up.
3

Perplexity: turning loss into a number you can feel

Foundations

Training minimizes the average negative log-probability the model assigns to the actual next token — the cross-entropy loss. That number in nats or bits isn't very intuitive on its own, so it's usually reported as perplexity, its exponential:

$$\text{loss} = -\frac{1}{T}\sum_{t=1}^{T} \log P(x_t \mid x_{

Perplexity has a clean reading: it's the effective number of equally-likely choices the model was uncertain between, on average, at each step. A perplexity of 1 means the model was completely certain every time (loss = 0); a perplexity equal to the vocabulary size means it was no better than guessing uniformly at random.

Try a sequence the bigram model has never seen the like of (e.g. the window read the robot) and watch perplexity spike — an ungrammatical or surprising continuation is, by definition, one the model assigned low probability to.

4

Before prediction: tokens and embeddings

Foundations

Models don't operate on raw characters or whole words — they operate on tokens from a fixed vocabulary, each looked up in a learned embedding table to become a vector. Both of those deserve real treatment rather than a paragraph, so Part 2 is entirely devoted to them: a byte-level BPE tokenizer we actually trained, and word embeddings we actually learned from co-occurrence statistics — not illustrative placeholders.

5

Putting it together: autoregressive generation

Foundations

Tokenize the prompt, look up embeddings, run them through the network to get a probability distribution over the next token, pick one, append it, and repeat. That loop — one forward pass per new token — is exactly how every response you've ever gotten from a chat model was produced (Part 11 covers the sampling choices — temperature, top-k, top-p — in full). Use the same bigram model from Step 1 to generate a few tokens yourself:

⚠️ Notice: a bigram model only remembers one word back, so it wanders and repeats quickly. Real transformer models condition on the entire preceding context (OLMo 2 supports contexts of several thousand tokens) via the attention mechanism covered in Part 3 — which is most of why they stay coherent, not a fundamentally different objective.
✓

Cheat sheet

Recap

ConceptWhat it isWhere it fits
Next-token distributionP(xt | x<t), a probability over the whole vocabularyWhat the model outputs at every position
n-gram sparsityExact-context counting tables need exponentially more data as context growsThe reason neural LMs replaced n-gram models
Cross-entropy lossAverage negative log-probability of the true next tokenWhat training actually minimizes
Perplexitye^loss — effective number of equally-likely choicesThe human-readable form of the loss curve (Part 5)
Autoregressive generationPredict → sample → append → repeatHow a distribution over one token becomes a whole response
📚

Further reading

References

?

Check your understanding

0/5 answered
Next: what tokens and embeddings actually are, built for real instead of illustrated. Continue: tokenization & embeddings →