Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

Numbers to know

OLMo 2 7B d_model
4,096
OLMo 2 7B layers
32
OLMo 2 7B attention heads
32
OLMo 2 vocab size
~100,278
Compute-optimal tokens/param
~20 (Chinchilla)
Production models' actual ratio
100s of tokens/param
Mixed-precision training memory
~16 bytes/param
OLMoE experts / active
64 experts, top-8
OLMoE total / active params
6.9B / 1.3B
Dolma v1.6 web share
74.6% Common Crawl
SFT vs. pretraining learning rate
~100× smaller
C = 6ND
FLOPs ≈ 6 × params × tokens
OLMo 3 context window
8,192 → 65,536
Extension token budget
50B (7B) / 100B (32B)
Midtraining mix
100B tokens (Dolmino 2)
OLMES base eval tasks
43 (was 11)
GRPO group size / pass@k
k = 16
OLMoRL throughput
2,949 tok/s (MBU 43.2%)
Preference dataset peak
~75–100K pairs
RL-Zero difficulty cutoff
62.5% pass rate
SimFC trajectories
200K / 42.6K functions

Terms

Autoregressive generation Part 1
Predict → sample → append → repeat, one token at a time.
Perplexity Part 1
exp(cross-entropy loss); the effective number of equally-likely choices per step.
n-gram sparsity Part 1
Exact-context counting tables need exponentially more data as context length grows.
Byte-level BPE Part 2
Tokenization by greedily merging the most frequent adjacent byte pair.
Static vs. contextual embedding Part 2
One fixed vector per token (static) vs. a vector reshaped per-instance by attention (contextual).
Self-attention Part 3
softmax(QKᵀ/√d_k + mask)V — lets tokens exchange information.
Grouped-query attention (GQA) Part 3
Several query heads share one key/value head, shrinking the KV cache.
RoPE Part 3
Rotary position embedding — rotates Q/K so their dot product depends only on relative distance.
Residual stream Part 3
The running sum each sublayer's output is added into, layer after layer.
RMSNorm / QK-norm Part 3
Normalize by root-mean-square rather than full LayerNorm; OLMo 2 also norms queries/keys directly.
SwiGLU Part 3
Gated MLP variant: a SiLU-gated projection multiplied elementwise into another before projecting down.
Mixture of experts (MoE) Part 3
Many parallel expert MLPs with a router selecting a few per token — more parameters, same per-token compute.
MinHash / LSH Part 4
Approximate, scalable near-duplicate document detection without all-pairs comparison.
Decontamination Part 4
Removing training text that overlaps with benchmark evaluation sets.
Data mixing / epochs Part 4
How much of each source to show the model, and how many times each source repeats.
Cross-entropy loss Part 5
Average negative log-probability the model assigned the true next token.
WSD schedule Part 5
Warmup–stable–decay learning-rate schedule; holds a constant peak, decays only near the end.
Gradient accumulation / packing Part 5
Simulating a larger batch, and filling sequences with concatenated documents instead of padding.
Gradient clipping Part 5
Capping gradient norm to prevent a bad batch from destabilizing training.
Mid-training anneal Part 5
Late-pretraining shift to a smaller, higher-quality mixture (OLMo 2's Dolmino).
Chinchilla / compute-optimal Part 6
The parameter/token split that minimizes loss for a fixed training compute budget.
C ≈ 6ND Part 6
Training FLOPs ≈ 6 × parameters × tokens.
MFU Part 6
Model FLOPs Utilization — achieved throughput divided by a GPU's peak FLOP/s.
FSDP / ZeRO Part 6
Sharding a model's parameters, gradients, and optimizer state across GPUs instead of replicating them.
LoRA / QLoRA Part 7
Freeze the base weights, train a small low-rank update BA instead; QLoRA also quantizes the frozen base.
Catastrophic forgetting / alignment tax Part 7
General capability lost from fine-tuning too narrowly or aggressively.
Reward model Part 8
A model trained to score responses, using the Bradley-Terry pairwise-preference loss.
PPO clipped objective Part 8
Caps how much a policy update can move the probability ratio versus the data-collecting policy.
KL–reward frontier Part 8
The trade-off β controls: reward-model score vs. divergence from the reference policy.
Reward hacking Part 8
A policy exploiting reward-model quirks in ways that raise score without raising real quality.
DPO Part 8
Direct Preference Optimization — trains the policy on preference pairs directly, no reward model or RL loop.
RLVR Part 9
Reinforcement Learning with Verifiable Rewards — reward from a deterministic checker, not a learned model.
GRPO Part 9
Group Relative Policy Optimization — advantage from a sampled group's own mean/std, no value model.
pass@k Parts 9, 10
Probability at least one of k sampled attempts is correct; unbiased estimator uses a binomial-coefficient formula.
Length normalization Part 10
Dividing summed log-likelihood by token count so shorter answers aren't unfairly favored.
LLM-as-judge / position bias Part 10
Using a model to grade responses; position bias is its tendency to favor whichever response is shown first.
Elo / Bradley-Terry arena Part 10
Fitting per-model skill scores from pairwise human-preference votes.
OLMES Part 10
Ai2's standardized evaluation protocol, fixing prompt format and scoring conventions across models.
KV cache Part 11
Cached per-token key/value vectors so decoding doesn't recompute attention over the whole sequence each step.
Prefill vs. decode Part 11
Compute-bound whole-prompt processing vs. memory-bandwidth-bound one-token-at-a-time generation.
Continuous batching Part 11
Swapping finished requests out for queued ones every decode step, instead of waiting for a fixed batch to finish.
Quantization Part 11
Storing/computing weights at lower precision (int8/int4) to save memory and increase throughput.
Speculative decoding Part 11
A small draft model proposes tokens; the large model verifies them in one batched pass.
Constitutional AI Part 12
The model critiques and revises its own responses against written principles, then trains on the revisions.
Prompt injection vs. jailbreaking Part 12
Third-party content hijacking the model's instructions, vs. the user themselves trying to bypass safety training.
Defense in depth Part 12
Layering independent guard mechanisms (input guard, system prompt, model, output guard) so no single failure is fatal.
Model card Part 12
Standardized documentation of a model's training data, intended uses, limitations, and eval results at release.
Long-context extension Part 13
A dedicated late stage that grows the model's context window after pretraining/midtraining (OLMo 3: 8,192 → 65,536).
YaRN Part 13
Per-frequency RoPE interpolation plus attention temperature; OLMo 3 applies it to full-attention layers only.
Position interpolation Part 13
Dividing positions by s = L'/L so a long sequence fits the pretrained positional range.
RoPE base-frequency scaling Part 13
Raising RoPE's base θ so long distances don't wrap past what the model learned.
Best-fit document packing Part 13
Binning whole documents into fixed-length sequences to minimize splits and padding.
Intra-document masking Part 13
A block-diagonal mask preventing packed documents from attending across their boundaries.
RULER Part 13
Synthetic long-context dev suite (needle-in-a-haystack variants and aggregation tasks).
HELMET Part 13
Broad held-out long-context suite covering retrieval, in-context learning, and summarization.
LongPPL Part 13
Long-range-dependency proxy based on perplexity changes under more context; tested but not used in the final OLMo 3 filter.
gzip compressibility filter Part 13
Discard the most- and least-compressible 20% of long documents.
Synthetic aggregation (CWE / REX) Part 13
Generated tasks — count a unigram exactly (CWE) or rewrite an aggregation into a vignette (REX) — answerable only from the full window.
Function calling / tool use Part 14
The model emits a structured call, a runtime executes it, and the result returns to the context.
Environment role Part 14
A dedicated message role carrying tool outputs back to the model.
MCP Part 14
Model Context Protocol — the standard transport connecting agents to tools and servers.
SimFC Part 14
LLM-simulated function-calling trajectories: 200K examples across 42.6K unique functions.
BFCL Part 14
Berkeley Function Calling Leaderboard — intrinsic function-choice and argument-correctness benchmark.
Delta learning Part 15
Preference signal comes from the chosen-minus-rejected contrast, not the chosen response's absolute quality.
Length bias Part 15
Preference data's tendency to favor longer responses; countered by capping the chosen–rejected length difference.
Multi-turn preference Part 15
Preference pairs over synthetic conversations, varying only the final turn so the ranking is unambiguous.
Zero-gradient filtering Part 15
Dropping groups whose rewards are identical, since they produce no advantage and no gradient.
Active sampling Part 15
Resampling prompts after filtering to keep a full batch of non-zero-gradient completions.
Token-level loss Part 15
Normalizing the RL loss by total tokens across the batch rather than per sample, removing length bias.
Clip-higher Part 15
Setting the upper PPO/GRPO clip bound above the lower one so probabilities can rise more than fall.
Truncated importance sampling Part 15
Correcting the log-probability gap between the inference engine and the trainer.
No std normalization Part 15
Advantage = reward − group mean (no division by std), avoiding a difficulty bias toward near-saturated prompts.
Offline difficulty filtering Part 15
Removing prompts the model solves more than 62.5% of the time across 8 rollouts.
Spurious rewards Part 15
Random signal-free rewards used as a contamination control: if benchmarks improve, eval data leaked in.
Model souping / merging Parts 5, 13
Averaging weights from independently trained checkpoints; used for both midtraining and long-context checkpoints.
Microanneal Part 5
A 5–10B-token anneal used to cheaply test a candidate dataset's midtraining impact.
Integration test Part 5
A full 100B-token midtraining run on a candidate mix, plus SFT, to judge combined data sources.
Fill-in-the-middle (FIM) Part 4
Training on prefix + suffix → concealed middle, teaching code infilling; OLMo 3 applies it to 50% of code documents.
Back to the series overview. Series hub →