Long-context & Context Extension
Every model in this series so far reads a window of a few thousand tokens. But real work — a whole codebase, a stack of papers, a long agent trajectory — needs far more. Training directly at long sequence lengths is prohibitively expensive, so frontier models pretrain short and then extend their context in a dedicated late stage. OLMo 3 stretches from 8,192 to 65,536 tokens this way. In this part we open up that stage: how RoPE gets stretched, what data teaches long-range attention, how sequences are packed and masked, and how you know it actually worked.
Why context is its own stage
Pretrain short, extend late
Attention cost grows with the square of sequence length, so a training run at 64K costs dramatically more per token than the same run at 8K. The standard compromise is to pretrain and midtrain at a shorter window — OLMo 3 uses 8,192 tokens — and then spend a comparatively small number of tokens adapting the already-trained model to a longer one. Because the recipe only touches a few tens of billions of tokens, it is cheap enough to iterate on, and it can be tuned on the same base model repeatedly.
The catch is that the recipes are wildly inconsistent across labs. SmolLM3 and GLM 4.5 spend about 100B extension tokens; DeepSeek V3 spends 123B; Apertus 225B; Kimi K2 400B; Llama 3.1 and DeepSeek V3.1 spend 800–840B. The outliers prove how much the budget is an engineering choice rather than a law: AFM and Nemotron Nano 2 reach 64K and 128K respectively on fewer than 20B tokens. Labs also disagree on when to extend — Llama 3.1 extends before midtraining, Qwen 2.5/3 after it, and GLM 4.5 only after supervised fine-tuning.
Extension-stage token budgets for open models (billions of tokens). Click a bar for the note. Bars are sorted largest-first; some models reach long windows far more cheaply.
Stretching RoPE
Positional embeddings
RoPE (Part 3) encodes position by rotating query and key vectors by an angle proportional to their position, with a different rotation frequency per dimension. A model pretrained on 8K positions has simply never seen the rotations that a 32K-token sequence produces, so its attention pattern breaks down past the training length. Context extension is mostly the art of re-mapping positions into the range the model already understands. Three families dominate:
- Base-frequency scaling — change RoPE's base θ so the whole spectrum of frequencies stretches; lower frequencies now cover much longer distances before wrapping.
- Position interpolation (PI) — divide each position by a scale factor s = L'/L, compressing the long sequence back into the original range. Cheap, but it lowers resolution at every distance.
- YaRN — combine interpolation with a per-frequency correction and an attention “temperature” factor, interpolating only the slow dimensions and leaving fast ones near their pretrained frequencies. This is what OLMo 3 uses.
OLMo 3's key finding is architectural: it uses sliding-window attention (SWA) in three of every four layers, with the last layer always full attention. YaRN works best when applied to the RoPE instances in the full-attention layers only. Touching the sliding-window layers as well hurts.
Illustrative attention score at a given relative distance for each method (shapes modeled, not measured). The dashed line marks the 8K pretraining window.
Data for long context
Longmino Mix
A longer window is only useful if the training data actually has long-range dependencies. OLMo 3 builds Longmino Mix from a 639B-token pool of long documents — principally scientific PDFs processed by olmocr — plus synthetic augmentation, then mixes in high-quality short-context data from midtraining so short tasks don't regress. The final 50B mix is roughly 34% long-context data (long docs plus synthetic) and 66% short-context midtraining data; a 66/34 long-heavy mix measurably hurt short-context evals.
Two filters shape the raw documents:
- gzip compressibility — compression ratio is a cheap proxy for redundancy. OLMo 3 throws away the most-compressible 20% (boilerplate, repetitive) and the least-compressible 20% (noise, garbled text), keeping the informative middle.
- LongPPL — a token is a “key” token if its perplexity drops a lot when given much more preceding context. OLMo 3 computed it over 10B tokens with Gemma 3 4B as reference, but found no thresholding strategy beat plain gzip, so it was tested and not used.
Illustrative compressibility distribution. The retention window keeps the middle; drag the cut to see how aggressively extremes are removed.
Synthetic aggregation tasks
Most long documents have no supervision for the tasks long-context models are actually used for — summarizing, counting, and synthesizing across a whole document. OLMo 3 injects that supervision synthetically, inspired by CLIPPER, by asking a model to build aggregation questions from snippets it can see. Two task types are used:
QA pairs whose answer is the exact number of times a common unigram occurs in a document partition. Only answerable by scanning the whole window.
The same aggregation task rendered as one of 12 vignettes — a dialogue, flashcards, a quiz, a game show, a debate, an ELI5 explainer — forcing synthesis over the noun phrase throughout.
Try a CWE task. The passage below is short enough to skim, but the same task at 32K tokens can only be answered by attending to the whole document. Count every occurrence of “enzyme”.
Packing & masking
Turning documents into training sequences
Training consumes fixed-length sequences, but real documents come in every length. The naive approach — concatenate all documents, then chop every N tokens — cuts documents mid-sentence and produces training instances shorter than the documents they came from. Best-fit document packing instead bins whole documents into fixed-length slots, so almost no document is split and almost no padding is added. On long-context benchmarks this alone is a large win.
Ten sample documents packed into fixed-length sequences. Dark red slivers mark documents cut across a boundary; the light strip at the end of a row is padding.
Intra-document masking
Packing documents next to each other creates a hazard: with only a causal mask, tokens in one document can attend into the previous document. That cross-document signal is spurious — during pretraining the model only ever saw coherent text — and it degrades long-range attention. Intra-document masking adds a block-diagonal structure so each position attends only within its own document (plus the usual causal constraint).
Attention mask for a packed sequence made of three documents (rows = query, columns = key).
OLMo 3 also splits the sequence across devices with 8-way context parallelism: each of eight devices processes 8K of the 65K sequence. The all-gather attention strategy it borrows from Llama 3 supports irregular masks, which is what makes sliding-window and intra-document masking practical at the same time.
Does it work?
RULER and HELMET
Long-context evaluation splits into a development suite and a held-out one. RULER bundles synthetic tasks — variants of needle-in-a-haystack and simple aggregation — and is the signal labs tune against. HELMET is broader (retrieval, in-context learning, summarization) and is kept unseen to check generalization; OLMo 3 is careful to note that RULER and HELMET overlap on the easier subsets, so it isn't a perfect held-out set. Toggle models below and watch how the spread widens as context grows.
RULER dev-suite score (0–100) versus context length. Click a model to toggle it.
Notice the shape of the deltas: at 4K almost every model is in the mid-90s, and the differences are uninformative. It is only at 32K and 65K that the curves separate. OLMo 3 also merges the last three extension checkpoints — a model soup — to get its final long-context 32B base checkpoint.
Cheat sheet
Recap
| Concept | What it is |
|---|---|
| Context extension stage | A short late-training phase that grows the window after pretraining/midtraining — OLMo 3: 8,192 → 65,536 tokens on 50B (7B) / 100B (32B) tokens. |
| Base-frequency scaling | Raise RoPE's base θ so long distances don't wrap past the learned range. |
| Position interpolation | Divide positions by s = L'/L to squeeze long sequences into the pretrained range. |
| YaRN | Per-frequency interpolation plus attention temperature; applied to full-attention layers only in OLMo 3. |
| Longmino Mix | 639B pool of long docs (olmocr PDFs) plus synthetic data; final mix ≈ 34% long / 66% short. |
| gzip filter | Drop the most- and least-compressible 20% of documents. |
| LongPPL | Long-range-dependency proxy using perplexity under more context; tested, not used. |
| CWE / REX | Synthetic aggregation tasks (count a word / rewrite into a vignette) that require the full window. |
| Best-fit packing | Bin whole documents into fixed sequences to minimize splits and padding. |
| Intra-document masking | Block-diagonal mask so packed documents can't attend across boundaries. |
| Context parallelism | Split a long sequence across devices; OLMo 3 uses 8-way CP (8K per device). |
| RULER / HELMET | Synthetic dev suite / broad held-out long-context suite. |
Further reading
References
- OLMo 3 Team (Ai2), "OLMo 3" (2025) — the long-context extension recipe (§3.6) and RULER/HELMET comparison.
- Peng et al., "YaRN: Efficient Context Window Extension of Large Language Models" (2023).
- Chen et al., "Extending Context Window of Large Language Models via Positional Interpolation" (2023).
- Xiong et al., "Effective Long-Context Scaling of Foundation Models" (2023) — base-frequency scaling.
- Ding et al., "Fewer Truncations Improve Language Modeling" (2024) — best-fit packing.
- Gao et al., "ProLong" (2024) and Wu et al., "LongAttn" (2025) — token-efficient extension.
- Hsieh et al., "RULER: What's the Real Context Size of Your Long-Context Language Models?" (2024).
- Yen et al., "HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly" (2024).
- Pham et al., "CLIPPER" (2025) — injected aggregation tasks.
- Serving consequence: LLM Serving, Part 15 takes the same window as given and asks what it costs per request.