The Transformer, Block by Block
Part 1 said a language model predicts a distribution over the next token. Part 2 turned text into token vectors. This part is about the function in between: the transformer block, stacked dozens of times, that turns a sequence of vectors into a better sequence of vectors — one that a final linear layer can read off as next-token logits. Every matrix on this page is small enough to compute by hand and large enough to show the real mechanism: nothing here is illustrative or hand-waved.
Why attention is permutation-invariant (for now)
Setup
A transformer block does two things to a sequence of token vectors, repeatedly: it lets tokens exchange information (self-attention), then it processes each position independently (a per-token MLP). Both are wrapped in a residual connection and a normalization step. Stack this block $L$ times and you have a full model — OLMo 2 7B stacks it 32 times.
Self-attention: queries, keys, values
Core mechanism
Each token embedding x is projected three ways by learned weight matrices: a query $Q=XW_Q$ ("what am I looking for"), a key $K=XW_K$ ("what do I contain"), and a value $V=XW_V$ ("what do I offer if attended to"). Every token's query is dot-producted against every token's key, scaled, masked (a token may only attend to itself and earlier tokens — that's what makes generation autoregressive), turned into a probability distribution with softmax, and used to take a weighted average of the values.
Edit the 4 toy token embeddings below (4 tokens × 4 dimensions) and watch every intermediate matrix update — projections, raw scores, the causal mask, the softmax, and the final output.
Multiple heads, running in parallel
Core mechanism
One attention pattern can only capture one notion of "relevant." Transformers instead split Q, K, V into h smaller heads, each with its own learned projections, run attention independently per head, then concatenate the results back to the full width. OLMo 2 7B uses 32 heads of dimension 128 each (32 × 128 = 4096 = dmodel). Some models (including many recent OLMo variants) use grouped-query attention (GQA), where several query heads share one key/value head — this shrinks the KV cache we'll meet in Part 11 with little quality loss.
Each row is one head's attention pattern over the same 4 tokens — different random projections, different pattern.
Positional information: RoPE
Core mechanism
Since raw attention can't see position, every transformer injects positional information somewhere. OLMo, like Llama and most current open models, uses Rotary Position Embeddings (RoPE): instead of adding a position vector, RoPE rotates each query/key vector by an angle proportional to its position, in 2D sub-planes of the embedding. Two vectors rotated by the same angle keep the same dot product — so what matters for the QK dot product is the relative rotation, i.e. the distance between the two positions, not their absolute positions.
Query token is fixed at position 0. As you drag the key token's position n away, both vectors keep rotating together (same relative angle for a given distance) — watch the dot-product-vs-distance curve on the right decay smoothly instead of jumping around, which is exactly the "nearby tokens matter more, and it's about distance, not absolute position" property RoPE is designed to give a model for free.
The residual stream
Architecture
Attention and the MLP don't replace a token's representation — they compute an update that gets added to it: $x \leftarrow x + \text{Attention}(x)$, then $x \leftarrow x + \text{MLP}(x)$. This running sum across all 32 (or however many) layers is the residual stream. It's the reason gradients can flow all the way back to layer 1 in a 32-layer network without vanishing, and it means every layer's job is only to compute a small correction, not to recompute the whole representation from scratch.
LayerNorm, RMSNorm, and where to put it
Architecture
Before feeding a vector into attention or the MLP, it's normalized. The original transformer used LayerNorm (subtract the mean, divide by the standard deviation, then apply a learned scale and shift). Most current LLMs, including OLMo, use RMSNorm — skip the mean-centering, just divide by the root-mean-square and rescale — which is cheaper and works about as well:
Where normalization sits also matters. The original ("post-norm") design normalizes after adding the sublayer output to the residual stream; most modern models use pre-norm (normalize the input to attention/MLP, then add the raw output to the residual stream), which trains far more stably at depth. OLMo 2 made a further, specific tweak: it normalizes the sublayer's output before adding it to the residual stream (sometimes called "post-norm on the branch"), plus adds RMSNorm directly to queries and keys (QK-norm) — both changes the OLMo 2 paper credits with fixing training-loss spikes that plagued OLMo 1 at scale.
The MLP: most of the parameters, half the story
Architecture
After attention mixes information across tokens, the MLP (feed-forward block) processes each token position independently — expand to a wider hidden dimension, apply a nonlinearity, project back down. Modern models use SwiGLU: two parallel "up" projections, one gated through a SiLU nonlinearity and multiplied elementwise into the other, then a "down" projection back to dmodel:
OLMo 2 7B uses an FFN hidden size of 11008 against dmodel=4096 — a ratio of about 2.7× per projection, but three projections (gate, up, down) instead of the plain MLP's two, which is why the MLP block, not attention, holds roughly two-thirds of a dense transformer's parameters.
Parameter-count calculator
Putting it together
Per transformer layer: attention contributes roughly $4d^2$ (Q, K, V, and output projections, each d×d), the MLP contributes roughly $3 d \cdot d_{ffn}$ (gate, up, down), plus two small RMSNorm vectors. Add the embedding table ($V \times d$, doubled if input/output embeddings aren't tied) once, multiply the per-layer cost by L, and you get the whole model. Try the preset for OLMo 2 7B — the calculator should land within a percent or two of the real, published 7B parameter count.
Mixture of experts: more parameters, same compute
Scaling the architecture
A dense model runs every parameter on every token. A mixture-of-experts (MoE) model replaces the single MLP in each block with many parallel "expert" MLPs and a small learned router that picks a handful of them per token. Total parameters grow (more experts to store), but compute per token doesn't (only the chosen experts run) — this is how you get a bigger, more capable model at roughly the same serving cost. Ai2's open MoE, OLMoE, uses 64 experts per layer across 16 MoE layers, routing each token to just 8 of them: 6.9B total parameters, only 1.3B active per token.
The load-balance meter shows how evenly tokens spread across experts — real MoE training adds an auxiliary loss specifically to discourage the router from collapsing onto a favorite few experts.
Cheat sheet
Recap
| Concept | What it is | OLMo 2 7B value |
|---|---|---|
| Self-attention | Softmax(QKᵀ/√dk + mask)V — tokens exchange information | 32 heads × 128 dim |
| RoPE | Rotates Q/K by position so their dot product depends on relative distance | θ base 10,000 |
| Residual stream | Running sum each sublayer adds a correction to | 32 layers deep |
| RMSNorm | Cheaper LayerNorm variant; OLMo 2 also reorders it and adds QK-norm | ε = 1e-5 |
| SwiGLU MLP | Gated expansion → nonlinearity → projection back down | d_ffn = 11,008 |
| MoE | Many experts, few active per token — more params, same compute | OLMoE: 64 experts, top-8, 6.9B/1.3B |
Further reading
References
- Vaswani et al., "Attention Is All You Need" (2017).
- Su et al., "RoFormer: Enhanced Transformer with Rotary Position Embedding" (2021) — RoPE.
- Shazeer, "GLU Variants Improve Transformer" (2020) — SwiGLU.
- OLMo 2 Team (Ai2), "2 OLMo 2 Furious" (2024) — norm reordering, QK-norm, z-loss.
- Muennighoff et al. (Ai2), "OLMoE: Open Mixture-of-Experts Language Models" (2024).