Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

Why attention is permutation-invariant (for now)

Setup

A transformer block does two things to a sequence of token vectors, repeatedly: it lets tokens exchange information (self-attention), then it processes each position independently (a per-token MLP). Both are wrapped in a residual connection and a normalization step. Stack this block $L$ times and you have a full model — OLMo 2 7B stacks it 32 times.

💡 The shape of this part: we'll build the block piece by piece — attention, multiple heads, positional encoding (attention alone can't tell position i from position j!), the residual stream those pieces write into, normalization, the MLP, and finally how to count and route parameters at scale.
1

Self-attention: queries, keys, values

Core mechanism

Each token embedding x is projected three ways by learned weight matrices: a query $Q=XW_Q$ ("what am I looking for"), a key $K=XW_K$ ("what do I contain"), and a value $V=XW_V$ ("what do I offer if attended to"). Every token's query is dot-producted against every token's key, scaled, masked (a token may only attend to itself and earlier tokens — that's what makes generation autoregressive), turned into a probability distribution with softmax, and used to take a weighted average of the values.

$$\text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}} + M\right)V$$

Edit the 4 toy token embeddings below (4 tokens × 4 dimensions) and watch every intermediate matrix update — projections, raw scores, the causal mask, the softmax, and the final output.

scores = QKᵀ/√dk (masked)
→ softmax →
attention weights
× V =
Hover a cell in the score heatmap to see its dot product.
⚠️ Notice: nothing above the mask/softmax step knows token order — swap two rows of the input and you get exactly the same set of dot products, just relabeled. Attention by itself is permutation-invariant. We fix that in the RoPE section below.
2

Multiple heads, running in parallel

Core mechanism

One attention pattern can only capture one notion of "relevant." Transformers instead split Q, K, V into h smaller heads, each with its own learned projections, run attention independently per head, then concatenate the results back to the full width. OLMo 2 7B uses 32 heads of dimension 128 each (32 × 128 = 4096 = dmodel). Some models (including many recent OLMo variants) use grouped-query attention (GQA), where several query heads share one key/value head — this shrinks the KV cache we'll meet in Part 11 with little quality loss.

Each row is one head's attention pattern over the same 4 tokens — different random projections, different pattern.

3

Positional information: RoPE

Core mechanism

Since raw attention can't see position, every transformer injects positional information somewhere. OLMo, like Llama and most current open models, uses Rotary Position Embeddings (RoPE): instead of adding a position vector, RoPE rotates each query/key vector by an angle proportional to its position, in 2D sub-planes of the embedding. Two vectors rotated by the same angle keep the same dot product — so what matters for the QK dot product is the relative rotation, i.e. the distance between the two positions, not their absolute positions.

$$q_m^\top k_n \text{ depends only on } (m-n) \text{ after rotating } q_m, k_n \text{ by angles } m\theta, n\theta$$

Query token is fixed at position 0. As you drag the key token's position n away, both vectors keep rotating together (same relative angle for a given distance) — watch the dot-product-vs-distance curve on the right decay smoothly instead of jumping around, which is exactly the "nearby tokens matter more, and it's about distance, not absolute position" property RoPE is designed to give a model for free.

4

The residual stream

Architecture

Attention and the MLP don't replace a token's representation — they compute an update that gets added to it: $x \leftarrow x + \text{Attention}(x)$, then $x \leftarrow x + \text{MLP}(x)$. This running sum across all 32 (or however many) layers is the residual stream. It's the reason gradients can flow all the way back to layer 1 in a 32-layer network without vanishing, and it means every layer's job is only to compute a small correction, not to recompute the whole representation from scratch.

⚠️ Disable the residual connection and each layer's random, untrained-looking update simply overwrites the last — the representation's norm drifts and the original token identity is lost by layer 3–4. This is a real, not merely didactic, failure mode: very deep networks without residual connections are notoriously hard to train.
5

LayerNorm, RMSNorm, and where to put it

Architecture

Before feeding a vector into attention or the MLP, it's normalized. The original transformer used LayerNorm (subtract the mean, divide by the standard deviation, then apply a learned scale and shift). Most current LLMs, including OLMo, use RMSNorm — skip the mean-centering, just divide by the root-mean-square and rescale — which is cheaper and works about as well:

$$\text{RMSNorm}(x) = \frac{x}{\sqrt{\frac{1}{d}\sum_i x_i^2 + \epsilon}} \cdot \gamma$$

Where normalization sits also matters. The original ("post-norm") design normalizes after adding the sublayer output to the residual stream; most modern models use pre-norm (normalize the input to attention/MLP, then add the raw output to the residual stream), which trains far more stably at depth. OLMo 2 made a further, specific tweak: it normalizes the sublayer's output before adding it to the residual stream (sometimes called "post-norm on the branch"), plus adds RMSNorm directly to queries and keys (QK-norm) — both changes the OLMo 2 paper credits with fixing training-loss spikes that plagued OLMo 1 at scale.

💡 OLMo hook: if you've read Part 5's loss-spike section, this is the architectural half of the fix; the optimizer/learning-rate half lives there.
6

The MLP: most of the parameters, half the story

Architecture

After attention mixes information across tokens, the MLP (feed-forward block) processes each token position independently — expand to a wider hidden dimension, apply a nonlinearity, project back down. Modern models use SwiGLU: two parallel "up" projections, one gated through a SiLU nonlinearity and multiplied elementwise into the other, then a "down" projection back to dmodel:

$$\text{SwiGLU}(x) = \big(\text{SiLU}(xW_{\text{gate}}) \odot xW_{\text{up}}\big)\,W_{\text{down}}$$

OLMo 2 7B uses an FFN hidden size of 11008 against dmodel=4096 — a ratio of about 2.7× per projection, but three projections (gate, up, down) instead of the plain MLP's two, which is why the MLP block, not attention, holds roughly two-thirds of a dense transformer's parameters.

7

Parameter-count calculator

Putting it together

Per transformer layer: attention contributes roughly $4d^2$ (Q, K, V, and output projections, each d×d), the MLP contributes roughly $3 d \cdot d_{ffn}$ (gate, up, down), plus two small RMSNorm vectors. Add the embedding table ($V \times d$, doubled if input/output embeddings aren't tied) once, multiply the per-layer cost by L, and you get the whole model. Try the preset for OLMo 2 7B — the calculator should land within a percent or two of the real, published 7B parameter count.

8

Mixture of experts: more parameters, same compute

Scaling the architecture

A dense model runs every parameter on every token. A mixture-of-experts (MoE) model replaces the single MLP in each block with many parallel "expert" MLPs and a small learned router that picks a handful of them per token. Total parameters grow (more experts to store), but compute per token doesn't (only the chosen experts run) — this is how you get a bigger, more capable model at roughly the same serving cost. Ai2's open MoE, OLMoE, uses 64 experts per layer across 16 MoE layers, routing each token to just 8 of them: 6.9B total parameters, only 1.3B active per token.

The load-balance meter shows how evenly tokens spread across experts — real MoE training adds an auxiliary loss specifically to discourage the router from collapsing onto a favorite few experts.

✓

Cheat sheet

Recap

ConceptWhat it isOLMo 2 7B value
Self-attentionSoftmax(QKᵀ/√dk + mask)V — tokens exchange information32 heads × 128 dim
RoPERotates Q/K by position so their dot product depends on relative distanceθ base 10,000
Residual streamRunning sum each sublayer adds a correction to32 layers deep
RMSNormCheaper LayerNorm variant; OLMo 2 also reorders it and adds QK-normε = 1e-5
SwiGLU MLPGated expansion → nonlinearity → projection back downd_ffn = 11,008
MoEMany experts, few active per token — more params, same computeOLMoE: 64 experts, top-8, 6.9B/1.3B
📚

Further reading

References

?

Check your understanding

0/6 answered
The architecture is fixed once training starts. What varies — and what makes or breaks the run — is the data it's trained on. Continue: building the pretraining dataset →