Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Where the whole series was heading

This guide began with a vector as an arrow, a list of numbers, and an object you can add and scale. It end here, in a stack of matrix multiplies that predicts the next token. That is not a metaphor. A transformer layer is a short program written in exactly the operations of this series: project a token's embedding with a learned matrix (Part 3), measure similarity between two tokens with an inner product (Part 8), normalise a row of scores so it sums to one (Part 8), and mix the value vectors with those weights (Part 4). Repeat, and add nonlinearity between the blocks.

The applied act of the series built up to this. Part 21 turned the derivative of a vector-valued function into a Jacobian and the Gauss–Newton step into a least-squares problem; that is exactly the machinery used to train the matrices you will see here. Part 17 proved that truncating the SVD gives the closest rank-r matrix; that theorem is, almost word for word, the definition of LoRA. This part is the payoff: the ideas are not toy examples for a textbook, they are what runs.

It is worth saying why a linear algebra guide ends here rather than with a survey of models. The architecture is novel, but the algebra is not. Almost every operation that makes a transformer work is one this series has already named and drawn: a change of basis, a projection onto a subspace, a matrix product that does not commute, a rank that controls how much information survives, a decomposition that sorts directions by how much of the data they carry. Learning to see those in a large model is the same as learning to see them in a small one; only the numbers get bigger.

💡 By the end of this part you'll see why an attention head is three learned matrices producing queries, keys and values, why the score matrix QKᵀ/√d is a table of inner products measuring token similarity, why a row of softmax weights is a recipe for mixing the values, and why fine-tuning a giant model with a rank-r update BA is the same low-rank approximation you already know — just stored as two skinny matrices instead of one big one.
2

An attention head is three matrix multiplies

Similarity, read off an inner product

Collect the token embeddings of a short sentence into a matrix X, one row per token, d columns per embedding. A single attention head learns three d×d weight matrices and multiplies X by each of them:

$$Q = XW_Q, \qquad K = XW_K, \qquad V = XW_V$$

The name of each result says what it is for. A query is what a token is looking for, a key is what a token advertises, and a value is what a token will hand over if it is chosen. The three matrices are learned, so the projection from “meaning” to “looking for” and “advertising” is chosen by training, not by us. But the scoring rule between a query and a key is fixed, and it is the one from Part 8: the inner product. A query matches a key when their dot product is large, so the score matrix is

$$S = \frac{QK^\top}{\sqrt{d}}$$

Row i, column j of S is the similarity of token i's query to token j's key. The division by √d is bookkeeping, not geometry: as the embedding dimension grows, so does the typical size of a dot product, and dividing keeps the numbers from blowing up before the softmax. In the demo below, the head starts with near-identity projections so the first thing you see is raw embedding similarity — the score table is XXᵀ/√d. Edit any entry of W_Q, W_K or W_V to change what each token looks for and advertises, and nudge a single embedding coordinate to watch that token's whole row of scores move.

Rows are queries, columns are keys. The outlined row is the token whose embedding you are nudging.

WQ
WK
WV

Read the score table the way you would read a similarity chart. A dark cell means the row token's query points nearly along the column token's key; a pale cell means the two projections are close to perpendicular. When you nudge one embedding, only the rows and columns that touch that token change, because the score is built from that token's query and key and nothing else. This is the entire mechanism by which a transformer decides “which other tokens is this one about?” — a matrix of inner products, one per pair.

The shapes are worth tracking, because they explain what the block can and cannot do. If X is n×d then each of Q, K and V is n×d, the score matrix is n×n, and the output is n×d again. The one place the sequence length n enters twice is the score matrix, and that is exactly the object that grows with the square of the context in the naive forward pass. Everything else is a fixed-width projection. When people say attention is expensive at long context, they are pointing at the n×n table you just drew and the softmax that has to normalise each of its n rows. A real layer does not use one head but several, splitting the d columns into groups, running a head on each group, and concatenating the results back into one n×d output. The heads are independent attention tables looking at different projected features, and the arithmetic per head is precisely the arithmetic in this demo.

3

Softmax turns scores into a mixing recipe

From similarities to weights that sum to one

The score table is a set of raw numbers, one per pair, and they can be any size and either sign. To use them as weights you need them to be positive and to sum to one along each row, so that “how much of each value to take” is a genuine proportion. The softmax does exactly that, row by row: exponentiate every score so it becomes positive, then divide by the row's total. With S the scored matrix and M the mask,

$$A = \operatorname{softmax}(S + M), \qquad \text{out} = AV$$

where the mask M sets the scores of forbidden pairs to a very large negative number, which the exponential flattens to zero. That is a causal mask: at position i a token is allowed to attend to itself and everything before it, but not to tokens that come later. Turn the toggle on and the upper triangle of the weight matrix goes to zero; that is how a decoder-only model is prevented from reading the answer it is about to predict. The temperature slider divides the scores before exponentiating: a low temperature makes the largest score dominate and the distribution spike, a high temperature flattens it toward a uniform average over the row.

Each row sums to one: how much attention the query token pays to every key.

The output row is a weighted sum of the value vectors, which is why it is the same width as an embedding. A single head produces one such mixed vector per token; a real layer runs several heads side by side and concatenates the results.

Watch the outlined query row. With the mask on, the leading entries survive and the rest are deleted; with the temperature low, the recipe concentrates on one column and the output snaps toward that token's value. Both changes are the same operation — a re-weighting of a linear combination.

This is the point of contact with the rest of the series. The output is literally a linear combination of the rows of V, with coefficients taken from a row of A. Attention is not a special new operation; it is a way of letting the model choose the coefficients of a linear combination from the data, and then applying them. Everything downstream — the feed-forward block, the residual stream, the unembedding — is more matrix multiplies on top of that choice.

Why exponentiate at all, rather than simply normalising the raw scores? Two reasons, and both are visible in the demo. First, scores can be negative and a weight cannot, so the exponential is a smooth way to make every entry positive while preserving the ranking of the scores. Second, it is a soft version of picking the single largest score: a low temperature drives the row toward a one-hot vector on the argmax, while a high temperature drives it toward the uniform average. That single dial runs from hard nearest-neighbour lookup to a plain mean over the row, and every setting in between is a legitimate weighted combination. The mask fits the same frame. Adding a very large negative number to a forbidden score sends its weight to zero without changing the arithmetic of the allowed entries, so the causal constraint costs nothing but one addition before the exponential. In training you can mask and score all positions at once; at generation time the same head runs on one new token against the cached keys and values, which is why the serving guide cares so much about the sizes of K and V.

4

LoRA is the low-rank update you already proved

Eckart–Young, applied to a weight matrix

Fine-tuning a large model by updating every weight is expensive for a reason you can count: a weight matrix with d_{in} columns and d_{out} rows has d_{in}d_{out} numbers, and optimiser state multiplies that again. LoRA freezes the pretrained weight W and learns only a correction ΔW, factored into two skinny matrices:

$$W' = W + BA, \qquad B \in \mathbb{R}^{d_{out}\times r}, \quad A \in \mathbb{R}^{r \times d_{in}}$$

Only B and A are trained, and they hold r(d_{in}+d_{out}) numbers instead of d_{in}d_{out}. For a square weight and a small rank, that is a large saving: at r equal to a few percent of the dimension you store a few percent of the parameters. The reason this is not merely a compression trick is Part 17. When you truncate the SVD of ΔW at rank r,

$$\Delta W_r = U_r\,\Sigma_r\,V_r^\top = B A,$$

you get the closest rank-r matrix to ΔW, with error equal to the energy of the singular values you threw away. Real fine-tuning updates are close to low rank — the direction of adaptation is a few dominant patterns — which is why a small r reconstructs most of the update. The demo below builds a seeded target update and lets you dial the rank, watching the reconstruction fill in.

The target update ΔW, seeded as a low-rank pattern plus a little noise.

The best rank-r approximation BA = U_rΣ_rV_rᵀ.

W + BA

Slide the rank up and two things move together. The picture on the right gains the fine-grained cells it was missing, and the relative error falls toward the floor set by the noise that was never low-rank to begin with. The parameter count climbs linearly, r times the two side lengths, while the full matrix costs their product. There is a rank where the picture already looks like the target and the count is still a small fraction of the whole; finding that rank is the same skill Part 17 taught for images and point clouds, now spent on a model's adaptation.

There is a second reason the factored form is more than storage. Once training is done you can multiply B and A together and add the result to the frozen weight, folding the adapter back into W with no extra matrices at inference time. The product BA is an ordinary d_{out}×d_{in} matrix; it is only during training that you need the two factors, because that is what lets the gradient reach a small number of parameters instead of all of them. LoRA is the low-rank approximation idea used not as a compressor but as a parametrisation: you constrain the update to live in the set of rank-r matrices, and the truncated SVD tells you exactly what that constraint costs in error.

5

Embedding geometry: meaning as angle

Back to the dot product

The scores in the attention head are inner products, and an inner product is a length and a cosine in disguise. For two embedding vectors u and v,

$$\cos\theta = \frac{u\cdot v}{\lVert u\rVert\,\lVert v\rVert}$$

so the angle between two embeddings — not their distance, and not their individual lengths — is what “similar meaning” means when people say a language model represents words as directions. Two tokens with nearly parallel embeddings have a cosine near one; unrelated tokens have a cosine near zero; a cosine near minus one would mean they point opposite. The demo below uses a small seeded set of embeddings built as two clusters, and shows their pairwise cosine similarity. Pick a token and compare its angle to each of the others: the similar tokens cluster on the diagonal blocks, and the odd one out sits in the pale band.

Cosine similarity between token embeddings; the diagonal is 1 because every vector is parallel to itself.

This is the picture Part 8 drew, running at full scale: an attention score is a scaled cosine between two learned projections, and a softmax over those scores is a soft nearest-neighbour lookup over directions.

Seen this way the whole head is geometry. Each token projects its embedding into a query direction and a key direction; the score measures the angle between one token's query and another's key; the softmax picks the nearest few. The value vectors are the payload that gets averaged, and the output is a new direction per token that the next layer will treat as its embedding. Stack the layers and the geometry keeps composing, one projection and one weighted average at a time.

It matters that the measure is an angle and not a distance, and the reason is the normalisation. An embedding's length is not something the token controls in any meaningful way; what carries meaning is the direction it points. Cosine similarity throws the lengths away by construction, so a token cannot become “more important” merely by having a long embedding vector — only by pointing in a direction the query is looking for. That normalisation also makes the quantity bounded between minus one and one, which is what lets you read a cosine matrix as a correlation chart without worrying about scale. The score in the attention head is this same number, except that the query and key are already projections of the embedding, so the model first chooses which directions to compare and only then measures the angle between them. Similarity in a transformer is angle, measured after a learned change of coordinates.

6

Where this shows up

One idea, two worlds

The four moves in this part — project with a matrix, score with an inner product, normalise with a softmax, mix with a weighted sum — are not unique to attention. The same four moves, with different names and different shapes, are what a robot uses to reconcile a thousand noisy measurements into one coherent pose.

ML / AI

The transformer and its cache

Every block of a language model is the matrix multiplies you just edited: the architecture chapter lays out the tensors and how many of them there are, and the KV cache is simply the keys and values from this head kept in memory so a new token does not recompute them. Generation is, in the end, a stream of inner products.

Robotics

Jacobians and least squares

The same operations fit a robot. A pose-graph optimiser linearises its residuals into a Jacobian, forms the normal equations, and solves a sparse least-squares problem — the subject of Part 21, running on real sensors in the pose-graph chapter. Attention and bundle adjustment are different problems wearing the same linear algebra.

7

Further reading

8

Check your understanding

0/4 answered