Linear algebra in ML and AI
Every claim in this series has been about a picture you could draw or a decomposition you could compute. The last part of the guide points all of it at the machinery of a modern language model. A transformer attention head is three learned matrices multiplying the same token embeddings three different ways, an inner product that scores how alike two tokens are, and a softmax that turns those scores into a mixing recipe. Fine-tuning those weights with LoRA is the low-rank approximation you built a few parts ago, applied to a frozen matrix. Nothing new is introduced; the same four operations, stacked deep.
The question
Where the whole series was heading
This guide began with a vector as an arrow, a list of numbers, and an object you can add and scale. It end here, in a stack of matrix multiplies that predicts the next token. That is not a metaphor. A transformer layer is a short program written in exactly the operations of this series: project a token's embedding with a learned matrix (Part 3), measure similarity between two tokens with an inner product (Part 8), normalise a row of scores so it sums to one (Part 8), and mix the value vectors with those weights (Part 4). Repeat, and add nonlinearity between the blocks.
The applied act of the series built up to this. Part 21 turned the derivative of a vector-valued function into a Jacobian and the Gauss–Newton step into a least-squares problem; that is exactly the machinery used to train the matrices you will see here. Part 17 proved that truncating the SVD gives the closest rank-r matrix; that theorem is, almost word for word, the definition of LoRA. This part is the payoff: the ideas are not toy examples for a textbook, they are what runs.
It is worth saying why a linear algebra guide ends here rather than with a survey of models. The architecture is novel, but the algebra is not. Almost every operation that makes a transformer work is one this series has already named and drawn: a change of basis, a projection onto a subspace, a matrix product that does not commute, a rank that controls how much information survives, a decomposition that sorts directions by how much of the data they carry. Learning to see those in a large model is the same as learning to see them in a small one; only the numbers get bigger.
An attention head is three matrix multiplies
Similarity, read off an inner product
Collect the token embeddings of a short sentence into a matrix X, one row per token, d columns per embedding. A single attention head learns three d×d weight matrices and multiplies X by each of them:
The name of each result says what it is for. A query is what a token is looking for, a key is what a token advertises, and a value is what a token will hand over if it is chosen. The three matrices are learned, so the projection from “meaning” to “looking for” and “advertising” is chosen by training, not by us. But the scoring rule between a query and a key is fixed, and it is the one from Part 8: the inner product. A query matches a key when their dot product is large, so the score matrix is
Row i, column j of S is the similarity of token i's query to token j's key. The division by √d is bookkeeping, not geometry: as the embedding dimension grows, so does the typical size of a dot product, and dividing keeps the numbers from blowing up before the softmax. In the demo below, the head starts with near-identity projections so the first thing you see is raw embedding similarity — the score table is XXᵀ/√d. Edit any entry of W_Q, W_K or W_V to change what each token looks for and advertises, and nudge a single embedding coordinate to watch that token's whole row of scores move.
Rows are queries, columns are keys. The outlined row is the token whose embedding you are nudging.
Read the score table the way you would read a similarity chart. A dark cell means the row token's query points nearly along the column token's key; a pale cell means the two projections are close to perpendicular. When you nudge one embedding, only the rows and columns that touch that token change, because the score is built from that token's query and key and nothing else. This is the entire mechanism by which a transformer decides “which other tokens is this one about?” — a matrix of inner products, one per pair.
The shapes are worth tracking, because they explain what the block can and cannot do. If X is n×d then each of Q, K and V is n×d, the score matrix is n×n, and the output is n×d again. The one place the sequence length n enters twice is the score matrix, and that is exactly the object that grows with the square of the context in the naive forward pass. Everything else is a fixed-width projection. When people say attention is expensive at long context, they are pointing at the n×n table you just drew and the softmax that has to normalise each of its n rows. A real layer does not use one head but several, splitting the d columns into groups, running a head on each group, and concatenating the results back into one n×d output. The heads are independent attention tables looking at different projected features, and the arithmetic per head is precisely the arithmetic in this demo.
Softmax turns scores into a mixing recipe
From similarities to weights that sum to one
The score table is a set of raw numbers, one per pair, and they can be any size and either sign. To use them as weights you need them to be positive and to sum to one along each row, so that “how much of each value to take” is a genuine proportion. The softmax does exactly that, row by row: exponentiate every score so it becomes positive, then divide by the row's total. With S the scored matrix and M the mask,
where the mask M sets the scores of forbidden pairs to a very large negative number, which the exponential flattens to zero. That is a causal mask: at position i a token is allowed to attend to itself and everything before it, but not to tokens that come later. Turn the toggle on and the upper triangle of the weight matrix goes to zero; that is how a decoder-only model is prevented from reading the answer it is about to predict. The temperature slider divides the scores before exponentiating: a low temperature makes the largest score dominate and the distribution spike, a high temperature flattens it toward a uniform average over the row.
Each row sums to one: how much attention the query token pays to every key.
Watch the outlined query row. With the mask on, the leading entries survive and the rest are deleted; with the temperature low, the recipe concentrates on one column and the output snaps toward that token's value. Both changes are the same operation — a re-weighting of a linear combination.
This is the point of contact with the rest of the series. The output is literally a linear combination of the rows of V, with coefficients taken from a row of A. Attention is not a special new operation; it is a way of letting the model choose the coefficients of a linear combination from the data, and then applying them. Everything downstream — the feed-forward block, the residual stream, the unembedding — is more matrix multiplies on top of that choice.
Why exponentiate at all, rather than simply normalising the raw scores? Two reasons, and both are visible in the demo. First, scores can be negative and a weight cannot, so the exponential is a smooth way to make every entry positive while preserving the ranking of the scores. Second, it is a soft version of picking the single largest score: a low temperature drives the row toward a one-hot vector on the argmax, while a high temperature drives it toward the uniform average. That single dial runs from hard nearest-neighbour lookup to a plain mean over the row, and every setting in between is a legitimate weighted combination. The mask fits the same frame. Adding a very large negative number to a forbidden score sends its weight to zero without changing the arithmetic of the allowed entries, so the causal constraint costs nothing but one addition before the exponential. In training you can mask and score all positions at once; at generation time the same head runs on one new token against the cached keys and values, which is why the serving guide cares so much about the sizes of K and V.
LoRA is the low-rank update you already proved
Eckart–Young, applied to a weight matrix
Fine-tuning a large model by updating every weight is expensive for a reason you can count: a weight matrix with d_{in} columns and d_{out} rows has d_{in}d_{out} numbers, and optimiser state multiplies that again. LoRA freezes the pretrained weight W and learns only a correction ΔW, factored into two skinny matrices:
Only B and A are trained, and they hold r(d_{in}+d_{out}) numbers instead of d_{in}d_{out}. For a square weight and a small rank, that is a large saving: at r equal to a few percent of the dimension you store a few percent of the parameters. The reason this is not merely a compression trick is Part 17. When you truncate the SVD of ΔW at rank r,
you get the closest rank-r matrix to ΔW, with error equal to the energy of the singular values you threw away. Real fine-tuning updates are close to low rank — the direction of adaptation is a few dominant patterns — which is why a small r reconstructs most of the update. The demo below builds a seeded target update and lets you dial the rank, watching the reconstruction fill in.
The target update ΔW, seeded as a low-rank pattern plus a little noise.
The best rank-r approximation BA = U_rΣ_rV_rᵀ.
Slide the rank up and two things move together. The picture on the right gains the fine-grained cells it was missing, and the relative error falls toward the floor set by the noise that was never low-rank to begin with. The parameter count climbs linearly, r times the two side lengths, while the full matrix costs their product. There is a rank where the picture already looks like the target and the count is still a small fraction of the whole; finding that rank is the same skill Part 17 taught for images and point clouds, now spent on a model's adaptation.
There is a second reason the factored form is more than storage. Once training is done you can multiply B and A together and add the result to the frozen weight, folding the adapter back into W with no extra matrices at inference time. The product BA is an ordinary d_{out}×d_{in} matrix; it is only during training that you need the two factors, because that is what lets the gradient reach a small number of parameters instead of all of them. LoRA is the low-rank approximation idea used not as a compressor but as a parametrisation: you constrain the update to live in the set of rank-r matrices, and the truncated SVD tells you exactly what that constraint costs in error.
Embedding geometry: meaning as angle
Back to the dot product
The scores in the attention head are inner products, and an inner product is a length and a cosine in disguise. For two embedding vectors u and v,
so the angle between two embeddings — not their distance, and not their individual lengths — is what “similar meaning” means when people say a language model represents words as directions. Two tokens with nearly parallel embeddings have a cosine near one; unrelated tokens have a cosine near zero; a cosine near minus one would mean they point opposite. The demo below uses a small seeded set of embeddings built as two clusters, and shows their pairwise cosine similarity. Pick a token and compare its angle to each of the others: the similar tokens cluster on the diagonal blocks, and the odd one out sits in the pale band.
Cosine similarity between token embeddings; the diagonal is 1 because every vector is parallel to itself.
Seen this way the whole head is geometry. Each token projects its embedding into a query direction and a key direction; the score measures the angle between one token's query and another's key; the softmax picks the nearest few. The value vectors are the payload that gets averaged, and the output is a new direction per token that the next layer will treat as its embedding. Stack the layers and the geometry keeps composing, one projection and one weighted average at a time.
It matters that the measure is an angle and not a distance, and the reason is the normalisation. An embedding's length is not something the token controls in any meaningful way; what carries meaning is the direction it points. Cosine similarity throws the lengths away by construction, so a token cannot become “more important” merely by having a long embedding vector — only by pointing in a direction the query is looking for. That normalisation also makes the quantity bounded between minus one and one, which is what lets you read a cosine matrix as a correlation chart without worrying about scale. The score in the attention head is this same number, except that the query and key are already projections of the embedding, so the model first chooses which directions to compare and only then measures the angle between them. Similarity in a transformer is angle, measured after a learned change of coordinates.
Where this shows up
One idea, two worlds
The four moves in this part — project with a matrix, score with an inner product, normalise with a softmax, mix with a weighted sum — are not unique to attention. The same four moves, with different names and different shapes, are what a robot uses to reconcile a thousand noisy measurements into one coherent pose.
The transformer and its cache
Every block of a language model is the matrix multiplies you just edited: the architecture chapter lays out the tensors and how many of them there are, and the KV cache is simply the keys and values from this head kept in memory so a new token does not recompute them. Generation is, in the end, a stream of inner products.
Jacobians and least squares
The same operations fit a robot. A pose-graph optimiser linearises its residuals into a Jacobian, forms the normal equations, and solves a sparse least-squares problem — the subject of Part 21, running on real sensors in the pose-graph chapter. Attention and bundle adjustment are different problems wearing the same linear algebra.
Further reading
- Denim Patel, LLM training guide — the architecture this part describes, built up tensor by tensor.
- 3Blue1Brown, "The SVD" and "Attention in transformers, visually explained" — the two visuals this part is in conversation with.
- Gilbert Strang, 18.06 Linear Algebra, MIT OpenCourseWare — Lectures 29–30, the SVD and linear transformations.
- Jay Alammar, "The Illustrated Transformer" — the standard picture-first walkthrough of Q, K, V and multi-head attention.
- Edward Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models", 2021 — the paper that made W + BA a default.
- The series reference card: Glossary and identities to know.