Same map, different coordinates
Everything so far has quietly assumed one thing: that when you write down a vector or a matrix, you and the reader agree on which axes you are using. Part 3 showed that a linear map is determined by where it sends the basis vectors, and Part 11 ended with a map wearing a matrix. This part removes the assumption. A linear map is a physical act — it stretches, shears and rotates the plane — and the act does not know about axes. Coordinates are a choice you lay on top of the act. Change the choice and the act does not move, but every number attached to it does, including the matrix. The goal is to work out exactly how, and to see that the new matrix is $P^{-1}AP$.
The question
One act, many costumes
Think about what a transformation actually is before anyone writes it down. Take the plane, grab every point, and move it: double each point's distance from the horizontal axis, slide the top of the picture sideways, spin the whole sheet by a fixed angle. That is a physical event. It happens to the plane, not to a coordinate system. If you and a friend stand on opposite sides of the sheet and describe the same event, you are describing the same thing.
Now hand each of you a set of axes. You both watch the same event, but you record it differently, because your axes point in different directions. You write down a matrix, your friend writes down a different matrix, and both are correct. This is the tension the whole part lives in: the transformation is one object, but its matrix depends on the basis. If the transformation were a piece of music, the matrix would be the score in a particular key — transposing the key changes every note on the page and not one note you hear.
There is a second half to the story, and it is about vectors. A vector is an arrow, and the arrow does not change when you change your axes. Its coordinates change, because coordinates are answers to the question "how far along each axis", and the axes moved. So the same physical arrow is (3, 1) in one basis and something else entirely in another, while the tip of the arrow stays exactly where it was. Two things change under a change of basis — the matrix of a map, and the coordinates of a vector — and they change in opposite directions so that the equation $A\mathbf{x} = \mathbf{b}$ keeps telling the same physical truth.
The map does not know about axes
Two frames, one transformation
Fix a linear map $A$, written in the standard basis, and choose a new basis $\mathcal{B}$ with vectors $\mathbf{b}_1$ and $\mathbf{b}_2$. Stack those basis vectors side by side as the columns of a matrix:
The matrix $P$ is a translator. If a vector has coordinates $(c_1, c_2)$ in the new basis, then its physical position — the arrow you could draw in the room — is $c_1\mathbf{b}_1 + c_2\mathbf{b}_2$, which is exactly $P$ times the column $[c_1, c_2]^{\mathsf{T}}$. In symbols, $\mathbf{x} = P\,[\mathbf{x}]_{\mathcal{B}}$. Going the other way, from a physical arrow back to $\mathcal{B}$-coordinates, is the inverse translation $P^{-1}$.
That is all you need to find the new matrix. Suppose you want to apply the map $A$ to a vector that arrives in $\mathcal{B}$-coordinates. Three steps, in order: translate to physical coordinates with $P$, apply the map with $A$, translate the result back with $P^{-1}$. Composition of maps is multiplication of matrices, so the whole round trip is one matrix:
Read the order right to left the way the arrows act: $P$ first, then $A$, then $P^{-1}$. This is called a similarity transformation, and two matrices related this way are similar. They are not equal — they rarely are — but they describe the same map and therefore share everything that is intrinsic to the map: the determinant, the trace, the eigenvalues, the rank, whether the map is invertible. Those quantities are properties of the act, not of the score, which is why they survive the change of basis untouched.
The demo below makes the translation visible. On the left is the world: a fixed map $A$ drawn as the deformed blue grid, with your chosen basis vectors $\mathbf{b}_1$ and $\mathbf{b}_2$ overlaid as draggable arrows. On the right is the same map, but every point is now recorded in $\mathcal{B}$-coordinates, so the deformed grid is drawn with the matrix $P^{-1}AP$. Drag the basis vectors and watch both the pink grid and the matrix change while the physical map on the left stays exactly where it was.
World coordinates: the deformed blue grid is the fixed map $A$. Drag b₁ and b₂ to choose the new basis.
The same map in $\mathcal{B}$-coordinates: the pink grid is $P^{-1}AP$. The arrows b₁ and b₂ are now the ordinary unit vectors.
The probe arrow is the consistency check. On the left, x is drawn in world coordinates and Ax is the map applied to it. On the right, the very same physical arrow is drawn at $[\mathbf{x}]_{\mathcal{B}} = P^{-1}\mathbf{x}$, and its image under the new matrix is at $P^{-1}A\mathbf{x}$, which is exactly $[\mathbf{Ax}]_{\mathcal{B}}$. Two different pairs of numbers, one physical result. If the basis vectors become collinear the matrix $P$ stops being invertible, the right-hand picture has no basis to draw in, and the readout says so — a basis has to be a genuine basis, not a collapsed one, which is the lesson of Part 2.
A point's coordinates change; the arrow does not
Solving for the coefficients
The matrix of a map is one half of the story; the coordinates of a vector are the other. They are two sides of the same equation. A vector $\mathbf{x}$ is a physical arrow, and its $\mathcal{B}$-coordinates are the unique coefficients that build that arrow out of the basis vectors:
Finding those coefficients is a linear solve. You know the physical arrow $\mathbf{x}$, you know the basis columns of $P$, and you want the coordinate column, so you solve $P\,[\mathbf{x}]_{\mathcal{B}} = \mathbf{x}$ for $[\mathbf{x}]_{\mathcal{B}}$. In the standard basis the answer is immediate, because $P$ is the identity and the coordinates are the entries. In any other basis the coefficients are what a solver returns, and they can be surprising: an arrow that sits comfortably at (1, 1) can have coordinates like (0.3, 0.7) once the axes are skewed.
The demo below puts the two records side by side. On the left, the world: the standard grid, the basis vectors $\mathbf{b}_1$ and $\mathbf{b}_2$, and a draggable point x decomposed into a dashed parallelogram whose edges are the two scaled basis vectors. On the right, the $\mathcal{B}$-coordinate plane: the same arrow, now drawn at the coordinates the solver found. Move the point on the left and both pictures update together — the arrow moves, and its description moves, but they never disagree about where the arrow is.
World coordinates. The dashed edges are $c_1\mathbf{b}_1$ and $c_2\mathbf{b}_2$; they meet at the tip of x. Drag the point.
The $\mathcal{B}$-coordinate plane. The same arrow is now plotted at $[\mathbf{x}]_{\mathcal{B}}$ on the ordinary unit axes.
Watch the last line of the readout. It computes $P[\mathbf{x}]_{\mathcal{B}} - \mathbf{x}$, the gap between the arrow rebuilt from its new coordinates and the arrow you dragged. It stays at zero, because that is what it means for coordinates to be correct: they reconstruct the vector exactly. This is the sense in which a vector's coordinates are a representation and not the vector itself. The arrow is the invariant object; the list of numbers is one basis's opinion about it. When you change the basis you are not moving the arrow, you are re-answering the question "how far along each axis", and the answer changes because the axes did.
There is a tidy way to remember which way each quantity transforms. Coordinates carry the inverse of the basis matrix, $[\mathbf{x}]_{\mathcal{B}} = P^{-1}\mathbf{x}$, while the map's matrix carries the whole sandwich, $A_{\mathcal{B}} = P^{-1}AP$. The two inverses cancel in the equation they serve. Apply the map in $\mathcal{B}$-coordinates, then convert back, and you get the world answer: $P\,(P^{-1}AP)\,(P^{-1}\mathbf{x}) = A\mathbf{x}$. That cancellation is why physics does not care which basis you compute in, and it is the algebraic heart of this part.
A basis that makes the matrix diagonal
Choosing axes the map already likes
If the matrix of a map depends on the basis, a natural question follows: is there a basis in which it looks as simple as possible? The simplest matrix is a diagonal one, because it scales each axis independently and mixes nothing. So we are looking for basis vectors that the map does not rotate, only stretches. Those directions have a name — eigenvectors — and the stretch factors are their eigenvalues. They are the subject of Part 13; here they appear only as the answer to the change-of-basis question.
The reason eigenvectors diagonalise a matrix is exactly the formula from the last section. If you build $P$ from eigenvectors, then $P^{-1}AP$ is diagonal, with the eigenvalues on the diagonal:
You can read the algebra straight off the geometry. In the eigenvector basis, the map sends the first basis arrow to $\lambda_1$ times itself and the second to $\lambda_2$ times itself. A map that only stretches its own basis vectors has a diagonal matrix, by definition. The off-diagonal entries measure how much each basis direction leaks into the other, so they vanish precisely when the basis vectors are the directions the map refuses to turn. This is the whole content of diagonalisation, and it is a change of basis in a very particular, very useful costume.
The demo lets you hunt for those directions by hand. The purple dashed lines are the eigen-directions of the fixed map $A$, drawn faintly so you can aim at them. Drag the handles $\mathbf{b}_1$ and $\mathbf{b}_2$ until the off-diagonal entries in the readout fall to zero, and the matrix $A_{\mathcal{B}}$ snaps to a diagonal. The dashed pink arrows show where the map sends each basis vector; when a dashed arrow lies along its own solid arrow, that basis vector is an eigenvector. If aiming by hand is tedious, the button places the basis on the eigenvectors in one step.
Drag b₁ and b₂ toward the purple eigen-directions. The pink dashed arrows are the images Ab₁ and Ab₂; when one lies along its own basis arrow, the matrix is diagonal.
Not every map rewards this search. A rotation has no real eigen-directions at all, so no real basis makes it diagonal; the best you can do is a block that still contains a rotation. That is why the button can only snap when the eigenvectors are real, and why the theory of eigenvalues has to admit complex numbers to stay complete. When the eigenvectors do exist and are independent, they form the friendliest possible basis for the map, and every power of the matrix becomes trivial to compute — the story Part 14 tells. When the map is also symmetric, the eigenvectors can even be chosen perpendicular, which is what makes Part 15 so clean.
Where this shows up
One change of basis, two worlds
Change of basis sounds like bookkeeping, and in a sense it is: it is the bookkeeping that lets two people describe one physical situation without arguing. That is why it is everywhere. A robot and a language model both spend most of their linear algebra moving between frames that are convenient for different parts of a computation, and the matrix $P^{-1}AP$ is the exchange rate between them.
Body frame versus world frame
A robot's sensors measure motion in its own body frame, while its map and its goals live in the world frame. The same physical motion has different coordinates in the two frames, and the rotation matrix between them is exactly a change of basis. Composing those frames down a chain of links is how a controller knows where its hand is: the pose it builds is precisely these translations and rotations composed. The pinhole camera part uses the same idea to move between camera-centred and world-centred coordinates.
Embeddings and attention in a learned basis
A transformer's embedding space has no preferred axes — the model is free to learn whichever basis makes its computation easy, and the numbers attached to a token depend on that learned frame. Attention is computed after projecting the embeddings into query, key and value bases, and each projection is a change of coordinates in which the same semantic arrow wears different numbers. The architecture chapter introduces those projections; Part 22 reads attention as exactly these matrices.
The common thread is invariance. In both fields, some quantity is physically real — a joint angle, a distance between embeddings — and the computation is free to pick whichever basis is cheapest, because the answer does not depend on the choice. Recognising that freedom is what lets a robotics stack factor a long chain of transforms, and what lets a model project its activations into a dozen subspaces without losing track of the underlying object. Change of basis is the licence to choose your coordinates to suit the problem.
Cheat sheet
| Idea | Formula | What it means |
|---|---|---|
| Change-of-basis matrix | $P = [\mathbf{b}_1\;\mathbf{b}_2]$ | The new basis vectors, stacked as columns |
| Coordinates of a vector | $\mathbf{x} = P\,[\mathbf{x}]_{\mathcal{B}}$ | Build the arrow from its coefficients |
| Finding the coordinates | $[\mathbf{x}]_{\mathcal{B}} = P^{-1}\mathbf{x}$ | Solve $P[\mathbf{x}]_{\mathcal{B}} = \mathbf{x}$ |
| Matrix of the same map | $A_{\mathcal{B}} = P^{-1}AP$ | Translate in, apply, translate back |
| Similar matrices | $A_{\mathcal{B}} \sim A$ | Same determinant, trace, eigenvalues, rank |
| Consistency | $P\,(P^{-1}AP)\,(P^{-1}\mathbf{x}) = A\mathbf{x}$ | The inverses cancel; the physics is unchanged |
| Diagonal basis | $P^{-1}AP = \operatorname{diag}(\lambda_1, \lambda_2)$ | Chosen when the columns are eigenvectors |
The first three rows are the mechanics of coordinates; the next two are the mechanics of the matrix; the last two are the reason the mechanics are worth knowing. Read the table as one sentence: a map and a vector both have intrinsic meaning, a basis is a choice, and $P$ is the dictionary between the choice and the meaning.
Further reading
- Grant Sanderson, "Change of basis", Essence of Linear Algebra, 3Blue1Brown — the chapter this part is in conversation with, including the translation picture and the notation $A_{\mathcal{B}} = P^{-1}AP$.
- Gilbert Strang, 18.06 Linear Algebra, MIT OpenCourseWare — Lectures 27–28 on change of basis, similar matrices and diagonalisation, with the eigenvalue connection.
- Immersive Math, Chapter 8: Change of Basis — the same translation between frames as a draggable figure.
- Sheldon Axler, Linear Algebra Done Right, chapter 3 and chapter 5 — change of basis treated through operators, and the diagonalisation that follows.