Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The one picture that makes matrices inevitable

A rule, not a table of numbers

Most introductions to matrices start with a rectangular block of numbers and a bizarre multiplication rule, leaving you to wonder who decided it. Turn it around. Suppose you want a rule that moves points around while staying compatible with vector addition and scaling. Compatibility forces the grid of integer lines to stay a grid: lines that were parallel stay parallel, evenly spaced lines stay evenly spaced, and the origin stays put. A map with that behaviour is called linear, and a linear map is completely described by two pieces of information in the plane — where the basis vector î = (1, 0) lands, and where ĵ = (0, 1) lands.

The reason two arrows are enough is the linear-combination idea from the previous parts. Every vector is a combination of the basis vectors,

$$\mathbf{p} = x\,\hat{\mathbf{\imath}} + y\,\hat{\mathbf{\jmath}}.$$

Apply the transformation to both sides. Linearity lets you pull the map through the scalars and the sum, so the image of p is the same combination of the images of the basis vectors. If you know where the two basis vectors go, you know where everything goes. That is the entire content of the next four demos.

Write the image of î as a column and the image of ĵ as another column, and stack them side by side. The resulting array is the matrix of the transformation, and the rule above becomes matrix-vector multiplication. A matrix is not an arbitrary table; it is a compact record of where a linear map sends the basis, arranged so that evaluating the map is a single combination.

Spelled out in symbols, the whole procedure is one line. Suppose î goes to the column (a, c) and ĵ goes to the column (b, d). Then a general input (x, y), which is x î + y ĵ, must land on

$$\begin{bmatrix} a & b \\ c & d \end{bmatrix}\begin{bmatrix} x \\ y \end{bmatrix} = x\begin{bmatrix} a \\ c \end{bmatrix} + y\begin{bmatrix} b \\ d \end{bmatrix} = \begin{bmatrix} ax + by \\ cx + dy \end{bmatrix}.$$

Read the middle expression as an instruction and the rightmost as arithmetic. The instruction says: use x as a weight on the first column and y as a weight on the second, then add. The arithmetic says the same thing entry by entry. There is no mysterious rule being imposed from outside; the familiar row-times-column formula is simply the coordinate form of “combine the transformed basis vectors”.

It is worth pausing on why the origin is special. A linear map has no freedom to translate; if it moved the origin to some point q, then transforming the zero vector would give q, and transforming 0 · v would give 0 · T(v) = 0, a contradiction unless q is zero. Linearity alone forces the origin to stay put, which is why every matrix sends (0, 0) to (0, 0) and why the grid lines of the picture all pass through the same fixed point. Translations need homogeneous coordinates precisely because they are not linear; they are affine, and affine maps are linear maps followed by a shift. Keeping that distinction straight will matter when we meet camera projections and robot poses.

💡 By the end of this part you'll see why a linear map is determined by just two arrows, why those arrows are literally the columns of its matrix, and why multiplying a matrix by a vector is the same act as combining the transformed basis vectors by that vector's coordinates.
2

Moving the grid with two arrows

Drag the images of î and ĵ

Below, the two handles are the destinations of î and ĵ. Drag either one and the entire grid deforms with it, because the grid is just the images of all the integer-spaced lines, and each of those lines is dragged along by linearity. The straight solid arrows are the transformed basis vectors, and the faded square is the image of the unit square, whose sides are exactly those two arrows. The probe vector p, chosen by the sliders, is drawn dashed in its original position, and its image Mp is drawn solid: it is the combination x along the first column plus y along the second.

The live matrix on the right is the same information in numeric form. Its first column is the coordinates of the first handle and its second column is the coordinates of the second. Nothing is hidden: editing the geometry edits the numbers, and the readout reports where the probe actually lands. Watch what stays true no matter where you drag. The origin never moves. Straight lines through the origin stay straight. Parallel lines stay parallel. If the two handles line up, the grid collapses onto a single line and the image of the unit square flattens to zero area — the matrix has lost a dimension, which is the rank collapse from Part 2 reappearing inside a map.

Drag the î and ĵ handles; the grid and the unit square follow. Move the sliders to send a probe vector through the map.

A useful habit is to name the columns by what they do rather than by their entries. Column one is “where î goes” column two is “where ĵ goes”. When you later meet a matrix in someone else's code, the fastest way to understand it is to feed it (1, 0) and (0, 1) and look at the two outputs. Those outputs are the columns, and the columns are the transformation.

It is just as useful to notice what linearity rules out. A translation — sliding every point one unit to the right — is not linear, because it moves the origin and it does not preserve vector addition. A map that bends straight lines into curves is not linear either. The grid picture is a strict test, not a stylistic preference: if the origin moves, or if parallel evenly spaced lines lose their spacing, no matrix can represent the map. This is why so much of applied mathematics spends its effort approximating curved maps by linear ones near a point. The Jacobian of calculus is exactly that: the best linear transformation matching a nonlinear map locally, which is the bridge to the optimization parts of this series.

The two-dimensional case is the one you can draw, but the recipe does not care about dimension. In three dimensions a linear map is determined by where it sends the three basis vectors, and its matrix has three columns; in a thousand dimensions, by a thousand images. The picture stops being drawable but the rule “columns are the images of the basis” remains literally true, and it is the rule that makes matrix multiplication, determinants and eigenvalues meaningful in every dimension at once.

Once you trust that rule, a great deal can be read straight off the columns. Two columns that are scalar multiples of one another mean a direction has been lost, so the map crushes the plane onto a line and the determinant is zero. Two columns of equal length meeting at a right angle mean the map is a rotation or a reflection, and areas are preserved. A column of length greater than one means that direction is being stretched; less than one, compressed. None of these facts require any calculation beyond looking at the two arrows, and they are the same facts the numeric matrix reports. Learning to see them is the point of drawing the columns rather than the entries.

3

Typing a matrix is describing a transformation

The same object, entered as numbers

This demo runs the same map as the one above, but the controls are four number inputs instead of two handles. Change any entry and the transformed grid, the columns and the magenta ellipse all update together. The ellipse is the image of the unit circle, and it is a faithful fingerprint of the map: its axes tell you the directions the transformation stretches most and least. Typing 2 in the top-left cell widens the first column, which stretches everything in the x direction; typing in the off-diagonal cells shears the grid sideways. The gap between “a matrix” and “a transformation” closes the moment you notice there is only one object here with two notations.

This is also why the standard basis is a convenience and not a law. A matrix stores a map relative to the basis you happened to use when you wrote down its columns. Change the basis — the subject of a later part — and the same transformation earns a different matrix, while the arrows it moves are unchanged. The numbers are a description; the map is the thing described.

Edit the four entries. The columns of the matrix are the images of î and ĵ, drawn as arrows; the ellipse is the image of the unit circle.

Watch the determinant in the readout as you type. It is the signed area of the image of the unit square, and it changes continuously with the entries. Drive it to zero by making the two columns collinear and the ellipse degenerates into a line segment: the map has crushed the plane onto a lower-dimensional image, and the matrix is singular. The determinant is not a separate topic bolted on later; it is the area scaling you can read straight off this picture, and it is the next part's main character.

There is a second way to read the same multiplication, and it is the one textbooks usually teach first. Instead of combining columns, you take the dot product of the input with each row: the first output coordinate is the first row dotted with the input, and the second coordinate is the second row dotted with the input. The two views are equal because the arithmetic is associative, but they tell different stories. The row view is local — one number per output, a weighted average of the input's coordinates — and it is how a neural network layer is implemented in practice. The column view is global, and it explains what the transformation does to the space. Fluency means being able to switch between them without thinking.

5

Why the columns tell you everything

Multiplying is combining

Here is the whole multiplication rule, stripped of notation. Take the input vector's coordinates. Use the first coordinate as a weight on the first column and the second coordinate as a weight on the second column. Add the two weighted columns. That is the output, and it is drawn below as a tip-to-tail walk in the image space: first a step of a along column one, then a step of b along column two, landing on Mu. The sliders set a and b; the readout highlights those two weights, because they are the only things that vary. The columns are fixed by the transformation; the input chooses how much of each to use.

$$M\begin{bmatrix} a \\ b \end{bmatrix} = a\,\mathbf{col}_1 + b\,\mathbf{col}_2.$$

The unit square makes the same point from the other direction. Its four corners are (0, 0), (1, 0), (1, 1) and (0, 1), so their images are the origin, the first column, the sum of the columns, and the second column. The image of the square is therefore built entirely out of the two columns, which is why it is a parallelogram whose sides are the columns. A second test vector, drawn in the darker ink, shows the same mechanism with different weights.

Slide the weights a and b. The solid arrow is Mu, assembled tip-to-tail from the two columns.

This is the bridge to everything that follows. Matrix multiplication is composition of transformations, determinants measure how the map scales area, and eigenvector hunting is the search for inputs whose two weights are a single number. All of it lives inside the columns. When a neural network computes y = Wx, it is not doing anything more exotic than the tip-to-tail walk drawn above: each output coordinate is a combination of the input's coordinates, weighted by a row of W.

Try it on a concrete pair of weights. With the matrix shown and a = 1, b = 1, the output is the sum of the two columns, which is where the corner (1, 1) of the unit square lands. Halve one weight and the output slides halfway back along that column; set a weight to zero and the corresponding column drops out of the combination entirely. Negative weights run along a column backwards, which is why the image can extend past the origin even when the input does not. Every one of these behaviours is visible in the demo as the two dashed steps rearrange themselves before the solid arrow lands.

The same bookkeeping will let us multiply two matrices. If a second transformation N acts after M, then the composite map sends the basis vector î to N applied to the first column of M, and similarly for the second column. Those two results become the columns of a new matrix, and that matrix is the product NM. The order matters because the transformations are applied in order; running the walk above and then running another map is not the same as doing it the other way round. The next part makes that composition rule explicit and shows why the usual row-times-column formula is exactly this statement written out.

6

Where this shows up

One idea, two worlds

Robotics & vision

Cameras and image warps are matrices

A pinhole camera projects a 3D point to a 2D image through a linear map, and its intrinsic matrix is exactly the columns-are-images object drawn above: it turns coordinates in the camera frame into pixel coordinates. Rectifying a lens, aligning a stereo pair and warping an image are all matrices acting on homogeneous coordinates — the same picture, one dimension up. The pinhole camera part builds the intrinsics explicitly.

ML / AI

A linear layer is this transformation

The dense linear layer at the heart of a transformer is y = Wx, with W a learned matrix. Its columns are the directions the layer can send the basis vectors, and training adjusts those columns. Activations, attention projections and the value/output maps are all instances of the same operation, applied at a much larger scale. The architecture chapter places them inside the model; a later part of this series returns to attention as three matrices.

Further reading

7

Check your understanding

0/4 answered