Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

Two moves, one matrix

A linear transformation takes a vector and returns a vector. Do two of them in a row and you have a new rule: feed a vector to the first, take its output, feed that to the second. The question is whether this two-step rule is itself linear — whether it can be written as a single matrix — and if so, which matrix.

It is linear, because both steps are: scaling the input scales each step, and adding inputs adds each step, so the combined rule preserves both operations. A linear rule is determined by where it sends the basis vectors, so the combined rule has a matrix, and we are free to define the product of two matrices to be exactly that matrix. Matrix multiplication is not an arbitrary arithmetic ritual. It is composition, written down in coordinates.

One convention is worth fixing now, because everything downstream depends on it. When you see a product written as AB, the rightmost factor acts first. This matches ordinary function notation, where (f\circ g)(x) = f(g(x)) applies g first, and it means the letters read from last-applied to first-applied. Read a matrix equation right to left and the geometry comes out in the correct sequence; read it left to right and you will compose the maps backwards.

💡 By the end of this part you'll see why applying B and then A is a single linear map whose matrix is AB, why the order of the factors changes the picture so that AB is usually not BA, and why a chain of transformations can be grouped however you like — the property called associativity.
2

Doing B then A is a matrix

The product is defined to make it true

Fix two maps, B first and A second. A vector x goes to Bx, and that goes to A(Bx). If the combined map is AB, then for every x

$$(AB)\,x = A\,(Bx)$$

This is the whole definition. The product AB is the matrix that applies B first and A second, so the factors are written in the opposite order to the words: the matrix nearest the vector acts first. Read the equation right to left and it says exactly what you do — B touches x before A does.

The formula you were taught is a consequence of this definition, not a separate fact. Read the columns of B one at a time. Column j of B is where B sends the j-th basis vector, and A then sends that vector on to A times it. So column j of the product is just A applied to column j of B:

$$AB = A\begin{bmatrix} \vert & & \vert \\ b_1 & \cdots & b_n \\ \vert & & \vert \end{bmatrix} = \begin{bmatrix} \vert & & \vert \\ Ab_1 & \cdots & Ab_n \\ \vert & & \vert \end{bmatrix}$$

For a pair of two-by-two matrices this expands to the familiar pattern, each entry a row of A against a column of B:

$$\begin{bmatrix} a_{11} & a_{12} \\ a_{21} & a_{22} \end{bmatrix}\begin{bmatrix} b_{11} & b_{12} \\ b_{21} & b_{22} \end{bmatrix} = \begin{bmatrix} a_{11}b_{11}+a_{12}b_{21} & a_{11}b_{12}+a_{12}b_{22} \\ a_{21}b_{11}+a_{22}b_{21} & a_{21}b_{12}+a_{22}b_{22} \end{bmatrix}$$

The rows-against-columns rule is exactly "apply A to each column of B and stack the results". Nothing about it is arbitrary; it is composition written in coordinates. That is also why a matrix times a column vector is just a special case: a vector is a matrix with one column, and Ax is the composition of A with the map that takes the single number 1 to x.

The demo below makes the definition visible. The pale grid is the plane after B, the strong grid is the plane after B and then A. The pink arrow tracks a vector x through both steps; the matrix AB is computed live from your two matrices, and multiplying it by x lands in the same place as doing the two moves in sequence.

Drag the round handle to move x. The faint grid is the plane after B; the strong grid is the plane after B then A.

Two things to notice. First, the strong grid is not the pale grid moved by A on its own; it is the pale grid moved by A, which is a different set of lines because B already bent the plane. Second, the arithmetic in the readout never has to be believed on faith: the coordinates of A(Bx) and of (AB)x agree because the product was built to make them agree.

If the product is defined by composing maps, it can only be formed when the shapes line up: the map A must accept vectors of the dimension that B produces. For two-by-two matrices that is automatic, but the same rule explains why an m\times n matrix can multiply an n\times p matrix and why the middle dimensions must match. Shape compatibility is not an arbitrary typing rule; it is the statement that the output of one map is a legal input to the next.

Checking a product by hand is easier by columns. To get the first column of AB, apply A to the first column of B alone; to get the second, apply A to the second column. Each column is a small matrix-vector multiplication you already know how to do, and the full product is those results placed side by side. The row-times-column algorithm is what you get when you compute those same columns entry by entry.

3

Order matters

AB is almost never BA

Numbers multiply in any order: 3×5 = 5×3. Transformations do not. Rotating a picture and then shearing it leaves different lines from shearing it and then rotating it, and the two products AB and BA are the coordinate versions of those two different pictures. Composition is not commutative, and neither is the product that encodes it.

Flip between the two orders below. The grid for the order you selected is drawn strong; the grid for the other order is left faint underneath so you can see they are genuinely different shapes, not the same shape described twice. The readout prints both products so you can compare them entry by entry.

Order is not a technicality. Put on a sock and then a shoe and you get a wearable result; reverse the order and you get a nonsense one. Rotations behave the same way — a quarter turn about one axis followed by a quarter turn about another does not equal the reverse sequence — which is why a robot's end-effector pose depends on the exact order of its joints, and why the exponential map of Part 20 has to be careful about how it composes rotations.

You can see the failure of commutativity in the simplest possible formulas. A rotation by angle \theta in the plane is the matrix with cosine and sine in a fixed pattern, and a shear along the horizontal axis has ones on its diagonal and a single off-diagonal entry. Multiplying them in the two orders produces different off-diagonal entries, so the two products genuinely differ. The pictures differ too: one order shears the already-rotated axes, the other rotates the already-sheared axes, and a sheared square is not the same shape as a rotated one.

Scalars are the one family that escapes the rule. Multiplying by a number stretches every direction equally, so a uniform scale can be applied before or after any other map without changing the result; a scalar matrix commutes with everything. That exception explains why scalar factors can be pulled out of a product at will, k(AB) = (kA)B = A(kB), and why scaling is so much tamer than the matrix multiplication it rides inside.

The faint grid is the order you did not pick. Where the two grids disagree is where order matters.

⚠ Do not read AB = BA into the picture. The only everyday cases where they agree are special: when one factor is the identity, when one is a plain scaling, or when the two happen to share the same eigenvectors. For generic matrices the two products differ, and a robot arm or a stack of network layers depends on that difference.
4

Composition as a path

First B, then A, in one sweep

It helps to watch the plane travel rather than jump. Start at the identity — nothing has happened. Move the slider: for the first half the plane morphs from the identity into the state after B alone; for the second half it morphs from B's state into the state after B and then A. The midpoint of the slider is B, and the endpoint is the product AB.

That is what composition means as a process: a sequence of two maps is a path through two intermediate pictures, and the product is the label on the far end. The pale grids mark the two landmarks — B halfway, AB at the end — so you can always see which stage the moving grid is between.

The straight-line interpolation between matrices is a choice made for the demo, not part of the mathematics; any continuous path from the identity to B and on to AB tells the same story. What matters is the reading order. The first half of the slider is the statement "B has not finished being applied"; the second half is "B is done and A is now at work". At every instant the moving grid is a legal linear map of the plane, and the slider is walking through a family of them from the do-nothing map to the product.

Once you are comfortable reading a product as a sequence, a power like A^2 stops looking like an algebraic oddity and becomes a description of a process: apply A, then apply it again. Repeating A over and over is how a discrete dynamical system moves forward in time, and asking what happens after many repetitions — does the process settle, cycle, or blow up — is a question about the eigenvalues of A. The next act of this series is built entirely from iterating the composition you are looking at here.

One subtlety about the slider: it shows matrices, not a single moving point, so what travels is the whole grid. When the moving grid crosses B at the halfway mark, every point of the plane is at its post-B location; when it reaches the end, every point is at its post-A location. Watching whole grids rather than single arrows is the habit this series wants to build, because a transformation is defined by what it does to everything at once.

Slider at 0 is the identity, at 0.5 the plane has been moved by B, at 1 by B then A.

5

Chains of any length

Associativity is grouping, not reordering

Three maps in a row, C then B then A, can be combined in two ways. Group the first two and then apply A, A(BC), or apply C first and then combine B with A, (AB)C. Both describe the identical sequence of moves — nothing has been reordered, only the brackets have moved — so both give the identical result. That is associativity:

$$(AB)\,C = A\,(BC)$$

It matters because it lets you compute a long chain in whatever grouping is cheapest, without changing the answer. Edit the three matrices below and watch both groupings agree to the last digit; the difference shown is floating-point noise at worst.

The same three maps applied in the same order, drawn once through each grouping. They land on the same grid.

Associativity is why matrix multiplication scales to deep stacks: a hundred layers is a hundred factors, and you may fold them in any order while you build up the computation. It is also why the notation ABCD needs no brackets at all.

Two related laws round out the algebra of composition. There is an identity: the do-nothing map has the identity matrix, and composing with it changes nothing, so IA = AI = A. And composition distributes over addition, A(B + C) = AB + AC, because adding two maps and then applying a third is the same as applying the third to each and adding the results. Notice which law is missing. There is no AB = BA in general, and there is no division, only an inverse when the determinant permits one.

Associativity has a very practical consequence for computation. If you need the product of a long chain of matrices, the order in which you multiply pairs does not change the answer, only the cost. Multiplying a chain of matrices of mismatched shapes can be done cheaply or expensively depending on how you bracket it, and because associativity guarantees the result is the same, a program is free to choose the cheap bracketing. This is the classic matrix-chain problem, and it is a pure consequence of the fact that composition is associative.

There is a limit to what associativity buys you. It lets you regroup, but it never lets you reorder. The sequence of maps is fixed by the problem, and swapping two neighbours in the chain changes the result unless those two maps happen to commute. So a product of many factors has exactly one meaning — the composed map — and the brackets are pure bookkeeping that the mathematics lets you ignore and the implementation is free to choose.

Composition and inversion are two sides of the same coin. If a chain called C, then B, then A can be undone, its inverse undoes each step from last to first, so (ABC)^{-1} = C^{-1}B^{-1}A^{-1}. Undoing a sequence means walking it backwards, and the reversal is not optional: the same non-commutativity that makes AB differ from BA is what forces the order to flip when you invert. A path that goes out through several transformations is retraced in reverse.

6

Where this shows up

One idea, two worlds

Robotics

Kinematic chains

A robot arm is a chain of joints, and each joint contributes a transform from one link's frame to the next. The pose of the end-effector is the product of those transforms, in order from base to tip, so a single matrix captures where the tool is. The 3D rotations part builds the rotation factors, and the same chain of transforms is what a motion planner composes. Because the factors do not commute, swapping two joints' transforms describes a different arm.

ML / AI

Stacked layers

A deep network is a chain of linear maps punctuated by nonlinearities, and each layer's weights are a matrix multiplying the activations from the layer before. Associativity is what lets a compiler fuse adjacent linear layers into one, and the architecture chapter shows the attention and feed-forward blocks where those products live. Training then adjusts each factor by a gradient, which is calculus applied on top of the composition.

Composition is the thread that ties the rest of the guide together. A change of basis is a composition, P^{-1}AP, and the whole question of eigenvectors is the question of which directions a composition leaves alone. The singular value decomposition describes a matrix as a rotation composed with a stretch composed with another rotation. Once you can read a product as a sequence of moves, each of those decompositions becomes a sentence about a journey rather than a formula to memorise.

This is also why a matrix factorization is worth doing at all. When a computation is awkward in one coordinate system, you compose a change into a friendlier one, do the easy version there, and compose the change back. The pattern is always the same — insert a pair of inverse changes, let them cancel around the hard part — and the second half of this guide is largely a catalogue of good choices for what to insert.

Further reading

The sources below all approach composition from the same geometric starting point. Watch the 3Blue1Brown chapter first if the rows-against-columns rule still feels arbitrary; the point of this part is that it never had to.

Strang's Lecture 3 develops the product once from columns and again from elimination, and the two derivations meet in the middle. Immersive Math lets you drag the factors and watch the grid move, which is the quickest way to feel the non-commutativity. Axler proves the composition rule straight from the definition of a linear map, which is the cleanest logical route if you want to see that nothing was assumed.

Cheat sheet

StatementGeometric reading
(AB)x = A(Bx)B acts first, A second; the product encodes the sequence
AB = BAFalse in general; true only in special cases such as identity or shared eigenvectors
(AB)C = A(BC)Grouping is free; the order of the maps is fixed
IA = AI = AComposing with the do-nothing map changes nothing
A(B + C) = AB + ACApply a map to a sum by applying it to each part
column j of ABColumn j of B, then acted on by A
7

Check your understanding

0/4 answered