Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

A shape that remembers everything

A matrix is a linear map, and a linear map is determined by very little: send the arrow (1, 0) somewhere, send (0, 1) somewhere, and everything else is fixed. So the fastest way to understand a matrix is to feed it a shape that contains every direction at once. The cleanest such shape is the unit circle. It is round, it points every way, and it is easy to draw. Whatever comes out is a picture of the map.

What comes out is always an ellipse (possibly squashed flat into a segment, if the matrix throws information away). The reason is geometric: a linear map takes straight lines to straight lines and the origin to the origin, so the round circle deforms smoothly into a closed convex curve with two perpendicular axes of symmetry. Those two axes are special, because along them the map does nothing but stretch. Every other direction gets both stretched and turned.

Compare this with eigenvalues, the subject of Part 13. Eigenvectors are the directions a matrix leaves un-turned, but they exist only for square matrices and even then may be complex or may collapse into a shortage of independent directions. The ellipse picture has no such caveats: it works for any shape of matrix, the axes are always real and always perpendicular, and the directions are always available. This is why the SVD is the workhorse where diagonalisation fails.

That gives us three objects to name. The two directions on the circle that get sent to the ellipse's axes are the right singular vectors, collected as the columns of V. The half-lengths of the resulting axes, in descending order, are the singular values σ₁ ≥ σ₂, collected on the diagonal of Σ. The axis directions on the ellipse are the left singular vectors, the columns of U. Put together, A = UΣVᵀ, and the geometry of the matrix is completely exposed.

💡 By the end of this part you'll see why every matrix can be read as a rotation, a stretch along perpendicular axes, and another rotation — and why the two stretch factors, the singular values, measure how much the matrix magnifies space in the two directions where it does not turn at all.
2

Circle in, ellipse out

The image of the unit circle, and the four arrows that describe it

The static equation is short. For each singular value, the matching circle direction lands on the matching ellipse direction, scaled by that singular value:

$$\mathbf{A}\mathbf{v}_1 = \sigma_1 \mathbf{u}_1, \qquad \mathbf{A}\mathbf{v}_2 = \sigma_2 \mathbf{u}_2, \qquad \sigma_1 \ge \sigma_2 \ge 0$$

Read the first identity carefully, because it is the entire idea. The unit vector v₁ sits on the circle. The matrix sends it to σ₁u₁, a point on the ellipse, and u₁ is the direction of the ellipse's long axis. The second identity does the same for the short axis. The two v arrows are perpendicular, the two u arrows are perpendicular, and the σ numbers are the only lengths involved. Nothing about A is left unexplained: it is a turn, a pair of stretches, and a turn.

Edit the matrix below and watch the ellipse change. The pale circle is the input, the curve is its image, and the four arrows are drawn exactly as the two identities describe them: the v arrows start on the circle, the σu arrows end on the ellipse. Drag the angle slider to carry a single direction around the circle and follow where it lands.

Unit circle in, ellipse out. v₁, v₂ are the circle directions that land on the axes; σ₁u₁, σ₂u₂ are the axes themselves. The slider walks one test vector around the circle without touching the matrix.

v circle directions σu ellipse axes test vector

Some outcomes are worth hunting for. Make the matrix symmetric and the v and u arrows line up: a symmetric matrix stretches along its own eigen-directions, so the two rotations cancel. Make the matrix a pure rotation and the ellipse is a circle, because a rotation stretches every direction by exactly one. Make one column a multiple of the other and the ellipse flattens, which is where the collapse section goes.

Two facts about the picture are worth stating plainly before any more algebra arrives. First, the ellipse is not an approximation or a coincidence: every linear map sends the unit circle to an exact ellipse, and its two axes are exactly the directions in which the map only stretches. Second, the decomposition is canonical but not always unique. The singular values are ordered by size and that ordering is fixed, but when the two singular values are equal the ellipse is a circle and any perpendicular pair of directions qualifies as the axes; when one singular value is zero, the direction paired with it is ambiguous too. Software never promises a particular sign for an arrow either, because an axis is a line, not a directed segment. Every comparison in this guide therefore treats v and −v as the same answer, which is the honest thing to do.

There is a third way to write the same fact, and it is the form the applications actually use. Instead of grouping the moves into matrices, add up one rank-one piece per singular value:

$$\mathbf{A} = \sigma_1 \mathbf{u}_1 \mathbf{v}_1^{\mathsf{T}} + \sigma_2 \mathbf{u}_2 \mathbf{v}_2^{\mathsf{T}} = \sum_i \sigma_i\, \mathbf{u}_i \mathbf{v}_i^{\mathsf{T}}$$

Each term is an outer product: a single direction on the circle paired with a single direction on the ellipse, scaled by one stretch factor. These pieces are as simple as matrices get — every column of the first term is a multiple of u₁, so it has rank one — and they add up to the full matrix. The singular values are the weights on those pieces, which is why the largest one dominates the picture and the smallest one can often be dropped without anyone noticing.

3

Rotate, stretch, rotate

A = U Σ Vᵀ, one move at a time

Factorise the picture itself. Because V and U are built from unit vectors at right angles, both are orthogonal matrices, and an orthogonal matrix is exactly a rigid rotation or reflection — it preserves lengths and angles. So the decomposition

$$\mathbf{A} = \mathbf{U}\,\boldsymbol{\Sigma}\,\mathbf{V}^{\mathsf{T}}, \qquad \mathbf{V}^{\mathsf{T}}\mathbf{V} = \mathbf{I}, \qquad \mathbf{U}^{\mathsf{T}}\mathbf{U} = \mathbf{I}$$

says that applying A is three easy moves in a row. Multiplications act on the right first, so a vector x is touched by Vᵀ first: a rigid rotation that lines the singular directions up with the coordinate axes. Then Σ stretches the first axis by σ₁ and the second by σ₂ — this is the step that turns the circle into an axis-aligned ellipse. Finally U rotates that ellipse to its true orientation. The two outer matrices can only turn things; all the deformation is locked inside the middle one.

Press the three buttons below to apply the moves to the shape. Start with the unit circle (the identity), and watch the grid carry along with it so you can see the transformation of the whole plane, not just the outline. The demo uses whichever A you set in the matrix above, so edit that first if you want a different windmill.

The three moves act on the right first: Vᵀ rotates, Σ stretches, U rotates. The faint circle is where the shape started.

When the last button is pressed the accumulated matrix is UΣVᵀ, which the readout checks against the original A to the last decimal. The order matters and it is not arbitrary: the axes come from V, the lengths come from Σ, and the final orientation comes from U. Swap the two rotations and you get a different matrix with the same singular values — a good reason the singular values, not the whole list of entries, are what people quote.

Nothing in the recipe is special to two dimensions, and that is the real payoff. A matrix with m rows and n columns still has an SVD: the unit sphere in n dimensions goes to an ellipsoid in m dimensions, whose axes are the min(m, n) singular values. In three dimensions the unit sphere becomes an ellipsoid with three semi-axes; in the high-dimensional spaces that show up in vision and machine learning, it becomes an ellipsoid with hundreds or thousands of axes described by a list of singular values. You cannot draw it, but you can sort it, and the sorted list is the matrix's size profile.

Right-to-left, always. Matrix multiplication applies the rightmost factor first. In UΣVᵀ, the vector is rotated by Vᵀ, then stretched by Σ, then rotated by U. Reading the product left-to-right and imagining U acts first is the single most common way to get the picture backwards. The buttons in the demo fire in the order the vector experiences them, which is why the sequence is labelled Vᵀ, then Σ, then U.

One more reading helps fix the geometry. The matrices U and V can only rotate or reflect, so they cannot change a length at all; they merely choose which directions are called "the axes". All of the stretching lives in the diagonal Σ. If you ever want to know how much a matrix can magnify a vector, you are asking for the largest diagonal entry of Σ, and the answer is σ₁. That quantity has a name, the spectral norm, and it is the reason singular values rather than eigenvalues are the right notion of a matrix's size.

4

When the ellipse collapses

Rank is the number of surviving axes

Push the matrix toward singularity and one of the singular values heads for zero. The corresponding axis of the ellipse shrinks until the closed curve becomes a flat segment traced out and back along a single direction. Two things happen at once, and they are the same thing. Geometrically, the image loses a dimension: the unit circle no longer covers an area, only a line. Algebraically, the columns of the matrix become parallel, so the matrix has rank one, its determinant is zero, and it can no longer be inverted because two different inputs now land on the same point.

The number of nonzero singular values is the rank. That is the cleanest definition of rank there is, and it is the one the numerics actually use: count the singular values above a tolerance. Before the collapse, rank two and a fat ellipse; after it, rank one and a segment; if both singular values vanish, rank zero and everything lands at the origin.

$$\operatorname{rank}(\mathbf{A}) = \#\{\,i : \sigma_i > 0 \,\}, \qquad \mathbf{A}\mathbf{v}_2 = \mathbf{0} \ \text{ when } \ \sigma_2 = 0$$

That second identity is the null space in disguise. The vector v₂ that gets squashed to nothing is a direction the matrix cannot see at all, and it is a genuine, unit-length vector — not a rounding artefact. If you wanted to recover x from Ax you would have to divide by σ₂, and dividing by zero is precisely the failure. Part 7 built the null space from row operations; the SVD hands it to you as a direction, which is a much more satisfying answer.

In floating point the zero is rarely exact, so "nonzero" means "above a tolerance", and choosing that tolerance is a judgement call that the readout makes for you here. A singular value of 10⁻¹⁶ relative to the largest is noise; the same value relative to a matrix whose largest singular value is 10⁻¹⁰ is signal. This is why serious software lets you pass a cutoff rather than pretending the question has one answer.

Drag the collapse slider and watch the second singular value die. The readout tracks the rank, the determinant, the condition number, and the angle between the two columns. Long before the columns are exactly parallel they are already almost parallel, and that is the situation that makes numerical work dangerous.

The matrix interpolates from a healthy one to a rank-one one as you drag. Solid arrows are the singular-value axes; the two faint arrows are the columns of A, which swing toward each other.

The ratio κ = σ₁/σ₂ is the condition number. While the ellipse is roundish, κ is close to one and the matrix treats all directions about equally. As the ellipse flattens, κ climbs without bound: at the far end of the slider one direction is magnified by a healthy factor and the other is nearly annihilated. A linear solve Ax = b divides by those singular values, so a small error in b along the tiny axis gets amplified by κ in x. Part 19 takes that amplification apart; here it is enough to notice that the geometry of the ellipse predicts the trouble.

5

The symmetric matrix hiding inside

AᵀA has the singular values squared

The right singular vectors can be found without ever mentioning the SVD. Multiply A = UΣVᵀ by its transpose and almost everything cancels, because U is orthogonal:

$$\mathbf{A}^{\mathsf{T}}\mathbf{A} = \mathbf{V}\,\boldsymbol{\Sigma}^{2}\,\mathbf{V}^{\mathsf{T}}, \qquad \mathbf{A}^{\mathsf{T}}\mathbf{A} = (\mathbf{A}^{\mathsf{T}}\mathbf{A})^{\mathsf{T}}$$

The matrix AᵀA is symmetric — its transpose is itself, by construction — and the displayed identity is precisely the eigen-decomposition of a symmetric matrix from Part 15. Its eigen-directions are the columns of V, the same right singular vectors, and its eigenvalues are the squares of the singular values, σ₁² and σ₂², in the same order. So the ellipse's axis directions are the eigen-directions of AᵀA, and its axis half-lengths are the square roots of its eigenvalues.

That is how many people first compute an SVD, and the demo below verifies it live. Move the entries of A and compare the two answers: the eigenvectors of AᵀA produced by an eigen-routine against the right singular vectors produced by the SVD, and the eigenvalues against the squared singular values. They should agree every time, up to the overall sign of an eigenvector, which is never meaningful (an eigen-direction is a line, not an arrow).

v arrows are the right singular vectors, dashed lines are the eigen-directions of AᵀA. They should sit on top of each other whatever you choose.

The identity also explains the phrase "size of a matrix". The squared Frobenius norm — the sum of the squares of every entry, the total energy of the matrix — is the sum of the squared singular values, because the rotations U and V do not change any length. So the singular values are a complete budget of how much the matrix does, split into independent channels. Drop the smallest channel and you have accounted for only a small slice of the energy, which is exactly the bargain Part 17 strikes.

$$\|\mathbf{A}\|_F^2 = \sum_{i,j} a_{ij}^2 = \sigma_1^2 + \sigma_2^2, \qquad \|\mathbf{A}\|_2 = \sigma_1$$

One warning, since it matters in practice. Forming AᵀA squares the condition number: if κ = 10⁶, then AᵀA has condition number 10¹² and the small singular value can drown in rounding error. The SVD routines used by real software work directly on A and never square anything. The identity above is the right way to understand the singular vectors; it is not always the right way to compute them.

6

Where this shows up

One decomposition, two worlds

The ellipse is not a teaching device that gets retired once the chapter ends. It is the shape that shows up, unnamed, inside the failures and the shortcuts of real systems. Here are two places where the singular values are the thing being measured, one in robotics and one in machine learning.

Robotics

A near-singular Jacobian

A robot's Jacobian maps joint velocities to end-effector velocity, and its singular values say how much motion each direction buys you. When σ₂ shrinks toward zero the arm is near a singular configuration: the ellipse collapses, one Cartesian direction becomes almost unreachable no matter how the joints move, and the inverse step demands huge joint speeds to make up the difference. Pose-graph solvers feel the same effect in their normal equations, where a flat ellipse in the information matrix is the signature of an under-constrained gauge — a direction the data simply does not pin down. The pose-graph part of the optimization guide is where this becomes a practical headache.

ML / AI

Low-rank weight matrices and LoRA

A weight matrix in a transformer has an SVD, and a striking number of them have a spectrum that decays fast: a few large singular values carry most of the action and the rest are small. That is what low-rank approximation exploits — keep the biggest singular values and the matching directions, throw the rest away, and you have most of the map for a fraction of the storage. Fine-tuning methods such as LoRA freeze the full matrix and learn only a small low-rank correction, which is the same rank-versus-energy trade promoted into a training trick. The scaling chapter of the LLM-training guide picks up the story.

7

Further reading

8

Check your understanding

0/4 answered