Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

A matrix you can rotate to face you

In Part 13 and Part 14 we asked when a matrix has special directions that only get stretched. The answer was often messy: a generic matrix can rotate its eigenvectors off the axes, and its eigenvalues can come in complex pairs that describe rotation rather than stretch. A symmetric matrix removes both problems at once. If a matrix equals its own transpose — the entry in row one, column two equals the entry in row two, column one — then its story is forced to be simple.

The symmetry has a geometric meaning too. It means the matrix acts the same way along every pair of opposite directions: measuring the output along x after feeding it y gives the same number as measuring along y after feeding it x. That is a kind of balance, and balance always buys structure. The structure is the spectral theorem: a real symmetric matrix can be written as a rotation, a diagonal scaling, and the rotation back.

That last sentence is worth pausing on, because it is the difference between a definition and a promise. An arbitrary matrix may shear, rotate complexly, or refuse to settle into eigenvectors at all. A symmetric one always behaves like a set of independent springs along perpendicular rails. The springs may pull in opposite directions — one stiff, one floppy, one pulling back — but each acts alone, and you can always find the rails by turning your head. Once you know the rails, every question about the matrix becomes a question about the springs.

💡 By the end of this part you'll see why a symmetric matrix is the one matrix you can always rotate to a diagonal form, why its eigenvectors arrive already perpendicular, and why the shape of xᵀAx — ellipse, saddle, or a flat degenerate direction — is decided purely by the signs of its eigenvalues.
2

The signature of a quadratic form

Level sets you can read off the eigenvalues

Multiply a symmetric matrix A by a vector on both sides and you get a single number: the quadratic form xᵀAx. In two dimensions, writing the matrix with entries a, b, d and using the fact that the two off-diagonal entries are equal, the form expands to a plain polynomial in the two coordinates:

$$\mathbf{A} = \begin{bmatrix} a & b \\ b & d \end{bmatrix}, \qquad x^{\mathsf{T}}\mathbf{A}x = a\,x^2 + 2b\,xy + d\,y^2$$

Because xᵀAx is a number attached to each direction, we can ask for its level sets: the points where the form equals a fixed constant. Those level sets are conic sections, and the symmetry decides which kind. Diagonalise the matrix and the polynomial turns into a weighted sum of squares, λ₁u² + λ₂v², in the rotated coordinates. Every level set is then easy to classify from the two eigenvalues alone:

Why must the eigenvalues be real? For a two-by-two matrix you can see it directly. The characteristic polynomial is λ² − (a + d)λ + (ad − b²), and its discriminant works out to (a − d)² + 4b², a sum of squares. A sum of squares cannot be negative, so the quadratic has two real roots — end of story. The same argument generalises: a symmetric matrix's eigenvalues are always real because symmetry forces the quadratic form xᵀAx to be a real number for every real input, and a complex eigenvalue would require a direction where the form is complex. The perpendicularity of the eigenvectors follows from the same balance: if two eigenvectors had different eigenvalues, their inner product would have to equal both itself and its own negative, which only works when it is zero.

The demo below lets you choose the three entries and watch all three regimes appear. It reads the eigenvalues with a numerical routine, shades the level sets of the form, and draws the eigenvector axes so you can see how the conic lines up with them. Push b past the point where it overcomes the diagonal entries and the ellipse splits into a saddle without any warning beyond the sign flip of one eigenvalue.

Level sets of xᵀAx for the symmetric matrix on the right. Solid arrows are the scaled eigenvectors; dashed lines are the eigen-directions.

Two of the eigenvalues are ordered for you: the routine always reports the larger one first. That ordering is not cosmetic. It tells you the direction in which the form grows fastest, which is exactly the direction a minimiser of the form is least keen to move. If you picture a ball rolling in the bowl, it wants to settle at the bottom, and the curvature it feels depends on which way it is nudged; the largest eigenvalue is the stiffest direction and the smallest is the softest. When the smallest eigenvalue is only barely positive, the bowl is very flat along that direction, and a numerical method will wander across the floor instead of committing to the minimum. That is the practical face of a nearly-degenerate quadratic form.

3

Rotating it to diagonal

QᵀAQ = Λ, and Q is a rotation

The spectral theorem is the promise that the friendly picture is not wishful thinking. Write the eigenvectors as the columns of a matrix Q and the eigenvalues as the diagonal entries of a matrix Λ. Because the eigenvectors of a symmetric matrix are perpendicular, Q is an orthogonal matrix: its columns are unit vectors at right angles, so QᵀQ is the identity and Qᵀ is the inverse of Q. The theorem then says:

$$\mathbf{A} = \mathbf{Q}\,\boldsymbol{\Lambda}\,\mathbf{Q}^{\mathsf{T}}, \qquad \mathbf{Q}^{\mathsf{T}}\mathbf{A}\mathbf{Q} = \boldsymbol{\Lambda}, \qquad \mathbf{Q}^{\mathsf{T}}\mathbf{Q} = \mathbf{I}$$

Read the first equation as a recipe: rotate by Qᵀ, stretch along the coordinate axes by the eigenvalues, then rotate back. The middle equation is that recipe reversed — stand in the eigenbasis and the matrix is diagonal. Nothing about the matrix's action changed; only the frame you describe it in.

The demo makes the frame itself move. A grid starts aligned with the ordinary axes and turns until its lines run along the eigen-directions. At the end of the turn the rotated grid is exactly the eigenbasis, and the matrix expressed in that grid has zeros everywhere except the diagonal. The ellipse is the image of the unit circle under A; its principal axes are the eigenvectors scaled by the eigenvalues, which is why the solid arrows always touch the ellipse where it is widest or narrowest.

Slide frame to turn the grid from the standard axes to the eigenbasis Q. The ellipse stays put; only the coordinates change.

Do not miss the check in the readout. The product QᵀQ has ones on its diagonal and zeros off it, and QᵀAQ has the eigenvalues on its diagonal and nothing else. Those are not approximations that happen to look tidy; they are the two claims of the theorem, recomputed every time you move a slider.

It helps to say out loud what has not changed. The ellipse is the same set of points before and after the frame turns; the vectors are the same vectors. All that moves is the ruler you measure them with. In the standard frame the matrix looks dense and its action couples the two coordinates. In the eigenframe the same action decouples completely: the first coordinate is multiplied by λ₁, the second by λ₂, and never the twain shall meet. This is the whole value of the decomposition. A coupled problem becomes two independent one-dimensional problems the moment you agree to describe it in the right basis, and because the basis is orthogonal, switching into it costs nothing more than a rotation.

4

The form on the unit circle

Eigenvalues are the extremes of xᵀAx

There is a second way to see the eigenvalues that needs no matrix diagonalisation at all. Restrict the input to the unit circle, so x and y are the cosine and sine of an angle, and evaluate the form as the angle sweeps around. The result is a smooth curve whose radius in each direction is the value of xᵀAx for that direction. This curve is the Rayleigh quotient of the unit circle, and its highest and lowest points are the eigenvalues.

$$f(\theta) = (\cos\theta,\ \sin\theta)^{\mathsf{T}}\mathbf{A}\,(\cos\theta,\ \sin\theta), \qquad \max_\theta f = \lambda_1, \qquad \min_\theta f = \lambda_2$$

The reason is the same diagonalisation, read as a weighted average. In the eigenbasis the unit circle stays a circle, and f becomes λ₁u² + λ₂v² with u² + v² = 1. A weighted average of two numbers can never beat the larger or fall below the smaller, and it hits the larger exactly when the input points along the first eigenvector and the smaller when it points along the second. That is why the maximum and minimum below always land on the dashed eigen-directions and always agree with the eigenvalues in the readout.

The thin circle is the unit inputs; the coloured curve plots f(θ) as a radius. The marked points sit on the eigen-directions.

When the matrix is indefinite, the curve dips to a negative radius, which draws the point on the opposite side of the origin: the form is negative there, so the surface is below the floor. The maximum stays positive and the minimum stays negative, and the gap between them — the spread of the curve — is what makes the saddle steep. When the two eigenvalues are equal the curve is a perfect circle, the matrix is a multiple of the identity, and every direction is an eigen-direction. That is the only case where the eigenvectors are ambiguous, and the numerics will happily hand you any perpendicular pair.

This circle view also gives you a test you can run in your head. For a two-by-two symmetric matrix, check the top entry a and the determinant ad − b². If both are positive, the eigenvalues are both positive and the form is a bowl. If the determinant is positive but a is negative, both eigenvalues are negative and the form is a dome. If the determinant is negative, the eigenvalues have opposite signs and the form is a saddle. A determinant of zero is the degenerate case, where the curve passes through the origin and one direction of the plane costs nothing. That is the same Sylvester test the finite-element and robotics codes apply, just written in two dimensions where you can see why it works.

5

Where this shows up

One object, two worlds

Robotics

Hessians decide whether a step is safe

A robot's cost function usually looks like a bowl near its minimum, and the bowl's shape is a symmetric matrix: the Hessian, whose entries are the second derivatives. If that Hessian is positive definite, the local model curves upward in every direction and a gradient step decreases the cost. If it has a negative eigenvalue, the model curves downward along that direction and the optimiser can slide away. The gradient descent part of the optimization guide spends its time checking exactly this signature.

ML / AI

Covariances and attention scores

A covariance matrix is symmetric by construction, so it always has real eigenvalues and a perpendicular eigenbasis; those eigenvectors are the principal axes of the data cloud, and the eigenvalues are the variance along them. Attention is often described with a symmetric score matrix too, where the symmetry means token i weighing token j is the same relationship read in reverse. The architecture chapter of the LLM-training guide uses both objects when it explains how a transformer mixes context.

6

Cheat sheet

What to check, and what it buys you

ObjectWhat to checkWhat it means
SymmetryA = Aᵀ, so the two off-diagonal entries agreeReal eigenvalues and perpendicular eigenvectors, guaranteed
EigenvaluesBoth positivePositive definite: level sets are ellipses, the surface is a bowl
EigenvaluesOne positive, one negativeIndefinite: level sets are hyperbolas, the surface is a saddle
EigenvaluesEither one is zeroSemidefinite: a flat direction, degenerate level sets
Basis QQᵀQ = IThe eigenvectors are orthonormal; Qᵀ is the inverse rotation
Diagonal formQᵀAQ = ΛIn the eigenframe the matrix is just the eigenvalues on the diagonal
Quadratic formλ₁u² + λ₂v² with u² + v² = 1A weighted average, so its extremes are exactly λ₁ and λ₂
Ellipse axesImage of the unit circle under APrincipal axes along the eigenvectors, lengths |λ₁| and |λ₂|

If you remember one line, remember the middle of the table: the signs of the eigenvalues are the entire classification. Magnitudes tell you how steep the bowl or saddle is, and the eigenvectors tell you which way it leans, but the signs decide which of the three pictures you are looking at. Every algorithm in the next part that hunts for a minimum is, underneath, testing those signs.

7

Further reading

8

Check your understanding

0/4 answered