Glossary and identities to know
This is the series' back matter: the notation fixed in Part 1, a filterable index that links every term to the part that introduces it, and a cheat sheet of the decompositions and identities the later acts lean on. When a symbol in a page is unfamiliar, find it here first.
Notation, fixed once
The conventions every part follows
Vectors are columns and are written bold: v. The transpose vᵀ is the same numbers in a row, so a column vector is an n×1 matrix. The i-th coordinate is vi. Matrices are capital letters, A, with entry Aij in row i, column j. In code a matrix is an array of rows, matching the toolkit. The identity is I, the zero matrix 0, the inverse A⁻¹, the transpose Aᵀ, the norm ‖v‖, eigenvalues λ, singular values σ, and the pseudoinverse A⁺.
Sizes
A: m×n · x: n×1 · Ax: m×1
J: m×n · JᵀJ: n×n · σ: min(m,n)
Reading order
Parts build on each other, but each stands alone. Acts I–III are the machinery; Act IV is where it shows up. Every core part ends with a "Where this shows up" panel linking into the site's other guides.
Cheat sheet
Decompositions, and what each one is for
| Decomposition | Form | What it is for | Leading cost | Part |
|---|---|---|---|---|
| LU | PA = LU | Solving Ax=b for a square, non-singular A; reusable across many right-hand sides. | 2n³/3 | 6 |
| QR | A = QR | Least squares and stable solving; Q orthonormal, R upper triangular. | 2mn² − 2n³/3 | 9–10 |
| Cholesky | A = LLᵀ | Symmetric positive-definite systems; half the work of LU and always stable. | n³/3 | 15 |
| Eigendecomposition | A = QΛQ⁻¹, or QΛQᵀ if symmetric | Dynamics, stability, quadratic forms; diagonalises the map. | iterative, ~10n³ | 14 |
| SVD | A = UΣVᵀ | Rank, conditioning, best low-rank approximation, pseudoinverse — the complete picture. | iterative, ~20n³ | 16–18 |
Identities worth keeping
The SVD and the pseudoinverse it produces. Σ⁺ inverts the nonzero singular values and zeroes the rest.
For a 2×2 the eigenvalues are two numbers that account for the determinant and the trace; the discriminant decides real or complex.
Rank–nullity counts the dimensions a map preserves and destroys; the condition number measures how close σmin is to zero.
Orthonormal columns; the cross product as a skew matrix; a rotation as the exponential of a skew matrix.
Glossary
Every term, linked to the part that introduces it
| Term | What it means | Introduced in |
|---|---|---|
| A | ||
| Angle | The angle θ between two vectors satisfies cos θ = u·v / (‖u‖‖v‖); the dot product in disguise. | Part 8 |
| Attention | Token similarity computed as QKᵀ, normalised by softmax, and used to mix the value matrix V. | Part 22 |
| B | ||
| Basis | A linearly independent set that spans the space; coordinates in it are unique. | Part 2 |
| C | ||
| Characteristic polynomial | det(A − λI) = 0; its roots are the eigenvalues. | Part 13 |
| Column | One column of A is the image of a basis vector under the map; the matrix is its columns side by side. | Part 3 |
| Column space | Every vector Ax can reach; the range of the map. | Part 7 |
| Composition | Applying B then A is the single map AB. | Part 4 |
| Condition number | κ(A) = σmax/σmin; the factor by which a relative error in b can be amplified in x. | Part 19 |
| Cosine similarity | u·v divided by the product of norms; a scale-free measure of alignment. | Part 8 |
| Covariance matrix | A symmetric matrix whose entries are feature variances along the diagonal and pairwise covariances off it. | Part 17 |
| Cross product | a×b, perpendicular to both, with length equal to the parallelogram's area. | Part 11 |
| D | ||
| Determinant | The signed factor by which A scales area (2D) or volume (3D); zero exactly when A is singular. | Part 5 |
| Diagonalization | Writing A = PDP⁻¹ with D diagonal; possible when there are enough independent eigenvectors. | Part 14 |
| Dimension | The number of vectors in any basis of the space. | Part 2 |
| Dot product | The sum of entrywise products; measures projection and angle. | Part 8 |
| E | ||
| Eigenvalue | The scalar λ with Av = λv; the stretch factor along an eigen-direction. | Part 13 |
| Eigenvector | A direction that A only stretches or flips, never turns. | Part 13 |
| Exponential map | exp sends a skew matrix θ[û]× to the rotation of angle θ about û. | Part 20 |
| Outer product | uvᵀ; a rank-1 matrix, the building block of low-rank approximations. | Part 7 |
| G | ||
| Gauss–Newton | Iterating the linear solve JᵀJ δ = −Jᵀr to fit a nonlinear least-squares model. | Part 21 |
| Gaussian elimination | Row operations that turn A triangular while preserving the solution set. | Part 6 |
| Gimbal lock | The degeneracy of Euler angles when two rotation axes align, losing a degree of freedom. | Part 20 |
| Gram–Schmidt | Turning independent vectors into an orthonormal set by subtracting successive projections. | Part 9 |
| H | ||
| Hessian | The matrix of second derivatives of a function; JᵀJ approximates it for least-squares costs. | Part 21 |
| I | ||
| Identity | The matrix I that leaves every vector unchanged; the multiplicative unit. | Part 3 |
| Linear independence | No vector in the set is a linear combination of the others. | Part 2 |
| Inverse | A⁻¹ undoes A; it exists exactly when det A ≠ 0. | Part 5 |
| J | ||
| Jacobian | The matrix of first partial derivatives; one row per residual, one column per parameter. | Part 21 |
| K | ||
| KV cache | The stored key and value matrices a decoder reuses at every generation step. | Part 22 |
| L | ||
| Least squares | The best fit when Ax=b has no exact solution; minimise ‖Ax−b‖². | Part 10 |
| Least-norm solution | Among infinitely many solutions, the one with the smallest norm, produced by A⁺b. | Part 18 |
| Left null space | The solutions of Aᵀy = 0; the part of the output space the map never reaches. | Part 7 |
| Linear combination | A sum of scaled vectors, a₁v₁ + … + akvk. | Part 1 |
| Linear transformation | A map that respects addition and scaling; equivalently, a matrix. | Part 3 |
| LoRA | Fine-tuning a frozen weight by adding a low-rank product BA. | Part 22 |
| Low-rank update | Adding a rank-r product to a matrix; cheap to store when r is small. | Part 22 |
| LU | The factorisation PA = LU; a reusable machine for solving square systems. | Part 6 |
| M | ||
| Markov chain | A system advanced by a column-stochastic transition matrix; converges to its stationary distribution. | Part 14 |
| Matrix | An array of numbers that represents a linear map once a basis is fixed. | Part 3 |
| Moore–Penrose pseudoinverse | The unique A⁺ satisfying the four Penrose conditions; it generalises the inverse to every matrix. | Part 18 |
| N | ||
| Norm | The length of a vector, ‖v‖ = √(v·v). | Part 8 |
| Normal equations | AᵀA x = Aᵀb; the condition for a least-squares solution. | Part 10 |
| Null space | Every x with Ax = 0; the directions the map destroys. | Part 7 |
| Nullity | The dimension of the null space; rank plus nullity equals the number of columns. | Part 7 |
| O | ||
| Orientation | Whether a transformation preserves handedness; a reflection flips it, changing the sign of the determinant. | Part 5 |
| Orthogonal | Perpendicular: the dot product is zero. | Part 8 |
| Orthonormal | Mutually orthogonal unit vectors; the columns of Q satisfy QᵀQ = I. | Part 9 |
| P | ||
| PCA | Principal component analysis: project data onto the top right singular vectors (the top covariance eigenvectors). | Part 17 |
| Pivot | The leading nonzero entry used to eliminate below it during Gaussian elimination. | Part 6 |
| Positive definite | xᵀAx > 0 for every nonzero x; equivalently all eigenvalues are positive and the quadratic form is a bowl. | Part 15 |
| Projection | The component of one vector along another; the closest point on the other's span. | Part 8 |
| Pseudoinverse | A⁺ = VΣ⁺Uᵀ; gives the least-squares least-norm solution for any shape. | Part 18 |
| Q | ||
| QR | A = QR with Q orthonormal and R upper triangular; the stable route to least squares. | Part 9 |
| Quadratic form | xᵀAx; a scalar-valued function whose level sets the eigenvalues shape. | Part 15 |
| Quaternion | Four numbers representing a rotation; avoids gimbal lock at the cost of a double cover. | Part 20 |
| R | ||
| Rank | The dimension of the column space; the number of independent directions the map preserves. | Part 7 |
| Rank-k approximation | The closest rank-k matrix to A, obtained by keeping the k largest singular values. | Part 17 |
| Residual | What is left over, b − Ax; perpendicular to the column space at a least-squares solution. | Part 10 |
| Rotation | An orthonormal matrix with determinant +1; length- and angle-preserving. | Part 20 |
| Row | One row of A is a linear functional: it dots a vector to produce one output coordinate. | Part 3 |
| Row echelon form | The triangular form left by elimination, ready for back-substitution. | Part 6 |
| Row space | The span of the rows; the map sends it isomorphically onto the column space. | Part 7 |
| S | ||
| Scalar | A single number; it scales a vector, changing its length and possibly its sign. | Part 1 |
| Singular | Not invertible; the determinant is zero and the columns are dependent. | Part 5 |
| Singular value | σ, a nonnegative stretch factor in the SVD, equal to √(eigenvalue of AᵀA). | Part 16 |
| Singular vector | The right vectors v (circle directions) and left vectors u (ellipse axes) of the SVD. | Part 16 |
| Skew-symmetric | Aᵀ = −A; the matrix [a]× that performs a×. | Part 11 |
| SO(3) | The group of 3D rotations: orthonormal matrices with determinant +1. | Part 20 |
| Softmax | Exponentiate and normalise a vector so its entries are positive and sum to one. | Part 22 |
| Span | Every vector reachable by linear combination; the space a set generates. | Part 2 |
| Spectral theorem | A symmetric matrix factorises as A = QΛQᵀ with real λ and orthogonal eigenvectors. | Part 15 |
| Stationary distribution | A probability vector π with Pπ = π; the λ = 1 eigenvector of a Markov chain. | Part 14 |
| Subspace | A subset closed under addition and scaling; every subspace contains the origin. | Part 7 |
| SVD | A = UΣVᵀ: a rotation, a stretch, and another rotation. Every matrix has one. | Part 16 |
| Symmetric matrix | A = Aᵀ; real eigenvalues and perpendicular eigenvectors. | Part 15 |
| T | ||
| Trace | The sum of the diagonal entries; it equals the sum of the eigenvalues. | Part 13 |
| Transpose | Flipping a matrix across its diagonal; (Aᵀ)ij = Aji. | Part 3 |
| U | ||
| Unit vector | A vector of length one; used to carry a direction. | Part 9 |
| V | ||
| Vector | An arrow, a list of numbers, or an abstract object you can add and scale — three views of one thing. | Part 1 |
| Vector space | A collection with addition and scaling obeying the usual rules. | Part 1 |