Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Notation, fixed once

The conventions every part follows

Vectors are columns and are written bold: v. The transpose vᵀ is the same numbers in a row, so a column vector is an n×1 matrix. The i-th coordinate is vi. Matrices are capital letters, A, with entry Aij in row i, column j. In code a matrix is an array of rows, matching the toolkit. The identity is I, the zero matrix 0, the inverse A⁻¹, the transpose Aᵀ, the norm ‖v‖, eigenvalues λ, singular values σ, and the pseudoinverse A⁺.

Sizes

A: m×n  ·  x: n×1  ·  Ax: m×1
J: m×n  ·  JᵀJ: n×n  ·  σ: min(m,n)

Reading order

Parts build on each other, but each stands alone. Acts I–III are the machinery; Act IV is where it shows up. Every core part ends with a "Where this shows up" panel linking into the site's other guides.

2

Cheat sheet

Decompositions, and what each one is for

DecompositionFormWhat it is forLeading costPart
LUPA = LUSolving Ax=b for a square, non-singular A; reusable across many right-hand sides.2n³/36
QRA = QRLeast squares and stable solving; Q orthonormal, R upper triangular.2mn² − 2n³/39–10
CholeskyA = LLᵀSymmetric positive-definite systems; half the work of LU and always stable.n³/315
EigendecompositionA = QΛQ⁻¹, or QΛQᵀ if symmetricDynamics, stability, quadratic forms; diagonalises the map.iterative, ~10n³14
SVDA = UΣVᵀRank, conditioning, best low-rank approximation, pseudoinverse — the complete picture.iterative, ~20n³16–18

Identities worth keeping

$$A = U\Sigma V^\top, \qquad A^{+} = V\Sigma^{+}U^\top, \qquad A^\top A = V\Sigma^2 V^\top$$

The SVD and the pseudoinverse it produces. Σ⁺ inverts the nonzero singular values and zeroes the rest.

$$\lambda_1\lambda_2 = \det A, \qquad \lambda_1+\lambda_2 = \operatorname{tr}A$$

For a 2×2 the eigenvalues are two numbers that account for the determinant and the trace; the discriminant decides real or complex.

$$\operatorname{rank}A + \operatorname{nullity}A = n, \qquad \kappa(A) = \frac{\sigma_{\max}}{\sigma_{\min}}$$

Rank–nullity counts the dimensions a map preserves and destroys; the condition number measures how close σmin is to zero.

$$Q^\top Q = I, \qquad [\mathbf{a}]_\times \mathbf{b} = \mathbf{a}\times\mathbf{b}, \qquad R = e^{\theta[\hat{\mathbf{u}}]_\times}$$

Orthonormal columns; the cross product as a skew matrix; a rotation as the exponential of a skew matrix.

3

Glossary

Every term, linked to the part that introduces it

💡 Filter the list. Type any fragment — a term, a symbol, an idea — and the table hides non-matching rows and reports how many remain. Or narrow by category with the buttons. Matching is case-insensitive and looks at both the term and its definition.
TermWhat it meansIntroduced in
A
AngleThe angle θ between two vectors satisfies cos θ = u·v / (‖u‖‖v‖); the dot product in disguise.Part 8
AttentionToken similarity computed as QKᵀ, normalised by softmax, and used to mix the value matrix V.Part 22
B
BasisA linearly independent set that spans the space; coordinates in it are unique.Part 2
C
Characteristic polynomialdet(A − λI) = 0; its roots are the eigenvalues.Part 13
ColumnOne column of A is the image of a basis vector under the map; the matrix is its columns side by side.Part 3
Column spaceEvery vector Ax can reach; the range of the map.Part 7
CompositionApplying B then A is the single map AB.Part 4
Condition numberκ(A) = σmax/σmin; the factor by which a relative error in b can be amplified in x.Part 19
Cosine similarityu·v divided by the product of norms; a scale-free measure of alignment.Part 8
Covariance matrixA symmetric matrix whose entries are feature variances along the diagonal and pairwise covariances off it.Part 17
Cross producta×b, perpendicular to both, with length equal to the parallelogram's area.Part 11
D
DeterminantThe signed factor by which A scales area (2D) or volume (3D); zero exactly when A is singular.Part 5
DiagonalizationWriting A = PDP⁻¹ with D diagonal; possible when there are enough independent eigenvectors.Part 14
DimensionThe number of vectors in any basis of the space.Part 2
Dot productThe sum of entrywise products; measures projection and angle.Part 8
E
EigenvalueThe scalar λ with Av = λv; the stretch factor along an eigen-direction.Part 13
EigenvectorA direction that A only stretches or flips, never turns.Part 13
Exponential mapexp sends a skew matrix θ[û]× to the rotation of angle θ about û.Part 20
Outer productuvᵀ; a rank-1 matrix, the building block of low-rank approximations.Part 7
G
Gauss–NewtonIterating the linear solve JᵀJ δ = −Jᵀr to fit a nonlinear least-squares model.Part 21
Gaussian eliminationRow operations that turn A triangular while preserving the solution set.Part 6
Gimbal lockThe degeneracy of Euler angles when two rotation axes align, losing a degree of freedom.Part 20
Gram–SchmidtTurning independent vectors into an orthonormal set by subtracting successive projections.Part 9
H
HessianThe matrix of second derivatives of a function; JᵀJ approximates it for least-squares costs.Part 21
I
IdentityThe matrix I that leaves every vector unchanged; the multiplicative unit.Part 3
Linear independenceNo vector in the set is a linear combination of the others.Part 2
InverseA⁻¹ undoes A; it exists exactly when det A ≠ 0.Part 5
J
JacobianThe matrix of first partial derivatives; one row per residual, one column per parameter.Part 21
K
KV cacheThe stored key and value matrices a decoder reuses at every generation step.Part 22
L
Least squaresThe best fit when Ax=b has no exact solution; minimise ‖Ax−b‖².Part 10
Least-norm solutionAmong infinitely many solutions, the one with the smallest norm, produced by A⁺b.Part 18
Left null spaceThe solutions of Aᵀy = 0; the part of the output space the map never reaches.Part 7
Linear combinationA sum of scaled vectors, a₁v₁ + … + akvk.Part 1
Linear transformationA map that respects addition and scaling; equivalently, a matrix.Part 3
LoRAFine-tuning a frozen weight by adding a low-rank product BA.Part 22
Low-rank updateAdding a rank-r product to a matrix; cheap to store when r is small.Part 22
LUThe factorisation PA = LU; a reusable machine for solving square systems.Part 6
M
Markov chainA system advanced by a column-stochastic transition matrix; converges to its stationary distribution.Part 14
MatrixAn array of numbers that represents a linear map once a basis is fixed.Part 3
Moore–Penrose pseudoinverseThe unique A⁺ satisfying the four Penrose conditions; it generalises the inverse to every matrix.Part 18
N
NormThe length of a vector, ‖v‖ = √(v·v).Part 8
Normal equationsAᵀA x = Aᵀb; the condition for a least-squares solution.Part 10
Null spaceEvery x with Ax = 0; the directions the map destroys.Part 7
NullityThe dimension of the null space; rank plus nullity equals the number of columns.Part 7
O
OrientationWhether a transformation preserves handedness; a reflection flips it, changing the sign of the determinant.Part 5
OrthogonalPerpendicular: the dot product is zero.Part 8
OrthonormalMutually orthogonal unit vectors; the columns of Q satisfy QᵀQ = I.Part 9
P
PCAPrincipal component analysis: project data onto the top right singular vectors (the top covariance eigenvectors).Part 17
PivotThe leading nonzero entry used to eliminate below it during Gaussian elimination.Part 6
Positive definitexᵀAx > 0 for every nonzero x; equivalently all eigenvalues are positive and the quadratic form is a bowl.Part 15
ProjectionThe component of one vector along another; the closest point on the other's span.Part 8
PseudoinverseA⁺ = VΣ⁺Uᵀ; gives the least-squares least-norm solution for any shape.Part 18
Q
QRA = QR with Q orthonormal and R upper triangular; the stable route to least squares.Part 9
Quadratic formxᵀAx; a scalar-valued function whose level sets the eigenvalues shape.Part 15
QuaternionFour numbers representing a rotation; avoids gimbal lock at the cost of a double cover.Part 20
R
RankThe dimension of the column space; the number of independent directions the map preserves.Part 7
Rank-k approximationThe closest rank-k matrix to A, obtained by keeping the k largest singular values.Part 17
ResidualWhat is left over, b − Ax; perpendicular to the column space at a least-squares solution.Part 10
RotationAn orthonormal matrix with determinant +1; length- and angle-preserving.Part 20
RowOne row of A is a linear functional: it dots a vector to produce one output coordinate.Part 3
Row echelon formThe triangular form left by elimination, ready for back-substitution.Part 6
Row spaceThe span of the rows; the map sends it isomorphically onto the column space.Part 7
S
ScalarA single number; it scales a vector, changing its length and possibly its sign.Part 1
SingularNot invertible; the determinant is zero and the columns are dependent.Part 5
Singular valueσ, a nonnegative stretch factor in the SVD, equal to √(eigenvalue of AᵀA).Part 16
Singular vectorThe right vectors v (circle directions) and left vectors u (ellipse axes) of the SVD.Part 16
Skew-symmetricAᵀ = −A; the matrix [a]× that performs a×.Part 11
SO(3)The group of 3D rotations: orthonormal matrices with determinant +1.Part 20
SoftmaxExponentiate and normalise a vector so its entries are positive and sum to one.Part 22
SpanEvery vector reachable by linear combination; the space a set generates.Part 2
Spectral theoremA symmetric matrix factorises as A = QΛQᵀ with real λ and orthogonal eigenvectors.Part 15
Stationary distributionA probability vector π with Pπ = π; the λ = 1 eigenvector of a Markov chain.Part 14
SubspaceA subset closed under addition and scaling; every subspace contains the origin.Part 7
SVDA = UΣVᵀ: a rotation, a stretch, and another rotation. Every matrix has one.Part 16
Symmetric matrixA = Aᵀ; real eigenvalues and perpendicular eigenvectors.Part 15
T
TraceThe sum of the diagonal entries; it equals the sum of the eigenvalues.Part 13
TransposeFlipping a matrix across its diagonal; (Aᵀ)ij = Aji.Part 3
U
Unit vectorA vector of length one; used to carry a direction.Part 9
V
VectorAn arrow, a list of numbers, or an abstract object you can add and scale — three views of one thing.Part 1
Vector spaceA collection with addition and scaling obeying the usual rules.Part 1