The complete picture of Ax = b
You have met A x = b four times already. Part 6 solved the square case by elimination, Part 7 named the four subspaces that decide reachability, Part 10 projected when no solution exists, and Part 17 kept only the largest singular values. This part closes the loop. There is a single matrix, built straight from the singular value decomposition, that answers every version of the question at once: it returns an exact solution when one exists, the closest solution when none does, and among all the solutions that fit it always returns the smallest one. It is the Moore–Penrose pseudoinverse, written A⁺, and it is the reason the phrase “the solution to A x = b” can be said out loud without first checking the shape of A.
The question
Four cases, one object
Write down a matrix and a target and there are really only four things that can happen. The columns may be independent and b may be reachable, so there is exactly one solution. The columns may be independent but b may lie off the column space, so there is no solution and the honest task is to minimise the residual. The columns may be dependent while b is still reachable, so there are infinitely many solutions and the interesting task is to pick the smallest. Or both failures can happen at once: b is unreachable and the null space is nontrivial, so there is an infinite family of equally good least-squares answers.
Textbooks usually split these into separate chapters — invertible matrices here, least squares there, minimum-norm problems somewhere else. The geometry says they are one problem. A linear map sends its row space one-to-one onto its column space and crushes the null space to zero. So the only sensible answer to A x = b is the one built by reversing the map on the part it did not destroy: take b, keep its component in the column space, map that component back into the row space, and ignore everything else. The pseudoinverse is exactly that procedure written as a matrix.
The reason it always works is that the two halves of the domain and the two halves of the codomain are locked together by orthogonality. The row space is perpendicular to the null space, and the column space is perpendicular to the left null space. Reversing the map is therefore not a guess about which solution to prefer; it is the unique choice that stays inside the row space, and staying inside the row space is the same as having no null-space component, which is the same as having the smallest norm.
One widget, every case
Edit the matrix, watch the answer change character
Here is the whole subject in a single panel. Pick a shape with the buttons: a tall matrix with more equations than unknowns, a square one, a wide one with more unknowns than equations, or a rank-deficient one whose columns are secretly dependent. Then edit the entries of A and b directly. The widget computes A⁺, forms x̂ = A⁺b, and reports which of the four cases you are in, together with the two numbers that matter: the residual ‖A x̂ − b‖, and the size ‖x̂‖ of the answer.
The bars on the left explain what the pseudoinverse is doing. Every matrix has a list of singular values, and A⁺ inverts the ones that are meaningfully nonzero while quietly zeroing the rest. A rank-deficient matrix has a singular value at the numerical floor; that direction is exactly the one the pseudoinverse refuses to amplify. When the columns are independent all the bars stand above the cutoff, and when they are dependent one bar falls below it and the rank drops.
Singular values of the current A. The pseudoinverse reciprocates the bars above the dashed cutoff and discards the ones below.
Try each preset and read the case label. The tall preset is inconsistent, so A⁺b is the unique least-squares answer and the residual is positive. The square preset is invertible, so A⁺b equals the ordinary solution solve(A, b) and the residual is zero. The wide preset is consistent but underdetermined, so the pseudoinverse returns the least-norm member of an infinite family. The rank-deficient preset is consistent with a dependent column, so again there is a family, and the pseudoinverse picks the shortest member. The cross-check line at the bottom confirms the agreement wherever an ordinary solver applies.
The family slider is the part worth playing with. Once the null space is nontrivial, add t times the displayed null-space direction to x̂ and watch two things stay true at once. The residual does not move, because A n = 0 for every null vector, so every point on that line solves the same problem equally well. But the norm grows the moment you step away from x̂, which is the geometric content of “least norm”: the pseudoinverse answer is the one point of the family closest to the origin.
The pseudoinverse formula
Reciprocate the singular values, and only those
The construction is short once you have the SVD from Part 16. Every matrix factors as A = U Σ Vᵀ, where U and V are orthogonal and Σ is diagonal with the singular values σ₁ ≥ σ₂ ≥ … ≥ 0 on its diagonal. An inverse would want to divide by those numbers, but a zero singular value has no reciprocal and a tiny one produces an enormous, meaningless answer. The pseudoinverse resolves both problems with one rule: reciprocate every singular value that is not effectively zero, and set the rest to zero.
Read the formula as a set of instructions on the four subspaces. Multiplication by Uᵀ resolves a vector into the output directions; Σ⁺ reverses the ones the map genuinely stretched, by the reciprocal amount, and annihilates the directions beyond the rank; multiplication by V rebuilds the answer in the input directions. Restricted to the column space, A⁺ is a true inverse of A, sending each column-space direction back to the row-space direction it came from. On the left null space it is the zero map, because there is nothing to undo. And it never produces a null-space component, because its output lives entirely in the row space. Those three sentences are the entire geometry.
The definition is usually pinned down by four identities, the Moore–Penrose conditions. They look like bookkeeping, but each one names a projection:
The symmetry conditions say that A A⁺ and A⁺ A are symmetric, hence honest orthogonal projections. The first is the projection onto the column space, the second the projection onto the row space. So A A⁺ b is the component of b the map can reach, and A⁺ A x strips any null-space part from x. The remaining two conditions say the reversal is consistent: reversing, then applying, then reversing again changes nothing.
Everything you need about A x = b now follows in two lines. Split b into its column-space part and its left-null part. The left-null part can never be matched, so the best residual is its length, and A A⁺ b is exactly the column-space part, which is why the residual is always orthogonal to the column space. Then x̂ = A⁺ b is the unique preimage of that column-space part inside the row space, and it is the smallest preimage because every other one differs by a null-space vector. A consistent system has zero left-null part, so the residual vanishes and x̂ is the least-norm exact solution; an inconsistent system has no exact solution, so x̂ is the least-norm least-squares solution. One formula, both stories.
One practical note carries into the next section. The cutoff that decides which singular values count is a choice, not a fact. At exactly zero rank the matrix is singular and the pseudoinverse is unambiguous in exact arithmetic, but in floating point a singular value may be 10⁻¹⁶ rather than zero, and a naive reciprocal would produce an answer in the 10¹⁰ range. The standard fix is a relative tolerance, which is what the toolkit uses by default and what the dashed line in the widget shows. Choosing that line generously is the first step toward regularisation.
The big picture
Two rooms, four subspaces, two maps
Put the whole story in one diagram. On the left is the input space ℝn, split into the row space, where A acts faithfully, and the null space, which it destroys. On the right is the output space ℝm, split into the column space, which is reachable, and the left null space, which is not. The map A runs left to right: it carries the row space onto the column space and sends the null space to zero. The pseudoinverse runs right to left: it carries the column space back onto the row space and sends the left null space to zero.
Look at what A⁺ chooses to ignore. It has a perfectly good definition on the left null space — it is zero there — but it never uses that definition when you ask it to solve A x = b, because b’s left-null component is the part that cannot be matched and is simply left as residual. And it never adds a null-space component to its answer, even though the null space is where all the other solutions live. Those two refusals are not limitations. They are exactly what makes the answer canonical.
Select a subspace to highlight it and list the basis the current matrix actually has.
Step through the four buttons and the diagram becomes a map of the algebra. The row-space basis comes from colSpace(Aᵀ), the null-space basis from nullSpace(A), the column-space basis from colSpace(A), and the left-null basis from nullSpace(Aᵀ). The dimensions always add up the way Part 7 promised: the two pieces of the domain sum to n, and the two pieces of the codomain sum to m.
The dashed perpendicular marks are the load-bearing part of the picture. Row space and null space are orthogonal complements, so any input splits uniquely into a part the map keeps and a part it discards. Column space and left null space are orthogonal complements, so any target splits uniquely into a part that is reachable and a part that is not. The pseudoinverse pairs the two faithful pieces together and zeroes the two lost pieces, which is why it is the natural inverse even when no inverse exists.
Damped least squares
Trading residual for a smaller answer
The exact pseudoinverse is brittle near the cutoff. A singular value just above the line gets reciprocated to a huge number, and a small change in b along that direction swings the solution wildly — the conditioning story of Part 19. The standard remedy is to stop inverting singular values exactly and instead apply a soft filter that fades them out. Add a multiple of the identity to AᵀA and you get the damped or Tikhonov-regularised pseudoinverse:
The two forms say the same thing from two directions. Algebraically, Aλ⁺ is the minimiser of a penalised loss, ‖A x − b‖2 + λ2‖x‖2, so the damping term charges you for a large answer and the residual term charges you for a bad fit. In the SVD picture, each singular value is replaced by the filter factor σi / (σi2 + λ2), which equals the exact reciprocal when σi is much larger than λ and fades smoothly to zero when it is much smaller. No direction is ever amplified without limit.
The slider below sweeps λ and shows both sides of the bargain at once. At λ = 0 you recover the ordinary pseudoinverse: smallest residual, largest answer. As λ grows, the residual rises — you are deliberately fitting less well — while the solution norm falls, because the penalty pushes the answer toward the origin. The two curves always move in opposite directions. Choosing λ is choosing a point on that trade-off, and the usual heuristic is the bend of the L-shaped curve the two quantities trace out.
Residual and solution size as functions of the damping λ. The marker follows the slider; λ = 0 is the exact pseudoinverse.
Damping is also what makes the pseudoinverse usable on matrices whose columns are nearly dependent. In that regime the exact inverse is a division by a number that is almost zero, and the answer is dominated by noise. The filter factor caps the division at 1/(2λ), so the pseudoinverse stays bounded and the solution stays stable. You give up a little residual to buy a lot of robustness, and for a robot or a regression that is almost always the right trade.
Where this shows up
One inverse, two worlds
Redundant inverse kinematics
A robot arm with more joints than task dimensions has a wide Jacobian, so a desired end-effector twist usually has a whole family of joint-velocity solutions. The damped pseudoinverse picks the least-norm member and, because it caps the near-singular directions, keeps the commanded velocities finite as the arm approaches a singularity. The pose-graph part of the optimization guide uses the same machinery to handle gauge freedom in a map.
Regression and fine-tuning updates
A linear regression layer fits its weights by the normal equation, and when features are correlated that system is ill-conditioned; ridge regression adds exactly the λ2I term above and is the damped pseudoinverse under another name. The same least-norm choice appears in fine-tuning, where a small update is preferred among all the updates that fit a batch. The architecture chapter sets up the matrices these updates act on.
Further reading
- Grant Sanderson, "Inverse matrices, column space and null space", Essence of Linear Algebra, 3Blue1Brown — the geometric picture of what an inverse can and cannot do.
- Gilbert Strang, 18.06 Linear Algebra, MIT OpenCourseWare — Lectures 32 and 33 develop the pseudoinverse and the four subspaces it acts on.
- Immersive Math, Chapter 11: The singular value decomposition — the decomposition the pseudoinverse is built from, in a draggable textbook.
- Åke Björck, Numerical Methods for Least Squares Problems — the reference for the damped and truncated pseudoinverses and how to pick the damping.