Regularization is a prior
Least squares will happily hand you enormous coefficients if that shaves a little residual error off a small or collinear sample. Adding a penalty for coefficient size tames that, and the penalty looks arbitrary until you read it as a probability: a quadratic penalty is a Gaussian prior, an absolute penalty is a Laplace prior. Once you see that, regularization becomes the oldest idea in Bayesian inference wearing numerical clothing — maximum a posteriori estimation is just penalised maximum likelihood, and the penalty strength is the prior precision.
The question
Why would you deliberately bias an estimate?
Ordinary least squares has one job: choose the coefficient vector $\beta$ that makes the residuals as small as possible in sum of squares. A fit that minimises the residual sum of squares is not always the fit you want. When the predictors are numerous, correlated with each other, or outnumbered by the observations, the least-squares coefficients can be huge and wildly unstable: refit on a slightly different sample and they jump. The model has memorised the noise instead of the signal.
The fix is to add a price for the size of the coefficients, and there are two canonical prices. The ridge (or $L_2$) penalty charges for the sum of squares of the coefficients, so the objective becomes
The lasso (or $L_1$) penalty charges for the sum of absolute values instead:
In both, $\lambda \ge 0$ is the dial. At $\lambda = 0$ you get ordinary least squares; as $\lambda \to \infty$ every coefficient is dragged to zero. Everything in between is a compromise between fitting the data and keeping the coefficients small. The two penalties look almost the same written down and behave very differently.
The need to constrain
The unregularized fit and its instability
Least squares solves the normal equations $(X^{T}X)\,\beta = X^{T}y$. Everything hinges on the matrix $X^{T}X$, the Gram matrix of the predictors. If the columns of $X$ are nearly parallel — the predictors are collinear — then $X^{T}X$ is nearly singular, its smaller eigenvalue is tiny, and inverting it multiplies any noise in $y$ by a huge factor. The coefficients that come out are large, of opposite signs, and exquisitely sensitive to the data. This is not a subtle numerical accident; it is the same conditioning problem the numerics part describes, and it is exactly what regularisation repairs.
Adding $\lambda \lVert\beta\rVert_2^2$ to the objective adds $\lambda I$ to that matrix, so ridge solves $(X^{T}X + \lambda I)\beta = X^{T}y$. Every eigenvalue of the system is lifted by $\lambda$, the smallest one no longer sits near zero, and the inverse becomes tame. The price is bias: ridge estimates are no longer centred on the true coefficients, because the penalty pulls them toward the origin. What you buy with that bias is a large reduction in variance, and if the reduction outweighs the bias the predictions get better. That is the bias–variance trade-off of the previous part, now with a knob attached.
Larger $\lambda$ means more shrinkage, more bias, and less variance; smaller $\lambda$ moves back toward the jumpy unbiased fit. The useful value is chosen by holding out data — cross-validation and the information criteria of the next part — because the objective being minimised is deliberately not the quantity you care about.
Two details matter before the pictures make sense. The penalty is not scale invariant — doubling one predictor's units halves its coefficient, so predictors are standardised first. And the intercept is left unpenalised, so shrinking the slopes does not force the fit through the origin. The demo centres every column, which removes the intercept and lets the geometry live in two dimensions.
Two shapes, two solutions
Where a loss contour first touches the constraint
Every penalised problem has an equivalent constrained twin. Minimising residuals plus a penalty is the same as minimising residuals subject to the coefficients lying inside a region, with $\lambda$ playing the role of a Lagrange multiplier. For ridge the region is the ball $\lVert\beta\rVert_2 \le t$, a circle in two dimensions; for lasso it is the ball $\lVert\beta\rVert_1 \le t$, a diamond with vertices on the axes. The Lagrange-multiplier part derives exactly why the penalty and the constraint describe the same family of solutions; here the picture is the argument.
In coefficient space, constant values of the residual sum of squares form concentric ellipses centred on the ordinary least-squares solution. Minimising a loss is a walk inward, downhill toward that centre. The constrained problem asks for the smallest loss ellipse that still touches the allowed region: the answer is the point where the ellipse first kisses the boundary.
Now the shape decides everything. A circle is smooth, and a generic downhill path meets it at a point where both coordinates are nonzero; the ridge solution shrinks both coefficients together but almost never removes one. The diamond has corners sitting on the coordinate axes, and a shrinking ellipse often reaches a corner first, because a corner sticks out toward the centre. Landing on a corner means one coordinate is exactly zero — the geometric accident that makes the lasso a selection method and ridge not.
The demo uses a genuinely small two-predictor dataset so the ellipse lives on the page. The $\lambda$ slider changes the size of the allowed region and the highlighted contour, and the readout reports the ridge and lasso coefficient vectors side by side. Push $\lambda$ up and watch the diamond shrink until the loss ellipse touches it at a vertex: the lasso coordinate snaps to a hard zero while the ridge coordinate only slides toward it.
Faint dashed ellipses are RSS contours centred on the OLS solution; the magenta curve is the contour the solution touches. The blue region is the active constraint — a circle for ridge, a diamond for lasso.
Two refinements are visible here. When predictors are highly correlated the ellipse is thin and elongated: ridge spreads the shrinkage evenly across both, while the lasso keeps whichever predictor wins the race and zeroes the other. Neither is better; the shape of the region tells you which behaviour you are buying.
Read the numbers against the picture. The OLS point is fixed; for every $\lambda$ the ridge dot sits on the circle and the lasso dot on the diamond, each on its own touching contour, and only the lasso dot may arrive exactly on an axis.
A prior is a penalty in disguise
MAP is penalised maximum likelihood
So far $\lambda$ is an inexplicable dial. Bayesian reasoning gives it a meaning. Suppose you believe, before seeing data, that the coefficients are drawn from a distribution: a prior. Bayes' rule combines the likelihood of the residuals with that prior into a posterior over $\beta$, and the maximum a posteriori estimate is its most probable coefficient vector:
Now take the two standard priors. A Gaussian prior $\beta_j \sim \mathcal N(0, \tau^2)$ contributes $-\log p(\beta) = \lVert\beta\rVert_2^2/(2\tau^2)$ up to a constant, a quadratic penalty — ridge. A Laplace prior with density $\tfrac{b}{2}e^{-b\lvert\beta\rvert}$ contributes $b\lVert\beta\rVert_1$, an absolute penalty — lasso. Writing $\textrm{RSS}/(2\sigma^2)$ for the negative log likelihood of Gaussian noise, the penalised objectives appear with $\lambda = \sigma^2/\tau^2$ for ridge and $\lambda = 2\sigma^2 b$ for lasso. The dial is the prior precision: a narrow prior means a large $\lambda$ means heavy shrinkage, and a flat prior recovers maximum likelihood.
This is the same Gaussian prior that powered the conjugacy part of this series. There, a Gaussian prior on a mean fused with a Gaussian likelihood to give a Gaussian posterior by precision-weighted averaging — the update that later becomes the Kalman filter. Ridge regression is that story in higher dimensions: with a Gaussian prior and Gaussian noise the posterior over $\beta$ is Gaussian, its mean is the ridge estimate, and its covariance is the familiar $(\tfrac{1}{\sigma^2}X^{T}X + \tfrac{1}{\tau^2}I)^{-1}$.
The lasso is the more interesting case. A Laplace prior has a sharp peak at zero and exponential tails, and that peak is what produces sparse MAP estimates. But the Laplace is not conjugate to the Gaussian likelihood, so the posterior has no closed form — the toolbox of Parts 19 and 20, MCMC and Laplace or variational approximation, is what you reach for to get the whole posterior rather than just its mode. The MAP itself is exactly the penalised problem above, so "lasso = MAP with a Laplace prior" is a precise statement even though the full Bayesian treatment is harder.
Top: the Gaussian and Laplace prior densities on one coefficient. Bottom: their negative logs — the penalties themselves. Drag λ and watch the Gaussian narrow while the Laplace grows a sharper spike; the corner at zero in the $L_1$ penalty is the source of sparsity.
Where this shows up
One penalty, two worlds
Weight decay is a Gaussian prior
"Weight decay" in the pretraining chapter is this page's $L_2$ penalty applied to every weight of a neural network. The loss is training error plus $\lambda\lVert\theta\rVert_2^2$, which is a Gaussian prior on the parameters and a shrunken, better-conditioned optimum. Decoupled variants such as AdamW exist precisely because the interaction between the adaptive gradient and the penalty matters when the prior is being applied per-parameter at scale.
The regularised normal equations
The least-squares part solves $X\beta = y$ in the least-squares sense; ridge is the same problem with $\lambda I$ added to the Gram matrix, trading exactness for a bounded, stable solution. The same move appears whenever an inverse is nearly singular, and the rank and subspaces part explains the geometry: a penalty suppresses the directions where the data carry little information.
The pattern recurs wherever a model has more parameters than the data can pin down. The optimizers part shows the penalty as gradient flow: ridge is a constant force pulling weights toward the origin, while the lasso's kink calls for a proximal step that sets small coordinates exactly to zero. Sparse coding and feature selection are the diamond doing its work; smooth interpolation and stable deep nets are the circle doing its. Choosing between them is choosing which prior you are willing to assert about the world.
Further reading
The references below connect the geometric and the Bayesian readings, where the constraint region and the prior are two views of one object.
- Trevor Hastie, Robert Tibshirani and Jerome Friedman, The Elements of Statistical Learning, chapter 3 — ridge and lasso, the constraint picture, and the coefficient paths as $\lambda$ varies.
- Robert Tibshirani, "Regression Shrinkage and Selection via the Lasso", 1996 — the original lasso paper, including the geometric argument for sparsity.
- Christopher Bishop, Pattern Recognition and Machine Learning, section 3.3 — the Bayesian linear model, where the Gaussian prior gives the ridge posterior in closed form.
- Andrew Gelman et al., Bayesian Data Analysis, chapter 3 — priors, MAP versus posterior mean, and why a mode is only one summary of a posterior.
Cheat sheet
| Object | Reading here |
|---|---|
| Ridge objective | $\lVert y-X\beta\rVert^2 + \lambda\lVert\beta\rVert_2^2$; shrinks all coefficients smoothly |
| Lasso objective | $\lVert y-X\beta\rVert^2 + \lambda\lVert\beta\rVert_1$; can set coefficients exactly to zero |
| Constraint region | Circle $\lVert\beta\rVert_2 \le t$ for ridge, diamond $\lVert\beta\rVert_1 \le t$ for lasso |
| Solution | The smallest RSS ellipse that touches the boundary of the region |
| Why lasso is sparse | The region's corners sit on the axes, and a shrinking ellipse tends to reach a corner first |
| $\lambda = 0$ | Ordinary least squares: unbiased, high variance, unstable when collinear |
| $\lambda \to \infty$ | Every coefficient is driven to zero; the prior overwhelms the data |
| Gaussian prior | $\beta_j \sim \mathcal N(0,\tau^2)$ gives the quadratic penalty; $\lambda = \sigma^2/\tau^2$ |
| Laplace prior | $p(\beta_j)\propto e^{-b\lvert\beta_j\rvert}$ gives the absolute penalty; $\lambda = 2\sigma^2 b$ |
| MAP | Penalised maximum likelihood; a mode of the posterior, not the whole posterior |