Lagrange multipliers & KKT
An unconstrained minimum asks where the gradient vanishes. A constrained one asks a different question, because the answer is no longer allowed to be anywhere: it has to sit on a surface. Our running example is the two-link arm choosing the least-joint-effort configuration that still reaches a target. The optimum turns out to be a statement about shapes — at the answer, an objective level set is tangent to the constraint — and the number that records that tangency is the multiplier $\lambda$, which also measures what a little more reach costs you.
Tangency is the condition
Why the level set must touch the constraint
Write the arm's two joint angles as $\mathbf{x}=(\theta_1,\theta_2)$ and let the effort be a smooth objective $f(\mathbf{x})$, smallest near the rest pose. The reach requirement confines the arm to a curve $g(\mathbf{x})=c$ — the set of joint configurations that satisfy one scalar task. The problem is
Suppose $\mathbf{x}(t)$ is any smooth path that stays on the constraint, so $g(\mathbf{x}(t))=c$ for all $t$. Differentiating gives $\nabla g\cdot\mathbf{x}'(t)=0$: every feasible velocity is orthogonal to $\nabla g$. At a constrained minimum the objective cannot decrease in any feasible direction, so $\tfrac{d}{dt}f(\mathbf{x}(t)) = \nabla f\cdot\mathbf{x}'(t) = 0$ for all such directions. The only way a vector can be orthogonal to the whole tangent space and still be felt is if it points along the normal:
The implicit function theorem is what makes this legitimate. At a point where $\nabla g\neq 0$ the constraint $g=c$ is locally the graph of a function, so the feasible set really is a one-dimensional curve with a well-defined tangent line — the tangent line the optimizer slides along. Drag the point around the constraint: while $\nabla f$ and $\nabla g$ disagree, their non-parallel parts push $f$ downhill; they line up exactly at the optimum.
Top: the Plotly surface $f$ with the constraint curve lifted onto it. Bottom: the same picture in the plane — shaded level sets of $f$, the constraint $g=c$ (pink), the feasible tangent line (dashed), and the two gradients at the draggable point. The ring marks the constrained optimum.
Move the constraint, watch the optimum slide
The envelope, and what the multiplier means
Now let the requirement change: replace $c$ by a slider and the constraint curve sweeps through the plane, the optimum sliding along with it. Record the best objective value at each level and you get the value function $f^*(c)=\min_{g=c} f$. Its slope is the multiplier — the single most useful fact about $\lambda$:
This is why $\lambda$ is called the shadow price. It answers a marginal question without re-solving anything: if the target is moved a little, the least achievable effort changes by $\lambda$ times that move. In the arm it is the extra joint effort bought by one more unit of reach. Move the slider and compare the envelope's local slope with the $\lambda$ read off the two gradients at the optimum; they agree to finite-difference accuracy.
Top: the constraint at the current level with the unconstrained minimum and the sliding constrained optimum. Bottom: the envelope $f^*(c)$; the dashed tangent has slope $\lambda\approx df^*/dc$.
Drag the constraint curve
Change the geometry and re-solve
The multiplier belongs to the constraint, not to the objective, so changing the constraint changes both the answer and its price. Grab the two handles: the centre moves the reach constraint through the plane, the radius handle stretches it. The optimum re-solves on every drag and the readout reports the new configuration and the new $\lambda$. The slider changes the level $c$ as before.
Drag the dark centre handle and the pink radius handle to reshape $g$; the optimum (ring) and the reported $\lambda$ update live.
Inequalities and the KKT conditions
Feasible regions, active and inactive constraints
Real tasks are usually one-sided: reach at least this far, stay inside this region. Write the feasible set as $g(\mathbf{x})\le 0$ and the shaded region below is it. The unconstrained minimum sits at the centre; slide the region so it contains that point and the constraint is doing nothing — it is inactive, the multiplier is zero. Slide it away and the optimum is forced onto the boundary, the constraint is active, and a positive multiplier appears. The Karush–Kuhn–Tucker conditions capture both cases at once:
The four lines are, in order: stationarity, primal feasibility, dual feasibility, and complementary slackness. The last one is the switch — either $g=0$ (the constraint is active and may charge a price) or $\lambda=0$ (it is inactive and charges nothing). Move the sliders to cross the boundary and watch the checklist flip.
Shaded: the feasible region $g\le 0$. The ring is the constrained optimum; when it is interior the constraint is inactive ($\lambda=0$), when it is on the boundary it is active ($\lambda>0$).
The multiplier as a penalty weight
One unconstrained problem hiding inside the constrained one
At the end of Step 1, $\nabla f = \lambda\nabla g$. Rearranged, that says the gradient of $f + \lambda g$ vanishes at the constrained optimum. So the constrained answer is also an unconstrained minimum of the penalised objective
provided the penalty weight $\mu$ is exactly the multiplier $\lambda$. Slide $\mu$ toward $\lambda$: the penalised minimum walks from the unconstrained centre to the constrained optimum and the gap closes. This is the seed of penalty and augmented-Lagrangian methods — solve a sequence of easy unconstrained problems whose target shifts with the price of violating the constraint.
The constrained optimum $x^*$ (ring) and the unconstrained minimiser of $f+\mu g$ (dot). At $\mu=\lambda$ they coincide.
Where this shows up
The geometry behind every constrained solver
Anything that must be optimal and legal is this page. A robot reaching a target, a controller respecting an actuator limit, a network trained under a norm budget — each is an objective plus one or more constraints, and each has a KKT system at its solution. That is exactly the territory of Nonlinear Optimization, where constrained problems, penalty methods and the KKT conditions are the working toolkit; the Lagrange picture here is the geometry those algorithms are approximating. And the moment the objective depends on a matrix of parameters, differentiating those stationarity conditions needs the layout rules of the next part, which is why the two are usually learned together.
Notation to carry forward
| Object | Meaning | Where it comes up |
|---|---|---|
∇f = λ∇g | Lagrange condition for an equality constraint: the gradients are parallel | Constrained optima, this part |
λ = (∇f·∇g)/‖∇g‖² | The multiplier as a projection of the objective gradient onto the constraint normal | Readout in Steps 1–3 |
r = ∇f − λ∇g | Stationarity residual; zero exactly when tangency holds | Numerical optimality tests |
df*/dc = λ | Shadow price: marginal change in the optimum per unit of constraint | Sensitivity, envelope theorem |
∇f + λ∇g = 0, g ≤ 0, λ ≥ 0, λg = 0 | The KKT conditions for an inequality constraint | Step 4; constrained solvers |
f + μ g | Penalty / embedded unconstrained objective whose optimum matches at μ = λ | Step 5; penalty methods |
Further reading
- Wikipedia, Lagrange multiplier — the tangent argument, the implicit-function proof, and worked examples.
- Wikipedia, Karush–Kuhn–Tucker conditions — the inequality theory, constraint qualifications and complementary slackness.
- MIT 18.02, Multivariable Calculus — the course companion for Lagrange multipliers and constrained extrema.