Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Tangency is the condition

Why the level set must touch the constraint

Write the arm's two joint angles as $\mathbf{x}=(\theta_1,\theta_2)$ and let the effort be a smooth objective $f(\mathbf{x})$, smallest near the rest pose. The reach requirement confines the arm to a curve $g(\mathbf{x})=c$ — the set of joint configurations that satisfy one scalar task. The problem is

$$ \min_{\mathbf{x}}\; f(\mathbf{x}) \quad\text{subject to}\quad g(\mathbf{x}) = c. $$

Suppose $\mathbf{x}(t)$ is any smooth path that stays on the constraint, so $g(\mathbf{x}(t))=c$ for all $t$. Differentiating gives $\nabla g\cdot\mathbf{x}'(t)=0$: every feasible velocity is orthogonal to $\nabla g$. At a constrained minimum the objective cannot decrease in any feasible direction, so $\tfrac{d}{dt}f(\mathbf{x}(t)) = \nabla f\cdot\mathbf{x}'(t) = 0$ for all such directions. The only way a vector can be orthogonal to the whole tangent space and still be felt is if it points along the normal:

$$ \nabla f(\mathbf{x}^*) \;=\; \lambda\,\nabla g(\mathbf{x}^*), \qquad \lambda = \frac{\nabla f\cdot\nabla g}{\lVert\nabla g\rVert^{2}}. $$

The implicit function theorem is what makes this legitimate. At a point where $\nabla g\neq 0$ the constraint $g=c$ is locally the graph of a function, so the feasible set really is a one-dimensional curve with a well-defined tangent line — the tangent line the optimizer slides along. Drag the point around the constraint: while $\nabla f$ and $\nabla g$ disagree, their non-parallel parts push $f$ downhill; they line up exactly at the optimum.

Top: the Plotly surface $f$ with the constraint curve lifted onto it. Bottom: the same picture in the plane — shaded level sets of $f$, the constraint $g=c$ (pink), the feasible tangent line (dashed), and the two gradients at the draggable point. The ring marks the constrained optimum.

💡 Read the residual. The residual $\nabla f-\lambda\nabla g$ is the part of the objective's gradient that the constraint cannot absorb. It is zero precisely at the multiplier that makes tangency hold — the residual is a numerical stationarity test.
2

Move the constraint, watch the optimum slide

The envelope, and what the multiplier means

Now let the requirement change: replace $c$ by a slider and the constraint curve sweeps through the plane, the optimum sliding along with it. Record the best objective value at each level and you get the value function $f^*(c)=\min_{g=c} f$. Its slope is the multiplier — the single most useful fact about $\lambda$:

$$ \frac{df^*}{dc} \;=\; \lambda. $$

This is why $\lambda$ is called the shadow price. It answers a marginal question without re-solving anything: if the target is moved a little, the least achievable effort changes by $\lambda$ times that move. In the arm it is the extra joint effort bought by one more unit of reach. Move the slider and compare the envelope's local slope with the $\lambda$ read off the two gradients at the optimum; they agree to finite-difference accuracy.

Top: the constraint at the current level with the unconstrained minimum and the sliding constrained optimum. Bottom: the envelope $f^*(c)$; the dashed tangent has slope $\lambda\approx df^*/dc$.

3

Drag the constraint curve

Change the geometry and re-solve

The multiplier belongs to the constraint, not to the objective, so changing the constraint changes both the answer and its price. Grab the two handles: the centre moves the reach constraint through the plane, the radius handle stretches it. The optimum re-solves on every drag and the readout reports the new configuration and the new $\lambda$. The slider changes the level $c$ as before.

Drag the dark centre handle and the pink radius handle to reshape $g$; the optimum (ring) and the reported $\lambda$ update live.

4

Inequalities and the KKT conditions

Feasible regions, active and inactive constraints

Real tasks are usually one-sided: reach at least this far, stay inside this region. Write the feasible set as $g(\mathbf{x})\le 0$ and the shaded region below is it. The unconstrained minimum sits at the centre; slide the region so it contains that point and the constraint is doing nothing — it is inactive, the multiplier is zero. Slide it away and the optimum is forced onto the boundary, the constraint is active, and a positive multiplier appears. The Karush–Kuhn–Tucker conditions capture both cases at once:

$$ \nabla f(\mathbf{x}^*) + \lambda\,\nabla g(\mathbf{x}^*) = 0, \qquad g(\mathbf{x}^*) \le 0, \qquad \lambda \ge 0, \qquad \lambda\,g(\mathbf{x}^*) = 0. $$

The four lines are, in order: stationarity, primal feasibility, dual feasibility, and complementary slackness. The last one is the switch — either $g=0$ (the constraint is active and may charge a price) or $\lambda=0$ (it is inactive and charges nothing). Move the sliders to cross the boundary and watch the checklist flip.

Shaded: the feasible region $g\le 0$. The ring is the constrained optimum; when it is interior the constraint is inactive ($\lambda=0$), when it is on the boundary it is active ($\lambda>0$).

5

The multiplier as a penalty weight

One unconstrained problem hiding inside the constrained one

At the end of Step 1, $\nabla f = \lambda\nabla g$. Rearranged, that says the gradient of $f + \lambda g$ vanishes at the constrained optimum. So the constrained answer is also an unconstrained minimum of the penalised objective

$$ \min_{\mathbf{x}} \; f(\mathbf{x}) + \mu\,g(\mathbf{x}), $$

provided the penalty weight $\mu$ is exactly the multiplier $\lambda$. Slide $\mu$ toward $\lambda$: the penalised minimum walks from the unconstrained centre to the constrained optimum and the gap closes. This is the seed of penalty and augmented-Lagrangian methods — solve a sequence of easy unconstrained problems whose target shifts with the price of violating the constraint.

The constrained optimum $x^*$ (ring) and the unconstrained minimiser of $f+\mu g$ (dot). At $\mu=\lambda$ they coincide.

6

Where this shows up

The geometry behind every constrained solver

Anything that must be optimal and legal is this page. A robot reaching a target, a controller respecting an actuator limit, a network trained under a norm budget — each is an objective plus one or more constraints, and each has a KKT system at its solution. That is exactly the territory of Nonlinear Optimization, where constrained problems, penalty methods and the KKT conditions are the working toolkit; the Lagrange picture here is the geometry those algorithms are approximating. And the moment the objective depends on a matrix of parameters, differentiating those stationarity conditions needs the layout rules of the next part, which is why the two are usually learned together.

7

Notation to carry forward

ObjectMeaningWhere it comes up
∇f = λ∇gLagrange condition for an equality constraint: the gradients are parallelConstrained optima, this part
λ = (∇f·∇g)/‖∇g‖²The multiplier as a projection of the objective gradient onto the constraint normalReadout in Steps 1–3
r = ∇f − λ∇gStationarity residual; zero exactly when tangency holdsNumerical optimality tests
df*/dc = λShadow price: marginal change in the optimum per unit of constraintSensitivity, envelope theorem
∇f + λ∇g = 0, g ≤ 0, λ ≥ 0, λg = 0The KKT conditions for an inequality constraintStep 4; constrained solvers
f + μ gPenalty / embedded unconstrained objective whose optimum matches at μ = λStep 5; penalty methods
8

Further reading

9

Check your understanding

0/4 answered