Nonlinear Optimization with Rotations
Part 1 found a robot's position from noisy distances alone. Real robots also have a heading — which way they're facing — and real sensors usually report a bearing (an angle) as well as a range. Adding that one extra unknown, θ, brings in rotation matrices, angle-wrapping, and a new kind of ambiguity that pure distance measurements didn't have. Same four methods, same running robot — now with rotations.
What's new: heading and bearing
Recap
In part 1 the robot's unknown was a position p = (x, y), and every landmark gave one equation: a measured distance. Now the robot also has an unknown heading θ — the direction it's facing — so the unknown is a full pose:
And each landmark can now give two equations instead of one: a measured range (how far), exactly like before, plus a measured bearing (which direction, relative to the robot's own nose) from an onboard sensor like a camera or lidar. That second measurement is where rotation enters the picture.
Orientation and rotation matrices
Foundations
If the robot is facing world-angle θ, its "forward" and "left" directions in world coordinates are:
⎣sin θ cos θ⎦
A landmark's position relative to the robot, expressed in the robot's own body frame (as its sensor would report it, forward/left instead of east/north), is R(θ)ᵀ · (L − p) — rotate the world-frame offset backward by the robot's own heading. Drag the heading slider and watch the landmark's body-frame coordinates change even though nothing in the world actually moved — only the robot's idea of "forward" did.
Bearing measurements & the wrap-around trap
A new kind of nonlinearity
Fix the robot's position and vary only its heading guess, θ. The residual is the (wrapped) difference between the bearing we'd predict at that heading and the bearing we actually measured:
Below, the solid curve is the cost using the correctly-wrapped residual — smooth and periodic, repeating every 360°. The dashed curve is what happens if you skip the wrap and just subtract angles directly: it keeps climbing forever, even though θ and θ+360° describe the exact same heading. Uncheck "handle wrap-around" and run gradient descent from a starting point past the seam — it happily marches off in the wrong direction, chasing a phantom slope that correct angle math would never produce.
Click the chart to set a starting heading guess.
Why one landmark still isn't enough
Observability
Given a measured range and bearing to a single landmark, any heading guess θ has a matching position that satisfies both measurements exactly — just walk backward from the landmark by the measured range, in the direction the bearing implies. Drag θ below and watch the "perfectly consistent" position trace out a full circle around the landmark. Every point on it, paired with the right heading, explains the one landmark's measurement exactly — the cost there is zero. There's no unique answer yet.
Newton-Raphson for the full pose
Method 2, three unknowns now
p ← p − H⁻¹∇C — it's just a 3×3 system now instead of 2×2, and each landmark contributes two rows (range and bearing) to the gradient and Hessian instead of one.Two landmarks, each giving range + bearing: 4 equations for 3 unknowns — enough to pin the pose down, with one equation to spare. Click the map to place a starting position guess, then use the heading slider for a starting orientation guess, and step through Newton's method. The triangle marker shows both position and heading at once — watch it rotate into alignment as much as it slides into place.
Click the map to set the starting position.
Gauss-Newton for the full pose
Method 3, three unknowns now
JᵀJ, and skip every second derivative.A third landmark now (three landmarks × 2 measurements = 6 equations for 3 unknowns — comfortably over-determined, which also helps average out sensor noise). Toggle the comparison to overlay Newton's path on the same scenario — Gauss-Newton reaches essentially the same answer without ever differentiating atan2 a second time.
Click the map to set the starting position.
Levenberg-Marquardt for the full pose
Method 4, three unknowns now
λI (a 3×3 identity now) to JᵀJ, and adapt λ the same way as before.Same three-landmark scenario, but this time try a starting guess that's both far away and badly oriented (drag the heading slider to point the robot almost backward). Plain Gauss-Newton tends to struggle from here; adaptive λ reins it in automatically. This combination — range+bearing measurements, an adaptively-damped least-squares solve — is essentially what real-world visual/lidar SLAM back-ends run at every keyframe.
Click the map to set the starting position.
Playground: race all four, full pose
Put it all together
Four landmarks, range+bearing to each, one shared starting pose — every method runs at once. This time the results table tracks two kinds of error: how far off the estimated position is, and how far off the estimated heading is.
Click the map to set the shared starting position.
| Method | Iters | Pos. err | Heading err | Status |
|---|
Cheat sheet
Recap
| Concept | What changed from part 1 |
|---|---|
| Unknown | Position (x, y) → full pose (x, y, θ) |
| Rotation matrix | R(θ) converts between the world frame and the robot's own body frame |
| New measurement | Bearing (angle), alongside range — each landmark now gives 2 equations, not 1 |
| New nonlinearity | Angles wrap: residuals must be folded into (−π, π] or the cost function is simply wrong |
| New failure mode | A single landmark leaves a whole circle of equally-good poses (unobservable); need ≥2 |
| Update rules | Identical in form — gradient, Jacobian and Hessian are just 3-wide instead of 2-wide |
This is, essentially, a miniature version of what real robots run continuously: fusing range-and-bearing (or purely visual) observations of known landmarks into an estimate of the robot's full pose, refined every time a new observation comes in. Bigger versions of exactly this least-squares machinery — usually solved with Gauss-Newton or Levenberg-Marquardt — sit at the core of visual-inertial odometry, lidar SLAM, and bundle adjustment in 3D. Going further from here means moving from a flat plane to full 3D rotations, which stop being a single angle and start living on a genuine rotation group (SO(3)) instead of a vector space — the update rule grows a "retraction" step, but the gradient-descent-vs-Newton-vs-Gauss-Newton-vs-Levenberg-Marquardt story stays exactly the same.