Triangulation & Recovering Pose
Epipolar geometry (Part 6) tells you where a match could be; homographies (Part 8) only work on flat scenes. This part closes the loop for the general case: given two cameras and a correspondence, actually compute the 3D point — the first real moment of 3D reconstruction in this series.
Two problems, in order
Setup
Reconstructing a scene from two images is really two separate problems, solved in order: (1) figure out where camera 2 is relative to camera 1 (its relative pose), then (2) given both poses, work out where each matched point actually is in 3D. Part 6 already got us the essential matrix E from correspondences; this page turns E into an actual pose, and then turns rays into points.
From E to (R, t): four candidates, one makes sense
Foundations
The essential matrix E = [t]×R can be factored back into a rotation and a translation direction via its singular value decomposition, E = UΣVᵀ. The algebra always produces exactly four candidate poses — two possible rotations, each paired with +t or −t — and all four satisfy the epipolar constraint equally well, because E only fixes t up to an unknown overall scale and sign. (The full worked recipe — the W matrix and all — is in Section 4 below, not deferred to a textbook.) A minimal calibrated solver for E — the five-point algorithm, which needs only five correspondences instead of eight — is treated in Part 13.
Only one of the four puts a triangulated point in front of both cameras (positive depth in both) — the other three put it behind one or both. That test is called the cheirality check: triangulate a single correspondence with each candidate and keep the one where both depths come out positive. This is also exactly why triangulation (next) and pose recovery are tangled together in practice — you need a candidate pose to triangulate, and triangulating is how you rule the bad candidates out.
Two ways to triangulate
Foundations
Once both camera poses are known, a correspondence (x₁, x₂) defines two rays in 3D that, with perfect noise-free data, cross at exactly one point. With real (noisy) pixel measurements, the rays almost never exactly intersect — so "triangulation" really means finding the point that best explains both observations, and there are two standard ways to define "best":
Linear (DLT) triangulation. Stack the two cameras' projection equations into one linear system and solve it directly — fast, no iteration, but it's minimizing an algebraic error that doesn't correspond to anything visually meaningful.
Nonlinear (reprojection-error) triangulation. Start from the linear solution, then refine it by minimizing the actual pixel distance between where X reprojects and where it was observed, in both images — the same kind of nonlinear least-squares problem, and the same Gauss-Newton machinery, as the earlier Nonlinear Optimization series.
The nonlinear version is slower (it iterates) but minimizes the error you actually care about — how far off the predicted pixel is — and is noticeably more accurate whenever the two camera rays are close to parallel (a narrow baseline) or the noise is large.
Play with it: triangulate a noisy point
Interactive
Two cameras (fixed poses) each get one noisy pixel observation of the same true 3D point. Compare where linear DLT vs. nonlinear (Gauss-Newton on reprojection error) place the reconstructed point.
Drag to orbit, scroll to zoom.
Image 1 observation
Image 2 observation
The actual inputs: ring = where the true point projects, dot = the noisy pixel the triangulators receive. Every estimation error traces back to these two offsets.
The SVD recipe, worked through
Foundations
E = [t]×R is built from a rotation and a translation, so it has to be possible to pull those two pieces back out — and unlike most factorizations in this series, the recipe is short enough to write out in full rather than wave at a textbook. Because [t]× is skew-symmetric with rank 2, E always has two equal nonzero singular values and a third that is exactly zero. (With real, noisy correspondences the two nonzero ones only come out approximately equal — the standard fix is to force them, replacing whatever the raw SVD returns with Σ = diag(1,1,0) before doing anything else.)
W = [[0,−1,0], [1,0,0], [0,0,1]] (a 90° rotation about z — det(W)=1, WWᵀ=I)
W is chosen because UWVᵀ and UΣVᵀ both reuse the same U and V — swapping Σ for a rotation-shaped matrix turns "the rank-2 essential-matrix basis" into "a candidate rotation," in the same basis E was already decomposed in. Plugging U and V into W two different ways (W or Wᵀ), and pairing each with either sign of the translation, produces exactly four candidate poses:
Why u₃ in particular: since E = [t]×R, Eᵀt = Rᵀ[t]×ᵀt = −Rᵀ(t×t) = 0 — so t lives in the left null space of E, exactly the singular vector paired with the zero singular value, i.e. u₃. That's the same vector that shows up as the epipole in image 2 (e₂ᵀE = 0 ⇔ Eᵀe₂ = 0): the translation direction and the epipole direction are the same null space, so recovering one recovers the other, up to the scale and sign E never carried in the first place. All four candidates satisfy the epipolar constraint identically well — E genuinely cannot distinguish between them.
Four candidates, one correspondence: which one is real? That's a separate, much simpler question than the algebra above — cheirality, next.
Cheirality: reconstruct once, check the sign, done
Interactive
Below: a ground-truth two-camera rig, decomposed live via the exact SVD recipe above, all four candidates triangulating the same point simultaneously. Drag the sliders — the test point and the relative rotation between the cameras — and watch which candidate wins update in real time.
Drag to orbit, scroll to zoom. Camera 1 is fixed at the origin.
Midpoint vs. optimal: a third and fourth way to triangulate
Foundations
Section 2 already compared linear DLT to nonlinear reprojection refinement. Two more members of the same family are worth knowing, because they sit at opposite ends of "cheap and biased" vs. "correct and expensive."
Midpoint triangulation. Back-project each pixel into a 3D ray, ray = C + λ·d (camera center C, unit direction d toward the pixel). With noisy pixels the two rays almost never meet, so take the two points — one on each ray — that are mutually closest, and report their midpoint. Solving for the two closest points is a small 2×2 linear system:
t* = (b·e′ − c·d′) / (ac − b²) s* = (a·e′ − b·d′) / (ac − b²)
M = ½[(C₁+t*d₁) + (C₂+s*d₂)]
This is fast, closed-form, and geometrically intuitive — and systematically wrong in a specific way. Worked example: two cameras a unit apart, C₁=(0,0,0), C₂=(1,0,0), true point X=(0.5, 0.3, 4). Camera 1's ray is exact; camera 2's ray is perturbed by a tiny angle (≈ the effect of a couple of pixels of noise at typical focal length) away from the true direction. Solving the system above gives a midpoint M ≈ (0.480, 0.288, 3.843) — error vector M − X ≈ (−0.020, −0.012, −0.157). A single small perturbation on one ray produces an error that is eight times larger along depth (z) than laterally (x, y). That's not a coincidence of this example: minimizing 3D distance to two rays weights every direction in space equally, but the rays themselves are far more sensitive to angular error along the depth axis than across it (a ray sweeps a huge distance in z for a tiny angular nudge, once it's traveled a few units out) — so the point that minimizes 3D ray distance is not the point that minimizes 2D pixel reprojection error. It systematically isn't the maximum-likelihood estimate under pixel noise.
Optimal (Hartley–Sturm) triangulation. Reframe the problem entirely in pixel space, where the noise actually lives. Instead of asking "which 3D point best explains these rays," ask: find corrected points x̂₁, x̂₂ that (a) satisfy the epipolar constraint exactly, x̂₂ᵀFx̂₁ = 0, and (b) are as close as possible to the actually-observed x₁, x₂:
Once x̂₁, x̂₂ exactly satisfy the epipolar constraint, their rays are guaranteed to intersect exactly — so triangulating them is trivial, any linear method gives the identical answer. All the work is in finding the corrected pair, and Hartley & Zisserman show the constrained minimization reduces to a single-variable function of one parameter along the epipolar line, whose minimum is a root of a degree-6 polynomial (solvable in closed form, no iteration, no local minima). The point of the machinery isn't the polynomial's coefficients — it's that this is provably the pixel-space maximum-likelihood point under Gaussian pixel noise, which is exactly the error the midpoint method above ignores.
Drag to orbit. Diamond-open = truth; gray = midpoint; teal = linear DLT; pink = Gauss–Newton; purple = epipolar-consistent optimal.
3D reconstruction error per method, split into depth (z) and lateral (x,y) components. Watch midpoint’s depth bar dwarf its lateral one — the 8:1 bias from the worked example, live.
How good is a triangulated point? The depth-uncertainty ellipsoid
Foundations
Every triangulated point carries uncertainty inherited from pixel noise, and that uncertainty is not the same in every direction — it forms an ellipsoid stretched along the viewing (depth) axis. The small-baseline stereo approximation makes the size of the stretch concrete. For two cameras separated by baseline B, focal length f (pixels), and disparity d (pixels) between the two observations of the same point:
Differentiate with respect to disparity to see how a pixel-sized error in d turns into an error in Z:
⇒ ΔZ ≈ σ · Z² / (fB) for pixel noise σ
Compare that to the lateral (x, y) uncertainty, which comes from a single camera's angular resolution at that depth and has no baseline in it at all: ΔX ≈ σ·Z/f — linear in depth, no benefit from a wider rig. Depth uncertainty, by contrast, is quadratic in Z and only inversely linear in B. Spelled out:
Geometrically, this is the same story from the other direction: as a point moves farther from the rig (or the baseline shrinks), the angle between the two rays converging on it narrows, and a fixed pixel-sized wobble in either ray sweeps a much larger span along the now-nearly-parallel viewing direction than across it — exactly the elongated-ellipsoid picture, and exactly why the worked midpoint example above showed an 8:1 depth-to-lateral error ratio.
Play with it: the uncertainty ellipse
Interactive
Two cameras B apart, both looking at a point at depth Z on their shared perpendicular bisector (keeping the point centered isolates the depth-vs-baseline effect — an off-center point tilts the ellipse instead of just stretching it). The ellipse is drawn 20× actual scale for visibility; the numbers are the true, unscaled values.
Depth axis is vertical; baseline axis is horizontal. Faint dots = 60 repeated triangulations of noisy observations (Monte-Carlo); the ellipse is the analytic prediction.
Scale ambiguity: the metric you never measured
Foundations
Trace the scale of every quantity on this page back to its source. E is defined only up to an overall scale — it's built from correspondences alone, and E and λE satisfy the identical epipolar constraint for any λ ≠ 0. The SVD recipe above hands back a unit-length t by construction. Every triangulated point inherits that same unknown scale. Nowhere in this entire pipeline — Part 6's essential matrix, this page's pose recovery, this page's triangulation — does an actual physical measurement (meters, millimeters, anything) ever enter. Two uncalibrated-scale images simply cannot tell you whether you're looking at a dollhouse or a house.
The size of that ambiguity depends on exactly one thing: whether the cameras are calibrated. A two-view reconstruction from an uncalibrated pair — K unknown, so the fundamental matrix is all you have — is defined only up to a general projective transformation H of 3D space, the full 15-DOF group of P³ (an arbitrary invertible 4×4 matrix, modulo overall scale). Applying any such H to every camera and every 3D point leaves every image measurement bit-for-bit unchanged: send X → HX and P → PH⁻¹ and the projection becomes P′(HX) = PH⁻¹HX = PX — literally the same number, for any H. An uncalibrated reconstruction is therefore only projectively correct: straight lines stay straight and cross-ratios survive, but angles, length ratios, and parallelism do not.
Knowing K removes the projective (and affine) part of that freedom. Calibration is what lets you promote F to the essential matrix E = KᵀFK, whose decomposition carries a genuine rotation and a metric translation direction; fold K in and only a similarity transform remains — 7 DOF (3 rotation + 3 translation + 1 uniform scale), i.e. a Euclidean reconstruction up to a single unknown scale. The similarity group is the subgroup of H that additionally preserves angles and length ratios, which is precisely the extra structure calibration buys you. So: uncalibrated pair → projective H, 15 DOF; calibrated pair → similarity (sR, t), 7 DOF.
The intermediate step of that promotion has a name worth keeping. Recovering the plane at infinity (equivalently, the absolute dual quadric, or the image of the absolute conic) is what upgrades a projective reconstruction to an affine one; imposing the absolute conic’s constraints on that object is what then upgrades the affine reconstruction to metric. Calling this the “metric upgrade” is exact — it is not more triangulation, it is the extra geometric object that calibration pins down.
In the calibrated case, what you get instead is a reconstruction correct up to a similarity transform — a global rotation, a global translation, and a single global scale factor applied to every camera and every point simultaneously, with no way to distinguish the true reconstruction from any rescaled copy of it:
That's 7 degrees of freedom (3 rotation + 3 translation + 1 scale) that no amount of additional two-view geometry can pin down — a gauge freedom, not a solvable unknown. The only way out is an external reference: a known physical size somewhere in the scene, a second sensor with its own metric ground truth (IMU, GPS, LIDAR), or a stereo rig with a physically-measured baseline (which is exactly why Section 7's B has to be a real, known number — that's where the metric scale that monocular two-view geometry can't supply comes from). This same 7-DOF freedom reappears, unchanged, as the gauge freedom that has to be fixed (usually by anchoring one camera and one distance) in bundle adjustment when reconstructing from many views in Structure from Motion.
Drag to orbit. The whole reconstruction (cube + cameras) is scaled by the slider about the origin.
Orthographic 2D view; the framing fits both the ghost and the transformed scene, so any motion you see is real. Ghost = original reconstruction, solid = transformed (and its cameras).
Cheat sheet
Recap
| Step | Input | Output | Ambiguity / cost |
|---|---|---|---|
| Decompose E | Essential matrix | 4 candidate (R,t) poses | Resolved by the cheirality (positive-depth) check |
| Linear triangulation | Both poses + a correspondence | 3D point, one linear solve | Minimizes algebraic error, not pixel error |
| Nonlinear triangulation | Linear estimate as a starting point | Refined 3D point | Minimizes true reprojection error via Gauss-Newton |
| SVD recipe | E = UΣVᵀ, W = [[0,−1,0],[1,0,0],[0,0,1]] | R∈{UWVᵀ,UWᵀVᵀ}, t=±u₃ | u₃ = U's 3rd column (null vector of Eᵀ, since Eᵀt=0) |
| Cheirality test | 1 correspondence, all 4 candidates | The correct (R,t) | Keep the candidate with positive depth in both cameras |
| Midpoint triangulation | Two back-projected rays | Midpoint of common perpendicular | Minimizes 3D ray distance, not pixel error — biased along depth |
| Optimal (Hartley–Sturm) triangulation | F + noisy x₁, x₂ | Corrected x̂₁, x̂₂ on the epipolar constraint | ML-optimal under Gaussian pixel noise; degree-6 polynomial, closed-form |
| Depth uncertainty | Pixel noise σ, focal f, baseline B, depth Z | ΔZ ≈ σ·Z²/(fB) | Quadratic in distance, only linear (inverse) in baseline |
| Scale ambiguity | Two views; uncalibrated vs. calibrated (K known) | Reconstruction up to H, or up to a similarity once K is known | Uncalibrated: 15-DOF projective H in P³. Calibrated: 7-DOF similarity gauge (3 rot + 3 trans + 1 scale). Both need an external metric reference |