Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

Two problems, in order

Setup

Reconstructing a scene from two images is really two separate problems, solved in order: (1) figure out where camera 2 is relative to camera 1 (its relative pose), then (2) given both poses, work out where each matched point actually is in 3D. Part 6 already got us the essential matrix E from correspondences; this page turns E into an actual pose, and then turns rays into points.

1

From E to (R, t): four candidates, one makes sense

Foundations

The essential matrix E = [t]×R can be factored back into a rotation and a translation direction via its singular value decomposition, E = UΣVᵀ. The algebra always produces exactly four candidate poses — two possible rotations, each paired with +t or −t — and all four satisfy the epipolar constraint equally well, because E only fixes t up to an unknown overall scale and sign. (The full worked recipe — the W matrix and all — is in Section 4 below, not deferred to a textbook.) A minimal calibrated solver for E — the five-point algorithm, which needs only five correspondences instead of eight — is treated in Part 13.

(R₁,t), (R₁,−t), (R₂,t), (R₂,−t)  —  only one is physically valid

Only one of the four puts a triangulated point in front of both cameras (positive depth in both) — the other three put it behind one or both. That test is called the cheirality check: triangulate a single correspondence with each candidate and keep the one where both depths come out positive. This is also exactly why triangulation (next) and pose recovery are tangled together in practice — you need a candidate pose to triangulate, and triangulating is how you rule the bad candidates out.

2

Two ways to triangulate

Foundations

Once both camera poses are known, a correspondence (x₁, x₂) defines two rays in 3D that, with perfect noise-free data, cross at exactly one point. With real (noisy) pixel measurements, the rays almost never exactly intersect — so "triangulation" really means finding the point that best explains both observations, and there are two standard ways to define "best":

Linear (DLT) triangulation. Stack the two cameras' projection equations into one linear system and solve it directly — fast, no iteration, but it's minimizing an algebraic error that doesn't correspond to anything visually meaningful.

x₁×(P₁X) = 0, x₂×(P₂X) = 0  →  linear system AX = 0, solved directly

Nonlinear (reprojection-error) triangulation. Start from the linear solution, then refine it by minimizing the actual pixel distance between where X reprojects and where it was observed, in both images — the same kind of nonlinear least-squares problem, and the same Gauss-Newton machinery, as the earlier Nonlinear Optimization series.

minₓ Σₓ∈{1,2} ‖π(Pₓ, X) − xₓ‖²

The nonlinear version is slower (it iterates) but minimizes the error you actually care about — how far off the predicted pixel is — and is noticeably more accurate whenever the two camera rays are close to parallel (a narrow baseline) or the noise is large.

3

Play with it: triangulate a noisy point

Interactive

🎯 Learning goal: both methods agree almost exactly at low noise. Crank up pixel noise (or narrow the baseline) and watch the linear estimate drift further from the truth than the nonlinear, reprojection-error-minimizing one.

Two cameras (fixed poses) each get one noisy pixel observation of the same true 3D point. Compare where linear DLT vs. nonlinear (Gauss-Newton on reprojection error) place the reconstructed point.

Drag to orbit, scroll to zoom.

true point linear (DLT) nonlinear (refined)

Image 1 observation

Image 2 observation

The actual inputs: ring = where the true point projects, dot = the noisy pixel the triangulators receive. Every estimation error traces back to these two offsets.

4

The SVD recipe, worked through

Foundations

E = [t]×R is built from a rotation and a translation, so it has to be possible to pull those two pieces back out — and unlike most factorizations in this series, the recipe is short enough to write out in full rather than wave at a textbook. Because [t]× is skew-symmetric with rank 2, E always has two equal nonzero singular values and a third that is exactly zero. (With real, noisy correspondences the two nonzero ones only come out approximately equal — the standard fix is to force them, replacing whatever the raw SVD returns with Σ = diag(1,1,0) before doing anything else.)

E = UΣVᵀ    Σ = diag(1,1,0)
W = [[0,−1,0], [1,0,0], [0,0,1]]  (a 90° rotation about z — det(W)=1, WWᵀ=I)

W is chosen because UWVᵀ and UΣVᵀ both reuse the same U and V — swapping Σ for a rotation-shaped matrix turns "the rank-2 essential-matrix basis" into "a candidate rotation," in the same basis E was already decomposed in. Plugging U and V into W two different ways (W or Wᵀ), and pairing each with either sign of the translation, produces exactly four candidate poses:

R ∈ { UWVᵀ, UWᵀVᵀ }    t ∈ { +u₃, −u₃ }    (u₃ = U's third column)

Why u₃ in particular: since E = [t]×R, Eᵀt = Rᵀ[t]×ᵀt = −Rᵀ(t×t) = 0 — so t lives in the left null space of E, exactly the singular vector paired with the zero singular value, i.e. u₃. That's the same vector that shows up as the epipole in image 2 (e₂ᵀE = 0 ⇔ Eᵀe₂ = 0): the translation direction and the epipole direction are the same null space, so recovering one recovers the other, up to the scale and sign E never carried in the first place. All four candidates satisfy the epipolar constraint identically well — E genuinely cannot distinguish between them.

⚠ One bookkeeping detail the formula above glosses over: UWVᵀ is only guaranteed to be a proper rotation (det = +1, not a reflection) if det(U) = det(V) = +1. The SVD doesn't promise that on its own — the fix is to flip the sign of whichever singular vector has the unconstrained sign (the one paired with the zero singular value) until both determinants come out positive.

Four candidates, one correspondence: which one is real? That's a separate, much simpler question than the algebra above — cheirality, next.

5

Cheirality: reconstruct once, check the sign, done

Interactive

🎯 Learning goal: this is not a subtle test. Triangulate one correspondence under each of the four candidates and look at the sign of its depth in both cameras — three candidates put the point behind a camera (physically impossible, since a camera can't photograph what's behind its own lens), and exactly one puts it in front of both. That one is (R, t).

Below: a ground-truth two-camera rig, decomposed live via the exact SVD recipe above, all four candidates triangulating the same point simultaneously. Drag the sliders — the test point and the relative rotation between the cameras — and watch which candidate wins update in real time.

Drag to orbit, scroll to zoom. Camera 1 is fixed at the origin.

camera 1 (fixed) correct candidate impossible candidates
⚠ The candidate translations are unit vectors (that's all four singular-vector decomposition can give you — recall t's scale is unrecoverable from E alone). To make the picture legible, each candidate camera 2 here is drawn at the true baseline distance along its recovered direction — real pipelines don't get to do that, they just keep the unit vector and fix the scale later (Section 9, below).
6

Midpoint vs. optimal: a third and fourth way to triangulate

Foundations

Section 2 already compared linear DLT to nonlinear reprojection refinement. Two more members of the same family are worth knowing, because they sit at opposite ends of "cheap and biased" vs. "correct and expensive."

Midpoint triangulation. Back-project each pixel into a 3D ray, ray = C + λ·d (camera center C, unit direction d toward the pixel). With noisy pixels the two rays almost never meet, so take the two points — one on each ray — that are mutually closest, and report their midpoint. Solving for the two closest points is a small 2×2 linear system:

r = C₁ − C₂,  a = d₁·d₁,  b = d₁·d₂,  c = d₂·d₂,  d′ = d₁·r,  e′ = d₂·r
t* = (b·e′ − c·d′) / (ac − b²)    s* = (a·e′ − b·d′) / (ac − b²)
M = ½[(C₁+t*d₁) + (C₂+s*d₂)]

This is fast, closed-form, and geometrically intuitive — and systematically wrong in a specific way. Worked example: two cameras a unit apart, C₁=(0,0,0), C₂=(1,0,0), true point X=(0.5, 0.3, 4). Camera 1's ray is exact; camera 2's ray is perturbed by a tiny angle (≈ the effect of a couple of pixels of noise at typical focal length) away from the true direction. Solving the system above gives a midpoint M ≈ (0.480, 0.288, 3.843) — error vector M − X ≈ (−0.020, −0.012, −0.157). A single small perturbation on one ray produces an error that is eight times larger along depth (z) than laterally (x, y). That's not a coincidence of this example: minimizing 3D distance to two rays weights every direction in space equally, but the rays themselves are far more sensitive to angular error along the depth axis than across it (a ray sweeps a huge distance in z for a tiny angular nudge, once it's traveled a few units out) — so the point that minimizes 3D ray distance is not the point that minimizes 2D pixel reprojection error. It systematically isn't the maximum-likelihood estimate under pixel noise.

Optimal (Hartley–Sturm) triangulation. Reframe the problem entirely in pixel space, where the noise actually lives. Instead of asking "which 3D point best explains these rays," ask: find corrected points x̂₁, x̂₂ that (a) satisfy the epipolar constraint exactly, x̂₂ᵀFx̂₁ = 0, and (b) are as close as possible to the actually-observed x₁, x₂:

min ‖x̂₁ − x₁‖² + ‖x̂₂ − x₂‖²  subject to  x̂₂ᵀFx̂₁ = 0

Once x̂₁, x̂₂ exactly satisfy the epipolar constraint, their rays are guaranteed to intersect exactly — so triangulating them is trivial, any linear method gives the identical answer. All the work is in finding the corrected pair, and Hartley & Zisserman show the constrained minimization reduces to a single-variable function of one parameter along the epipolar line, whose minimum is a root of a degree-6 polynomial (solvable in closed form, no iteration, no local minima). The point of the machinery isn't the polynomial's coefficients — it's that this is provably the pixel-space maximum-likelihood point under Gaussian pixel noise, which is exactly the error the midpoint method above ignores.

⚠ The Gauss-Newton reprojection refinement from Section 2 minimizes the same reprojection-error objective, and in practice lands extremely close to the Hartley–Sturm answer — but it's a local iterative solve seeded from the linear estimate, not a closed-form global optimum. For two views the gap rarely matters; it matters more once outliers or bad initialization are in the picture.
🎯 Learning goal: all four methods on the same noisy observation. Midpoint is the fastest and visibly biased along depth; DLT and Gauss–Newton agree closely; the epipolar-consistent optimal correction (Hartley–Sturm's objective, minimized numerically here) lands on essentially the same answer as GN — the bar chart makes the depth-bias of midpoint impossible to miss.

Drag to orbit. Diamond-open = truth; gray = midpoint; teal = linear DLT; pink = Gauss–Newton; purple = epipolar-consistent optimal.

3D reconstruction error per method, split into depth (z) and lateral (x,y) components. Watch midpoint’s depth bar dwarf its lateral one — the 8:1 bias from the worked example, live.

7

How good is a triangulated point? The depth-uncertainty ellipsoid

Foundations

Every triangulated point carries uncertainty inherited from pixel noise, and that uncertainty is not the same in every direction — it forms an ellipsoid stretched along the viewing (depth) axis. The small-baseline stereo approximation makes the size of the stretch concrete. For two cameras separated by baseline B, focal length f (pixels), and disparity d (pixels) between the two observations of the same point:

Z = fB / d

Differentiate with respect to disparity to see how a pixel-sized error in d turns into an error in Z:

dZ/dd = −fB / d² = −Z² / (fB)    (substituting d = fB/Z)
⇒  ΔZ ≈ σ · Z² / (fB)    for pixel noise σ

Compare that to the lateral (x, y) uncertainty, which comes from a single camera's angular resolution at that depth and has no baseline in it at all: ΔX ≈ σ·Z/f — linear in depth, no benefit from a wider rig. Depth uncertainty, by contrast, is quadratic in Z and only inversely linear in B. Spelled out:

🎯 The one intuition to keep: doubling your baseline only halves your depth error at a given distance. Doubling your distance quadruples it. Baseline is a linear knob; distance is a quadratic tax. That asymmetry — not sensor noise, not calibration error — is the fundamental reason long-range stereo and triangulation get hard fast, and why active range sensors (LIDAR, radar, ToF) exist for anything beyond a few tens of meters: they measure range directly instead of inferring it from a shrinking intersection angle.

Geometrically, this is the same story from the other direction: as a point moves farther from the rig (or the baseline shrinks), the angle between the two rays converging on it narrows, and a fixed pixel-sized wobble in either ray sweeps a much larger span along the now-nearly-parallel viewing direction than across it — exactly the elongated-ellipsoid picture, and exactly why the worked midpoint example above showed an 8:1 depth-to-lateral error ratio.

8

Play with it: the uncertainty ellipse

Interactive

🎯 Learning goal: narrow the ray-intersection angle — by shrinking the baseline or pushing the point farther away — and watch the ellipse stretch along the depth axis, exactly tracking the ΔZ ≈ σ·Z²/(fB) formula shown numerically alongside.

Two cameras B apart, both looking at a point at depth Z on their shared perpendicular bisector (keeping the point centered isolates the depth-vs-baseline effect — an off-center point tilts the ellipse instead of just stretching it). The ellipse is drawn 20× actual scale for visibility; the numbers are the true, unscaled values.

Depth axis is vertical; baseline axis is horizontal. Faint dots = 60 repeated triangulations of noisy observations (Monte-Carlo); the ellipse is the analytic prediction.

9

Scale ambiguity: the metric you never measured

Foundations

Trace the scale of every quantity on this page back to its source. E is defined only up to an overall scale — it's built from correspondences alone, and E and λE satisfy the identical epipolar constraint for any λ ≠ 0. The SVD recipe above hands back a unit-length t by construction. Every triangulated point inherits that same unknown scale. Nowhere in this entire pipeline — Part 6's essential matrix, this page's pose recovery, this page's triangulation — does an actual physical measurement (meters, millimeters, anything) ever enter. Two uncalibrated-scale images simply cannot tell you whether you're looking at a dollhouse or a house.

The size of that ambiguity depends on exactly one thing: whether the cameras are calibrated. A two-view reconstruction from an uncalibrated pair — K unknown, so the fundamental matrix is all you have — is defined only up to a general projective transformation H of 3D space, the full 15-DOF group of P³ (an arbitrary invertible 4×4 matrix, modulo overall scale). Applying any such H to every camera and every 3D point leaves every image measurement bit-for-bit unchanged: send X → HX and P → PH⁻¹ and the projection becomes P′(HX) = PH⁻¹HX = PX — literally the same number, for any H. An uncalibrated reconstruction is therefore only projectively correct: straight lines stay straight and cross-ratios survive, but angles, length ratios, and parallelism do not.

Knowing K removes the projective (and affine) part of that freedom. Calibration is what lets you promote F to the essential matrix E = KᵀFK, whose decomposition carries a genuine rotation and a metric translation direction; fold K in and only a similarity transform remains — 7 DOF (3 rotation + 3 translation + 1 uniform scale), i.e. a Euclidean reconstruction up to a single unknown scale. The similarity group is the subgroup of H that additionally preserves angles and length ratios, which is precisely the extra structure calibration buys you. So: uncalibrated pair → projective H, 15 DOF; calibrated pair → similarity (sR, t), 7 DOF.

The intermediate step of that promotion has a name worth keeping. Recovering the plane at infinity (equivalently, the absolute dual quadric, or the image of the absolute conic) is what upgrades a projective reconstruction to an affine one; imposing the absolute conic’s constraints on that object is what then upgrades the affine reconstruction to metric. Calling this the “metric upgrade” is exact — it is not more triangulation, it is the extra geometric object that calibration pins down.

In the calibrated case, what you get instead is a reconstruction correct up to a similarity transform — a global rotation, a global translation, and a single global scale factor applied to every camera and every point simultaneously, with no way to distinguish the true reconstruction from any rescaled copy of it:

(R, t, {Xⁱ})  and  (R, s·t, {s·Xⁱ})  are indistinguishable, for any s > 0

That's 7 degrees of freedom (3 rotation + 3 translation + 1 scale) that no amount of additional two-view geometry can pin down — a gauge freedom, not a solvable unknown. The only way out is an external reference: a known physical size somewhere in the scene, a second sensor with its own metric ground truth (IMU, GPS, LIDAR), or a stereo rig with a physically-measured baseline (which is exactly why Section 7's B has to be a real, known number — that's where the metric scale that monocular two-view geometry can't supply comes from). This same 7-DOF freedom reappears, unchanged, as the gauge freedom that has to be fixed (usually by anchoring one camera and one distance) in bundle adjustment when reconstructing from many views in Structure from Motion.

🎯 Learning goal: drag the global scale slider. The reconstruction grows and shrinks like a dollhouse, but the reprojection error — the only thing the images can actually measure — stays exactly at zero for every scale. Nothing in the two images moves. That is what "7-DOF gauge freedom" feels like.

Drag to orbit. The whole reconstruction (cube + cameras) is scaled by the slider about the origin.

🎯 Learning goal: switch the ambiguity group and roll a random transform. Under the projective H the cube shears and its edge lengths stop being equal; under the similarity (sR, t) it only rotates, translates, and uniformly rescales. In both cases the reprojected pixels are untouched: the readout’s max reprojection difference stays at float noise while the 3D coordinates visibly move. That gap between “the images cannot tell” and “the 3D scene clearly changed” is the ambiguity.

Orthographic 2D view; the framing fits both the ghost and the transformed scene, so any motion you see is real. Ghost = original reconstruction, solid = transformed (and its cameras).

original transformed camera 1 camera 2
✓

Cheat sheet

Recap

StepInputOutputAmbiguity / cost
Decompose EEssential matrix4 candidate (R,t) posesResolved by the cheirality (positive-depth) check
Linear triangulationBoth poses + a correspondence3D point, one linear solveMinimizes algebraic error, not pixel error
Nonlinear triangulationLinear estimate as a starting pointRefined 3D pointMinimizes true reprojection error via Gauss-Newton
SVD recipeE = UΣVᵀ, W = [[0,−1,0],[1,0,0],[0,0,1]]R∈{UWVᵀ,UWᵀVᵀ}, t=±u₃u₃ = U's 3rd column (null vector of Eᵀ, since Eᵀt=0)
Cheirality test1 correspondence, all 4 candidatesThe correct (R,t)Keep the candidate with positive depth in both cameras
Midpoint triangulationTwo back-projected raysMidpoint of common perpendicularMinimizes 3D ray distance, not pixel error — biased along depth
Optimal (Hartley–Sturm) triangulationF + noisy x₁, x₂Corrected x̂₁, x̂₂ on the epipolar constraintML-optimal under Gaussian pixel noise; degree-6 polynomial, closed-form
Depth uncertaintyPixel noise σ, focal f, baseline B, depth ZΔZ ≈ σ·Z²/(fB)Quadratic in distance, only linear (inverse) in baseline
Scale ambiguityTwo views; uncalibrated vs. calibrated (K known)Reconstruction up to H, or up to a similarity once K is knownUncalibrated: 15-DOF projective H in P³. Calibrated: 7-DOF similarity gauge (3 rot + 3 trans + 1 scale). Both need an external metric reference
Triangulation assumed you already had two calibrated camera poses. In a real reconstruction, every new camera has to be located against a 3D map that already exists — that's PnP, camera resection. Continue: PnP & camera resection →