Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

When a plane makes the world flat again

Setup

Every point on a flat surface — a wall, a floor, a book cover, a whiteboard — shares one property: it lies in a single 3D plane. Restricting attention to points on that plane removes exactly the degree of freedom (depth "off the plane") that made epipolar geometry only a line-constraint. The result is a direct linear map between the two images of that plane, with no depth left as an unknown.

💡 The other special case: even for a non-planar scene, if the camera only rotates in place (no translation — think of pivoting a phone camera to shoot a panorama), every point in the world behaves as if it were on a plane at infinity. Same math, same matrix.
1

The homography matrix

Foundations

For a plane with unit normal n (in camera 1's frame) at distance d from camera 1's center, and the relative pose (R, t) from camera 1 to camera 2, the induced homography (in normalized/calibrated coordinates) is:

H = R − t nᵀ / d

and every point X on that plane satisfies, exactly (no approximation, no residual):

x₂ ≅ H x₁    (≅ = equal up to overall scale, since these are homogeneous coordinates)

In the pure-rotation case (t = 0), this collapses to H = K₂ R K₁⁻¹ — no plane, no depth, no t at all; it works for every point in the scene, which is exactly why panorama-stitching software can align photos taken by pivoting a camera in place.

Given four or more point correspondences (x₁ᵕ, x₂ᵕ) known to lie on the same plane, H's 8 free parameters (9 entries, minus 1 for scale) can be recovered by solving a linear system — the Direct Linear Transform (DLT). Each correspondence contributes two linear equations in H's entries; four points give exactly eight equations for eight unknowns.

2

Play with it: warp a face of the cube

Interactive

🎯 Learning goal: a homography fit from four corners of a plane predicts every other point on that same plane exactly — but says nothing reliable about points that aren't on it.

Treat the cube's front face as a book cover lying in a plane. Camera 1's view of that face (left) is fit to camera 2's view (right) using only its four corners — the dot in the middle is a fifth, held-out point on the same face. Toggle "off-plane point" to add a point that isn't on the plane (a back corner of the cube) and watch the homography's prediction for it (hollow) drift away from where that point actually lands (filled) as you increase the baseline.

Drag to orbit. Shaded face = the plane the homography assumes.

Image 1 (camera 1)

Image 2 — actual vs. H-predicted

actual projection H-predicted
⚠️ Weakness: a homography estimated from one plane is silently wrong for anything not on that plane — the error (called parallax) grows with how far a point is off the plane and how wide the baseline between cameras is. This is also exactly why phone panorama stitching works great for distant scenery (nearly flat / nearly pure rotation) and falls apart on close-up scenes with real depth variation.
3

Estimating H: the DLT

Estimation

Everything so far assumed H was handed to you. In practice you only have pixel correspondences (x₁ᵕ, x₂ᵕ) from feature matching, and need to recover H from them — the same problem the eight-point algorithm solves for F, just with a different constraint. x₂ is only equal to Hx₁ up to scale (they're homogeneous coordinates), so the usable equation isn't x̃₂ = Hx̃₁ directly — it's that x̃₂ and Hx̃₁ point the same direction, i.e. their cross product vanishes:

x̃₂ × (H x̃₁) = 0

Writing that out for a single correspondence x̃₁=(x,y,1), x̃₂=(x′,y′,1) gives three equations in the 9 unknown entries of H (stacked into a vector h = vec(H)), only two of which are independent (the third is a linear combination of the first two once you eliminate the shared scale) — the standard DLT row pair:

[ −x, −y, −1,  0, 0, 0,  x′x, x′y, x′ ] h = 0 [ 0, 0, 0,  −x, −y, −1,  y′x, y′y, y′ ] h = 0

Four correspondences stack into an 8×9 matrix A, and H (8 free parameters — 9 entries minus 1 for overall scale) is the vector h with Ah = 0. This is exactly the same homogeneous least-squares pattern used to recover the fundamental matrix from correspondences: build A, then find its null vector — the right singular vector of A with the smallest singular value, equivalently the eigenvector of AᵀA with the smallest eigenvalue. With more than 4 correspondences the system is overdetermined (no exact null vector), and the same solve — same A, more rows, smallest-eigenvalue eigenvector of AᵀA — becomes the least-squares fit that minimizes algebraic error over all the noisy matches at once.

⚠ The same conditioning problem, again: run this directly on raw pixel coordinates and it's numerically unstable for exactly the reason unnormalized eight-point is — A mixes products of pixel values (hundreds to thousands) with the constant −1 in the same rows, so AᵀA is badly scaled and the smallest-eigenvalue direction is dominated by rounding error rather than geometry. The fix is the same Hartley normalization from the eight-point page: translate each image's points to centroid 0, scale so the average distance from the origin is √2, solve for the normalized Ĥ, then denormalize H = T₂⁻¹ĤT₁.
🎯 Learning goal: watch the DLT null-space solve run on live correspondences — drag the destination markers and see A, its smallest eigenvalue, and the resulting warp update in real time. Add a fifth correspondence and the exact 4-point solve becomes an overdetermined least-squares fit, using the identical code.

Image 1 — fixed correspondences

Image 2 — drag the matching points

The warped grid is the source checkerboard pushed through the live-fit H; it should pass exactly through every filled marker.

4

How good is the fit? Three error metrics

Evaluation

The DLT above minimizes algebraic error — how far Ah is from 0, which has no direct pixel-space meaning and weights correspondences unevenly. Fitting or scoring a homography for real (e.g. inside RANSAC, or as a final bundle-adjustment cost) uses one of three geometric alternatives instead:

Symmetric transfer error  =  ‖x̃₂ − Hx̃₁‖² + ‖x̃₁ − H⁻¹x̃₂‖²

Symmetric transfer error pushes each point through H forward and its match back through H⁻¹, and sums both squared pixel distances. Unlike F — where x₂ᵀFx₁=0 is already symmetric in the two images — H maps one image onto the other, so a one-directional error silently rewards a degenerate H that maps everything to a single point in image 2 (zero forward error) while doing arbitrarily badly in reverse; summing both directions closes that loophole.

Reprojection error is the "real" cost used when jointly refining H and the correspondences themselves: instead of trusting the measured, noisy (x₁, x₂) as ground truth, it optimizes over both H and a set of "true" plane points x̂₁, x̂₂ satisfying x̂₂=Hx̂₁ exactly, and minimizes ‖x₁−x̂₁‖² + ‖x₂−x̂₂‖² — the distance from each observation to its nearest point actually consistent with some homography. It's the statistically correct (maximum-likelihood, under Gaussian pixel noise) cost, and the most expensive: it adds two new unknowns per correspondence.

Sampson error is the cheap first-order approximation to reprojection error, the same trick used for F's Sampson distance: instead of solving the nonlinear correction, linearize the constraint g(x₁,x₂) = x̃₂×Hx̃₁ around the measured points and divide its magnitude by the local sensitivity (the Jacobian norm) to convert an algebraic residual into an approximate geometric one — one evaluation of H and a small Jacobian, no extra unknowns, and close to reprojection error when the residual is already small (which is the only regime it's meant for).

🎯 Learning goal: the three costs computed on the same fitted H and the same noisy correspondences. Symmetric transfer is what you can draw as segments; Sampson tracks it closely at small residuals; reprojection (small Gauss–Newton fit of each point pair) is the statistically right number the other two are approximating. Drag σ up and watch all three grow — and watch the DLT's own algebraic fit get left behind.

Image 1 — x₁ (filled) vs H⁻¹x₂ (ring)

Image 2 — x₂ (filled) vs Hx₁ (ring)

Segments = the two one-directional transfer errors whose squares sum to the symmetric transfer error. H is the DLT fit on the noisy points (Hartley-normalized).

5

Decomposing H back into pose

Recovering (R, t, n) from a fitted H

A fitted H is 8 numbers with no obvious physical meaning. But §1 showed exactly what those numbers are built from — for calibrated cameras, x̃₂ ≅ K₂Hx̃₁ reduces to a plain rotation-plus-plane relation once the intrinsics are divided out. Writing the rigid motion as X₂ = RX₁ − t (the same (R,t) convention as §1) and substituting the plane constraint nᵀX₁=d as 1 = nᵀX₁/d:

X₂ = RX₁ − t·(nᵀX₁/d) = (R − t nᵀ/d) X₁

— exactly the calibrated H from §1, and for uncalibrated pixel coordinates, the full homography is that same matrix sandwiched between the two intrinsics:

H = K₂ (R − t nᵀ/d) K₁⁻¹

Going the other way — given a fitted (uncalibrated) H and known K₁, K₂ — first strip the intrinsics to recover the calibrated Ĥ = K₂⁻¹HK₁, then decompose Ĥ into (R, t/d, n). The standard method (Faugeras & Lustman 1988; Zhang 1998) runs on the SVD of Ĥ: with singular values σ₁≥σ₂≥σ₃, the middle one is always exactly σ₂=1 for a genuine calibrated homography (a real check on whether an estimated H is even consistent with a rigid motion), and the two extreme singular vectors combine — with a case-splitting sign choice — into a closed-form R, t/d, and n.

⚠ Two solutions, not one: that sign choice isn't cosmetic — the algebra produces four raw candidates, and they come in two mirror-image pairs: (R, t/d, n) and its twin (R, −t/d, −n), both of which satisfy Ĥ = R − (t/d)nᵀ exactly (flipping the sign of both t/d and n leaves their outer product unchanged). Within each twin pair, cheirality — the identical positive-depth check used to prune E's four candidate poses down to one during essential-matrix decomposition — kills the physically backwards twin: the correct one is the one where the plane's normal faces toward the camera it was measured from, so every observed point ends up in front of both cameras. That removes exactly two of the four raw candidates, but — unlike E — it does not get you down to one: it leaves two algebraically distinct, both-cameras-positive-depth solutions standing. Nothing in a single homography can tell them apart; that needs extra information — a known vertical direction, a non-planar point, or a second frame.

This is exactly how planar visual SLAM initializes when the scene starts out flat — ORB-SLAM, for instance, fits both a homography and an essential matrix to the first two frames in parallel and picks whichever model explains the data better; if it's the homography (near-planar scene, or a pure-rotation start), decomposing H — with the same two-solution ambiguity above — is how the initial map and pose come out.

🎯 Learning goal: construct H from a known ground-truth (R, t, n, d), decompose it back, and watch cheirality narrow four raw candidates to the two that survive — one of which is the ground truth and one of which is an equally valid impostor.

Gray = true plane & camera 1. Colored arrows = candidate plane normals, drawn from the true plane's center.

true (R, t/d, n) surviving candidate A surviving candidate B
⚠ What decomposition can't give back: only the ratio t/d comes out, never t and d separately — the same monocular scale ambiguity as essential-matrix pose recovery. A homography fit from image measurements alone can't tell a nearby plane and a small motion from a distant plane and a large one.
6

What homographies are actually used for

Applications

Panorama stitching. Successive frames from a camera that only rotates (no translation — the case that makes E degenerate because the baseline is zero) are related by exactly the pure-rotation homography from §1, H=K RK⁻¹: no depth, no baseline, one matrix per frame pair. Stitching software estimates that H between overlapping frames, warps each onto a shared cylindrical or spherical surface (a flat mosaic stretches area without bound away from the image center — a sphere doesn't), then blends the seams (feathering or multi-band blending) and corrects exposure differences between shots so the seam doesn't show as a brightness jump.

Planar AR-marker pose (ArUco / AprilTag). A printed marker of known real-world size defines a plane with a known local coordinate frame. Detect its four corners in the image, DLT-fit the homography from the marker's canonical corners to the observed ones, then run exactly the decomposition above (using the camera's known K) to get the marker's (R,t) relative to the camera directly — this is the whole pose-estimation step in ArUco and AprilTag, correspondence-to-pose with nothing else involved.

Inverse perspective mapping (bird's-eye view). A forward-facing road camera sees the ground plane in perspective — lane markings converge toward a vanishing point. Given the homography that maps that ground plane onto a top-down virtual camera, warping the whole image through it turns the road into a rectangle, which is what parking-assist and lane-detection systems actually run their geometry on.

⚠ Only the ground plane is real: IPM's homography is fit assuming everything in the image lies on the ground — true for lane paint and flat pavement, false for anything with height. A curb, a parked car, another vehicle's body: none of it is on the assumed plane, so IPM smears and distorts it exactly the way §2's off-plane point drifted away from its prediction. The demo below makes that failure visible on purpose.
🎯 Learning goal: the panorama case. The camera only rotates (zero baseline — the exact case that kills E), so image 2 is a pure-rotation homography of image 1, valid for every depth. Warping image 2 into image 1's frame through H = K R K⁻¹ aligns the two photos pixel-perfectly in the overlap — drag the rotation slider and watch the blended overlay stay seamless.

Photo 1 (yaw 0)

Photo 2 (rotated camera)

Photo 2 warped into photo 1's frame

Right canvas: photo 1 in its original pixels (left half) cross-faded into the H-warped photo 2 (right half) — the checkerboard and the colored disks must line up across the fade with no ghosting, because pure rotation has no parallax.

Oblique camera view

Bird's-eye warp (ground-plane H)

The checkerboard road is warped correctly; the box (which stands above the ground plane) smears, because the single ground-plane H is only exact for z=0.

✓

Cheat sheet

Recap

CaseHomographyValid for
Planar sceneH = R − tnᵀ/dOnly points on that one plane
Pure rotation (any scene)H = K₂RK₁⁻¹Every point — no plane or depth needed
Estimating HDLT from ≥4 correspondencesA linear solve, no calibration required
DLT row pair[−x,−y,−1,0,0,0,x′x,x′y,x′]h=0 & the y′ rowStack ≥4 pts → null vector of AᵀA (exact if 4, least-squares if more)
Uncalibrated HH = K₂(R − tnᵀ/d)K₁⁻¹Strip intrinsics first (Ĥ=K₂⁻¹HK₁) before decomposing
Decomposing HSVD of Ĥ → 4 raw (R, t/d, n)Cheirality kills the backwards twins → 2 physically valid candidates remain
Fit qualitySymmetric transfer, reprojection, Sampson errorReprojection = ground truth; Sampson = its cheap first-order stand-in
Panorama stitchingH = K R K⁻¹ (pure rotation)Warp to cylinder/sphere, then blend seams & exposure
AR marker poseDLT-fit H → decompose with known KMarker size + K → (R,t) directly (ArUco/AprilTag)
Bird's-eye view (IPM)Warp by the ground-plane HExact only on z=0; anything with height smears
For a general (non-planar) scene, epipolar geometry alone can't recover 3D structure — you need to actually intersect rays. That's triangulation. Continue: triangulation & pose recovery →