Homography
Epipolar geometry (Part 6) narrows a match down to a line but never to a point — depth is still unknown. There's one important special case where that ambiguity disappears entirely: when the scene, or the camera motion, is flat. Then one image can predict the other exactly, pixel for pixel, with a single 3×3 matrix — the homography.
When a plane makes the world flat again
Setup
Every point on a flat surface — a wall, a floor, a book cover, a whiteboard — shares one property: it lies in a single 3D plane. Restricting attention to points on that plane removes exactly the degree of freedom (depth "off the plane") that made epipolar geometry only a line-constraint. The result is a direct linear map between the two images of that plane, with no depth left as an unknown.
The homography matrix
Foundations
For a plane with unit normal n (in camera 1's frame) at distance d from camera 1's center, and the relative pose (R, t) from camera 1 to camera 2, the induced homography (in normalized/calibrated coordinates) is:
and every point X on that plane satisfies, exactly (no approximation, no residual):
In the pure-rotation case (t = 0), this collapses to H = K₂ R K₁⁻¹ — no plane, no depth, no t at all; it works for every point in the scene, which is exactly why panorama-stitching software can align photos taken by pivoting a camera in place.
Given four or more point correspondences (x₁ᵕ, x₂ᵕ) known to lie on the same plane, H's 8 free parameters (9 entries, minus 1 for scale) can be recovered by solving a linear system — the Direct Linear Transform (DLT). Each correspondence contributes two linear equations in H's entries; four points give exactly eight equations for eight unknowns.
Play with it: warp a face of the cube
Interactive
Treat the cube's front face as a book cover lying in a plane. Camera 1's view of that face (left) is fit to camera 2's view (right) using only its four corners — the dot in the middle is a fifth, held-out point on the same face. Toggle "off-plane point" to add a point that isn't on the plane (a back corner of the cube) and watch the homography's prediction for it (hollow) drift away from where that point actually lands (filled) as you increase the baseline.
Drag to orbit. Shaded face = the plane the homography assumes.
Image 1 (camera 1)
Image 2 — actual vs. H-predicted
Estimating H: the DLT
Estimation
Everything so far assumed H was handed to you. In practice you only have pixel correspondences (x₁ᵕ, x₂ᵕ) from feature matching, and need to recover H from them — the same problem the eight-point algorithm solves for F, just with a different constraint. x₂ is only equal to Hx₁ up to scale (they're homogeneous coordinates), so the usable equation isn't x̃₂ = Hx̃₁ directly — it's that x̃₂ and Hx̃₁ point the same direction, i.e. their cross product vanishes:
Writing that out for a single correspondence x̃₁=(x,y,1), x̃₂=(x′,y′,1) gives three equations in the 9 unknown entries of H (stacked into a vector h = vec(H)), only two of which are independent (the third is a linear combination of the first two once you eliminate the shared scale) — the standard DLT row pair:
Four correspondences stack into an 8×9 matrix A, and H (8 free parameters — 9 entries minus 1 for overall scale) is the vector h with Ah = 0. This is exactly the same homogeneous least-squares pattern used to recover the fundamental matrix from correspondences: build A, then find its null vector — the right singular vector of A with the smallest singular value, equivalently the eigenvector of AᵀA with the smallest eigenvalue. With more than 4 correspondences the system is overdetermined (no exact null vector), and the same solve — same A, more rows, smallest-eigenvalue eigenvector of AᵀA — becomes the least-squares fit that minimizes algebraic error over all the noisy matches at once.
Image 1 — fixed correspondences
Image 2 — drag the matching points
The warped grid is the source checkerboard pushed through the live-fit H; it should pass exactly through every filled marker.
How good is the fit? Three error metrics
Evaluation
The DLT above minimizes algebraic error — how far Ah is from 0, which has no direct pixel-space meaning and weights correspondences unevenly. Fitting or scoring a homography for real (e.g. inside RANSAC, or as a final bundle-adjustment cost) uses one of three geometric alternatives instead:
Symmetric transfer error pushes each point through H forward and its match back through H⁻¹, and sums both squared pixel distances. Unlike F — where x₂ᵀFx₁=0 is already symmetric in the two images — H maps one image onto the other, so a one-directional error silently rewards a degenerate H that maps everything to a single point in image 2 (zero forward error) while doing arbitrarily badly in reverse; summing both directions closes that loophole.
Reprojection error is the "real" cost used when jointly refining H and the correspondences themselves: instead of trusting the measured, noisy (x₁, x₂) as ground truth, it optimizes over both H and a set of "true" plane points x̂₁, x̂₂ satisfying x̂₂=Hx̂₁ exactly, and minimizes ‖x₁−x̂₁‖² + ‖x₂−x̂₂‖² — the distance from each observation to its nearest point actually consistent with some homography. It's the statistically correct (maximum-likelihood, under Gaussian pixel noise) cost, and the most expensive: it adds two new unknowns per correspondence.
Sampson error is the cheap first-order approximation to reprojection error, the same trick used for F's Sampson distance: instead of solving the nonlinear correction, linearize the constraint g(x₁,x₂) = x̃₂×Hx̃₁ around the measured points and divide its magnitude by the local sensitivity (the Jacobian norm) to convert an algebraic residual into an approximate geometric one — one evaluation of H and a small Jacobian, no extra unknowns, and close to reprojection error when the residual is already small (which is the only regime it's meant for).
Image 1 — x₁ (filled) vs H⁻¹x₂ (ring)
Image 2 — x₂ (filled) vs Hx₁ (ring)
Segments = the two one-directional transfer errors whose squares sum to the symmetric transfer error. H is the DLT fit on the noisy points (Hartley-normalized).
Decomposing H back into pose
Recovering (R, t, n) from a fitted H
A fitted H is 8 numbers with no obvious physical meaning. But §1 showed exactly what those numbers are built from — for calibrated cameras, x̃₂ ≅ K₂Hx̃₁ reduces to a plain rotation-plus-plane relation once the intrinsics are divided out. Writing the rigid motion as X₂ = RX₁ − t (the same (R,t) convention as §1) and substituting the plane constraint nᵀX₁=d as 1 = nᵀX₁/d:
— exactly the calibrated H from §1, and for uncalibrated pixel coordinates, the full homography is that same matrix sandwiched between the two intrinsics:
Going the other way — given a fitted (uncalibrated) H and known K₁, K₂ — first strip the intrinsics to recover the calibrated Ĥ = K₂⁻¹HK₁, then decompose Ĥ into (R, t/d, n). The standard method (Faugeras & Lustman 1988; Zhang 1998) runs on the SVD of Ĥ: with singular values σ₁≥σ₂≥σ₃, the middle one is always exactly σ₂=1 for a genuine calibrated homography (a real check on whether an estimated H is even consistent with a rigid motion), and the two extreme singular vectors combine — with a case-splitting sign choice — into a closed-form R, t/d, and n.
This is exactly how planar visual SLAM initializes when the scene starts out flat — ORB-SLAM, for instance, fits both a homography and an essential matrix to the first two frames in parallel and picks whichever model explains the data better; if it's the homography (near-planar scene, or a pure-rotation start), decomposing H — with the same two-solution ambiguity above — is how the initial map and pose come out.
Gray = true plane & camera 1. Colored arrows = candidate plane normals, drawn from the true plane's center.
What homographies are actually used for
Applications
Panorama stitching. Successive frames from a camera that only rotates (no translation — the case that makes E degenerate because the baseline is zero) are related by exactly the pure-rotation homography from §1, H=K RK⁻¹: no depth, no baseline, one matrix per frame pair. Stitching software estimates that H between overlapping frames, warps each onto a shared cylindrical or spherical surface (a flat mosaic stretches area without bound away from the image center — a sphere doesn't), then blends the seams (feathering or multi-band blending) and corrects exposure differences between shots so the seam doesn't show as a brightness jump.
Planar AR-marker pose (ArUco / AprilTag). A printed marker of known real-world size defines a plane with a known local coordinate frame. Detect its four corners in the image, DLT-fit the homography from the marker's canonical corners to the observed ones, then run exactly the decomposition above (using the camera's known K) to get the marker's (R,t) relative to the camera directly — this is the whole pose-estimation step in ArUco and AprilTag, correspondence-to-pose with nothing else involved.
Inverse perspective mapping (bird's-eye view). A forward-facing road camera sees the ground plane in perspective — lane markings converge toward a vanishing point. Given the homography that maps that ground plane onto a top-down virtual camera, warping the whole image through it turns the road into a rectangle, which is what parking-assist and lane-detection systems actually run their geometry on.
Photo 1 (yaw 0)
Photo 2 (rotated camera)
Photo 2 warped into photo 1's frame
Right canvas: photo 1 in its original pixels (left half) cross-faded into the H-warped photo 2 (right half) — the checkerboard and the colored disks must line up across the fade with no ghosting, because pure rotation has no parallax.
Oblique camera view
Bird's-eye warp (ground-plane H)
The checkerboard road is warped correctly; the box (which stands above the ground plane) smears, because the single ground-plane H is only exact for z=0.
Cheat sheet
Recap
| Case | Homography | Valid for |
|---|---|---|
| Planar scene | H = R − tnᵀ/d | Only points on that one plane |
| Pure rotation (any scene) | H = K₂RK₁⁻¹ | Every point — no plane or depth needed |
| Estimating H | DLT from ≥4 correspondences | A linear solve, no calibration required |
| DLT row pair | [−x,−y,−1,0,0,0,x′x,x′y,x′]h=0 & the y′ row | Stack ≥4 pts → null vector of AᵀA (exact if 4, least-squares if more) |
| Uncalibrated H | H = K₂(R − tnᵀ/d)K₁⁻¹ | Strip intrinsics first (Ĥ=K₂⁻¹HK₁) before decomposing |
| Decomposing H | SVD of Ĥ → 4 raw (R, t/d, n) | Cheirality kills the backwards twins → 2 physically valid candidates remain |
| Fit quality | Symmetric transfer, reprojection, Sampson error | Reprojection = ground truth; Sampson = its cheap first-order stand-in |
| Panorama stitching | H = K R K⁻¹ (pure rotation) | Warp to cylinder/sphere, then blend seams & exposure |
| AR marker pose | DLT-fit H → decompose with known K | Marker size + K → (R,t) directly (ArUco/AprilTag) |
| Bird's-eye view (IPM) | Warp by the ground-plane H | Exact only on z=0; anything with height smears |