Stereo: Rectification & Disparity
Part 6 showed that a point in one image constrains its match in the other to lie on an epipolar line — but that line can point in any direction, and searching a whole image along an arbitrary line, at every candidate depth, is still awkward to make fast. Real stereo systems sidestep the problem: warp both images so every pair of corresponding epipolar lines becomes the same horizontal row in both images, and the search collapses to sliding a small window along one scanline. That trick — rectification — is what turns epipolar geometry into an actual depth sensor, and it's the piece this guide was missing: stereo depth cameras, robot stereo rigs, and every real-time disparity map run on exactly this pipeline. This page derives the rectifying rotation, turns disparity into depth with Z = fB/d, and builds a live block-matching stereo pair so you can watch the whole thing run.
From an arbitrary line to a single row
Motivation
The epipolar constraint from Part 6 says a match must lie somewhere on the line l₂ = Fx₁ — a real improvement over searching the whole image, but that line's slope depends on F, which depends on the two cameras' relative pose. In general it's tilted, so "search along it" means resampling pixels along a diagonal at every candidate position, for every pixel in the image, for every frame. That's expensive and fiddly to vectorize.
But there's nothing sacred about the original pixel grid. Camera rotation is free to choose, and Part 8 will show that rotating a calibrated camera about its own optical center — without moving it — is exactly a homography: a warp of the image that changes nothing about the physical rays it represents, only how they're indexed by pixel coordinates. So: choose new, virtual camera orientations for both cameras such that every epipolar line becomes horizontal, and lands on the same row in both images. Warp the two real images into those virtual cameras. Now the search for a match is one dimension: fix a row, slide a window along it, done.
Constructing the rectifying rotation
Derivation
Start from a calibrated stereo pair: two cameras with known rotations R₁, R₂ (world → camera) and known relative pose between them, summarized by the baseline vector t pointing from camera 1's center to camera 2's center, expressed in world coordinates. The goal is a single new rotation Rrect that both virtual cameras will share — same orientation, so their image planes are parallel and their rows already line up — chosen so that direction also happens to send the epipole to infinity along the horizontal axis.
An epipole is the projection of one camera's center into the other camera's image, so all epipolar lines fan out from it; if it sits at infinity in a fixed horizontal direction, every epipolar line is parallel to that direction — i.e. horizontal. The standard construction picks Rrect's three rows to be an orthonormal basis built directly from t:
new y‑axis: e₂ = (0,0,1) × e₁, normalized (perpendicular to baseline, roughly “up” — assumes t isn’t vertical)
new z‑axis: e₃ = e₁ × e₂ (completes a right‑handed frame, points “forward”)
Rrect = [ e₁ᵀ ; e₂ᵀ ; e₃ᵀ ] (stack as rows)
Here's why this specific choice sends the epipole to infinity. Camera 2's center, seen from camera 1's new orientation, is displaced by t; expressed in the new frame that displacement is Rrect·t. Because e₁ = t/‖t‖ by construction, and e₂, e₃ are both orthogonal to t by construction:
All of the displacement lands on the new x-axis; none of it on y or z. The epipole — the projection of that displaced center — sits infinitely far along x with zero vertical component, i.e. exactly at the ideal point in the horizontal direction (Part 0's w = 0 points). A line through the epipole with zero vertical offset from any point is, by definition, horizontal. That's the whole derivation: the construction doesn't just happen to work, it's reverse-engineered so that t lands entirely on the axis that becomes "sideways" in the new image.
The rectifying rotation applied to each real camera is Rrect·R₁ᵀ and Rrect·R₂ᵀ (composing "undo the camera's own rotation" with "rotate into the shared rectified frame"). Since both new cameras get the identical Rrect, they end up row-aligned by construction — and since t is now purely along the shared x-axis, the two virtual cameras differ from each other by nothing but a pure horizontal translation of length B = ‖t‖, the baseline.
x′ ≅ Knew·Rrect·Riᵀ·Ki⁻¹·x. Rectification in practice (e.g. OpenCV's stereoRectify + remap) is exactly this: two per-image homography warps, computed once from calibration, applied to every incoming frame.Disparity, and why Z = fB / d
Derivation
After rectification, both virtual cameras share a focal length f, sit a baseline B apart along a shared x-axis, and share every row. A 3D point at depth Z and horizontal offset X from the left camera projects, in each camera, by the ordinary pinhole rule u = f·(horizontal coordinate)/Z — the only difference between the two cameras is which point they call the horizontal origin, and that offset is exactly B:
uright = f·(X − B) / Z (the right camera's center sits B further along +x, so the point is B closer to its own origin)
disparity d ≡ uleft − uright = f·X/Z − f·(X−B)/Z = fB / Z
⇒ Z = fB / d
Every quantity on the right of Z = fB/d is either fixed by calibration (f, B) or directly measured from the two images (d, in pixels) — which is the whole appeal: once rectified, recovering depth is arithmetic on a single measured integer (or sub-pixel value) per pixel, no triangulation solve required.
That formula also says something uncomfortable about accuracy. Disparity is measured to some finite precision — call one pixel of disparity noise Δd = 1. How much does that same pixel of noise mean in meters, as a function of how far away the point already is? Differentiate Z(d) = fB/d:
substitute d = fB/Z (so d² = f²B²/Z²):
dZ/dd = −fB · Z²/(f²B²) = −Z² / (fB)
⇒ ΔZ ≈ Z² / (fB) for one pixel of disparity error
Depth error grows with the square of distance and shrinks only linearly with baseline. A stereo rig that's comfortably accurate at 2 m can be off by meters at 20 m — the same one pixel of disparity noise means 100× more depth error, because disparity itself has shrunk 10× and the derivative scales as its inverse square. This is precisely why LiDAR, not stereo, dominates long-range depth, and why wide-baseline rigs (further-apart cameras) trade a larger blind zone near the camera for meaningfully better depth at range.
Play: rectify a stereo pair, then match a pixel
Interactive
The rig is two cameras B apart along a shared baseline, both pitched down slightly to see a checkered floor with one solid sphere floating above it. In Unrectified mode each camera additionally toes inward to fixate a point ahead — a common (bad) way to point two cameras at the same subject — which makes their relative rotation R₂R₁ᵀ ≠ I and tilts the epipolar line. In Rectified mode both cameras share the same orientation exactly as derived above, so t is purely horizontal and the epipolar line is exactly horizontal too.
Left image (click to pick a pixel)
Right image (epipolar line + best match)
Dashed line = epipolar line for the clicked pixel, computed from the live F. In rectified mode it's the search row; the filled dot marks the best SAD match found by sliding a 7×7 window along it.
How disparity actually gets computed
Algorithms
Block matching. For every pixel in the rectified left image, slide a small window (e.g. 7×7 or 9×9) along the same row in the right image and score how well it matches at each candidate column. Three common scores:
SSD (sum of squared differences): ∑ (IL(x,y) − IR(x−d,y))² — penalizes large errors harder
NCC (normalized cross-correlation): ∑(IL−ȳL)(IR−ȳR) / (σLσR) — robust to brightness/contrast differences between the two cameras
Whichever score is used, the disparity that minimizes cost is only known to the nearest integer pixel by default; fitting a parabola to the three cost values around the minimum (d−1, d, d+1) and solving for its vertex gives a cheap sub-pixel refinement, which is what the demo above does.
The cost volume. Doing this for every pixel and every candidate disparity builds a 3D array — cost[row][col][d] — the object every stereo algorithm, however sophisticated, is ultimately searching over. Naive block matching just takes the per-pixel minimum along the d axis independently for each pixel, which is fast but noisy: textureless regions (a blank wall) have a flat, ambiguous cost curve, and the per-pixel minimum can jump around randomly from one pixel to its neighbor.
SGBM (semi-global block matching). Instead of picking each pixel's minimum in isolation, SGBM aggregates cost along several 1D paths through the cost volume (typically 8 or 16 directions radiating from each pixel), adding a penalty whenever the disparity jumps between neighboring pixels along a path — a small penalty for a 1-pixel jump (edges are allowed to have small disparity steps), a much larger one for jumps of more than 1 (discourages spurious noise). Summing the aggregated cost from all directions and then taking the per-pixel minimum gives a result that's dramatically smoother than naive block matching while still respecting real depth discontinuities, at a fraction of the cost of solving the fully global (2D smoothness) version of the same optimization. This is exactly the algorithm behind OpenCV's StereoSGBM.
Occlusion. Because the two cameras sit at different positions, some 3D surface visible to the left camera is hidden from the right one — most obviously at depth discontinuities (the strip of background right behind a foreground object's edge) and at the image borders (the baseline shift exposes a sliver of scene only one camera can see). Those pixels have no correct match at all, and forcing block matching to pick something anyway just injects noise. The standard fix is a left-right consistency check: match left→right independently of right→left; if pixel x in the left image matches x−d in the right image, but that pixel's own best right→left match doesn't land back within about a pixel of x, the two disagree — flag the disparity invalid rather than trust it.
Cost slice: SAD cost vs disparity (one pixel)
Click pixels in the Step-3 left image to move this slice. Dashed arc: the parabola through (d−1, d, d+1) whose vertex is the sub-pixel refinement.
Disparity map (naive block matching)
Play: how fast does stereo get bad at distance?
Interactive
This is the disparity-specific version of depth uncertainty — distinct from Part 10's general triangulation covariance ellipsoid, which applies to any two-view ray intersection regardless of whether the images are rectified. Here it's just Z = fB/d and its derivative, evaluated live as you move the sliders.
ΔZ(Z) for the current baseline (solid) and 2× the baseline (dashed), log scale on the vertical axis. Marker = your current Z, B.
Cheat sheet
Recap
| Object | Definition / formula | Notes |
|---|---|---|
| Rectification, goal | Warp both images so corresponding epipolar lines become the same horizontal row | Turns 2D/line search into 1D row search |
| Rectifying frame | e₁=t/‖t‖, e₂=(0,0,1)×e₁, e₃=e₁×e₂, Rrect=[e₁ᵀe₂ᵀe₃ᵀ] | Forces Rrect·t = (‖t‖,0,0) — epipole to infinity along x |
| Rectifying warp | x′ ≅ Knew·Rrect·Riᵀ·Ki⁻¹·x | A homography per image (Part 8) — centers never move |
| Disparity → depth | Z = fB / d | d in pixels, f shared focal length, B baseline |
| Depth uncertainty | ΔZ ≈ Z² / (fB) per pixel of disparity error | Quadratic in Z, linear improvement from B |
| Block matching | Per-pixel minimum of SAD/SSD/NCC along the row | Fast; noisy on low-texture regions |
| SGBM | Aggregates cost along multiple 1D paths with a jump penalty, then takes the minimum | OpenCV StereoSGBM; smoother, still local-ish cost |
| Left-right consistency | Match L→R and R→L independently; disagreement > ~1px ⇒ invalid | Flags occluded / unmatchable pixels |