Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

From an arbitrary line to a single row

Motivation

The epipolar constraint from Part 6 says a match must lie somewhere on the line l₂ = Fx₁ — a real improvement over searching the whole image, but that line's slope depends on F, which depends on the two cameras' relative pose. In general it's tilted, so "search along it" means resampling pixels along a diagonal at every candidate position, for every pixel in the image, for every frame. That's expensive and fiddly to vectorize.

But there's nothing sacred about the original pixel grid. Camera rotation is free to choose, and Part 8 will show that rotating a calibrated camera about its own optical center — without moving it — is exactly a homography: a warp of the image that changes nothing about the physical rays it represents, only how they're indexed by pixel coordinates. So: choose new, virtual camera orientations for both cameras such that every epipolar line becomes horizontal, and lands on the same row in both images. Warp the two real images into those virtual cameras. Now the search for a match is one dimension: fix a row, slide a window along it, done.

💡 Why it matters: this is the difference between "epipolar geometry" as a nice constraint and stereo vision as a working depth sensor. Every stereo camera you can buy — from a robot's stereo rig to a laptop webcam pair to classic Kinect-style stereo depth — rectifies first and matches second. Almost nothing downstream (block matching, SGBM, real-time disparity) works without it.
1

Constructing the rectifying rotation

Derivation

Start from a calibrated stereo pair: two cameras with known rotations R₁, R₂ (world → camera) and known relative pose between them, summarized by the baseline vector t pointing from camera 1's center to camera 2's center, expressed in world coordinates. The goal is a single new rotation Rrect that both virtual cameras will share — same orientation, so their image planes are parallel and their rows already line up — chosen so that direction also happens to send the epipole to infinity along the horizontal axis.

An epipole is the projection of one camera's center into the other camera's image, so all epipolar lines fan out from it; if it sits at infinity in a fixed horizontal direction, every epipolar line is parallel to that direction — i.e. horizontal. The standard construction picks Rrect's three rows to be an orthonormal basis built directly from t:

new x‑axis: e₁ = t / ‖t‖   (points along the baseline itself)
new y‑axis: e₂ = (0,0,1) × e₁, normalized   (perpendicular to baseline, roughly “up” — assumes t isn’t vertical)
new z‑axis: e₃ = e₁ × e₂   (completes a right‑handed frame, points “forward”)

Rrect = [ e₁ᵀ ; e₂ᵀ ; e₃ᵀ ]   (stack as rows)

Here's why this specific choice sends the epipole to infinity. Camera 2's center, seen from camera 1's new orientation, is displaced by t; expressed in the new frame that displacement is Rrect·t. Because e₁ = t/‖t‖ by construction, and e₂, e₃ are both orthogonal to t by construction:

Rrect·t = ( e₁·t,  e₂·t,  e₃·t ) = ( ‖t‖,  0,  0 )

All of the displacement lands on the new x-axis; none of it on y or z. The epipole — the projection of that displaced center — sits infinitely far along x with zero vertical component, i.e. exactly at the ideal point in the horizontal direction (Part 0's w = 0 points). A line through the epipole with zero vertical offset from any point is, by definition, horizontal. That's the whole derivation: the construction doesn't just happen to work, it's reverse-engineered so that t lands entirely on the axis that becomes "sideways" in the new image.

The rectifying rotation applied to each real camera is Rrect·R₁ᵀ and Rrect·R₂ᵀ (composing "undo the camera's own rotation" with "rotate into the shared rectified frame"). Since both new cameras get the identical Rrect, they end up row-aligned by construction — and since t is now purely along the shared x-axis, the two virtual cameras differ from each other by nothing but a pure horizontal translation of length B = ‖t‖, the baseline.

⚠️ What actually gets warped: the cameras' centers never move — only their orientation (and, for a clean result, their intrinsics K, usually reconciled to a shared value for both images) — so per Part 8, this rotation-about-the-center is exactly a homography of the original pixel grid: x′ ≅ Knew·Rrect·Riᵀ·Ki⁻¹·x. Rectification in practice (e.g. OpenCV's stereoRectify + remap) is exactly this: two per-image homography warps, computed once from calibration, applied to every incoming frame.
2

Disparity, and why Z = fB / d

Derivation

After rectification, both virtual cameras share a focal length f, sit a baseline B apart along a shared x-axis, and share every row. A 3D point at depth Z and horizontal offset X from the left camera projects, in each camera, by the ordinary pinhole rule u = f·(horizontal coordinate)/Z — the only difference between the two cameras is which point they call the horizontal origin, and that offset is exactly B:

uleft = f·X / Z
uright = f·(X − B) / Z   (the right camera's center sits B further along +x, so the point is B closer to its own origin)

disparity  d ≡ uleft − uright = f·X/Z − f·(X−B)/Z = fB / Z

⇒   Z = fB / d

Every quantity on the right of Z = fB/d is either fixed by calibration (f, B) or directly measured from the two images (d, in pixels) — which is the whole appeal: once rectified, recovering depth is arithmetic on a single measured integer (or sub-pixel value) per pixel, no triangulation solve required.

That formula also says something uncomfortable about accuracy. Disparity is measured to some finite precision — call one pixel of disparity noise Δd = 1. How much does that same pixel of noise mean in meters, as a function of how far away the point already is? Differentiate Z(d) = fB/d:

dZ/dd = −fB / d²
substitute  d = fB/Z  (so  d² = f²B²/Z²):
dZ/dd = −fB · Z²/(f²B²) = −Z² / (fB)

⇒   ΔZ ≈ Z² / (fB)    for one pixel of disparity error

Depth error grows with the square of distance and shrinks only linearly with baseline. A stereo rig that's comfortably accurate at 2 m can be off by meters at 20 m — the same one pixel of disparity noise means 100× more depth error, because disparity itself has shrunk 10× and the derivative scales as its inverse square. This is precisely why LiDAR, not stereo, dominates long-range depth, and why wide-baseline rigs (further-apart cameras) trade a larger blind zone near the camera for meaningfully better depth at range.

3

Play: rectify a stereo pair, then match a pixel

Interactive

🎯 Learning goal: toggle between an "unrectified" (toed-in) rig and a rectified one and watch the epipolar line snap horizontal. Then, in the rectified pair, click a pixel and watch a real block match run along that single row.
⚠️ Honest framing: there's no real camera hardware here. Both images below are ray-traced from a small synthetic scene (a textured ground plane plus one sphere) through two virtual pinhole cameras, purely so there's real, live-computable stereo geometry to click on. Real rectification also has to correct lens distortion and reconcile two slightly different K's; this demo skips both since the virtual cameras are already distortion-free and share K exactly.

The rig is two cameras B apart along a shared baseline, both pitched down slightly to see a checkered floor with one solid sphere floating above it. In Unrectified mode each camera additionally toes inward to fixate a point ahead — a common (bad) way to point two cameras at the same subject — which makes their relative rotation R₂R₁ᵀ ≠ I and tilts the epipolar line. In Rectified mode both cameras share the same orientation exactly as derived above, so t is purely horizontal and the epipolar line is exactly horizontal too.

Left image (click to pick a pixel)

Right image (epipolar line + best match)

Dashed line = epipolar line for the clicked pixel, computed from the live F. In rectified mode it's the search row; the filled dot marks the best SAD match found by sliding a 7×7 window along it.

Click a pixel in the left image.
4

How disparity actually gets computed

Algorithms

Block matching. For every pixel in the rectified left image, slide a small window (e.g. 7×7 or 9×9) along the same row in the right image and score how well it matches at each candidate column. Three common scores:

SAD (sum of absolute differences):  ∑ |IL(x,y) − IR(x−d,y)|  —  cheapest, sensitive to brightness offsets
SSD (sum of squared differences):  ∑ (IL(x,y) − IR(x−d,y))²  —  penalizes large errors harder
NCC (normalized cross-correlation): ∑(IL−ȳL)(IR−ȳR) / (σLσR)  —  robust to brightness/contrast differences between the two cameras

Whichever score is used, the disparity that minimizes cost is only known to the nearest integer pixel by default; fitting a parabola to the three cost values around the minimum (d−1, d, d+1) and solving for its vertex gives a cheap sub-pixel refinement, which is what the demo above does.

The cost volume. Doing this for every pixel and every candidate disparity builds a 3D array — cost[row][col][d] — the object every stereo algorithm, however sophisticated, is ultimately searching over. Naive block matching just takes the per-pixel minimum along the d axis independently for each pixel, which is fast but noisy: textureless regions (a blank wall) have a flat, ambiguous cost curve, and the per-pixel minimum can jump around randomly from one pixel to its neighbor.

SGBM (semi-global block matching). Instead of picking each pixel's minimum in isolation, SGBM aggregates cost along several 1D paths through the cost volume (typically 8 or 16 directions radiating from each pixel), adding a penalty whenever the disparity jumps between neighboring pixels along a path — a small penalty for a 1-pixel jump (edges are allowed to have small disparity steps), a much larger one for jumps of more than 1 (discourages spurious noise). Summing the aggregated cost from all directions and then taking the per-pixel minimum gives a result that's dramatically smoother than naive block matching while still respecting real depth discontinuities, at a fraction of the cost of solving the fully global (2D smoothness) version of the same optimization. This is exactly the algorithm behind OpenCV's StereoSGBM.

Occlusion. Because the two cameras sit at different positions, some 3D surface visible to the left camera is hidden from the right one — most obviously at depth discontinuities (the strip of background right behind a foreground object's edge) and at the image borders (the baseline shift exposes a sliver of scene only one camera can see). Those pixels have no correct match at all, and forcing block matching to pick something anyway just injects noise. The standard fix is a left-right consistency check: match left→right independently of right→left; if pixel x in the left image matches x−d in the right image, but that pixel's own best right→left match doesn't land back within about a pixel of x, the two disagree — flag the disparity invalid rather than trust it.

🎯 Learning goal: the cost volume isn't an abstraction. Left: the actual SAD cost curve over disparity for one pixel — sharp V on textured areas, flat and ambiguous on textureless ones — with the sub-pixel parabola fit overlaid. Right: a whole-image disparity map from naive block matching, with occluded pixels (left-right check failures) flagged in red and the textureless sky showing the streaky garbage the text warns about.

Cost slice: SAD cost vs disparity (one pixel)

Click pixels in the Step-3 left image to move this slice. Dashed arc: the parabola through (d−1, d, d+1) whose vertex is the sub-pixel refinement.

Disparity map (naive block matching)

Click “Compute disparity map” to run block matching over the whole image pair.
5

Play: how fast does stereo get bad at distance?

Interactive

🎯 Learning goal: watch ΔZ = Z²/(fB) explode quadratically as Z grows, and shrink only linearly when you double the baseline — directly from the formula, no simulated image involved this time.

This is the disparity-specific version of depth uncertainty — distinct from Part 10's general triangulation covariance ellipsoid, which applies to any two-view ray intersection regardless of whether the images are rectified. Here it's just Z = fB/d and its derivative, evaluated live as you move the sliders.

ΔZ(Z) for the current baseline (solid) and 2× the baseline (dashed), log scale on the vertical axis. Marker = your current Z, B.

✓

Cheat sheet

Recap

ObjectDefinition / formulaNotes
Rectification, goalWarp both images so corresponding epipolar lines become the same horizontal rowTurns 2D/line search into 1D row search
Rectifying framee₁=t/‖t‖, e₂=(0,0,1)×e₁, e₃=e₁×e₂, Rrect=[e₁ᵀe₂ᵀe₃ᵀ]Forces Rrect·t = (‖t‖,0,0) — epipole to infinity along x
Rectifying warpx′ ≅ Knew·Rrect·Riᵀ·Ki⁻¹·xA homography per image (Part 8) — centers never move
Disparity → depthZ = fB / dd in pixels, f shared focal length, B baseline
Depth uncertaintyΔZ ≈ Z² / (fB) per pixel of disparity errorQuadratic in Z, linear improvement from B
Block matchingPer-pixel minimum of SAD/SSD/NCC along the rowFast; noisy on low-texture regions
SGBMAggregates cost along multiple 1D paths with a jump penalty, then takes the minimumOpenCV StereoSGBM; smoother, still local-ish cost
Left-right consistencyMatch L→R and R→L independently; disagreement > ~1px ⇒ invalidFlags occluded / unmatchable pixels
Rectification is a special case of a more general idea: when a scene (or camera motion) is flat, one image can predict another completely, with a single 3×3 matrix — the homography. Continue: homography →