Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

Why two views aren't the whole story

Setup

Pairwise epipolar geometry between views (1,2), (1,3) and (2,3) certainly constrains a three-view correspondence — but treating three views as three independent pairs throws away information. A match in views 1 and 2 already determines a unique 3D point (via triangulation, Part 10); and once you know a 3D point and camera 3's pose, its image in view 3 is just a projection. Chaining those two steps is point transfer: predicting a third view's pixel from the other two, without ever writing the 3D point down explicitly.

💡 The idea for this page: point transfer works even without knowing any camera's calibration — the same way F replaced E for uncalibrated pairs, a single object called the trifocal tensor replaces the "triangulate-then-project" chain with one direct algebraic map, estimable from correspondences alone.
1

The trifocal tensor, briefly

Foundations

The trifocal tensor T is a 3×3×3 array of numbers (27 entries, 18 of them independent) built from the three cameras' projection matrices. It plays the same role for three views that the fundamental matrix plays for two — except instead of one constraint per correspondence, it directly transfers geometry between views: given a point in views 1 and 2, T predicts its exact location in view 3 (and, more generally, can transfer lines too — hence "point-line-line" and "line-line-line" relationships in the fuller treatment). Because it's linear in the unknowns, it can be estimated from as few as 6-7 point correspondences seen in all three views, with no calibration required — a three-view generalization of the eight-point algorithm from Part 6.

The full index-notation formula for point transfer is dense multilinear algebra and won't add much intuition at this level — what matters is the shape of the idea: redundancy across views is information. The more cameras that see a point, the more its location elsewhere is constrained (and eventually over-determined), which is exactly the leverage that structure-from-motion pipelines exploit with dozens or hundreds of views at once.

⚠️ The demo directly below still shows the effect of point transfer — predicting view 3 from views 1 & 2 — using the triangulate-then-project shortcut from Part 10, which needs known camera poses, purely as a warm-up/sanity check. The trifocal tensor T is built and used for real — for both point and line transfer, with no 3D triangulation anywhere in the transfer step — further down: jump to the real thing.
2

Play with it: predict the third view

Interactive

🎯 Learning goal: a correspondence in views 1 & 2 fully determines where it must appear in view 3 — click and watch the prediction land almost exactly on the true projection, until you add noise.

Click a pixel in image 1 to choose a point (depth slider sets where it actually sits in 3D — a stand-in for having already matched it in image 2 too). Image 3's prediction (hollow circle) comes purely from combining the noisy observations in images 1 & 2; the true projection (filled dot) is the answer key. This particular demo still fakes the transfer (triangulate + reproject) — the next one below builds and uses T for real.

Drag to orbit. Diamond = camera 1, square = camera 2, triangle-up = camera 3.

Image 1 (click to pick a point)

Image 2

Image 3: predicted vs. actual

actual projection predicted (transfer)
3

Building T for real

The actual tensor

Put camera 1 in canonical form, P=[I|0] (always possible by choosing world coordinates to coincide with camera 1's frame — exactly what the demo below does). Write the other two cameras as P′=[A|a₄] and P″=[B|b₄], where A,B are their leading 3×3 blocks and a₄,b₄ their last columns. Let ai be the i-th column of A, and bi the i-th column of B. The trifocal tensor is the set of three 3×3 matrices

Ti = aib₄ᵀ − a₄biᵀ    (i = 1, 2, 3)

Each Ti is an outer product of two 3-vectors minus another outer product — nothing more exotic than that; no SVD, no iteration, just six outer products and three subtractions. That's 27 numbers total (three 3×3 matrices), but they aren't all free: scale is arbitrary (as always in projective geometry) and the 27 entries satisfy internal algebraic constraints, leaving 18 genuine degrees of freedom — the same count you'd get adding up two calibration-free camera pairs' worth of relative-pose information, which is exactly what T encodes.

Why this particular formula? It falls straight out of eliminating the 3D point X from the three projection equations x̃₁≅[I|0]X̃, x̃₂≅[A|a₄]X̃, x̃₃≅[B|b₄]X̃ algebraically — T is what's left of "the third view" once the unknown depth and 3D position have been algebraically eliminated, which is exactly why it can predict view 3 without either.

The payoff is a single incidence relation that generalizes two-view epipolar coplanarity (Part 6's x₂ᵀFx₁=0) to three views. Take a point x̃₁ in view 1, and any two lines l₂, l₃ that happen to pass through the true corresponding points x̃₂, x̃₃ in views 2 and 3 (the simplest choice: a horizontal or vertical image-axis line through each). Then, writing x₁ⁱ for the three homogeneous components of x̃₁:

l₂ᵀ · ( ∑i x₁ⁱ Ti ) · l₃ = 0    for every l₂∋x̃₂, l₃∋x̃₃

Geometrically: back-project x̃₁ from camera 1 into a 3D ray, back-project l₂ from camera 2 into a 3D plane, and back-project l₃ from camera 3 into a 3D plane. If x̃₁, l₂, l₃ are truly images/silhouettes of the same 3D point, that ray and those two planes all meet at one common point in space — which is a strictly more rigid, more three-view-specific condition than any pairwise epipolar constraint, and it's exactly what the left-hand side above vanishing says algebraically.

4

True trifocal transfer: points and lines

The real demo

🎯 Learning goal: T, built once from the camera matrices, transfers a point exactly (matching the fake triangulate-then-project method to within numerical noise) and transfers an entire line exactly — something no pair of two-view fundamental matrices can ever do, because F alone only ever hands you a line-shaped constraint, never a full second-order transfer.

Point transfer, derived from the incidence relation above. Fix x̃₁ and define the 3×3 matrix M(x̃₁) = ∑i x₁ⁱ Ti (linear in x̃₁, nothing else). The incidence relation says l₂ᵀM(x̃₁)l₃=0 for every line l₂ through x̃₂ and every line l₃ through x̃₃. Hold l₂ fixed at any one line through x̃₂ and let l₃ range over every line through x̃₃: the only vector that's orthogonal to an entire pencil of lines through a point is that point itself. So:

x̃₃ ≅ M(x̃₁)ᵀ l₂    for any single line l₂ through x̃₂ (e.g. the vertical image line x=x₂ through it)

No triangulation, no camera centers, no depth — just one matrix built from x̃₁ and T, applied to one line through x̃₂. The demo below computes this and compares it, pixel for pixel, against the old triangulate-then-project shortcut.

Line transfer — the genuinely three-view-only trick. Given a line l₁ in view 1 (say, a cube edge) and its corresponding line l₂ in view 2, transfer two distinct points x̃₁ᵃ, x̃₁ᵇ that both lie on l₁ (any two points with l₁ᵀx=0 will do) using the exact point-transfer formula above, reusing the same whole line l₂ both times:

x̃₃ᵃ ≅ M(x̃₁ᵃ)ᵀ l₂,   x̃₃ᵇ ≅ M(x̃₁ᵇ)ᵀ l₂    ⇒     l₃ ≅ x̃₃ᵃ × x̃₃ᵇ

This works because x̃₁↦x̃₃ (for l₂ held fixed as the whole line, not a point-specific one) is linear in x̃₁ — a linear map sends the line l₁ to a line, and that image line has to be l₃, since every true correspondence on l₁ must land on the true l₃ by the incidence relation itself. Two-view geometry has no equivalent of this: from F alone, a point in view 1 only ever gives you a line constraint in view 2 — never a point, and never a whole transferred line in a third view. Three-view T categorically outruns any pair of two-view relations.

Click a pixel in image 1 below to pick a point (as before); its transfer to image 3 is now computed two independent ways — the old triangulate-then-project shortcut, and the real formula above — and the readout reports how far apart they land (it should be effectively zero, up to floating-point roundoff). A highlighted cube edge is also transferred as a line from views 1 & 2 into view 3 and drawn directly on top of that edge's ordinary projection, to show they coincide exactly.

Image 1 (click to pick a point)

Image 2

Image 3: transferred point & line vs. actual

actual projection fake transfer (triangulate + reproject) real transfer (T, point & line)

The orange dashed line is a second edge transferred independently through the same tensor; its intersection with the first transferred line is marked in view 3 — two line transfers composing into a point prediction, the trifocal analog of Part 6’s two-line epipolar trick.

⚠️ Still one simplification here: T is assembled from known camera poses (via P′,P″), the way this whole series has done for every earlier tensor/matrix. Real trifocal-tensor estimation in an uncalibrated setting instead solves for T's 27 numbers directly from ≈7 point correspondences seen in all three images — no poses required — the same spirit as the eight-point algorithm for F in Part 6. Once T exists, though, every formula above is used exactly as stated — the transfer step itself is fully real, not faked.
5

F and epipoles, straight out of T — and where the hierarchy stops

What T contains

T doesn't just generalize the two view-pairs (1,2) and (1,3) — it contains both of their fundamental matrices and epipoles as extractable byproducts, with no extra data needed. Each Ti is (generically) rank 2. Let ui be its left null vector (uiᵀTi=0) and vi its right null vector (Tivi=0). Stacking all three ui (or vi) must agree on a single common direction — and that direction is an epipole:

e₂ ≅ u₁ × u₂  (≅ u₁×u₃ ≅ u₂×u₃)    e₃ ≅ v₁ × v₂  (≅ v₁×v₃ ≅ v₂×v₃)

and the fundamental matrix between views 1 and 2 follows directly from e₂ and e₃:

F₂₁ ≅ [e₂]× · [ T₁e₃  |  T₂e₃  |  T₃e₃ ]    (three columns, each Tie₃, stacked into a 3×3 matrix)

F₃₁ (views 1,3) is built the same way with the roles of views 2 and 3 swapped. The upshot: T strictly contains more information than the pair {F₂₁, F₃₁} — it also fixes F₃₂ (views 2,3), which the two pairwise fundamental matrices alone leave underdetermined, plus the full point- and line-transfer maps this page has been building.

🎯 Seen live in the demo above: the tensor's left-null directions give e₂ (drawn on image 2) and its right-null directions give e₃ (drawn on image 3), and the readout checks both against the ground-truth epipoles — the residual e₂ᵀF₂₁e₁-style agreement is ~0 up to floating point. Drag the noise slider: transfer degrades gracefully while the epipoles stay pinned.

Add a fourth view and the same elimination trick produces a quadrifocal tensor Q — 34=81 components, encoding a four-view point-line-line-line incidence relation the same way T encodes point-line-line. But the hierarchy stops there: any geometric relation among five or more views reduces to combinations of the two-, three-, and four-view tensors (F, T, Q) applied to overlapping subsets — no fundamentally new type of constraint appears past four views. That's why real structure-from-motion systems (Part 15) never build a literal N-focal tensor for N>4; instead they work with pairwise and triplet relations plus bundle adjustment across however many views they have.

✓

Cheat sheet

Recap

ViewsConstraintWhat's determined
2 (calibrated)x₂ᵀEx₁=0A line in view 2 (the epipolar line)
2 (uncalibrated)x₂ᵀFx₁=0Same, without knowing intrinsics
3Trifocal tensor TAn exact point (or line) in view 3 — full transfer, not just a constraint
3Ti = aib₄ᵀ−a₄biᵀ (i=1,2,3)T itself, built from P′=[A|a₄], P″=[B|b₄] — 27 numbers, 18 DOF
3l₂ᵀ(∑ix₁ⁱTi)l₃=0Point-line-line incidence — three back-projected rays/planes meet at one 3D point
3x̃₃≅M(x̃₁)ᵀl₂, M=∑ix₁ⁱTiPoint transfer: view 3's point from x̃₁, x̃₂ and T — no 3D point ever built
3l₃≅M(x₁ᵃ)ᵀl₂ × M(x₁ᵇ)ᵀl₂Line transfer: an entire line handed to view 3 exactly — impossible from any pair of F's
3e₂≅u₁×u₂, F₂₁≅[e₂]×[T₁e₃|T₂e₃|T₃e₃]Epipoles and both pairwise F's, read directly out of T's null vectors
4Quadrifocal tensor Q (81 numbers)Point-line-line-line incidence; hierarchy of fundamental multi-view objects stops here (F, T, Q)
Real reconstructions use dozens or hundreds of views, not three. Making every camera and every point agree with every observation at once is bundle adjustment — the payoff of this whole series. Continue: bundle adjustment & structure from motion →