Three Views: the Trifocal Tensor
So far, everything has involved exactly two cameras. Add a third and something new becomes possible: given where a point lands in views 1 and 2, its location in view 3 isn't just constrained to a line (like the epipolar case) — it's pinned down exactly, without ever explicitly reconstructing a 3D point.
Why two views aren't the whole story
Setup
Pairwise epipolar geometry between views (1,2), (1,3) and (2,3) certainly constrains a three-view correspondence — but treating three views as three independent pairs throws away information. A match in views 1 and 2 already determines a unique 3D point (via triangulation, Part 10); and once you know a 3D point and camera 3's pose, its image in view 3 is just a projection. Chaining those two steps is point transfer: predicting a third view's pixel from the other two, without ever writing the 3D point down explicitly.
The trifocal tensor, briefly
Foundations
The trifocal tensor T is a 3×3×3 array of numbers (27 entries, 18 of them independent) built from the three cameras' projection matrices. It plays the same role for three views that the fundamental matrix plays for two — except instead of one constraint per correspondence, it directly transfers geometry between views: given a point in views 1 and 2, T predicts its exact location in view 3 (and, more generally, can transfer lines too — hence "point-line-line" and "line-line-line" relationships in the fuller treatment). Because it's linear in the unknowns, it can be estimated from as few as 6-7 point correspondences seen in all three views, with no calibration required — a three-view generalization of the eight-point algorithm from Part 6.
The full index-notation formula for point transfer is dense multilinear algebra and won't add much intuition at this level — what matters is the shape of the idea: redundancy across views is information. The more cameras that see a point, the more its location elsewhere is constrained (and eventually over-determined), which is exactly the leverage that structure-from-motion pipelines exploit with dozens or hundreds of views at once.
Play with it: predict the third view
Interactive
Click a pixel in image 1 to choose a point (depth slider sets where it actually sits in 3D — a stand-in for having already matched it in image 2 too). Image 3's prediction (hollow circle) comes purely from combining the noisy observations in images 1 & 2; the true projection (filled dot) is the answer key. This particular demo still fakes the transfer (triangulate + reproject) — the next one below builds and uses T for real.
Drag to orbit. Diamond = camera 1, square = camera 2, triangle-up = camera 3.
Image 1 (click to pick a point)
Image 2
Image 3: predicted vs. actual
Building T for real
The actual tensor
Put camera 1 in canonical form, P=[I|0] (always possible by choosing world coordinates to coincide with camera 1's frame — exactly what the demo below does). Write the other two cameras as P′=[A|a₄] and P″=[B|b₄], where A,B are their leading 3×3 blocks and a₄,b₄ their last columns. Let ai be the i-th column of A, and bi the i-th column of B. The trifocal tensor is the set of three 3×3 matrices
Each Ti is an outer product of two 3-vectors minus another outer product — nothing more exotic than that; no SVD, no iteration, just six outer products and three subtractions. That's 27 numbers total (three 3×3 matrices), but they aren't all free: scale is arbitrary (as always in projective geometry) and the 27 entries satisfy internal algebraic constraints, leaving 18 genuine degrees of freedom — the same count you'd get adding up two calibration-free camera pairs' worth of relative-pose information, which is exactly what T encodes.
Why this particular formula? It falls straight out of eliminating the 3D point X from the three projection equations x̃₁≅[I|0]X̃, x̃₂≅[A|a₄]X̃, x̃₃≅[B|b₄]X̃ algebraically — T is what's left of "the third view" once the unknown depth and 3D position have been algebraically eliminated, which is exactly why it can predict view 3 without either.
The payoff is a single incidence relation that generalizes two-view epipolar coplanarity (Part 6's x₂ᵀFx₁=0) to three views. Take a point x̃₁ in view 1, and any two lines l₂, l₃ that happen to pass through the true corresponding points x̃₂, x̃₃ in views 2 and 3 (the simplest choice: a horizontal or vertical image-axis line through each). Then, writing x₁ⁱ for the three homogeneous components of x̃₁:
Geometrically: back-project x̃₁ from camera 1 into a 3D ray, back-project l₂ from camera 2 into a 3D plane, and back-project l₃ from camera 3 into a 3D plane. If x̃₁, l₂, l₃ are truly images/silhouettes of the same 3D point, that ray and those two planes all meet at one common point in space — which is a strictly more rigid, more three-view-specific condition than any pairwise epipolar constraint, and it's exactly what the left-hand side above vanishing says algebraically.
True trifocal transfer: points and lines
The real demo
Point transfer, derived from the incidence relation above. Fix x̃₁ and define the 3×3 matrix M(x̃₁) = ∑i x₁ⁱ Ti (linear in x̃₁, nothing else). The incidence relation says l₂ᵀM(x̃₁)l₃=0 for every line l₂ through x̃₂ and every line l₃ through x̃₃. Hold l₂ fixed at any one line through x̃₂ and let l₃ range over every line through x̃₃: the only vector that's orthogonal to an entire pencil of lines through a point is that point itself. So:
No triangulation, no camera centers, no depth — just one matrix built from x̃₁ and T, applied to one line through x̃₂. The demo below computes this and compares it, pixel for pixel, against the old triangulate-then-project shortcut.
Line transfer — the genuinely three-view-only trick. Given a line l₁ in view 1 (say, a cube edge) and its corresponding line l₂ in view 2, transfer two distinct points x̃₁ᵃ, x̃₁ᵇ that both lie on l₁ (any two points with l₁ᵀx=0 will do) using the exact point-transfer formula above, reusing the same whole line l₂ both times:
This works because x̃₁↦x̃₃ (for l₂ held fixed as the whole line, not a point-specific one) is linear in x̃₁ — a linear map sends the line l₁ to a line, and that image line has to be l₃, since every true correspondence on l₁ must land on the true l₃ by the incidence relation itself. Two-view geometry has no equivalent of this: from F alone, a point in view 1 only ever gives you a line constraint in view 2 — never a point, and never a whole transferred line in a third view. Three-view T categorically outruns any pair of two-view relations.
Click a pixel in image 1 below to pick a point (as before); its transfer to image 3 is now computed two independent ways — the old triangulate-then-project shortcut, and the real formula above — and the readout reports how far apart they land (it should be effectively zero, up to floating-point roundoff). A highlighted cube edge is also transferred as a line from views 1 & 2 into view 3 and drawn directly on top of that edge's ordinary projection, to show they coincide exactly.
Image 1 (click to pick a point)
Image 2
Image 3: transferred point & line vs. actual
The orange dashed line is a second edge transferred independently through the same tensor; its intersection with the first transferred line is marked in view 3 — two line transfers composing into a point prediction, the trifocal analog of Part 6’s two-line epipolar trick.
F and epipoles, straight out of T — and where the hierarchy stops
What T contains
T doesn't just generalize the two view-pairs (1,2) and (1,3) — it contains both of their fundamental matrices and epipoles as extractable byproducts, with no extra data needed. Each Ti is (generically) rank 2. Let ui be its left null vector (uiᵀTi=0) and vi its right null vector (Tivi=0). Stacking all three ui (or vi) must agree on a single common direction — and that direction is an epipole:
and the fundamental matrix between views 1 and 2 follows directly from e₂ and e₃:
F₃₁ (views 1,3) is built the same way with the roles of views 2 and 3 swapped. The upshot: T strictly contains more information than the pair {F₂₁, F₃₁} — it also fixes F₃₂ (views 2,3), which the two pairwise fundamental matrices alone leave underdetermined, plus the full point- and line-transfer maps this page has been building.
Add a fourth view and the same elimination trick produces a quadrifocal tensor Q — 34=81 components, encoding a four-view point-line-line-line incidence relation the same way T encodes point-line-line. But the hierarchy stops there: any geometric relation among five or more views reduces to combinations of the two-, three-, and four-view tensors (F, T, Q) applied to overlapping subsets — no fundamentally new type of constraint appears past four views. That's why real structure-from-motion systems (Part 15) never build a literal N-focal tensor for N>4; instead they work with pairwise and triplet relations plus bundle adjustment across however many views they have.
Cheat sheet
Recap
| Views | Constraint | What's determined |
|---|---|---|
| 2 (calibrated) | x₂ᵀEx₁=0 | A line in view 2 (the epipolar line) |
| 2 (uncalibrated) | x₂ᵀFx₁=0 | Same, without knowing intrinsics |
| 3 | Trifocal tensor T | An exact point (or line) in view 3 — full transfer, not just a constraint |
| 3 | Ti = aib₄ᵀ−a₄biᵀ (i=1,2,3) | T itself, built from P′=[A|a₄], P″=[B|b₄] — 27 numbers, 18 DOF |
| 3 | l₂ᵀ(∑ix₁ⁱTi)l₃=0 | Point-line-line incidence — three back-projected rays/planes meet at one 3D point |
| 3 | x̃₃≅M(x̃₁)ᵀl₂, M=∑ix₁ⁱTi | Point transfer: view 3's point from x̃₁, x̃₂ and T — no 3D point ever built |
| 3 | l₃≅M(x₁ᵃ)ᵀl₂ × M(x₁ᵇ)ᵀl₂ | Line transfer: an entire line handed to view 3 exactly — impossible from any pair of F's |
| 3 | e₂≅u₁×u₂, F₂₁≅[e₂]×[T₁e₃|T₂e₃|T₃e₃] | Epipoles and both pairwise F's, read directly out of T's null vectors |
| 4 | Quadrifocal tensor Q (81 numbers) | Point-line-line-line incidence; hierarchy of fundamental multi-view objects stops here (F, T, Q) |