Affine Factorization: the Tomasi-Kanade Method
Everything so far has recovered geometry one error term at a time: triangulate a point, then estimate a pose, then reconcile the two with a nonlinear solver. This page does the opposite. When the camera is far from the scene, the entire set of tracked 2D features — every point in every frame — stacks into a single tall matrix that is provably rank three. One singular value decomposition splits that matrix into camera motion and 3D shape simultaneously, with no iteration, no initialization, and no alternation. It is the cleanest result in this whole series, and it comes with an honest price tag: the weak-perspective approximation.
Shape from tracked features, in one shot
Motivation
The classical structure-from-motion problem is: someone hands you P feature points that have been tracked through F frames of video, and you must produce both the 3D shape of the object and the camera motion — from the 2D tracks alone. Position-based, incremental pipelines solve this by matching point pairs, estimating pairwise geometry, triangulating, then running bundle adjustment to clean up the accumulated drift. That works, but it is a long chain of nonlinear steps, each needing a good starting guess.
Tomasi and Kanade's 1992 result says something startling: if the camera is far from the object relative to the object's depth, you can skip the whole chain. Write every tracked pixel coordinate into one big matrix, subtract the per-frame mean, and read off the answer from a single singular value decomposition. The shape and the motion come out together, linearly, with no iteration at all.
The catch is the approximation itself. "Far from the object" is a claim about the camera's distance relative to the depth variation of the scene; when that fails, perspective projection bends the tracks and the rank-three story degrades one singular value at a time. The interactive demo below lets you drive that failure with a slider and watch it happen.
Weak perspective: the imaging model and the measurement matrix
The model
Under full perspective, a 3D point X lands at u = f·x/z, which is a division and therefore not linear in X. Weak perspective (also called scaled orthographic) replaces that with an affine rule: for a camera whose center sits at t and whose orientation is the rotation R,
i.e. image coordinates are an affine function of the 3D point
Concretely, the image of a point is its camera-frame position with the depth coordinate simply discarded. Why is this a good approximation? Because when the camera distance D greatly exceeds the object's depth extent Δz, full perspective projects to u = f·x/(D + Δz), and expanding the division gives u ≈ (f/D)·x plus a correction of order Δz/D. Weak perspective keeps the leading term and throws away the correction. It is exact for an orthographic camera and accurate whenever Δz / D is small — telephoto lenses, distant objects, aerial imagery.
The measurement matrix. Stack all the tracked data. For frame i and point p, let (uᵢₚ, vᵢₚ) be the tracked pixel. Define the measurement matrix W of size 2F × P by writing each frame's two coordinate rows one after another:
Because the camera translation t is the same for every point in a frame, it contributes a constant offset to that frame's whole row. Subtracting each frame's centroid of tracked points therefore deletes the translation exactly:
From here on W means the centered matrix, with each frame's centroid subtracted. Optionally one also normalizes each frame by the RMS radius of its points, which removes any per-frame scale; it changes nothing conceptually and everything is cleaner without having to worry about the absolute scale, which is unrecoverable anyway.
The rank-3 theorem and the factorization
The core result
Write the shape of the object as a fixed 3 × P matrix S whose columns are the 3D points (after removing the mean point), and write the motion as a 2F × 3 matrix M whose frame-i block is the first two rows of that frame's rotation — the third row is simply discarded by the projection. Then the centered measurement matrix factors exactly:
Two structural facts about M are the key to everything that follows. Within frame i, the two rows m₂ᵢ₋₁ and m₂ᵢ are two rows of a rotation matrix, so they are orthonormal: each has unit norm and the two are perpendicular. And all frames share the same global scale.
Since W is the product of a 2F × 3 and a 3 × P matrix, it has rank at most three. That is the theorem. Every entry of W is a linear combination of just three underlying shape directions.
Factorization by truncated SVD. Compute the SVD W = U Σ Vᵀ. If the data were exactly weak perspective, only the first three singular values would be nonzero. Keep the top three and split them symmetrically between motion and shape:
Here Σ₃ is the 3 × 3 diagonal matrix of the three largest singular values, and U₃, V₃ their singular vectors. The split by the square root makes the two factors balanced, but it is not the answer yet.
The affine gauge ambiguity. The factorization is not unique. Pick any invertible 3 × 3 matrix A and you get another valid factorization, because
SVD can hand you any member of this family; it has no way to know which one is the physically correct rotation. This is the affine gauge freedom — the residual metric ambiguity of the problem — and it is exactly 9 unknowns' worth of it. What pins it down is the orthonormality of the motion rows, which M̂ does not in general satisfy but the true M must.
Solving for the gauge with a linear system. Let Q = A Aᵀ, a symmetric 3 × 3 matrix with six unknown entries. Look at a single frame's two rows of M̂, call them the row vectors aᵢ and bᵢ. The corresponding rows of the true motion are m₂ᵢ₋₁ = aᵢ A and m₂ᵢ = bᵢ A, so each orthonormality condition becomes a condition on Q alone:
‖m₂ᵢ‖² = bᵢ Q bᵢᵀ = 1
m₂ᵢ₋₁ · m₂ᵢ = aᵢ Q bᵢᵀ = 0
Each of these is linear in the six entries of Q. Writing them explicitly for one frame, with aᵢ = (a₁, a₂, a₃):
and the b and mixed rows are identical in form with b or the cross terms substituted. Stacking all three equations for all F frames gives 3F linear equations in only six unknowns:
Solve it in the least-squares sense, q = (GᵀG)⁻¹ Gᵀ1 — a single 6 × 6 inverse, no iteration. Then factor Q itself. Since Q = A Aᵀ, it must be symmetric positive definite, so the Cholesky factorization gives A directly: Q = L Lᵀ with L lower triangular, take A = L. Finally
and the columns of S are the recovered 3D points. Motion and shape are recovered together, in one SVD plus one 6×6 solve. The only residual ambiguity is a single global rotation and reflection (which SVD cannot resolve, since a rotation of the whole world while rotating every camera oppositely is observationally identical) — precisely the ambiguity a Procrustes alignment removes when validating against ground truth.
Play: factor a tracked object in one SVD
Interactive
A fixed synthetic point cloud is tracked through the chosen number of orthographic frames, each from a different viewpoint. The demo builds the centered measurement matrix W, computes a one-sided Jacobi SVD in your browser, solves the 3F × 6 system for Q, factors it by Cholesky, and reconstructs the 3D shape. On the canvas, the gray points are the ground-truth shape and the accent points are the reconstruction after an orthogonal Procrustes alignment (which removes the unavoidable global rotation/reflection/scale). Thin lines join each ground-truth point to its reconstruction. The readout reports the alignment RMS, the singular values of W (three large, the rest near zero), and the residual of the orthonormality constraints.
Ground truth (gray) vs. reconstructed shape (accent), aligned by Procrustes. Lines connect corresponding points.
With perspective distortion at 0 the model is exactly orthographic: the singular values collapse to three large ones and everything else at the level of floating-point noise, the orthonormality residual is ~10⁻¹⁵, and the reconstruction matches the ground truth to machine precision up to a global rotation and scale. As you raise the slider, the first three singular values stay put while the fourth, fifth, sixth… creep up out of the noise floor — that is perspective showing up as extra rank — and both the residual and the alignment error grow together. This is the honest limitation of the method: affine factorization is not "wrong" under perspective, it is simply operating in the wrong model, and the error is smooth and measurable rather than catastrophic.
Cheat sheet
Recap
| Object | Formula | Meaning / role |
|---|---|---|
| Weak perspective | x = R(X − t), drop the third coordinate | Affine imaging model; exact for orthographic, good when distance ≫ depth relief |
| Measurement matrix | W, size 2F × P | Stacked (u, v) tracks of every point in every frame, per-frame centroid subtracted |
| Centering | ũ = u − meanₚ(u) | Cancels the unknown camera translation exactly |
| Rank-3 theorem | W = M S, M is 2F×3, S is 3×P | The measurement matrix has rank at most three |
| Factorization | W ≈ U₃Σ₃V₃ᵀ, M̂=U₃Σ₃¹⁄², Ŝ=Σ₃¹⁄²V₃ᵀ | One SVD splits tracks into motion and shape |
| Affine gauge | M = M̂A, S = A⁻¹Ŝ | Any invertible 3×3 A is a valid factorization; SVD picks an arbitrary one |
| Orthonormality constraints | aᵢQaᵢᵀ=1, bᵢQbᵢᵀ=1, aᵢQbᵢᵀ=0 | Each frame's two motion rows are rows of a rotation; Q = AAᵀ |
| Linear solve | G q = 1, G is 3F×6 | Least-squares for the six entries of Q; then Cholesky Q = LLᵀ, A = L |
| Metric upgrade | M = M̂A, S = A⁻¹Ŝ | Turns the arbitrary factorization into metric motion and shape |
| Remaining ambiguity | one global rotation / reflection / scale | Unobservable; removed by Procrustes alignment for validation |
| Failure modes | σ₄⁄σ₁ no longer ≈ 0 | Perspective (near/large-depth scene), missing or mis-tracked data, outliers |