Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

0

Shape from tracked features, in one shot

Motivation

The classical structure-from-motion problem is: someone hands you P feature points that have been tracked through F frames of video, and you must produce both the 3D shape of the object and the camera motion — from the 2D tracks alone. Position-based, incremental pipelines solve this by matching point pairs, estimating pairwise geometry, triangulating, then running bundle adjustment to clean up the accumulated drift. That works, but it is a long chain of nonlinear steps, each needing a good starting guess.

Tomasi and Kanade's 1992 result says something startling: if the camera is far from the object relative to the object's depth, you can skip the whole chain. Write every tracked pixel coordinate into one big matrix, subtract the per-frame mean, and read off the answer from a single singular value decomposition. The shape and the motion come out together, linearly, with no iteration at all.

💡 The idea in one line: under weak perspective the measurement matrix is the product of a small motion matrix and a small shape matrix, so it is rank three — and rank-three structure is exactly what a truncated SVD exposes.

The catch is the approximation itself. "Far from the object" is a claim about the camera's distance relative to the depth variation of the scene; when that fails, perspective projection bends the tracks and the rank-three story degrades one singular value at a time. The interactive demo below lets you drive that failure with a slider and watch it happen.

1

Weak perspective: the imaging model and the measurement matrix

The model

Under full perspective, a 3D point X lands at u = f·x/z, which is a division and therefore not linear in X. Weak perspective (also called scaled orthographic) replaces that with an affine rule: for a camera whose center sits at t and whose orientation is the rotation R,

x = R (X − t)    then    u = x₁, v = x₂ (drop the third coordinate)
i.e. image coordinates are an affine function of the 3D point

Concretely, the image of a point is its camera-frame position with the depth coordinate simply discarded. Why is this a good approximation? Because when the camera distance D greatly exceeds the object's depth extent Δz, full perspective projects to u = f·x/(D + Δz), and expanding the division gives u ≈ (f/D)·x plus a correction of order Δz/D. Weak perspective keeps the leading term and throws away the correction. It is exact for an orthographic camera and accurate whenever Δz / D is small — telephoto lenses, distant objects, aerial imagery.

⚠️ This is where the method breaks. When the object is near the camera or has large depth relief, the discarded correction is no longer small — and the demo below shows precisely how it corrupts the reconstruction.

The measurement matrix. Stack all the tracked data. For frame i and point p, let (uᵢₚ, vᵢₚ) be the tracked pixel. Define the measurement matrix W of size 2F × P by writing each frame's two coordinate rows one after another:

W₂ᵢ₋₁,ₚ = uᵢₚ,    W₂ᵢ,ₚ = vᵢₚ    (row 2i−1 = the u's of frame i, row 2i = the v's)

Because the camera translation t is the same for every point in a frame, it contributes a constant offset to that frame's whole row. Subtracting each frame's centroid of tracked points therefore deletes the translation exactly:

ũᵢₚ = uᵢₚ − (1/P) ∑ₚ uᵢₚ    ṽᵢₚ = vᵢₚ − (1/P) ∑ₚ vᵢₚ

From here on W means the centered matrix, with each frame's centroid subtracted. Optionally one also normalizes each frame by the RMS radius of its points, which removes any per-frame scale; it changes nothing conceptually and everything is cleaner without having to worry about the absolute scale, which is unrecoverable anyway.

🎯 What centering buys you: the unknown camera translation is gone, so the only unknowns left are the rotation of each frame and the 3D shape. That is the whole reason the problem collapses to a rank condition.
2

The rank-3 theorem and the factorization

The core result

Write the shape of the object as a fixed 3 × P matrix S whose columns are the 3D points (after removing the mean point), and write the motion as a 2F × 3 matrix M whose frame-i block is the first two rows of that frame's rotation — the third row is simply discarded by the projection. Then the centered measurement matrix factors exactly:

W = M S    with   M (2F × 3),  S (3 × P)

Two structural facts about M are the key to everything that follows. Within frame i, the two rows m₂ᵢ₋₁ and m₂ᵢ are two rows of a rotation matrix, so they are orthonormal: each has unit norm and the two are perpendicular. And all frames share the same global scale.

m₂ᵢ₋₁ · m₂ᵢ₋₁ = 1,    m₂ᵢ · m₂ᵢ = 1,    m₂ᵢ₋₁ · m₂ᵢ = 0    for every frame i

Since W is the product of a 2F × 3 and a 3 × P matrix, it has rank at most three. That is the theorem. Every entry of W is a linear combination of just three underlying shape directions.

Factorization by truncated SVD. Compute the SVD W = U Σ Vᵀ. If the data were exactly weak perspective, only the first three singular values would be nonzero. Keep the top three and split them symmetrically between motion and shape:

W ≈ U₃ Σ₃ V₃ᵀ  ⇒   M̂ = U₃ Σ₃¹⁄²   Ŝ = Σ₃¹⁄² V₃ᵀ    so that  M̂ Ŝ = W

Here Σ₃ is the 3 × 3 diagonal matrix of the three largest singular values, and U₃, V₃ their singular vectors. The split by the square root makes the two factors balanced, but it is not the answer yet.

The affine gauge ambiguity. The factorization is not unique. Pick any invertible 3 × 3 matrix A and you get another valid factorization, because

W = M̂ Ŝ = (M̂ A)(A⁻¹ Ŝ)    ⇒    M = M̂ A,   S = A⁻¹ Ŝ

SVD can hand you any member of this family; it has no way to know which one is the physically correct rotation. This is the affine gauge freedom — the residual metric ambiguity of the problem — and it is exactly 9 unknowns' worth of it. What pins it down is the orthonormality of the motion rows, which M̂ does not in general satisfy but the true M must.

Solving for the gauge with a linear system. Let Q = A Aᵀ, a symmetric 3 × 3 matrix with six unknown entries. Look at a single frame's two rows of M̂, call them the row vectors aᵢ and bᵢ. The corresponding rows of the true motion are m₂ᵢ₋₁ = aᵢ A and m₂ᵢ = bᵢ A, so each orthonormality condition becomes a condition on Q alone:

‖m₂ᵢ₋₁‖² = aᵢ Q aᵢᵀ = 1
‖m₂ᵢ‖² = bᵢ Q bᵢᵀ = 1
m₂ᵢ₋₁ · m₂ᵢ = aᵢ Q bᵢᵀ = 0

Each of these is linear in the six entries of Q. Writing them explicitly for one frame, with aᵢ = (a₁, a₂, a₃):

a₁² q₁₁ + 2a₁a₂ q₁₂ + 2a₁a₃ q₁₃ + a₂² q₂₂ + 2a₂a₃ q₂₃ + a₃² q₃₃ = 1

and the b and mixed rows are identical in form with b or the cross terms substituted. Stacking all three equations for all F frames gives 3F linear equations in only six unknowns:

G q = 1  where  q = (q₁₁, q₁₂, q₁₃, q₂₂, q₂₃, q₃₃) and G is 3F × 6

Solve it in the least-squares sense, q = (GᵀG)⁻¹ Gᵀ1 — a single 6 × 6 inverse, no iteration. Then factor Q itself. Since Q = A Aᵀ, it must be symmetric positive definite, so the Cholesky factorization gives A directly: Q = L Lᵀ with L lower triangular, take A = L. Finally

M = M̂ A   (“the motion”)    S = A⁻¹ Ŝ   (“the shape”)

and the columns of S are the recovered 3D points. Motion and shape are recovered together, in one SVD plus one 6×6 solve. The only residual ambiguity is a single global rotation and reflection (which SVD cannot resolve, since a rotation of the whole world while rotating every camera oppositely is observationally identical) — precisely the ambiguity a Procrustes alignment removes when validating against ground truth.

🔑 Summary of the algorithm: center the tracks → SVD → truncate to three singular values → solve the 3F×6 linear system for Q → Cholesky for A → split into M and S. Every step is closed-form.
3

Play: factor a tracked object in one SVD

Interactive

🎯 Learning goal: watch the whole algorithm — SVD, rank-3 truncation, and the linear gauge solve — reconstruct an unknown 3D shape from nothing but its 2D tracks, and then watch a single slider (perspective distortion) break it in the honest, gradual way the theorem predicts.

A fixed synthetic point cloud is tracked through the chosen number of orthographic frames, each from a different viewpoint. The demo builds the centered measurement matrix W, computes a one-sided Jacobi SVD in your browser, solves the 3F × 6 system for Q, factors it by Cholesky, and reconstructs the 3D shape. On the canvas, the gray points are the ground-truth shape and the accent points are the reconstruction after an orthogonal Procrustes alignment (which removes the unavoidable global rotation/reflection/scale). Thin lines join each ground-truth point to its reconstruction. The readout reports the alignment RMS, the singular values of W (three large, the rest near zero), and the residual of the orthonormality constraints.

Ground truth (gray) vs. reconstructed shape (accent), aligned by Procrustes. Lines connect corresponding points.

ground truth reconstruction (rank-3 + gauge solve)

With perspective distortion at 0 the model is exactly orthographic: the singular values collapse to three large ones and everything else at the level of floating-point noise, the orthonormality residual is ~10⁻¹⁵, and the reconstruction matches the ground truth to machine precision up to a global rotation and scale. As you raise the slider, the first three singular values stay put while the fourth, fifth, sixth… creep up out of the noise floor — that is perspective showing up as extra rank — and both the residual and the alignment error grow together. This is the honest limitation of the method: affine factorization is not "wrong" under perspective, it is simply operating in the wrong model, and the error is smooth and measurable rather than catastrophic.

✓

Cheat sheet

Recap

ObjectFormulaMeaning / role
Weak perspectivex = R(X − t), drop the third coordinateAffine imaging model; exact for orthographic, good when distance ≫ depth relief
Measurement matrixW, size 2F × PStacked (u, v) tracks of every point in every frame, per-frame centroid subtracted
Centeringũ = u − meanₚ(u)Cancels the unknown camera translation exactly
Rank-3 theoremW = M S, M is 2F×3, S is 3×PThe measurement matrix has rank at most three
FactorizationW ≈ U₃Σ₃V₃ᵀ, M̂=U₃Σ₃¹⁄², Ŝ=Σ₃¹⁄²V₃ᵀOne SVD splits tracks into motion and shape
Affine gaugeM = M̂A, S = A⁻¹ŜAny invertible 3×3 A is a valid factorization; SVD picks an arbitrary one
Orthonormality constraintsaᵢQaᵢᵀ=1, bᵢQbᵢᵀ=1, aᵢQbᵢᵀ=0Each frame's two motion rows are rows of a rotation; Q = AAᵀ
Linear solveG q = 1, G is 3F×6Least-squares for the six entries of Q; then Cholesky Q = LLᵀ, A = L
Metric upgradeM = M̂A, S = A⁻¹ŜTurns the arbitrary factorization into metric motion and shape
Remaining ambiguityone global rotation / reflection / scaleUnobservable; removed by Procrustes alignment for validation
Failure modesσ₄⁄σ₁ no longer ≈ 0Perspective (near/large-depth scene), missing or mis-tracked data, outliers
Factorization is the elegant special case: closed-form, non-iterative, and exact under weak perspective. When perspective matters, or when data are noisy, occluded, or incomplete, the linear one-shot answer is only a starting point — and the general machinery that finishes the job is bundle adjustment. Continue: bundle adjustment & structure from motion →