Dot product, projection, duality
Adding arrows was easy; multiplying them takes a choice. The dot product is the useful choice: one number that measures how much of one arrow lies along another. It answers the question a projection asks — what is the closest point on this line? — and it hides a stranger fact underneath. A row matrix acting on a vector is nothing but a dot product with a fixed arrow, so the object you have been calling a matrix and the object you have been calling a vector turn out to be the same idea viewed from opposite sides. That reveal is called duality, and it is why half of machine learning can be written as “compare these two vectors”.
The question
One number that knows a direction
Two vectors live in the same plane, but they can point in wildly different directions. To compare them you want a single number that says how aligned they are: large and positive when they point the same way, zero when they are perpendicular, negative when they oppose. Lengths alone cannot do it, because length says nothing about direction. What you need is a measurement that mixes the two arrows together and returns one scalar.
The dot product is that measurement. Geometrically it is the length of one arrow times the length of the other times the cosine of the angle between them. Read that product in the other order and it becomes a projection: |v|cosθ is the length of the shadow that v casts on the line through w, and multiplying by |w| scales that shadow to the units of w. So the dot product is an overlap, and hidden in the overlap is a closest point.
That closest point matters more than it first appears. Projecting onto a line is the simplest least-squares problem there is: find the point on w's line that is nearest to v. The residual — the leftover piece of v that the line cannot reach — meets the line at a genuine right angle, and the whole calculation of regression, of fitting, of compressing data onto a direction, is this one picture repeated in higher dimensions.
Then comes the reveal. Write the same dot product as a matrix: a single row [ wx wy ] times the column x. The row matrix does nothing but compute w·x, so a row of numbers is a function on vectors, and a vector is what that function eats. Matrices map vectors to vectors; a row matrix maps vectors to numbers. Both are just dot products wearing different costumes, and once you see that, the shape 1×n stops looking like a degenerate matrix and starts looking like a vector holding a ruler.
It helps to say what the dot product is not, because the coordinate formula hides the choice behind it. Multiplying entry by entry and adding is a convention, and at first a strange one: nothing about the first coordinate of one arrow makes it naturally care about the first coordinate of the other, and rotating the axes changes both lists while the value of the dot product stays put. That invariance is the tell. The answer depends only on the two arrows, not on the coordinates you happened to write them in, which is why the operation counts as geometry and the formula counts as arithmetic.
Projection and the perpendicular drop
The shadow of one arrow on another
Fix a line through the origin — the line that w points along — and ask which of its points is closest to the tip of v. Slide a point along the line and the distance to v shrinks, bottoms out, then grows again. The minimising point is the foot of the perpendicular from v to the line, and the arrow from the origin to that foot is the projection of v onto w. The segment from the foot to the tip of v is the residual, and because it is perpendicular to the line it is the part of v the line cannot represent.
Two formulas fall out of the picture. The scalar projection is the signed length of the shadow, |v|cosθ, which is just (v·w)/|w|: divide the overlap by the length of the arrow doing the measuring. The vector projection restores the direction by multiplying that scalar by the unit vector along w, which is the same as scaling w itself by (v·w)/(w·w). Splitting v into that projection plus the residual gives the decomposition you will use everywhere from least squares to the SVD.
Drag the two arrows below. The number v·w in the readout is computed entrywise and compared against |v||w|cosθ: they agree to the last digit, because that is an identity and not an approximation. Watch what happens when v swings across the line and the shadow flips sides — the dot product changes sign at exactly the moment the angle passes a right angle, which is the next section.
Drag v and w. The heavy arrow is the projection of v onto w's line; the magenta dashed arrow is the perpendicular residual, and the small square marks the right angle at the foot.
Notice that the residual is perpendicular to w, not to v. That is the whole reason projection is a closest-point operation: at the minimiser, moving along the line changes the distance only to second order, which is the geometric restatement of “orthogonal”. When you meet the normal equations in the least-squares part, they are this right angle written in matrix algebra.
There is a bound hiding in the formula, and it is the reason the cosine is well defined at all. No projection can be longer than the vector casting it, so |v·w| ≤ |v||w|; divide through and cosθ lands in the interval from minus one to one, as a cosine must. Equality holds only when the two arrows are parallel, which is the algebraic way of saying that one is a multiple of the other. This inequality — Cauchy–Schwarz — is the first nontrivial theorem about inner products, and it is what lets you define an angle in a space where you cannot draw one, which is exactly what happens when the vectors are word embeddings with thousands of coordinates.
You can check the closest-point claim without any calculus. Pick another point on the line, so the arrow from the foot to it lies along w. Then the vector from that point to v is the residual minus that along-line step, and the residual is perpendicular to it, so Pythagoras says its squared length is the squared residual plus a positive squared step. The foot wins, and the winner is unique. Orthogonal projection is not merely a convenient construction; it is the exact minimiser of distance to the line, and the same argument works for a plane or any subspace.
When the overlap is zero
Perpendicularity, and why it is the core test
Set the overlap to zero and something clean happens. Since |v| and |w| are positive lengths, v·w = 0 forces cosθ = 0, so the angle is a right angle. The two arrows contribute nothing to each other; each one is invisible to the other's projection. This is the definition of orthogonal, and it is the single most-used test in the subject precisely because it costs one multiply-and-add per pair of coordinates and needs no trigonometry or square roots.
Rotate v below and watch the dot product pass through zero twice per turn. At those instants the projection collapses to the origin, the residual becomes all of v, and the right-angle marker appears at the common tail. Read the same fact from the algebra side: the equation wxx + wyy = 0 is a straight line through the origin, and every vector on that line is perpendicular to w. A right angle that looked like geometry has become a linear equation, which is the first hint that orthogonality will organise everything that follows.
Dividing the dot product by the two lengths strips out the lengths and leaves the pure direction comparison, the quantity called cosine similarity. It is invariant to how long the two vectors are, so two embeddings that point the same way but differ in magnitude score a perfect one. That invariance is why retrieval and attention compare normalized vectors: the magnitude of an embedding is often an artefact of how frequently a token appears, while its direction is what carries the meaning. When a model is said to have learned that two words are close, this number is what “close” usually means.
Rotate v with the slider, or drag either arrow. When the two arrows meet at a right angle both turn magenta and the dot product reads zero.
Orthogonality is what makes a basis convenient. If the basis arrows are mutually perpendicular and unit length, then the coordinates of any vector are just its dot products with the basis arrows, one projection per axis and no linear system to solve. Turning an arbitrary set of vectors into such a basis is Gram–Schmidt, and it runs on nothing but the projection you just dragged: subtract the part of each new vector that lies along the ones already chosen, and the residual is orthogonal by construction. That is the subject of the next part.
Spell the payoff out. If the two basis arrows e1 and e2 are unit length and perpendicular, then any vector x rebuilds from two numbers, x·e1 and x·e2. There is no matrix to invert and no elimination to run: the coefficients are the projections, read off one dot product at a time. In an orthonormal basis the dot product is not just a measurement, it is the inverse of the basis, which is why geometry pipelines go to the trouble of keeping their frames normalized.
A row matrix is a vector
Two sides of the same multiplication
Until now a matrix has been something that eats a vector and returns another vector. But a 1×n matrix — a single row — eats a vector and returns one number. Write out the multiplication and there is nowhere to hide: the row is a list of weights, you multiply it entry by entry against the input and add, and that is exactly a dot product with the row read as a vector.
Pull the two sides apart. From the matrix side, the row is a function that assigns a number to every input vector. From the vector side, that function is a projection and a scaling: its value is the length of x's shadow on w, times |w|. The level sets of the function — all the inputs sharing one output — are a family of parallel lines, and those lines are perpendicular to w, because moving perpendicular to w changes no shadow at all. Since vectors in the plane are pinned to the origin, the familiar picture is even simpler: the level sets are the striped contours of a hillside, and w points straight uphill.
This is duality, and it is worth stating plainly. A vector can be used as an input, and a vector can be used as a measuring device. The same list of numbers plays both roles, and the dot product is the bridge: it is the action of the vector-as-row on the vector-as-column. In n dimensions nothing changes except that the level sets become parallel hyperplanes, which is why a linear classifier is drawn as a line, a plane, or a hyperplane depending on the dimension of its input.
In the language of the subject, the row is a linear functional: a map from vectors to numbers that respects addition and scaling. In finite dimensions every linear functional arises this way, as a dot product with some fixed vector, so the collection of all functionals is itself just a copy of the space. That is the precise sense of the word duality — two spaces, vectors and measurements, mirroring each other, with the dot product as the pairing between them. The next time you transpose a matrix you are moving between those two mirror worlds, and the entries that look unchanged are doing a quiet change of role.
Drag w to re-aim the ruler and x to move the input. The parallel lines are the level sets of w·x and stay perpendicular to w; the number line below shows the single number the row matrix returns.
One consequence is worth carrying forward. A matrix with several rows is just several such measurements stacked: each row is a vector, each row asks its own dot-product question, and the output vector collects the answers. Multiplication by a matrix is therefore a list of projections onto a list of directions. When you reach the SVD and PCA, the directions the data is measured against are the singular vectors, chosen so those measurements carry as much of the variance as possible.
Duality also explains a notational habit that otherwise looks arbitrary. We write a vector as a column and its transpose as a row, and we are fussy about which side of a multiplication each one goes on. The reason is exactly the bridge above: a column is something a matrix acts on, while a row is something that acts. The dot product mixes the two roles, which is why it can be written as wTx, and why a formula that transposes one of its factors is usually saying “now measure instead of move”.
Where this shows up
One measurement, two worlds
Plane normals and signed distance
A plane through the origin is described by its normal vector n, and the quantity n·x tells you how far along that normal a point x lies; dividing by |n| turns it into the signed distance from the plane. This is the same projection, read as a measurement, and it is how image-to-world geometry checks whether a pixel ray meets a surface. The homography part of the geometry guide uses plane equations exactly this way, and a signed distance from a plane is how a controller keeps a robot on the right side of a surface.
Cosine similarity of embeddings
A token embedding is a vector, and the usual way to ask whether two tokens mean similar things is to compute the cosine of the angle between their embeddings — the dot product with lengths divided out. Attention scores, nearest-neighbour retrieval and contrastive training are all built from that one measurement. The tokens chapter builds the vectors that these comparisons run on.
Cheat sheet
Every formula in one place
| Idea | Arrow picture | Algebra |
|---|---|---|
| Dot product | Overlap: |v||w|cosθ | v·w = vxwx + vywy |
| Angle | The turn from one arrow to the other | cosθ = (v·w) / (|v||w|) |
| Scalar projection | Signed length of the shadow of v on w | (v·w) / |w| |
| Vector projection | Arrow from the origin to the foot of the drop | ((v·w)/(w·w)) w |
| Residual | The piece the line cannot reach | v − projwv, with w·v⊥ = 0 |
| Orthogonal test | A right angle at the common tail | v·w = 0 |
| Row matrix | A ruler laid along w | wTx = w·x |
| Level sets | Parallel lines cut across w | w·x = c, normal w |
Further reading
Where to go deeper
- Grant Sanderson, "Dot products and duality", Essence of Linear Algebra, 3Blue1Brown — the video this part is in conversation with, including the duality reveal.
- Gilbert Strang, 18.06 Linear Algebra, MIT OpenCourseWare — Lectures 14–16 on orthogonal vectors, projections onto subspaces, and the normal equations.
- Immersive Math, Chapter 3: The Dot Product — the same overlap and projection, in a draggable textbook.
- Sheldon Axler, Linear Algebra Done Right, chapter 6 — inner products and orthogonality, built abstractly from the axioms rather than from arrows.