Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The question

One object, three costumes

This series is about the machinery that every other guide on this site assumes: the geometry that turns images into 3D structure, the least-squares step that makes a solver converge, the matrix multiplies that train a language model. All of it is linear algebra, and all of it starts here, with the question of what a vector actually is.

The answer is deliberately slippery. A vector is not a particular kind of thing; it is a thing that behaves a particular way. If you can add two of them and get another one of the same kind, and you can stretch one by a number and still have one of the same kind, then you have a vector — whether it is an arrow, a pixel, or a list of a thousand embedding coordinates. Everything else in this series is about what those two operations let you build.

💡 By the end of this part you'll see why the arrow picture and the list-of-numbers picture are the same picture, why adding arrows is the same as adding columns of numbers, and why a robot pose, an RGB pixel and a word embedding are all the same kind of object wearing different costumes.
2

A vector is an arrow

Length and direction, and nothing else

The first answer is the one you can see. A vector is an arrow with a length and a direction: a displacement, a velocity, a force. It has no fixed home. The arrow from the origin to the point (2, 1) and the arrow from (3, 3) to (5, 4) are the same vector, because they have the same length and point the same way. Only the coordinates of the tail differ, and tail coordinates are not part of the vector.

Two operations make arrows do work. Adding two arrows stacks the second on the tip of the first, tip-to-tail, and the result is the arrow from the first tail to the last tip; because the order does not matter, the same sum is the diagonal of the parallelogram they span. Scaling an arrow by a number k keeps its direction and multiplies its length by k, flipping it if k is negative. Drag the two arrows below and watch both operations happen at once. The parallelogram is not decoration: its diagonal is the sum.

Drag the two handles to move v and w; the dashed arrow is the scalar multiple kv from the slider.

Nothing about that picture needed coordinates. If you rotate the page, the arrows rotate with it and their sum is still the diagonal. That is the arrow's whole case: addition and scaling are geometric facts first, and only later do we write them down as arithmetic.

3

A vector is a list of numbers

Coordinates, and the entrywise rules

The second answer is the one you compute with. Lay down a set of axes and a vector becomes an ordered list of coordinates: how far to walk along each axis to get from tail to tip. In two dimensions the arrow above is a pair of numbers, and by long convention we stack them in a column:

$$\mathbf{v} = \begin{bmatrix} v_x \\ v_y \end{bmatrix}, \qquad \mathbf{v}+\mathbf{w} = \begin{bmatrix} v_x + w_x \\ v_y + w_y \end{bmatrix}, \qquad k\mathbf{v} = \begin{bmatrix} k\,v_x \\ k\,v_y \end{bmatrix}$$

That is the whole rulebook. Addition is entrywise; scaling is entrywise. This is why the two answers are the same answer: tip-to-tail addition of arrows, once you choose axes, is exactly adding the pairs of coordinates that describe them. Drag the sliders to set the two vectors coordinate by coordinate and watch the arrow on the left move with the numbers — and, on the right, the sum computed entrywise.

The same two vectors as the demo above, now driven from their coordinates.

A vector in n dimensions is the same idea with n slots. The number of slots is the dimension of the space, and it can be large: a word embedding is a vector with hundreds or thousands of coordinates, and the arrow picture simply stops being drawable. The arithmetic does not stop working. That is the pattern of this whole series — the picture teaches the operation, then the list generalises it.

Notation, fixed once for the series. A vector is a column, v; its transpose vᵀ is the same numbers in a row; the i-th coordinate is vi. Matrices come next and will be written as arrays of rows in code, so a column vector is an n×1 matrix. The glossary keeps every symbol in one place.

4

A vector is anything you can add and scale

The abstract definition, which is the useful one

The third answer throws away both pictures and keeps only the two operations. A vector space is a collection of objects, plus a rule for adding any two of them and a rule for scaling any one of them by a number, such that the obvious bookkeeping holds: addition is commutative and associative, there is a zero vector, every vector has an opposite, scaling by 1 does nothing, and scaling distributes over addition.

Those eight rules sound pedantic until you notice how many things obey them. Polynomials obey them. Functions obey them. Solutions to a linear differential equation obey them. Once a collection obeys them, every theorem in this series applies to it, for free, because the theorems were proved from the rules and nothing else. The arrow and the list were never the point; the operations were.

The single most important operation is the linear combination: take any vectors v₁, …, vk, scale each by a number, and add the results. That is what the rules were for. Every question in Acts I–III — is a system solvable, what is the span of a set, does this transformation preserve dimension — is a question about which vectors a linear combination can reach.

Every value of the sliders lands somewhere in the plane. The shaded region is where linear combinations of v and w can reach.

When the two vectors point in different directions the combinations reach the whole plane, and the shading fills it. When they line up, the reachable set collapses to a line no matter how hard you push the coefficients. That collapse is the subject of Part 2, and it is the first place where linear algebra starts making claims about a picture.

5

Three costumes, one object

The running example starts here

Here is a set of two-dimensional points, drawn as arrows from the origin. Read it three ways. As geometry, it is a cloud of vectors, each with a direction and a length. As arithmetic, it is a list of coordinate pairs you could put in a spreadsheet. As data, it is a running example that this series will transform, fit, rotate and compress — the same cloud, gaining meaning in each act.

The running example: a seeded point cloud read as a set of vectors. Move the slider to pick one out.

A robot pose

A mobile robot's state is a vector: two metres east and one north is the arrow (2, 1), and a heading angle is one more coordinate. Planning is arithmetic on these vectors, and turning a goal into a path is exactly that arithmetic.

An RGB pixel

A colour is a vector (r, g, b): three coordinates, three axes, an arrow in a cube. Blending two images is a linear combination, and every filter in an image pipeline is a matrix acting on that vector — the story of multi-view geometry starts here too.

A word embedding

A token's embedding is a vector with hundreds of coordinates, and “similar meaning” becomes “small angle between two arrows”. The LLM training guide builds these vectors; Part 8 explains how the angle is measured.

A set of joint torques

A robot arm with n joints has an n-dimensional torque vector, and the nonlinear optimization guide searches over exactly such vectors. High-dimensional lists are not exotic; they are what a controller holds in memory every millisecond.

6

Where this shows up

One operation, two worlds

Robotics

Poses, twists and Jacobians

When a robot arm moves, its joint velocities form a vector and its end-effector velocity is a matrix times that vector. The pinhole camera part of the geometry guide uses the same object to describe where a camera is looking; Part 21 returns to it with derivatives.

ML / AI

Embeddings and attention

Every activation in a transformer is a vector, and the model's whole job is adding and scaling them in clever combinations. The architecture chapter introduces the tensors; LLM Serving Part 1 shows why moving them is the expensive part.

7

Cheat sheet

IdeaArrow pictureList picture
VectorLength and direction, no fixed homeAn ordered column of numbers, n×1
AdditionTip-to-tail; the parallelogram diagonalAdd corresponding entries
ScalingStretch, or flip if negativeMultiply every entry by the scalar
Zero vectorZero length, direction undefinedAll entries zero
Linear combinationReach a point by stepping along othersa₁v₁ + … + akvk, entrywise
DimensionDrawable only at 1, 2 or 3Number of entries; no limit

Further reading

8

Check your understanding

0/4 answered