Pixels into tokens
A language model reads a sequence of discrete tokens — separate, countable units, each one a whole integer id rather than a continuous value — and it expects them in order. A picture, by contrast, arrives as a dense grid of pixels, with no natural beginning, end, or word boundaries. The vision transformer closes that gap with the bluntest tool available: cut the image into patches, flatten each patch into a vector, and pretend the result is a sentence. Flattening here means reading the patch's pixels row by row and laying their colour values out in one long list, which is what turns a small square of an image into a single vector the model can treat like a token. Everything else — the position embeddings, the class token, the attention maps — exists to repair the two things that this trick destroys, which are order and locality. Order is lost because the transformer sees a set of tokens with no built-in sense of sequence, and locality is lost because nothing in the architecture knows that neighbouring pixels are related. This part builds that encoder patch by patch.
Cutting the image up
A patch is a token
A transformer wants a sequence, so the first job is to manufacture one out of an image. The simplest way to do that is to chop the image into non-overlapping $P \times P$ patches, flatten each patch row by row into a vector of $P^2 C$ numbers, and pass that vector through a learned linear projection into the model's width. Take a single concrete patch: a $16 \times 16$ square cut from a colour photograph has $16 \times 16 = 256$ pixels, and each pixel carries three colour channels (red, green and blue), which is the $C$ in $P^2 C$. Flattening lays those $256 \times 3 = 768$ numbers out in one long list. The linear projection — a single learned matrix, nothing more — then rescales that list to the transformer's width, meaning the fixed number of numbers in every vector the model passes around internally, say $1024$. The count of patches is fixed arithmetic: an $H \times H$ image becomes $(H/P)^2$ tokens. At the familiar $224 \times 224$ and $P = 16$ that is $14 \times 14 = 196$ tokens, roughly the length of a short paragraph. Nothing has been learned about the content yet; this step only decides how many pieces the model will reason over and how large each piece is.
The patch is the model's unit of meaning, and its size is a real trade. A small patch keeps fine detail, because each token covers a small area and can represent a small feature precisely. But a smaller patch means more of them: halving $P$ quadruples the token count, and attention compares every token with every other token, so its cost grows with the square of the sequence length. A large patch is cheap for the same reason in reverse, but it throws away exactly the high-frequency structure — sharp edges, texture, fine text — that a convolution would have kept, because any detail smaller than the patch is averaged away the moment the patch is flattened. The slider below changes $P$ on a fixed $32 \times 32$ test image so you can watch the grid and the token count move together.
A procedural test image drawn with GenMedia.raster, with the patch grid overlaid and the first patch outlined. The grid comes from Multimodal.vit.patchify.
Tokens, positions and the class token
Order the patches cannot see
Attention begins by comparing every token against every other token, and that comparison has no idea where any token sits. Formally, attention is permutation invariant: shuffle the patches of an image, and the unmodified transformer produces the same set of outputs, only reordered. That is fatal, because a photo is not a bag of patches — a face, a horizon and a skyline mean something only because of where their parts are relative to one another. The fix is to add a position embedding to every patch token, so that two otherwise identical patches in different places look different to the model. You can picture it as writing a seat number on each token after a shuffle: the token's content stays the same, but its address travels with it. The original ViT used fixed sinusoidal embeddings, the construction shown below, in which each address is a fixed pattern of sines and cosines so that nearby rows get similar patterns; later work learned them as parameters instead, and both work, because the signal the model needs is just "where am I" and there are many encodings of it.
The second repair is the class token. A classifier needs one vector to represent the whole image, and the natural choice is an extra learned token prepended to the sequence — placed at the front, before all the patch tokens. Attention lets that token gather from every patch, so after the last block its output is a global summary of the whole picture rather than a description of one square of it; a small head on top of it produces the logits, the raw scores, one per class, that a softmax later turns into probabilities. More recent encoders drop the class token and pool the patch tokens instead, meaning they average their final vectors into one, but the idea is the same: some vector has to stand for the whole picture.
Sinusoidal 2-D position embeddings from Multimodal.vit.positional, one row per patch token and one column per embedding dimension. The diagonal bands are the periodic encoding that tells the model where a patch sits.
Self-supervised encoders such as DINOv2 skip labels entirely: they train the same patch-token trunk with a teacher-student objective, and the resulting features transfer to segmentation and depth without fine-tuning.
What attention looks at
Weights you can read as a map
Every block computes attention weights: for each token, a softmax over how much it looks at every other token. In this operation each token plays three roles at once. It asks a question (the query), it advertises what it has to offer (the key), and it carries the information itself (the value). The weight between two tokens measures how well one token's query matches another's key. Because those weights form a row that sums to one, they can be drawn directly as a heatmap, and the picture is legible. A row is a query, a column is a key, and a bright cell means the query token leaned on that key. Deep in a trained network the maps often localise on objects — a patch sitting on a dog attends to the other patches on the dog — while early in training they are nearly uniform, which is what the temperature slider exaggerates here.
Temperature divides the logits before the softmax, and its effect is easiest to see at the extremes: dividing by a small number lets the largest scores dominate, so a small temperature sharpens each row onto a few patches, while dividing by a large number flattens the differences, so a large temperature spreads attention evenly. That knob is the same one that governs the contrastive objectives in the next part; it is the standard way to trade off "attend precisely" against "attend broadly", and it is why the map below changes character — from a few bright spots to a wash — rather than merely getting brighter.
Attention weights from Multimodal.vit.attention on the patch tokens of the test image. Rows are queries, columns are keys; each row is a softmax.
Convolutional bias versus learned bias
Built in against learned
A convolution is a small filter that slides across an image and responds to local patterns. It is born with two assumptions. The first is locality: nearby pixels matter to each other, so it is sensible to look at a small neighbourhood at a time. The second is translation equivariance: a feature is the same feature wherever it appears, so the same filter is reused at every position, and a corner learned in the top-left is recognised in the bottom-right. Stacking $3 \times 3$ convolutions grows the receptive field — the region of the input that can influence one output — by two pixels per layer, so an image-wide view of a $32 \times 32$ picture takes a dozen layers and an ImageNet-sized one takes far more. That slow growth is a restriction, and it is also a gift: the model does not have to learn from scratch that a corner is a corner in every location.
A patch transformer has the opposite profile. Self-attention connects every token to every other in the very first layer, so its receptive field is global immediately. The price is that locality and translation equivariance are not built in at all; they have to be discovered from data. In practice that makes ViTs data-hungry: they need far more examples than a comparable convolution to learn the same regularity, and they depend on strong augmentation or large-scale pretraining. Augmentation means manufacturing extra training variety from each image, with random crops, flips and colour shifts. It is also why hybrid designs exist — a convolutional stem in front of the transformer, attention restricted to local windows, or distillation from a trained convolution, in which the convolution's predictions become soft targets the transformer learns to match — all of them ways to hand the transformer a little of the prior it lacks.
Fraction of the image a token can see as a function of depth, for a $3 \times 3$ convolutional stack and for self-attention. The dashed line marks the layer depth on the slider.
Where this shows up
Every vision tower is this encoder with a different training signal
The vision half of a VLM
Every vision-language model begins here: a patch transformer produces image tokens, and a projector — a small learned module that turns vision vectors into vectors the language model accepts — feeds them in. The encoder's patch grid and its token count set the resolution budget, which is exactly the token arithmetic and tiling you meet in the shared embedding space and in the resolution discussion of the sibling volume. Tiling means splitting a large image into several fixed-size pieces and patchifying each of them, so that detail survives without letting the token count grow without limit.
Contrastive and self-supervised trunks
CLIP and SigLIP train this trunk — the shared backbone network, in this case the patch encoder — against text with a contrastive objective, in which matching image-caption pairs are pulled together and mismatched pairs are pushed apart. That is the next part. DINOv2 trains the same trunk with no labels at all, using a self-supervised recipe in which a student network learns to reproduce a teacher's features on different views of the same image. That is why its patch features segment objects so cleanly: each patch vector has learned to describe the object it belongs to rather than the scene as a whole. Either way the architecture is the one built above; only the loss changes.
Further out, the same patch encoder reappears in video models with an added temporal axis, so a token can stand for a small patch of a single frame gathered across time as well as space. It reappears in audio encoders, which turn a waveform into a spectrogram — a picture of how much energy each frequency carries over time — and treat each mel band, a coarsened range of frequencies, as a row of tokens. And it reappears in 3-D point encoders, which group a point cloud into local neighbourhoods the way patchify groups pixels. The modality zoo in the last part is mostly a list of ways to turn a signal into a patch sequence and then align it to a text tower.
Further reading
The vision transformer started as "an image is worth 16×16 words" and grew into the default vision backbone. These papers cover the original architecture, the data-efficient training that made it practical, the self-supervised version that produces general-purpose features, and the windowed hybrid that puts locality back. Read them in that order and the arc is clear: a clever idea that needed an unreasonable amount of data, the recipe that fixed the data problem, a label-free alternative, and a design that reintroduces the prior the first paper threw away.
- Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov and coauthors, "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale", 2020 — the patchify-and-attend vision transformer and the class token.
- Hugo Touvron, Matthieu Cord, Matthijs Douze and coauthors, "Training data-efficient image transformers & distillation through attention", 2020 — DeiT, the distillation token and the augmentation recipe that made ViTs trainable without giant datasets.
- Maxime Oquab, Timothée Darcet, Théo Moutakanni and coauthors, "DINOv2: Learning Robust Visual Features without Supervision", 2023 — the self-supervised trunk whose patch features transfer to segmentation and depth.
- Ze Liu, Yutong Lin, Yue Cao and coauthors, "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows", 2021 — locality reintroduced as a window prior.
Cheat sheet
| Term | Meaning here |
|---|---|
| Patchify | Cut an image into $P \times P$ patches and flatten each to a vector of $P^2 C$ numbers |
| Token count | $(H/P)^2$ for a square image; the sequence length the transformer sees |
| Patch size P | Smaller keeps detail and costs sequence length; larger is cheaper and blurs |
| Position embedding | Added so permutation-invariant attention knows where a patch came from; learned or sinusoidal |
| Class token | An extra learned token whose final output is the whole-image summary |
| Attention map | A row of softmax weights per query; readable as a heatmap over patches |
| Inductive bias | Locality and translation equivariance: a convolution has them, a ViT learns them |
| Data hunger | ViTs need large-scale pretraining or strong augmentation to match conv generalisation |