Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Cutting the image up

A patch is a token

A transformer wants a sequence, so the first job is to manufacture one out of an image. The simplest way to do that is to chop the image into non-overlapping $P \times P$ patches, flatten each patch row by row into a vector of $P^2 C$ numbers, and pass that vector through a learned linear projection into the model's width. Take a single concrete patch: a $16 \times 16$ square cut from a colour photograph has $16 \times 16 = 256$ pixels, and each pixel carries three colour channels (red, green and blue), which is the $C$ in $P^2 C$. Flattening lays those $256 \times 3 = 768$ numbers out in one long list. The linear projection — a single learned matrix, nothing more — then rescales that list to the transformer's width, meaning the fixed number of numbers in every vector the model passes around internally, say $1024$. The count of patches is fixed arithmetic: an $H \times H$ image becomes $(H/P)^2$ tokens. At the familiar $224 \times 224$ and $P = 16$ that is $14 \times 14 = 196$ tokens, roughly the length of a short paragraph. Nothing has been learned about the content yet; this step only decides how many pieces the model will reason over and how large each piece is.

The patch is the model's unit of meaning, and its size is a real trade. A small patch keeps fine detail, because each token covers a small area and can represent a small feature precisely. But a smaller patch means more of them: halving $P$ quadruples the token count, and attention compares every token with every other token, so its cost grows with the square of the sequence length. A large patch is cheap for the same reason in reverse, but it throws away exactly the high-frequency structure — sharp edges, texture, fine text — that a convolution would have kept, because any detail smaller than the patch is averaged away the moment the patch is flattened. The slider below changes $P$ on a fixed $32 \times 32$ test image so you can watch the grid and the token count move together.

A procedural test image drawn with GenMedia.raster, with the patch grid overlaid and the first patch outlined. The grid comes from Multimodal.vit.patchify.

💡 By the end of this part you'll be able to compute the token count (H/P)² for any image and patch size, explain what position embeddings (the addresses that tell the model where each patch sat) and the class token (the one vector asked to summarise the whole picture) are for, read an attention map, and say what a ViT has to learn that a convolution is born knowing.
2

Tokens, positions and the class token

Order the patches cannot see

Attention begins by comparing every token against every other token, and that comparison has no idea where any token sits. Formally, attention is permutation invariant: shuffle the patches of an image, and the unmodified transformer produces the same set of outputs, only reordered. That is fatal, because a photo is not a bag of patches — a face, a horizon and a skyline mean something only because of where their parts are relative to one another. The fix is to add a position embedding to every patch token, so that two otherwise identical patches in different places look different to the model. You can picture it as writing a seat number on each token after a shuffle: the token's content stays the same, but its address travels with it. The original ViT used fixed sinusoidal embeddings, the construction shown below, in which each address is a fixed pattern of sines and cosines so that nearby rows get similar patterns; later work learned them as parameters instead, and both work, because the signal the model needs is just "where am I" and there are many encodings of it.

The second repair is the class token. A classifier needs one vector to represent the whole image, and the natural choice is an extra learned token prepended to the sequence — placed at the front, before all the patch tokens. Attention lets that token gather from every patch, so after the last block its output is a global summary of the whole picture rather than a description of one square of it; a small head on top of it produces the logits, the raw scores, one per class, that a softmax later turns into probabilities. More recent encoders drop the class token and pool the patch tokens instead, meaning they average their final vectors into one, but the idea is the same: some vector has to stand for the whole picture.

Sinusoidal 2-D position embeddings from Multimodal.vit.positional, one row per patch token and one column per embedding dimension. The diagonal bands are the periodic encoding that tells the model where a patch sits.

Self-supervised encoders such as DINOv2 skip labels entirely: they train the same patch-token trunk with a teacher-student objective, and the resulting features transfer to segmentation and depth without fine-tuning.

⚠ Position embeddings are not free. They are fixed to a grid, so an encoder trained at $224 \times 224$ does not automatically accept a different resolution or aspect ratio: there is no embedding for patch 20 of 21 if the model only ever learned positions up to 14 of 14. Interpolating the embeddings is the usual fix — stretching the learned grid of addresses to fit the new one, like renumbering the seats when the room is rearranged — and it is one of the reasons high-resolution multimodal encoders are trained with variable-size inputs in the first place.
3

What attention looks at

Weights you can read as a map

Every block computes attention weights: for each token, a softmax over how much it looks at every other token. In this operation each token plays three roles at once. It asks a question (the query), it advertises what it has to offer (the key), and it carries the information itself (the value). The weight between two tokens measures how well one token's query matches another's key. Because those weights form a row that sums to one, they can be drawn directly as a heatmap, and the picture is legible. A row is a query, a column is a key, and a bright cell means the query token leaned on that key. Deep in a trained network the maps often localise on objects — a patch sitting on a dog attends to the other patches on the dog — while early in training they are nearly uniform, which is what the temperature slider exaggerates here.

Temperature divides the logits before the softmax, and its effect is easiest to see at the extremes: dividing by a small number lets the largest scores dominate, so a small temperature sharpens each row onto a few patches, while dividing by a large number flattens the differences, so a large temperature spreads attention evenly. That knob is the same one that governs the contrastive objectives in the next part; it is the standard way to trade off "attend precisely" against "attend broadly", and it is why the map below changes character — from a few bright spots to a wash — rather than merely getting brighter.

Attention weights from Multimodal.vit.attention on the patch tokens of the test image. Rows are queries, columns are keys; each row is a softmax.

💡 A heatmap you can read is a diagnostic you can act on. When a vision encoder is paired with a language model, inspecting which patches the image tokens attend to is the first thing to reach for when the model ignores a region or latches onto the wrong one. If a region is never attended to, the model cannot have used it, so the map tells you where to look next.
4

Convolutional bias versus learned bias

Built in against learned

A convolution is a small filter that slides across an image and responds to local patterns. It is born with two assumptions. The first is locality: nearby pixels matter to each other, so it is sensible to look at a small neighbourhood at a time. The second is translation equivariance: a feature is the same feature wherever it appears, so the same filter is reused at every position, and a corner learned in the top-left is recognised in the bottom-right. Stacking $3 \times 3$ convolutions grows the receptive field — the region of the input that can influence one output — by two pixels per layer, so an image-wide view of a $32 \times 32$ picture takes a dozen layers and an ImageNet-sized one takes far more. That slow growth is a restriction, and it is also a gift: the model does not have to learn from scratch that a corner is a corner in every location.

A patch transformer has the opposite profile. Self-attention connects every token to every other in the very first layer, so its receptive field is global immediately. The price is that locality and translation equivariance are not built in at all; they have to be discovered from data. In practice that makes ViTs data-hungry: they need far more examples than a comparable convolution to learn the same regularity, and they depend on strong augmentation or large-scale pretraining. Augmentation means manufacturing extra training variety from each image, with random crops, flips and colour shifts. It is also why hybrid designs exist — a convolutional stem in front of the transformer, attention restricted to local windows, or distillation from a trained convolution, in which the convolution's predictions become soft targets the transformer learns to match — all of them ways to hand the transformer a little of the prior it lacks.

Fraction of the image a token can see as a function of depth, for a $3 \times 3$ convolutional stack and for self-attention. The dashed line marks the layer depth on the slider.

⚠ Global from layer one is not automatically better. A model that can see everything immediately has more freedom to fit spurious long-range correlations — coincidences between distant pixels that happen to line up in the training set but mean nothing in the world. The inductive-bias trade is real in both directions: convolutions generalise from less data because their assumptions are good ones, and attention scales further when there is enough data for the model to learn those assumptions for itself.
5

Where this shows up

Every vision tower is this encoder with a different training signal

Encoders

The vision half of a VLM

Every vision-language model begins here: a patch transformer produces image tokens, and a projector — a small learned module that turns vision vectors into vectors the language model accepts — feeds them in. The encoder's patch grid and its token count set the resolution budget, which is exactly the token arithmetic and tiling you meet in the shared embedding space and in the resolution discussion of the sibling volume. Tiling means splitting a large image into several fixed-size pieces and patchifying each of them, so that detail survives without letting the token count grow without limit.

Pretraining

Contrastive and self-supervised trunks

CLIP and SigLIP train this trunk — the shared backbone network, in this case the patch encoder — against text with a contrastive objective, in which matching image-caption pairs are pulled together and mismatched pairs are pushed apart. That is the next part. DINOv2 trains the same trunk with no labels at all, using a self-supervised recipe in which a student network learns to reproduce a teacher's features on different views of the same image. That is why its patch features segment objects so cleanly: each patch vector has learned to describe the object it belongs to rather than the scene as a whole. Either way the architecture is the one built above; only the loss changes.

Further out, the same patch encoder reappears in video models with an added temporal axis, so a token can stand for a small patch of a single frame gathered across time as well as space. It reappears in audio encoders, which turn a waveform into a spectrogram — a picture of how much energy each frequency carries over time — and treat each mel band, a coarsened range of frequencies, as a row of tokens. And it reappears in 3-D point encoders, which group a point cloud into local neighbourhoods the way patchify groups pixels. The modality zoo in the last part is mostly a list of ways to turn a signal into a patch sequence and then align it to a text tower.

Further reading

The vision transformer started as "an image is worth 16×16 words" and grew into the default vision backbone. These papers cover the original architecture, the data-efficient training that made it practical, the self-supervised version that produces general-purpose features, and the windowed hybrid that puts locality back. Read them in that order and the arc is clear: a clever idea that needed an unreasonable amount of data, the recipe that fixed the data problem, a label-free alternative, and a design that reintroduces the prior the first paper threw away.

Cheat sheet

TermMeaning here
PatchifyCut an image into $P \times P$ patches and flatten each to a vector of $P^2 C$ numbers
Token count$(H/P)^2$ for a square image; the sequence length the transformer sees
Patch size PSmaller keeps detail and costs sequence length; larger is cheaper and blurs
Position embeddingAdded so permutation-invariant attention knows where a patch came from; learned or sinusoidal
Class tokenAn extra learned token whose final output is the whole-image summary
Attention mapA row of softmax weights per query; readable as a heatmap over patches
Inductive biasLocality and translation equivariance: a convolution has them, a ViT learns them
Data hungerViTs need large-scale pretraining or strong augmentation to match conv generalisation
7

Check your understanding

0/4 answered