What a shared space buys
Contrastive pretraining leaves behind a shared embedding space: a single coordinate system in which both images and sentences become vectors, so that an image vector and a text vector can be compared with a cosine — a measure of how closely two vectors point in the same direction, where a value near 1 means they point the same way, a value near 0 means they are unrelated, and a negative value means they point opposite ways. Picture a city map with neighbourhoods: similar things are placed on the same block, and the distance between two points says how related they are. That one property pays for a surprising amount of behaviour: classify an image into classes that were never in the training set, search a gallery with a sentence, and steer a model with a prompt instead of a retrained head. Ask for "a dog on a beach" and a retrieval system returns the photographs whose vectors sit nearest those words, even though no one ever labelled those particular photographs with that phrase. It also imports a specific set of failures, because a cosine flattens a whole vector down to one number, so it measures overall similarity and knows nothing about order, number or letters. This part spends the alignment, then catalogs what it cannot do.
One space
Images and sentences, measured with the same ruler
After training, the two towers have done something no classifier head does: they have placed pictures and sentences in the same coordinates. Picture a city map in which every image and every sentence has an address: a photograph of a red square sits on the same block as the words "a red square", and both are far from the block that holds "a blue circle". An image vector is close to the text vector of its caption and far from the text vector of an unrelated one. That is the entire interface. Everything this part demonstrates — zero-shot classification (sorting images into classes the model was never given labelled examples of), retrieval (ranking a gallery by similarity to a query), prompt steering — is a cosine and an argmax on top of it, with no additional training at all. Compare that with the usual classifier, which is a separate small network trained on labelled examples of exactly the classes you care about: there is no such head here, so a new class never requires a new round of training. The argmax is the operation that picks whichever option scored highest, and nothing about it needs to know what a picture is.
The matrix below is the aligned space made visible: rows are the image vectors of a set of concepts, columns are the text vectors of their captions, and the entry is their cosine similarity — the number that says how closely that image and that caption point in the same direction. The diagonal holds the matched pairs, so compare name tags at a party: each image should be closest to the caption that belongs to it, the way each guest should be closest to their own tag. A well-trained model has a bright diagonal and a dark off-diagonal, which is exactly the structure the contrastive loss was pushing for. The mean gap between the two is the margin the rest of this part gets to spend. That gap is worth pausing on, because every ability in this part is a decision made by ranking cosines, and a model with a large margin can rank correctly even when a question is phrased in a way it has never seen.
Cosine similarity of image against text embeddings (GenMedia.attn.simMatrix), with the diagonal outlined. Loaded from multimodal-embeddings.json.
Zero-shot classification
A classifier you build out of sentences
To classify an image with no training examples, write the class names down, embed them, and take the softmax over their cosines with the image. Softmax here is the same normalising step you have met before: it takes a row of raw scores — in this case the cosines, one per class — and turns them into positive numbers that sum to one, so the row can be read as a probability distribution over the classes. That is zero-shot classification, where "zero-shot" means the model sorts images into categories it was never given labelled examples of; it works because the text tower already placed "red" near the images that look red. The classes are not fixed when the model is trained; they are whatever strings you type at inference, which is why new classes need no new labels and no gradient step. That freedom has a cost worth naming up front: because the classifier was never calibrated on the classes you actually care about, its probabilities can be confidently wrong, and adding a class to the list changes the scores of the others.
Each class vector here is built from the embedding set by averaging the text vectors of the concepts that share an attribute value — every "red" concept is pooled into one direction, every "blue" concept into another — so the block of class names you select is a genuine partition of the data, with every image belonging to exactly one class. Change the concept being classified, the attribute block, and the softmax temperature, and watch the probabilities move. In a real system the class vector would come from a prompt such as "a photo of a red object": a prompt is the sentence you feed to the text tower to stand in for a class, and its exact wording is part of the input. That is the subject of the next step.
Class probabilities from Multimodal.contrastive.zeroShot: cosines of the query image against each class vector, softmaxed at the current temperature.
Retrieval and word order
Rank a gallery with a sentence
Retrieval is the same cosine run in the other direction. Embed a query — a caption, a sentence, a keyword — and rank a gallery of image vectors by similarity, returning the closest matches first. This is a soft keyword search over meaning rather than over literal text: the query and the gallery need not share a single word, they land near each other on the city map. There is no classifier and no index training; the aligned space already orders the gallery. The top-$k$ list below comes from Multimodal.retrieval.topK, and the score is just the cosine the training objective arranged to be meaningful. The k is how many results you keep: a search interface might show the best five, while a deduplication pass might keep everything above a fixed score. Ranking costs one cosine per gallery item, which is why this scales to millions of images with an approximate index, and it needs no labels for the gallery at all.
The second canvas is where the failure begins. Rewriting a caption into a different template — a template being the sentence frame around the attributes, such as "a photo of a ___" — or reordering its words produces a different text vector, but for a model trained on similarity the difference is often tiny. That is surprising at first: two sentences a reader would call clearly different come out as almost the same point. The bars compare the image against four caption variants: its own caption, the same attributes in a different word order, a template naming the wrong colour, and a bare shape template. The reordered caption sits almost on top of the original, which is the bag-of-words problem. A "bag of words" discards order and keeps only the collection of words, and that is what this space has learned: it encodes which attributes are present, not how they are related.
Top-5 gallery matches for the selected concept's caption, scored by Multimodal.retrieval.topK.
Cosine of the concept's image to several caption templates. The reordered caption lands almost exactly on the original.
What breaks
Four failures that follow from using a cosine
The compositionality failure above is the general pattern. Compositionality is the idea that the meaning of a phrase is built from its parts and the way they are combined, so "a horse riding a person" ought to mean something different from "a person riding a horse"; here relations between attributes are flattened into a sum-like code. Why a sum? Because pooling and averaging are the only operations the two-tower setup uses to combine information, and an average has no notion of which word modified which. "A horse riding a person" and "a person riding a horse" share almost all their words and land near each other, and a model that measures similarity cannot tell which relation is which. The same flatness costs counting: "two red circles" and "three red circles" differ in one token, so the vectors are close and the predicted count is unreliable. This counting failure is easy to see once you know where to look, because a count is a discrete, ordinal property — two is less than three and the step between them is exact — while a cosine is a continuous, symmetric one that has no slot for "how many".
Typographic attacks exploit the text tower directly. A typographic attack is printed text inside the picture, and it works because the vision tower was trained on images of signs and labels, so it can read: pasting the word "ipod" onto an apple can flip the label, because the image encoder sees letters it was trained to associate with text and the word fights with the fruit. Negation is a fourth casualty — "a photo without a dog" embeds near "a photo of a dog", so asking for the absence of a thing tends to retrieve the thing. None of these are bugs in the training code; they are the predictable consequence of squeezing two very different structures into one metric.
The failures are also why the encoder in a vision-language model is used differently from the encoder in a retrieval index. In both cases the same trunk produces the vectors, but what happens next differs. A VLM that has to answer a question about an image gets a language model after the vision tower, and the language model can reason about relations — who is riding whom, how many objects there are, whether something is absent — that the embedding space alone cannot represent. The shared space is a fast, cheap, untrained classifier and search index; it is not a reasoning module. Keep the two jobs separate in your head: one is a lookup, the other is a conversation.
Where this shows up
The reasons the aligned space is built at all
Text-to-image and image-to-text retrieval
Every photo search box is a nearest-neighbour query in a space like this one: embed the text you typed, then return the stored image vectors whose cosine to it is highest. The same geometry powers deduplication, dataset filtering and content moderation, all by thresholding a cosine — choosing a cutoff and keeping everything above it — rather than training a classifier. Near-duplicate photos, for instance, collapse onto almost the same point, so "too close to an existing item" becomes a quantity you can compare against a number.
The encoder inside a VLM
A vision-language model uses a contrastively trained encoder like SigLIP as its eyes, from the previous part. The alignment gives the projector a meaningful starting point: the image vectors already sit in a space the text tower understands, so the small learned module that bridges the two has a sensible direction to move in rather than a random one. The failures are then handled by the language model on top. The patch transformer underneath is the same trunk throughout.
The same recipe scales past images. Aligning audio to text (CLAP), video to text and 3-D shapes to text produces the modality zoo of the next part, where the shared space becomes a hub that every tower has to agree on. Nothing in the objective is specific to pictures: supply a matched pair in any two modalities, keep the same cosine and the same softmax, and the same map with neighbourhoods appears.
Further reading
The shared space is best understood through its uses and its documented failures, because the same alignment that makes the successes possible is also what produces the failures. These papers cover zero-shot transfer, the failure taxonomy, and the typographic attack that shows the image tower reading letters. Read them together and the picture is coherent: one mechanism with two faces.
- Alec Radford, Jong Wook Kim, Chris Hallacy and coauthors, "Learning Transferable Visual Models From Natural Language Supervision", 2021 — zero-shot transfer and the prompt templates that make it work.
- Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri and coauthors, "When and Why Vision-Language Models Behave like Bags-of-Words", 2022 — the compositionality failure made quantitative.
- Gabriel Goh, Nick Cammarata, Chelsea Voss and coauthors, "Multimodal Neurons in Artificial Neural Networks", 2021 — typographic attacks and the text-like features in the vision tower.
- Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov and Lucas Beyer, "Sigmoid Loss for Language Image Pre-Training", 2023 — the encoder family most current VLMs are built on.
Cheat sheet
| Term | Meaning here |
|---|---|
| Shared embedding space | One coordinate system where image and text vectors are compared directly |
| Zero-shot classification | Embed class names, take the softmax over cosines with the image; no training |
| Class vector | The text embedding of a class name or prompt; the only input you control |
| Prompt template | The wrapper around a class name — "a photo of a …" — which changes the vector |
| Retrieval | Rank a gallery by cosine to a query; top-k is the answer |
| Compositionality failure | Attribute order is flattened; "red square" and "square that is red" match alike |
| Counting failure | Number words are weak features; "two" and "three" embed close together |
| Typographic attack | Printed text in the image steers the vision tower's label |