Caption matches picture
Picture a party where every guest wears a name tag, and someone has to match each guest to the tag that belongs to them. The vision transformer from the previous part turns an image into tokens; a twin transformer turns a caption into tokens. Neither knows about the other — one has only ever seen pixels and the other only ever seen words — so on their own their outputs live in two unrelated worlds. Contrastive learning is the trick that makes them meet, and the name comes from the fact that it learns by contrasting right answers against wrong ones rather than by describing anything. Pull the vector of an image and the vector of its caption together, push every image away from every caption it does not belong to, and do it for hundreds of millions of pairs. If the same pair keeps being pulled together while the mismatched pairs keep being pushed apart, the two encoders are forced to agree on a common vocabulary of positions. The result is one embedding space — a single coordinate system in which both an image and a sentence become vectors, so that anything from either modality can be placed on the same map — where a picture and a sentence can be compared with a cosine, a measure of how closely two vectors point in the same direction. This part is about the loss that does the pulling, and the two knobs — temperature and batch size — that decide whether it works.
The training signal
No labels, only which caption goes with which picture
Contrastive learning uses the cheapest supervision there is: a matched pair. A matched pair is two things that belong together and are known to belong together — an image and its alt text, a video clip and its transcript, a photo and its caption. Notice that nobody has to write a description of what is in the picture. The label is only the fact that these two items go together, which is why this kind of supervision scales so cheaply: a web page's alt text is already there, and no human has to be paid to annotate it. Think of it as name tags at a party: you do not need a biography of each guest, only the knowledge of which tag belongs to whom. A batch of $N$ pairs — where $N$ is how many pairs are processed together at one step — is handed to two encoders, one for images and one for text, producing $N$ image vectors and $N$ text vectors. Both sets are mapped into the same $d$-dimensional space, meaning each image and each caption comes out as a list of $d$ numbers living in one shared coordinate system. The objective does not ask the model to describe anything; it only asks that the correct pairing be the most similar one.
The structure that makes this work is the in-batch negative. For each image, every other caption in the batch is a wrong answer; a positive pair is a genuinely matched image and caption, and a negative pair is any image and caption that do not go together. The trick is that you never have to go looking for negatives. Within a batch of $N$ pairs, each image has exactly one correct caption and $N-1$ wrong ones, so the batch supplies $N$ positives and $N(N-1)$ negatives for free, all of them already sitting in memory. The loss for an image is a softmax over its similarity to all $N$ captions — softmax being the operation that turns a row of raw similarity scores into a set of positive weights that sum to one, so it can be read as a probability distribution over "which caption is mine". The model is trained to put the probability mass on the true one. That is InfoNCE, short for noise-contrastive estimation, the objective that treats retrieval as a classification problem and rewards the model for ranking the true caption highest. Because the image→text and text→image directions are both summed, the objective is symmetric: the text encoder is judged on finding its image just as the image encoder is judged on finding its caption.
The similarity matrix
A square of predictions, one diagonal of truths
Arrange the batch as a matrix: rows are the $N$ images, columns are the $N$ captions, and the entry is the cosine similarity of the two vectors. A concrete entry might read: image 3 against caption 7, scored 0.82. The diagonal holds the matched pairs — image 1 against caption 1, image 2 against caption 2, and so on — and every off-diagonal cell is a negative, an image paired with a caption that is not its own. Training is the act of making that matrix as diagonal as possible: each row's softmax should put its mass on the cell it sits on. Read this way, contrastive learning is nothing more than $N$ classification problems on the same matrix, one per row, each asking "which of these $N$ captions belongs to this image?". Top-1 retrieval accuracy is then the fraction of rows whose largest entry is on the diagonal — the fraction of images whose own caption scored highest.
The demo below builds a mini-batch of synthetic pairs — each concept has one base vector, and the image and text versions are noisy copies of it, so the diagonal should be the brightest — then draws the cosine matrix. The temperature slider rescales the scores before the softmax. Here it helps to know what the softmax is actually fed: the raw similarity scores that go into it are called logits, and they are the numbers before any normalisation is applied. Temperature is a number that divides those logits first, and because it sits in the denominator it changes how peaked or flat the resulting distribution is. Drag the slider down and the diagonal dominates, with one bright cell per row; drag it up and the matrix goes soft, because a large divisor shrinks the gaps between the cells and every row is normalised into a flatter distribution.
Cosine similarity from GenMedia.attn.simMatrix for the current batch. Diagonal cells are outlined; the readout reports the InfoNCE loss from Multimodal.contrastive.infoNCE.
Temperature and batch size
Two knobs that decide how hard the task is
Temperature divides the logits before the softmax. A small τ makes the distribution sharp, so the loss punishes even small mistakes and the model is pushed to separate near-misses; a large τ softens the distribution and lets it get away with being sloppy, because the correct caption no longer has to stand out by much. Why is the sharp setting the useful one? Because captions in a batch are often similar to one another — "a dog on a wooden bench" and "a dog on a metal bench" share almost all their words — and a model that only ever has to beat plainly wrong answers never learns the fine distinctions that matter in practice. Those confusingly similar captions are the hard negatives: the captions that are almost right. CLIP learns this parameter rather than fixing it, and it settles around $0.07$ — small, because the useful signal is in the hard negatives. The problem with a very small τ is instability and saturation: as the divisor approaches zero the softmax collapses onto a single cell, its gradients stop changing, and training can wobble or stall. That is why τ is clamped rather than driven to zero.
Batch size is the other half. Batch size is how many pairs are processed in one step, and here it does far more than divide the work up. Every extra pair in the batch adds $2(N-1)$ new negatives, one in each direction, and those negatives are what teach fine distinctions: a model that has seen thousands of wrong captions for the same image has had thousands of chances to learn exactly what makes the right one different. That is why contrastive models are trained in enormous batches and why the batch size is described as a capacity parameter: it does not merely change the memory footprint, it changes the difficulty of the task itself, and with it how much the model can learn from each example. The first plot traces the two losses against temperature at the current batch; the second traces them against batch size at the current temperature.
InfoNCE and SigLIP loss against temperature, log-scaled. The dashed line is the current τ.
Loss against batch size at the current temperature. More pairs means more negatives, so the task is harder and the loss rises.
The SigLIP alternative
Drop the softmax, keep the pairs
InfoNCE is a softmax, and a softmax couples every cell in a row to every other. That coupling is the source of its power and of its cost. The power is that each row becomes a clean competition between the true caption and all the others. The cost is that the denominator has to be computed over the whole batch, so no device can finish its row on its own: it forces an all-gather of every embedding across every device, meaning all the machines training the model must send each other their vectors before any of them can compute a loss. SigLIP removes that step. Each cell gets an independent logistic loss on a binary decision — is this pair a match or not — with the diagonal labelled positive and everything else negative. The loss is the logistic, or sigmoid, loss, where the sigmoid is the S-shaped function that squashes any score into a number between 0 and 1 so it can be read as a probability of "match"; it pushes the score for a matching pair up and the score for a non-matching pair down, and nothing more. Because no term is normalised over the batch, the diagonal and the off-diagonal cells are judged individually rather than in competition, so training is cheaper and more parallel, and the loss behaves better at small batch sizes where a row has few negatives to compare against.
The trade is that the objective is no longer a proper probability over the batch; it is a pile of pairwise bets, each one placed without knowing what the other cells scored. Why is that acceptable? Because a proper probability was only ever a means to an end: the thing that matters is that matching pairs score higher than mismatched ones, and the pairwise loss optimises exactly that. What is lost is the mutual calibration between cells — a row no longer sums to one — but the encoders still end up separating matches from non-matches. In practice that is fine, and SigLIP's pairwise form has become the default for open vision encoders. The bars below compare the two losses at the current settings; note that they are different quantities and are not meant to match in value, only in what they are pushing the encoders to do.
The SigLIP generation did not stop there. SigLIP 2 (Tschannen et al., 2025) keeps the pairwise sigmoid loss demonstrated here and adds localization and dense-prediction objectives on top of it, so the same encoder that produces a global embedding also produces usable per-patch features. Localization means the encoder learns where a described thing sits, not only that it is present somewhere in the picture; dense prediction means it is trained to label patches individually rather than to emit one vector for the whole image. That combination is what a modern vision-language model wants: this generation of encoder is what a current open VLM such as Qwen3-VL continues training from, and once the encoder can report detail at the patch level, the next part is where resolution and dense detail become the binding constraint rather than the training objective.
InfoNCE against SigLIP's pairwise sigmoid loss at the current batch and temperature, with the chance-level InfoNCE value $\ln N$ for scale.
SigLIP computes one logistic loss per cell, so all $N^2$ pairs contribute independently; InfoNCE normalises each row across the batch.
Where this shows up
The objective behind the encoder you actually download
CLIP, SigLIP and their descendants
The two towers and the matrix above are CLIP — that is the whole recipe, with nothing else hiding in it. SigLIP swaps the loss for its pairwise version and becomes the encoder inside most current vision-language models. The vision trunk is the patch transformer from the previous part; the architecture did not change here, we only trained it. The contrastive objective is not a new network; it is a training signal wrapped around the encoder you already built.
Zero-shot, retrieval and rankers
Once a shared space exists, building a classifier reduces to a dot product against a class name, and building a retriever reduces to a nearest-neighbour search: embed the query, then find the stored vectors closest to it. That is what makes zero-shot classification possible, where "zero-shot" means the model sorts images into categories it was never given labelled examples of, because it has only ever seen each category as a piece of text. Both are the subject of the next part, where the same alignment that makes zero-shot classification possible also makes its failures predictable.
Beyond images, the identical machinery aligns audio to text (CLAP), video to text and 3-D shapes to text. Nothing about the objective is specific to pictures: all you need is a matched pair in two modalities and an encoder for each, and the same matrix and the same loss do the rest. That is why the modality zoo in the last part is mostly a list of towers plugged into this one loss.
Further reading
Contrastive image-text pretraining is the hinge between vision and language: it is the step that gives a picture and a sentence a common coordinate system, so everything downstream can compare the two. These papers introduce the two-tower recipe, scale it to noisy web data, and then replace the softmax with a simpler pairwise loss.
- Alec Radford, Jong Wook Kim, Chris Hallacy and coauthors, "Learning Transferable Visual Models From Natural Language Supervision", 2021 — CLIP, the symmetric InfoNCE objective and the learned temperature.
- Ting Chen, Simon Kornblith, Mohammad Norouzi and Geoffrey Hinton, "A Simple Framework for Contrastive Learning of Visual Representations", 2020 — SimCLR, where the temperature and the role of negatives were first laid out clearly.
- Chao Jia, Yinfei Yang, Ye Xia and coauthors, "Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision", 2021 — ALIGN, showing that noisy alt text at scale is enough.
- Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov and Lucas Beyer, "Sigmoid Loss for Language Image Pre-Training", 2023 — SigLIP, the pairwise sigmoid loss and the removal of the all-gather.
- Michael Tschannen, Alexey Gritsenko, Xiao Wang and coauthors, "SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features", 2025 — the pairwise loss plus the dense and localization objectives that current encoders inherit.
Cheat sheet
| Term | Meaning here |
|---|---|
| Positive pair | An image and its own caption; the diagonal cell of the matrix |
| In-batch negative | Every other caption in the batch; $N(N-1)$ negatives from $N$ pairs |
| InfoNCE | Row softmax over the similarity matrix, averaged over rows in both directions |
| Temperature τ | Divides the logits before the softmax; small sharpens, around $0.07$ in CLIP |
| Hard negative | A caption that is almost right; the pairs that a small τ forces the model to separate |
| Batch size | A capacity knob: more pairs means more negatives and a harder task |
| SigLIP | Pairwise logistic loss per cell, diagonal positive; no softmax, no all-gather |
| Chance | $\ln N$ for InfoNCE; the loss of a model that guesses uniformly |