The modality zoo
Contrastive pretraining — training that pulls matching pairs of examples together and pushes mismatched ones apart — is the recipe that aligned pictures to captions. A "pair" here means two things that belong together, such as a photograph and the sentence describing it. It turns out that the same recipe works for any pair of signals that can be tokenised and embedded, where tokenising means chopping a signal into small discrete units and embedding means mapping each unit to a vector of numbers. A spectrogram — a picture of how much energy each frequency carries over time, which is how sound is usually turned into an image — can be paired against a transcript. A video clip can be paired against a description, and a point cloud — a set of 3-D coordinates sampled from a shape's surface — against a shape name. The wilder version of the idea is to align all of them at once, into a single space where a photo of a barking dog, the sound of a bark and the sentence "a dog barking" all sit near one another. Picture a city map with neighbourhoods: similar things land on the same block, and the distance between two points tells you how related they are; here the map holds sounds and shapes as well as images and words. A modality, in this vocabulary, is one kind of signal — image, text, audio, video or 3-D — and a tower is the encoder network that a single modality is pushed through. This part is about how those towers are built, what a hub modality (one anchor that every other tower is trained against) is for, and the careful reading of the word "aligned".
More than pictures
Every signal becomes tokens, then a vector
The recipe is uniform even though the signals are not. Every modality — a modality being one kind of signal, such as an image, a sound, a video or a 3-D shape — is first turned into a sequence of units. For an image those units are patches, small square tiles cut from the grid. For audio they are mel-spectrogram frames: a mel spectrogram is a picture of how much energy each frequency band carries over time, and each frame is one thin vertical slice of it. For video they are sampled frames, and for a 3-D shape they are points or voxels, where a voxel is the 3-D equivalent of a pixel, a small cube of space. You then run a transformer over that sequence, and pool the output into one vector, pooling meaning combining the many per-unit vectors into a single summary vector, usually by averaging them. A text tower does exactly the same thing with subwords, the word pieces a tokeniser splits text into. Once every tower produces a vector of the same width — the same fixed number of numbers per vector — the contrastive loss from two parts back applies without modification: matched pairs together, everything else apart. That clause is the whole objective; everything else in this part is a change of front end, meaning the tokeniser that turns a raw signal into a sequence.
What changes from modality to modality is the tokeniser, meaning the rule that decides what counts as one unit, and the cost of running the transformer over the resulting sequence. An image is one grid, so its token count is fixed by the patch size. A video is a grid per frame, and it adds a second, harsher scaling problem in time: more frames means more tokens, and a long clip can easily produce a sequence too long to attend over cheaply. Audio is a time-frequency image whose frame rate sets the token count, so a higher sampling rate buys finer detail at a higher price. 3-D data arrives unordered and sparse: a point cloud has no natural reading order the way pixels do, and most of the space it describes is empty, so the tokeniser has to group nearby points into neighbourhoods before the transformer can use them. The cards below are the zoo as the site's data lists it: five encoders — one network per modality, each turning its signal into vectors — five contrastive pairings, and five different notions of what one token is.
These rows come from Multimodal.modalities. Each pair is a genuine contrastive objective in the literature; the point is that the architecture is the same and only the front end differs.
Audio and video towers
The same matrix, a different front end
CLAP is the audio analogue of CLIP — the model from the earlier parts that matched images to captions. CLAP swaps the image tower for an audio tower, so the pairing becomes sound against text. The audio tower consumes a mel spectrogram through a transformer, the text tower consumes a caption, and an InfoNCE loss — the specific contrastive objective that treats every other item in a batch as a negative example for a given pair — pulls matched audio-caption pairs together. The tokens are time-frequency cells, meaning each cell records the energy in one frequency band during one short slice of time. The caption might be "a dog barking" or the audio's tags, where tags are a few keywords describing the clip. Trained this way, the result classifies sound events, meaning it labels a clip as a bark, a siren or a doorbell, and retrieves audio with text, meaning it takes a typed phrase and returns the matching recordings. Video follows the same shape with a spatiotemporal torso, "spatiotemporal" meaning the model covers both space within a frame and time across frames. Concretely, you patchify each sampled frame, cutting it into patches the way the vision transformer does, add the temporal axis so the model can tell which frame a patch came from, and contrast clips against captions with VideoCLIP or InternVideo.
The demo builds one base vector per concept, then gives each modality its own noisy copy of that vector, and draws the cross-modal similarity matrix: a grid whose cell in row i and column j is the cosine between concept i in one modality and concept j in the other. The cosine is the standard similarity score for these spaces, measuring how closely two vectors point in the same direction rather than how far apart they are. The diagonal holds the matching pairs, the name-tags-at-a-party idea from the previous part, where each item should be closest to its own partner. The noise slider controls how well the towers agree. At low noise the diagonal is bright and the alignment is clean, so every matched pair is each other's nearest neighbour. Turn it up and the signal dissolves into a near-uniform matrix, which is what a badly trained or badly paired objective looks like: no better than guessing.
Cross-modal cosine similarity from GenMedia.attn.simMatrix for the selected pair. Diagonal cells are outlined; the readout reports top-1 alignment.
3-D and the hub problem
One anchor, many spokes
Aligning two modalities needs data in which the two co-occur, meaning examples where the same content appears in both at once, such as a clip and its transcript. Aligning five does not scale that way, because not every pair has such a dataset: there is no large corpus of paired point clouds and dog barks. Training every pair directly would need a dataset for each of the ten pairs, and most of those datasets do not exist. The workaround is a hub modality, a single anchor modality that every other tower is trained against instead of being trained against one another. Text is the usual choice, and the reason is largely practical: people have written captions, transcripts and descriptions for almost everything, so text can be paired with almost any other signal. Each tower is trained only against the hub, and transitivity does the rest: if audio lands near text and text lands near 3-D, then audio and 3-D are indirectly comparable, because both sit in the neighbourhood of the same text. The cost of the shortcut is that audio and 3-D never actually meet in training, so any error either one makes against text is inherited by that pair. ImageBind takes the extreme version, using images as the hub and binding six modalities through vision alone.
3-D encoders such as ULIP and OpenShape are the clearest case. A point cloud or voxel grid is encoded by a point transformer — a transformer whose tokens are groups of nearby 3-D points rather than image patches — then contrasted against shape captions or against rendered images of the same object. Through that objective the shape inherits a position in the same space as a sentence, so "a wooden chair" and the coordinates of a chair end up near each other. Notice that the comparison is not always point cloud against text directly: routing it through a rendered view of the shape is another instance of the hub idea, with images playing the anchor role. The diagram shows the arrangement: a central tower that every other one is pulled toward, and spokes that are trained one at a time.
A hub-modality diagram: each tower is trained to land near the hub, and no tower is trained directly against another.
A hub is a modelling choice, not a law. Text is convenient because captions exist for everything; images are convenient because vision models are strong. Either way, cross-modal skill that was never trained is inherited by transitivity, with all the error that implies.
What aligned does not mean
Four readings of one word
Aligned does not mean interchangeable. Two towers landing in one space does not let you feed a spectrogram to a text encoder or read an image vector as a sentence. The space is comparable, not shared in the sense of substitutable: both towers agree on how to measure similarity, but they do not agree on how to represent a signal, so a vector from one is meaningless to the other's decoder. There is also usually a measurable modality gap: embeddings of one modality occupy a slightly different region even for matched pairs, so a within-modality similarity — two audio clips compared with each other — is typically higher than a matched cross-modal one, such as an audio clip compared with its own caption. That gap is not a bug so much as the fingerprint of two encoders trained separately, which never had a reason to occupy identical coordinates. The bars below show that gap.
Aligned does not mean equivalent information. A point cloud and a caption both describe a chair, but the caption cannot tell you the chair's dimensions and the point cloud cannot tell you its material name. The two signals overlap without either containing the other, so the aligned vector can only carry the part they share. Alignment matches the parts that co-occur in the training pairs and ignores the rest, which is exactly why the shared space fails on the compositionality, counting and typographic cases catalogued in the previous part: those failures are the price of summarising each signal by the features it happens to share with its partner.
Aligned does not mean generative. A contrastive space gives you a similarity and a nearest neighbour; it does not produce a sound or a mesh, a mesh being a 3-D surface model built from vertices and faces. Turning one modality into another is a separate model — a decoder, a diffusion model or a language model — and the alignment is only a useful conditioning signal for it. The shared space tells you which things belong together; it cannot draw or speak, so the generation has to happen somewhere else. And aligned does not mean transitive in practice: the hub inherits every pairwise weakness, so audio and 3-D compare well only to the extent that each compares well to text. If the audio tower is sloppy about pitch, no amount of good 3-D training can repair the audio-to-3-D comparison.
Mean cosine within each modality, for matched cross-modal pairs, and for mismatched ones, at the current noise level. Matched cross-modal sits below within-modality: the modality gap.
Where this shows up
One trunk, many front ends
Early fusion instead of contrastive alignment
The frontier answer to the zoo is not a hub at all but a single trunk — one shared backbone network — that consumes text, image, audio and video tokens in one early-fused stream. Early fusion means the modalities are mixed together at the input, before any deep processing, rather than encoded separately and compared afterwards. Alignment by shared space becomes alignment by shared architecture, and the contrastive objective is replaced by next-token prediction over interleaved modalities, where interleaved means the sequence holds text tokens, then image tokens, then audio tokens in the order they occurred, much as a conversation alternates speakers. You can picture the difference as one shared language learned from birth, rather than two speakers raised apart who must be compared after the fact.
Search and grounding across signals
The practical wins are text-to-audio search, video moment retrieval — finding the few seconds of a clip that match a phrase — shape search by description and sound-event tagging. Each is a nearest-neighbour query in a space whose diagonal was made bright by the loss in the contrastive part, on top of the trunk from the first part. Concretely, you embed the thing you are looking for and return the stored vectors closest to it, which is why the same few lines of ranking code serve such different-sounding tasks.
This closes the volume's account of how a model is made to see. The rest of the series takes the aligned encoder and bolts it onto a language model, then pushes resolution, training and serving until one trunk reads every modality at once. The pattern worth carrying forward is that the architecture barely changes from one modality to the next: what changes is how the signal becomes tokens and which pairs the objective is trained on.
Further reading
These papers extend the two-tower recipe past images, then bind the towers to a single hub. Read together they make the point that alignment is one loss applied to many front ends, the same contrastive objective pointed in turn at audio, video, 3-D and finally at a whole set of modalities at once. Taken in that order, they also show the two ways the field has responded to the zoo: train everything against a shared anchor, or replace the anchor with one model that reads every signal directly.
- Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail and Huaming Wang, "CLAP: Learning Audio Concepts from Natural Language Supervision", 2022 — the audio-text analogue of CLIP.
- Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu and coauthors, "ImageBind: One Embedding Space To Bind Them All", 2023 — six modalities bound through the image as a hub.
- Le Xue, Mingfei Gao, Chen Xing and coauthors, "ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding", 2022 — 3-D shapes aligned to text through rendered images.
- Hu Xu, Gargi Ghosh, Po-Yao Huang and coauthors, "VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding", 2021 — the video-text version of the same objective.
Cheat sheet
| Term | Meaning here |
|---|---|
| Audio tower | Mel-spectrogram frames as tokens; CLAP aligns them to captions |
| Video tower | Patchified frames plus a temporal axis; contrasted with clip descriptions |
| 3-D tower | Points or voxels through a point transformer; ULIP and OpenShape align it to text |
| Hub modality | The anchor every tower is trained against, usually text; cross-modal skill is inherited by transitivity |
| Modality gap | Matched cross-modal similarity still sits below within-modality similarity |
| Aligned ≠ interchangeable | Vectors are comparable, not substitutable across encoders |
| Aligned ≠ generative | Similarity is not synthesis; producing one modality from another is a separate model |
| Early fusion | The omni-model alternative: one trunk over interleaved modality tokens |