Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

More than pictures

Every signal becomes tokens, then a vector

The recipe is uniform even though the signals are not. Every modality — a modality being one kind of signal, such as an image, a sound, a video or a 3-D shape — is first turned into a sequence of units. For an image those units are patches, small square tiles cut from the grid. For audio they are mel-spectrogram frames: a mel spectrogram is a picture of how much energy each frequency band carries over time, and each frame is one thin vertical slice of it. For video they are sampled frames, and for a 3-D shape they are points or voxels, where a voxel is the 3-D equivalent of a pixel, a small cube of space. You then run a transformer over that sequence, and pool the output into one vector, pooling meaning combining the many per-unit vectors into a single summary vector, usually by averaging them. A text tower does exactly the same thing with subwords, the word pieces a tokeniser splits text into. Once every tower produces a vector of the same width — the same fixed number of numbers per vector — the contrastive loss from two parts back applies without modification: matched pairs together, everything else apart. That clause is the whole objective; everything else in this part is a change of front end, meaning the tokeniser that turns a raw signal into a sequence.

What changes from modality to modality is the tokeniser, meaning the rule that decides what counts as one unit, and the cost of running the transformer over the resulting sequence. An image is one grid, so its token count is fixed by the patch size. A video is a grid per frame, and it adds a second, harsher scaling problem in time: more frames means more tokens, and a long clip can easily produce a sequence too long to attend over cheaply. Audio is a time-frequency image whose frame rate sets the token count, so a higher sampling rate buys finer detail at a higher price. 3-D data arrives unordered and sparse: a point cloud has no natural reading order the way pixels do, and most of the space it describes is empty, so the tokeniser has to group nearby points into neighbourhoods before the transformer can use them. The cards below are the zoo as the site's data lists it: five encoders — one network per modality, each turning its signal into vectors — five contrastive pairings, and five different notions of what one token is.

These rows come from Multimodal.modalities. Each pair is a genuine contrastive objective in the literature; the point is that the architecture is the same and only the front end differs.

💡 By the end of this part you'll be able to name the encoder and pairing for audio, video and 3-D, explain what a hub modality is and why text usually plays that role, demonstrate cross-modal alignment on a similarity matrix, and state precisely what "aligned" does and does not imply. Those four abilities are exactly the vocabulary you need to read a paper in this area without being misled by the word.
2

Audio and video towers

The same matrix, a different front end

CLAP is the audio analogue of CLIP — the model from the earlier parts that matched images to captions. CLAP swaps the image tower for an audio tower, so the pairing becomes sound against text. The audio tower consumes a mel spectrogram through a transformer, the text tower consumes a caption, and an InfoNCE loss — the specific contrastive objective that treats every other item in a batch as a negative example for a given pair — pulls matched audio-caption pairs together. The tokens are time-frequency cells, meaning each cell records the energy in one frequency band during one short slice of time. The caption might be "a dog barking" or the audio's tags, where tags are a few keywords describing the clip. Trained this way, the result classifies sound events, meaning it labels a clip as a bark, a siren or a doorbell, and retrieves audio with text, meaning it takes a typed phrase and returns the matching recordings. Video follows the same shape with a spatiotemporal torso, "spatiotemporal" meaning the model covers both space within a frame and time across frames. Concretely, you patchify each sampled frame, cutting it into patches the way the vision transformer does, add the temporal axis so the model can tell which frame a patch came from, and contrast clips against captions with VideoCLIP or InternVideo.

The demo builds one base vector per concept, then gives each modality its own noisy copy of that vector, and draws the cross-modal similarity matrix: a grid whose cell in row i and column j is the cosine between concept i in one modality and concept j in the other. The cosine is the standard similarity score for these spaces, measuring how closely two vectors point in the same direction rather than how far apart they are. The diagonal holds the matching pairs, the name-tags-at-a-party idea from the previous part, where each item should be closest to its own partner. The noise slider controls how well the towers agree. At low noise the diagonal is bright and the alignment is clean, so every matched pair is each other's nearest neighbour. Turn it up and the signal dissolves into a near-uniform matrix, which is what a badly trained or badly paired objective looks like: no better than guessing.

Cross-modal cosine similarity from GenMedia.attn.simMatrix for the selected pair. Diagonal cells are outlined; the readout reports top-1 alignment.

⚠ The front end is where the difficulty hides. A mel-spectrogram at 50 frames per second is a long, narrow image; how you patch it, window it and merge tokens decides both the cost and what the encoder can hear. A window here is a short span of frames treated as one unit before neighbouring spans are merged together, and the choices interact: a coarser merge is cheaper but blurs exactly the fast transients that make a bark sound like a bark. "Same loss, different tokeniser" is true and also where all the engineering lives.
3

3-D and the hub problem

One anchor, many spokes

Aligning two modalities needs data in which the two co-occur, meaning examples where the same content appears in both at once, such as a clip and its transcript. Aligning five does not scale that way, because not every pair has such a dataset: there is no large corpus of paired point clouds and dog barks. Training every pair directly would need a dataset for each of the ten pairs, and most of those datasets do not exist. The workaround is a hub modality, a single anchor modality that every other tower is trained against instead of being trained against one another. Text is the usual choice, and the reason is largely practical: people have written captions, transcripts and descriptions for almost everything, so text can be paired with almost any other signal. Each tower is trained only against the hub, and transitivity does the rest: if audio lands near text and text lands near 3-D, then audio and 3-D are indirectly comparable, because both sit in the neighbourhood of the same text. The cost of the shortcut is that audio and 3-D never actually meet in training, so any error either one makes against text is inherited by that pair. ImageBind takes the extreme version, using images as the hub and binding six modalities through vision alone.

3-D encoders such as ULIP and OpenShape are the clearest case. A point cloud or voxel grid is encoded by a point transformer — a transformer whose tokens are groups of nearby 3-D points rather than image patches — then contrasted against shape captions or against rendered images of the same object. Through that objective the shape inherits a position in the same space as a sentence, so "a wooden chair" and the coordinates of a chair end up near each other. Notice that the comparison is not always point cloud against text directly: routing it through a rendered view of the shape is another instance of the hub idea, with images playing the anchor role. The diagram shows the arrangement: a central tower that every other one is pulled toward, and spokes that are trained one at a time.

A hub-modality diagram: each tower is trained to land near the hub, and no tower is trained directly against another.

A hub is a modelling choice, not a law. Text is convenient because captions exist for everything; images are convenient because vision models are strong. Either way, cross-modal skill that was never trained is inherited by transitivity, with all the error that implies.

4

What aligned does not mean

Four readings of one word

Aligned does not mean interchangeable. Two towers landing in one space does not let you feed a spectrogram to a text encoder or read an image vector as a sentence. The space is comparable, not shared in the sense of substitutable: both towers agree on how to measure similarity, but they do not agree on how to represent a signal, so a vector from one is meaningless to the other's decoder. There is also usually a measurable modality gap: embeddings of one modality occupy a slightly different region even for matched pairs, so a within-modality similarity — two audio clips compared with each other — is typically higher than a matched cross-modal one, such as an audio clip compared with its own caption. That gap is not a bug so much as the fingerprint of two encoders trained separately, which never had a reason to occupy identical coordinates. The bars below show that gap.

Aligned does not mean equivalent information. A point cloud and a caption both describe a chair, but the caption cannot tell you the chair's dimensions and the point cloud cannot tell you its material name. The two signals overlap without either containing the other, so the aligned vector can only carry the part they share. Alignment matches the parts that co-occur in the training pairs and ignores the rest, which is exactly why the shared space fails on the compositionality, counting and typographic cases catalogued in the previous part: those failures are the price of summarising each signal by the features it happens to share with its partner.

Aligned does not mean generative. A contrastive space gives you a similarity and a nearest neighbour; it does not produce a sound or a mesh, a mesh being a 3-D surface model built from vertices and faces. Turning one modality into another is a separate model — a decoder, a diffusion model or a language model — and the alignment is only a useful conditioning signal for it. The shared space tells you which things belong together; it cannot draw or speak, so the generation has to happen somewhere else. And aligned does not mean transitive in practice: the hub inherits every pairwise weakness, so audio and 3-D compare well only to the extent that each compares well to text. If the audio tower is sloppy about pitch, no amount of good 3-D training can repair the audio-to-3-D comparison.

Mean cosine within each modality, for matched cross-modal pairs, and for mismatched ones, at the current noise level. Matched cross-modal sits below within-modality: the modality gap.

💡 When someone says two modalities are "aligned", ask three questions: aligned by what objective, through which hub, and on what pairs of data? The answers are the whole story, and they differ enormously between systems that use the same word. A tower aligned on ten thousand hand-labelled clips and one aligned on a million weakly captioned videos can both be called "aligned", and the word will not tell you which one you are looking at.
5

Where this shows up

One trunk, many front ends

Omni models

Early fusion instead of contrastive alignment

The frontier answer to the zoo is not a hub at all but a single trunk — one shared backbone network — that consumes text, image, audio and video tokens in one early-fused stream. Early fusion means the modalities are mixed together at the input, before any deep processing, rather than encoded separately and compared afterwards. Alignment by shared space becomes alignment by shared architecture, and the contrastive objective is replaced by next-token prediction over interleaved modalities, where interleaved means the sequence holds text tokens, then image tokens, then audio tokens in the order they occurred, much as a conversation alternates speakers. You can picture the difference as one shared language learned from birth, rather than two speakers raised apart who must be compared after the fact.

Applications

Search and grounding across signals

The practical wins are text-to-audio search, video moment retrieval — finding the few seconds of a clip that match a phrase — shape search by description and sound-event tagging. Each is a nearest-neighbour query in a space whose diagonal was made bright by the loss in the contrastive part, on top of the trunk from the first part. Concretely, you embed the thing you are looking for and return the stored vectors closest to it, which is why the same few lines of ranking code serve such different-sounding tasks.

This closes the volume's account of how a model is made to see. The rest of the series takes the aligned encoder and bolts it onto a language model, then pushes resolution, training and serving until one trunk reads every modality at once. The pattern worth carrying forward is that the architecture barely changes from one modality to the next: what changes is how the signal becomes tokens and which pairs the objective is trained on.

Further reading

These papers extend the two-tower recipe past images, then bind the towers to a single hub. Read together they make the point that alignment is one loss applied to many front ends, the same contrastive objective pointed in turn at audio, video, 3-D and finally at a whole set of modalities at once. Taken in that order, they also show the two ways the field has responded to the zoo: train everything against a shared anchor, or replace the anchor with one model that reads every signal directly.

Cheat sheet

TermMeaning here
Audio towerMel-spectrogram frames as tokens; CLAP aligns them to captions
Video towerPatchified frames plus a temporal axis; contrasted with clip descriptions
3-D towerPoints or voxels through a point transformer; ULIP and OpenShape align it to text
Hub modalityThe anchor every tower is trained against, usually text; cross-modal skill is inherited by transitivity
Modality gapMatched cross-modal similarity still sits below within-modality similarity
Aligned ≠ interchangeableVectors are comparable, not substitutable across encoders
Aligned ≠ generativeSimilarity is not synthesis; producing one modality from another is a separate model
Early fusionThe omni-model alternative: one trunk over interleaved modality tokens
7

Check your understanding

0/4 answered