Embeddings and vector similarity
BM25 matches the words you typed. An embedding is supposed to match the meaning: two texts with no word in common land near each other in a high-dimensional space. That promise is real but narrower than the marketing. Cosine similarity measures direction, not relevance; queries and documents are embedded differently by design; and the leaderboard that tells you which model to use is a shortlist, not an answer. This part takes the same 60-document corpus and looks at it as a space.
Meaning as a position
A point, a direction, and what similarity can and cannot mean
A text embedding is a fixed-length vector such that texts with similar meaning map to nearby points. "Near" is usually measured with cosine similarity — the cosine of the angle between the vectors, which ignores their lengths. Two documents about the same topic point in the same direction even if one is long and one is short. That is why cosine is the default, and why the magnitude of an embedding is usually normalised away.
The whole series has been using one embedding under the hood: the shared 60-document corpus has hand-fixed 2-D coordinates, coloured by guide, and those coordinates are the running example. A point's position is its meaning; closeness is the model's claim of similarity. Everything below happens in that same two-dimensional space, so the geometry stays visible.
The semantic space
Nearby points are topically similar
Plot all sixty documents by their coordinates and the topical structure is immediate: five clusters — LLM training, serving, applications, math, and vision & robotics. Type a query and the same 2-D space gets a point of its own, placed at the BM25-weighted centroid of the documents it lexically matches. Because the query lands in the space the documents live in, the nearest points by cosine are the model's nearest neighbours.
All 60 documents, coloured by guide. The star is the query's position; rings and lines mark the nearest neighbours by cosine.
Cosine is direction, not distance
Two vectors that rank in opposite orders
Cosine similarity is a·b / (|a| |b|) — the contribution of length cancels. Euclidean distance is |a − b|, which does not. So a vector that points almost exactly at the query can lose to a shorter vector that points slightly off, or win, depending on which measure you use. The two retrievers rank the same candidates differently, and neither is wrong; they are answering different questions.
Vector q points right. Drag A's length and B's length and angle. The readout re-ranks A and B under cosine and under euclidean distance — watch the two orders disagree.
Arrows are vectors from the origin. The dashed circle is the unit length: cosine compares angles, euclidean compares arrow tips.
Queries and documents are not the same text
Why asymmetric models exist
A query is a question; a document is an answer. They are different distributions of text, and a single encoder trained on both will place them in slightly different regions. Asymmetric embedding models handle this explicitly: E5 expects query: and passage: prefixes, and BGE takes an instruction on the query side only. Get the convention wrong — embed queries as passages, or forget the prefix — and the neighbour ordering shifts. The magnitude is small; the effect on the top few results is not.
Below, the same query is embedded two ways. On the left, the query is embedded as if it were a document (no prefix). On the right, it carries a query-side instruction, modelled as a fixed, seeded shift in the space. Lines follow each document from its rank on the left to its rank on the right.
Two rankings of the same five documents by the same model. Left: document-style query embedding. Right: instruction-prefixed query embedding.
Matryoshka truncation, and the MTEB trap
Smaller vectors, and a leaderboard that is a shortlist
Storage and search cost scale with dimension, so the interesting question is how much accuracy you give up by cutting the vector short. Matryoshka representation learning trains the embedding so that its most important information sits in the leading dimensions: a 1,536-dimension vector can be truncated to 768, 384 or 192 and still work, because each prefix is itself a usable embedding. The recall loss is small and, crucially, measurable — you plot it on your corpus rather than trusting a blog post.
The plot below is a seeded simulation of that tradeoff: exact nearest-neighbour search over the corpus scaled up to a few hundred points, compared against the same search after truncation. Drag the point count to change the density of the space.
Recall@1 against exact search as the dimension is truncated. The pale line is storage relative to the full 1,536 dimensions.
The other half of the trap is model selection. MTEB ranks hundreds of embedding models across dozens of tasks, and it is useful for exactly one thing: narrowing the field to a shortlist worth testing. The scores are dominated by a handful of benchmarks, the leaderboard is public enough to be overfitted, and a model that wins on a general web-retrieval task can lose on your corpus of support tickets or camera drivers. Choose from the shortlist, then measure on your own labelled set — the same rule the eval harness in Part 5 imposes on everything else.
Cheat sheet
| Question | The answer |
|---|---|
| What is an embedding? | A fixed-length vector where meaning is position; similar texts land nearby. |
| Why cosine and not L2? | Cosine measures direction and ignores magnitude, which is what "same topic, different length" needs. |
| Does higher similarity mean more relevant? | No. Similarity ranks; relevance needs the task and usually a reranker over a wide candidate set. |
| Why prefix the query? | Asymmetric models (E5, BGE) were trained with query and passage conventions; using them wrong shifts the ordering. |
| What is Matryoshka? | An embedding trained so prefixes are usable — truncate 1,536 → 768 → 384 for less storage at a small, plottable recall cost. |
| How do I pick a model? | Shortlist from MTEB, then measure recall on your own labelled queries. The leaderboard is not the answer. |
| What comes next? | Chunking: an embedding is per passage, so how you cut documents decides what can be retrieved at all. |
Further reading
- Reimers & Gurevych, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", EMNLP 2019 — the architecture that made sentence embeddings practical.
- Wang et al., "Text Embeddings by Weakly-Supervised Contrastive Pre-training" (E5), 2022 — the
query:/passage:prefixes behind the asymmetry in step 4. - Xiao et al., "C-Pack: Packed Resources for General Chinese Embeddings" (BGE), 2023 — the instruction-prefixed query convention and a widely used open model family.
- Kusupati et al., "Matryoshka Representation Learning", NeurIPS 2022 — nested, truncatable embeddings and the recall-versus-dimension tradeoff plotted in step 5.
- Muennighoff et al., "MTEB: Massive Text Embedding Benchmark", 2022 — the leaderboard, its task coverage, and why it shortlists rather than decides.
- Jégou, Douze & Schmid, "Product Quantization for Nearest Neighbor Search", IEEE TPAMI 2011 — the compression half of the storage question this part's truncation plot starts.