Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Meaning as a position

A point, a direction, and what similarity can and cannot mean

A text embedding is a fixed-length vector such that texts with similar meaning map to nearby points. "Near" is usually measured with cosine similarity — the cosine of the angle between the vectors, which ignores their lengths. Two documents about the same topic point in the same direction even if one is long and one is short. That is why cosine is the default, and why the magnitude of an embedding is usually normalised away.

The whole series has been using one embedding under the hood: the shared 60-document corpus has hand-fixed 2-D coordinates, coloured by guide, and those coordinates are the running example. A point's position is its meaning; closeness is the model's claim of similarity. Everything below happens in that same two-dimensional space, so the geometry stays visible.

💡 Carry two sentences forward. First: cosine measures direction, not distance — a short vector can outrank a long one on angle while losing on euclidean distance. Second: most similar is not most relevant — similarity ranks candidates, and relevance is decided by the task and the reranker.

⚠️ The trap: treating a similarity score as a relevance score. In a well-separated space the top similarity is usually on-topic, but among near-ties the ordering is noise. That is why production systems retrieve wide with a cheap similarity and then rerank narrow with a model that reads the query and the document together.
2

The semantic space

Nearby points are topically similar

Plot all sixty documents by their coordinates and the topical structure is immediate: five clusters — LLM training, serving, applications, math, and vision & robotics. Type a query and the same 2-D space gets a point of its own, placed at the BM25-weighted centroid of the documents it lexically matches. Because the query lands in the space the documents live in, the nearest points by cosine are the model's nearest neighbours.

All 60 documents, coloured by guide. The star is the query's position; rings and lines mark the nearest neighbours by cosine.

💡 Read the geometry, not just the list: the query point lands inside a cluster, and the neighbours it returns are mostly from that cluster. When a query has no lexical overlap at all it lands at the origin — this 2-D embedding is defined by lexical matching, which is the honest limitation of the running example.
3

Cosine is direction, not distance

Two vectors that rank in opposite orders

Cosine similarity is a·b / (|a| |b|) — the contribution of length cancels. Euclidean distance is |a − b|, which does not. So a vector that points almost exactly at the query can lose to a shorter vector that points slightly off, or win, depending on which measure you use. The two retrievers rank the same candidates differently, and neither is wrong; they are answering different questions.

Vector q points right. Drag A's length and B's length and angle. The readout re-ranks A and B under cosine and under euclidean distance — watch the two orders disagree.

Arrows are vectors from the origin. The dashed circle is the unit length: cosine compares angles, euclidean compares arrow tips.

⚠️ Practical consequence: if you switch a store from L2 to cosine, or normalise vectors after indexing but not before, you change the ranking. Pick one metric, record it in the index configuration, and never mix a normalised query with un-normalised documents.
4

Queries and documents are not the same text

Why asymmetric models exist

A query is a question; a document is an answer. They are different distributions of text, and a single encoder trained on both will place them in slightly different regions. Asymmetric embedding models handle this explicitly: E5 expects query: and passage: prefixes, and BGE takes an instruction on the query side only. Get the convention wrong — embed queries as passages, or forget the prefix — and the neighbour ordering shifts. The magnitude is small; the effect on the top few results is not.

Below, the same query is embedded two ways. On the left, the query is embedded as if it were a document (no prefix). On the right, it carries a query-side instruction, modelled as a fixed, seeded shift in the space. Lines follow each document from its rank on the left to its rank on the right.

Two rankings of the same five documents by the same model. Left: document-style query embedding. Right: instruction-prefixed query embedding.

5

Matryoshka truncation, and the MTEB trap

Smaller vectors, and a leaderboard that is a shortlist

Storage and search cost scale with dimension, so the interesting question is how much accuracy you give up by cutting the vector short. Matryoshka representation learning trains the embedding so that its most important information sits in the leading dimensions: a 1,536-dimension vector can be truncated to 768, 384 or 192 and still work, because each prefix is itself a usable embedding. The recall loss is small and, crucially, measurable — you plot it on your corpus rather than trusting a blog post.

The plot below is a seeded simulation of that tradeoff: exact nearest-neighbour search over the corpus scaled up to a few hundred points, compared against the same search after truncation. Drag the point count to change the density of the space.

Recall@1 against exact search as the dimension is truncated. The pale line is storage relative to the full 1,536 dimensions.

The other half of the trap is model selection. MTEB ranks hundreds of embedding models across dozens of tasks, and it is useful for exactly one thing: narrowing the field to a shortlist worth testing. The scores are dominated by a handful of benchmarks, the leaderboard is public enough to be overfitted, and a model that wins on a general web-retrieval task can lose on your corpus of support tickets or camera drivers. Choose from the shortlist, then measure on your own labelled set — the same rule the eval harness in Part 5 imposes on everything else.

💡 Two decisions, two measurements. Truncate dimensions and MTEB-rank models only on paper; then plot recall against your own queries and pick the smallest vector and the cheapest model that clears the bar. The leaderboard shortlists; your corpus decides.

Cheat sheet

QuestionThe answer
What is an embedding?A fixed-length vector where meaning is position; similar texts land nearby.
Why cosine and not L2?Cosine measures direction and ignores magnitude, which is what "same topic, different length" needs.
Does higher similarity mean more relevant?No. Similarity ranks; relevance needs the task and usually a reranker over a wide candidate set.
Why prefix the query?Asymmetric models (E5, BGE) were trained with query and passage conventions; using them wrong shifts the ordering.
What is Matryoshka?An embedding trained so prefixes are usable — truncate 1,536 → 768 → 384 for less storage at a small, plottable recall cost.
How do I pick a model?Shortlist from MTEB, then measure recall on your own labelled queries. The leaderboard is not the answer.
What comes next?Chunking: an embedding is per passage, so how you cut documents decides what can be retrieved at all.

Further reading

6

Check your understanding

0/5 answered