Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Two retrievers, two failure modes

Dense is fuzzy, lexical is brittle, and the failures do not overlap

Lexical retrieval scores documents by the terms they share with the query. It is unbeatable on the things that must match exactly — a part number, a function name, an error string, a surname, a rare acronym — because those tokens have high inverse document frequency and a sparse retriever finds them immediately. Its failure is vocabulary mismatch: ask for "how to make training faster" and it will not match a passage that only says "reducing step time".

Dense retrieval embeds query and passage and compares vectors, so it handles paraphrase, synonymy and abstraction. Its failure is the mirror image. A rare identifier that appears once in the corpus has a weak, uninformative embedding; a query that turns on one exact token ("error E1042", "section 4.2(a)") is exactly the case where similarity search tends to blur the distinction away. The two systems are wrong about different things, which is the whole argument for hybrid retrieval.

Learned sparse models occupy the middle. SPLADE learns a sparse vector over the vocabulary, with weights that can expand a query to related terms and contract away ones that do not help. It keeps the inverted index and the exact-match strength of BM25 while buying the synonymy of a dense model — a middle path, not a replacement for either side.

💡 The durable idea: hybrid retrieval is not about two mediocre retrievers averaging out. It is about covering two disjoint recall surfaces, then letting fusion decide which candidate survived both.
⚠️ The trap: combining scores as if they were comparable. A BM25 score is unbounded and corpus-dependent; a cosine similarity is bounded and means something else entirely. Adding them is a silent bug that looks like a tuning problem.
2

The pipeline, stage by stage

Lexical, vector, fused, reranked

A hybrid query produces two independent rankings, fuses them into one candidate list, and then hands that list to a reranker. AppSim.hybridRetrieve returns every stage, so the columns on the right show the same query moving through the pipeline: the BM25 order, the vector order, the reciprocal-rank-fusion of the two, and finally the cross-encoder's order over the wide set. The 2-D space on the left is the vector view, with the query's position and the top vector hit connected.

Left: the corpus in 2-D with the query point. Right: four rankings of the same query — lexical, vector, RRF-fused and reranked — with rank movement drawn from fused to reranked.

💡 Try “fragmentation” or “paged memory”. BM25 locks onto the document that literally contains the rare term; the vector centroid drifts toward a neighbouring cluster. Fusion keeps both candidates, and the cross-encoder then picks the one that actually answers the query — cover two recall surfaces, then let ranking decide.
3

Fusion: normalise the scores, or ignore them

Reciprocal rank fusion, k = 60

There are two families. Score normalisation tries to make the two score distributions comparable — min-max to [0, 1], or a z-score — and then sums. It can work, but it inherits the brittleness of the scores: min-max is set entirely by the largest value, so a single outlier flattens everything else toward zero, and the result depends on which retriever happened to produce a runaway score.

Reciprocal rank fusion throws the scores away and uses only the ranks. Each document's fused score is the sum, over every list it appears in, of 1 / (k + rank), with k = 60 in the original paper. A document at rank 1 in one list and rank 2 in another scores 1/61 + 1/62 ≈ 0.0325. RRF needs no tuning, cannot be broken by an incompatible scale, and is remarkably hard to beat — which is why it is the default in so many systems.

Below, the same two rankings are fused four ways while a slider magnifies one score into an outlier. Watch the raw sum chase the outlier, min-max flatten under it, and RRF not move at all.

Four fusion methods over the same lexical and vector rankings. The slider exaggerates one vector score; only the rank-based method is invariant.

⚠️ When normalisation is still worth it: when the two retrievers are the same kind of model and their scores are already calibrated to one another — two dense models, or a dense model and a learned-sparse one trained together. Otherwise, prefer rank fusion and stop tuning constants.
4

Retrieve wide, rerank narrow

Cheap recall first, expensive precision second

A bi-encoder embeds the query and every document independently, which is what makes it fast enough to search millions of vectors — the query and document never meet. A cross-encoder feeds the query and a single document through the model together, so it can attend to their interaction, which makes it far more accurate and far too slow to run over the corpus. The pipeline shape follows directly: use the cheap retriever to fetch a wide set of candidates, then use the expensive model to rerank a narrow slice of them.

The two sliders below are the whole design. Fetching more candidates raises the chance the answer is somewhere in the wide set; reranking deeper raises the chance the cross-encoder sees it. Both cost latency, and past a point both stop helping — the heatmap is the tradeoff, seeded rather than measured.

End-to-end accuracy over the wide-set size (columns) and the rerank depth (rows). The ring is your current setting; the readout prices it in latency.

💡 The canonical numbers: retrieve roughly 50–100 candidates with the cheap retriever, rerank the top 5–20 with the cross-encoder, and put only the final few in the context. Wider than that buys recall you cannot afford to re-score; deeper than that buys precision the model has already given you.
5

Bi-encoders, cross-encoders and late interaction

Accuracy against latency against memory

Three architectures sit on a spectrum. The bi-encoder produces one vector per text and compares them with a dot product: fastest, cheapest in memory, and unable to see the query and document together. The cross-encoder reads the pair jointly: the most accurate and the most expensive, with no index at all — it must score each candidate at query time. ColBERT's late interaction sits between them: it keeps one vector per token and scores a pair by a cheap max-similarity operation, which is much more expressive than a single vector and much cheaper than a full cross-encoder.

ColBERT's cost is storage, because a document is many vectors rather than one. MUVERA attacks exactly that: it turns multi-vector similarity into a single-vector maximum-inner-product search, recovering most of the quality with the index footprint of an ordinary dense system.

The bars below compare the three at illustrative operating points. They are a shape, not a benchmark: exact numbers depend on the model, the corpus and the hardware.

Relative accuracy, latency and memory for each architecture. Select one to read its role and its cost.

💡 They are not alternatives. A production stack uses a bi-encoder or a hybrid lexical+dense retriever to fetch wide, an optional late-interaction stage to re-rank cheaply, and a cross-encoder on the narrow set. Each stage's cost buys the next stage a better candidate list.

Cheat sheet

QuestionThe answer
Why hybrid?Dense and lexical fail on different inputs; the union of their candidates is better than either alone.
When does BM25 win?Exact identifiers, rare tokens, error codes, names — high-IDF terms a dense vector blurs.
When does dense win?Paraphrase and synonymy, where the query shares no vocabulary with the answer.
What is SPLADE?Learned sparse retrieval: an inverted index with learned term weights and query expansion — the middle path.
How do you fuse scores?Carefully. Add unlike scales and you get a silent bug; normalise, or avoid scores entirely.
What is RRF?Σ 1/(k + rank) with k = 60 by default. Rank-only, robust, no tuning, hard to beat.
Bi- vs cross-encoder?Bi-encoders embed independently and index; cross-encoders read the pair jointly and cannot.
What is ColBERT?Late interaction: per-token vectors scored by cheap max-similarity — between bi- and cross-encoders, at a storage cost.
What is MUVERA?Compresses multi-vector late interaction into single-vector MIPS, recovering quality at dense-index cost.
What is the pipeline shape?Retrieve ~50–100 wide, rerank ~5–20 narrow, put the final few in the context.

Further reading

6

Check your understanding

0/4 answered