Hybrid search, fusion and reranking
No single retriever is right. Dense vectors miss the rare identifier, the error code and the exact phrase; BM25 misses the paraphrase that shares no words at all. So production systems run both and combine them — which raises the real question: how do you combine two rankings that do not share a scale? This part answers that with reciprocal rank fusion, then follows the pipeline to its canonical shape: retrieve wide, rerank narrow, and put only the survivors in the context.
Two retrievers, two failure modes
Dense is fuzzy, lexical is brittle, and the failures do not overlap
Lexical retrieval scores documents by the terms they share with the query. It is unbeatable on the things that must match exactly — a part number, a function name, an error string, a surname, a rare acronym — because those tokens have high inverse document frequency and a sparse retriever finds them immediately. Its failure is vocabulary mismatch: ask for "how to make training faster" and it will not match a passage that only says "reducing step time".
Dense retrieval embeds query and passage and compares vectors, so it handles paraphrase, synonymy and abstraction. Its failure is the mirror image. A rare identifier that appears once in the corpus has a weak, uninformative embedding; a query that turns on one exact token ("error E1042", "section 4.2(a)") is exactly the case where similarity search tends to blur the distinction away. The two systems are wrong about different things, which is the whole argument for hybrid retrieval.
Learned sparse models occupy the middle. SPLADE learns a sparse vector over the vocabulary, with weights that can expand a query to related terms and contract away ones that do not help. It keeps the inverted index and the exact-match strength of BM25 while buying the synonymy of a dense model — a middle path, not a replacement for either side.
The pipeline, stage by stage
Lexical, vector, fused, reranked
A hybrid query produces two independent rankings, fuses them into one candidate list, and then hands that list to a reranker. AppSim.hybridRetrieve returns every stage, so the columns on the right show the same query moving through the pipeline: the BM25 order, the vector order, the reciprocal-rank-fusion of the two, and finally the cross-encoder's order over the wide set. The 2-D space on the left is the vector view, with the query's position and the top vector hit connected.
Left: the corpus in 2-D with the query point. Right: four rankings of the same query — lexical, vector, RRF-fused and reranked — with rank movement drawn from fused to reranked.
Fusion: normalise the scores, or ignore them
Reciprocal rank fusion, k = 60
There are two families. Score normalisation tries to make the two score distributions comparable — min-max to [0, 1], or a z-score — and then sums. It can work, but it inherits the brittleness of the scores: min-max is set entirely by the largest value, so a single outlier flattens everything else toward zero, and the result depends on which retriever happened to produce a runaway score.
Reciprocal rank fusion throws the scores away and uses only the ranks. Each document's fused score is the sum, over every list it appears in, of 1 / (k + rank), with k = 60 in the original paper. A document at rank 1 in one list and rank 2 in another scores 1/61 + 1/62 ≈ 0.0325. RRF needs no tuning, cannot be broken by an incompatible scale, and is remarkably hard to beat — which is why it is the default in so many systems.
Below, the same two rankings are fused four ways while a slider magnifies one score into an outlier. Watch the raw sum chase the outlier, min-max flatten under it, and RRF not move at all.
Four fusion methods over the same lexical and vector rankings. The slider exaggerates one vector score; only the rank-based method is invariant.
Retrieve wide, rerank narrow
Cheap recall first, expensive precision second
A bi-encoder embeds the query and every document independently, which is what makes it fast enough to search millions of vectors — the query and document never meet. A cross-encoder feeds the query and a single document through the model together, so it can attend to their interaction, which makes it far more accurate and far too slow to run over the corpus. The pipeline shape follows directly: use the cheap retriever to fetch a wide set of candidates, then use the expensive model to rerank a narrow slice of them.
The two sliders below are the whole design. Fetching more candidates raises the chance the answer is somewhere in the wide set; reranking deeper raises the chance the cross-encoder sees it. Both cost latency, and past a point both stop helping — the heatmap is the tradeoff, seeded rather than measured.
End-to-end accuracy over the wide-set size (columns) and the rerank depth (rows). The ring is your current setting; the readout prices it in latency.
Bi-encoders, cross-encoders and late interaction
Accuracy against latency against memory
Three architectures sit on a spectrum. The bi-encoder produces one vector per text and compares them with a dot product: fastest, cheapest in memory, and unable to see the query and document together. The cross-encoder reads the pair jointly: the most accurate and the most expensive, with no index at all — it must score each candidate at query time. ColBERT's late interaction sits between them: it keeps one vector per token and scores a pair by a cheap max-similarity operation, which is much more expressive than a single vector and much cheaper than a full cross-encoder.
ColBERT's cost is storage, because a document is many vectors rather than one. MUVERA attacks exactly that: it turns multi-vector similarity into a single-vector maximum-inner-product search, recovering most of the quality with the index footprint of an ordinary dense system.
The bars below compare the three at illustrative operating points. They are a shape, not a benchmark: exact numbers depend on the model, the corpus and the hardware.
Relative accuracy, latency and memory for each architecture. Select one to read its role and its cost.
Cheat sheet
| Question | The answer |
|---|---|
| Why hybrid? | Dense and lexical fail on different inputs; the union of their candidates is better than either alone. |
| When does BM25 win? | Exact identifiers, rare tokens, error codes, names — high-IDF terms a dense vector blurs. |
| When does dense win? | Paraphrase and synonymy, where the query shares no vocabulary with the answer. |
| What is SPLADE? | Learned sparse retrieval: an inverted index with learned term weights and query expansion — the middle path. |
| How do you fuse scores? | Carefully. Add unlike scales and you get a silent bug; normalise, or avoid scores entirely. |
| What is RRF? | Σ 1/(k + rank) with k = 60 by default. Rank-only, robust, no tuning, hard to beat. |
| Bi- vs cross-encoder? | Bi-encoders embed independently and index; cross-encoders read the pair jointly and cannot. |
| What is ColBERT? | Late interaction: per-token vectors scored by cheap max-similarity — between bi- and cross-encoders, at a storage cost. |
| What is MUVERA? | Compresses multi-vector late interaction into single-vector MIPS, recovering quality at dense-index cost. |
| What is the pipeline shape? | Retrieve ~50–100 wide, rerank ~5–20 narrow, put the final few in the context. |
Further reading
- Cormack, Clarke & Buettcher, "Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods", SIGIR 2009 — the k = 60 fusion rule.
- Formal et al., "SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking", SIGIR 2021 — learned sparse retrieval with query expansion.
- Khattab & Zaharia, "ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT", SIGIR 2020 — late interaction.
- Santhanam et al., "ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction", 2021 — the storage-savings follow-up.
- Jégou et al. / Cormack et al. as cited; and "MUVERA: Multi-Vector Retrieval via Fixed Dimensional Encodings", 2024 — compressing multi-vector search into single-vector MIPS.
- Nogueira & Cho, "Passage Re-ranking with BERT", 2019 — the cross-encoder reranker that fixed the retrieve-wide-rerank-narrow shape.