Why retrieve, and lexical search first
The reflex answer is that the model has not seen your data. That is the weakest of the reasons. The durable ones are systems properties: the data has changed since training, the answer has to cite a source, different users may be allowed to see different documents, and sending everything "just in case" costs more than retrieving the two thousand tokens that matter. This part makes that case, then starts where production systems actually start — BM25, the lexical baseline that is still hard to beat.
The real reasons are systems reasons
Not "the model doesn't know things"
A model's knowledge is parametric: it is baked into weights at training time and then frozen. That makes it broad but stale, unauditable and ungated. Retrieval is not a way to top up intelligence; it is a way to give the system five properties it cannot otherwise have.
- Freshness. The document you need may have been written after the model's training cutoff — or five minutes ago. No amount of model quality fixes that.
- Provenance. An answer you can trace to a document is checkable; an answer from weights is a claim. If the output has to be cited, you need the source in the context and in the logs.
- Per-user access control. The model has no notion of who is asking. Retrieval is where the permission check belongs, because documents carry ACLs and the model does not.
- Cost. A 200k-token prompt full of mostly-irrelevant text is expensive, slow, and worse than a focused 2k-token context. Retrieving the right tokens is a cost and accuracy win at once.
- Auditability. "Why did the system say that?" has an answer when you can replay the query, the retriever, and the passages that were in the context.
Anthropic's contextual-retrieval numbers make the point that retrieval quality is an engineering curve, not a single step: on a large corpus, the top-20 failure rate fell from 5.7% at baseline to 3.7% with contextual embeddings, 2.9% when BM25 was added, and 1.9% with reranking. The widely quoted 35%→49%→67% figures are the relative reductions of that same sequence. Every step is a retrieval-side change, not a model change.
Lexical search first: BM25
The baseline that keeps winning
BM25 scores a document by the query terms it contains, weighted three ways. A term that is rare across the corpus contributes more — that is inverse document frequency, IDF = ln(1 + (N − n + 0.5)/(n + 0.5)). Repeated occurrences of a term help, but with diminishing returns — the term-frequency factor saturates, controlled by k1. And longer documents are not automatically better, so the score is normalised by length, controlled by b. Lucene's defaults are k1 = 1.2, b = 0.75, and they are a good place to start.
It has no parameters to train, no embedding model, no index flush, and no drift. It is interpretable — you can point at the term that caused the hit — and it fails predictably on the one case it cannot represent: vocabulary mismatch, where the query and the document mean the same thing with no shared word.
Type a query and watch the ranking. Each bar is a document's BM25 score against the corpus this whole series runs on. Then try the query with no lexical overlap and watch BM25 return nothing at all — that is the gap Part 13's embeddings are for.
Ranked BM25 scores over the shared 60-document corpus. The term-IDF line at the top explains each contribution.
BM25's two knobs
Term saturation and length normalisation
k1 controls how quickly extra occurrences of a term stop helping. At low k1 the score is nearly binary — the document either contains the term or it does not. At high k1, repetition keeps adding. b controls length normalisation: b = 0 ignores document length entirely, b = 1 fully divides it out. In a corpus where long documents accumulate terms by accident, b is the knob that stops them from winning on sheer size.
Drag both sliders. The query is a single common term, so what you are watching is which documents win once the length penalty changes.
The top six documents for one query, with each document's term count (dl) shown. Change b to see the ordering move.
Retrieve, or just prompt?
A decision matrix, not a reflex
Retrieval adds latency, a second system to operate, a new failure surface, and a security boundary. It earns its place when at least one of the systems reasons is genuinely in play. Turn the reasons on and off and watch the verdict change: the one that never justifies retrieval on its own is "the model might not know it".
Toggle the reasons; the verdict is computed from the toggles, not the model.
Once retrieval is in, the metrics that tell you whether it worked are retrieval-side. RAGAS separates context precision (were the retrieved passages relevant?) from context recall (did they contain what was needed?), and faithfulness (did the answer stay inside them?). A single end-to-end "quality" score cannot tell you which of the three failed, which is exactly the distinction you need to fix one.
Retrieval is a security boundary
Filter before the model, and before the ranking you show
Once retrieval exists, a query from one user can pull a document that another user is not allowed to see. A naive retriever searches the whole index and hands the top documents to the model; the leak is already in the context by the time anyone reads the answer. The fix is to filter candidates by the caller's ACL before they enter the ranking, not after: the result list, the scores, and any snippets shown to the user are all places where the existence of a document leaks.
Two users, one query. The analyst has public access; the engineer also has the internal corpus. Toggle enforcement and watch a restricted document sit at the top of the analyst's list — then disappear.
Ranked results for the query "safety training". Red rows are documents the current user is not permitted to see.
Cheat sheet
| Question | The answer |
|---|---|
| Why retrieve, really? | Freshness, provenance, per-user access, cost, auditability — systems properties, not model ignorance. |
| Why not just a bigger prompt? | Cost, latency and context rot: sending everything is slower, pricier and often less accurate. |
| What is BM25? | A bag-of-words ranker: rare terms count more (IDF), repeats saturate (k1), length is normalised (b). |
| Default BM25 parameters? | k1 = 1.2, b = 0.75 (Lucene). b = 0 disables length normalisation. |
| Why start lexical? | No training, interpretable, cheap, and hard for dense retrieval to beat without tuning. |
| Where does BM25 fail? | Vocabulary mismatch: same meaning, no shared term. That is the embeddings gap. |
| Where does the ACL go? | Before ranking, so restricted documents never enter the candidate set or the UI. |
Further reading
- Robertson & Zaragoza, "The Probabilistic Relevance Framework: BM25 and Beyond", Foundations and Trends in Information Retrieval, 2009 — the definitive account of BM25 and its parameterisation.
- Apache Lucene, BM25Similarity — the k1 = 1.2, b = 0.75 defaults this part's playground starts from.
- Anthropic, "Introducing Contextual Retrieval", 19 September 2024 — the failure-rate sequence that shows BM25 still contributing on top of contextual embeddings.
- Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", NeurIPS 2020 — the paper that named the pattern and tied retrieval to provenance.
- Greshake et al., "Not what you've signed up for", 2023 — why retrieved content is untrusted input, the security framing this part's ACL section sets up.