Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The real reasons are systems reasons

Not "the model doesn't know things"

A model's knowledge is parametric: it is baked into weights at training time and then frozen. That makes it broad but stale, unauditable and ungated. Retrieval is not a way to top up intelligence; it is a way to give the system five properties it cannot otherwise have.

💡 Reframe: retrieval is a systems decision about freshness, provenance, access and cost. Model ignorance is a symptom you can fix with a fine-tune or a bigger prompt; the other four cannot be fixed that way.

Anthropic's contextual-retrieval numbers make the point that retrieval quality is an engineering curve, not a single step: on a large corpus, the top-20 failure rate fell from 5.7% at baseline to 3.7% with contextual embeddings, 2.9% when BM25 was added, and 1.9% with reranking. The widely quoted 35%→49%→67% figures are the relative reductions of that same sequence. Every step is a retrieval-side change, not a model change.

⚠️ The trap: shipping retrieval as a capability patch, with no labelled set and no measurement, replaces a known model failure with an unknown retrieval failure. If you cannot state whether the answer needs freshness, provenance or an ACL, you cannot tell whether retrieval helped.
2

Lexical search first: BM25

The baseline that keeps winning

BM25 scores a document by the query terms it contains, weighted three ways. A term that is rare across the corpus contributes more — that is inverse document frequency, IDF = ln(1 + (N − n + 0.5)/(n + 0.5)). Repeated occurrences of a term help, but with diminishing returns — the term-frequency factor saturates, controlled by k1. And longer documents are not automatically better, so the score is normalised by length, controlled by b. Lucene's defaults are k1 = 1.2, b = 0.75, and they are a good place to start.

It has no parameters to train, no embedding model, no index flush, and no drift. It is interpretable — you can point at the term that caused the hit — and it fails predictably on the one case it cannot represent: vocabulary mismatch, where the query and the document mean the same thing with no shared word.

Type a query and watch the ranking. Each bar is a document's BM25 score against the corpus this whole series runs on. Then try the query with no lexical overlap and watch BM25 return nothing at all — that is the gap Part 13's embeddings are for.

Ranked BM25 scores over the shared 60-document corpus. The term-IDF line at the top explains each contribution.

💡 Start lexical, not dense: BM25 is cheap, auditable and strong. Dense retrieval has to be tuned and measured to beat it on a specific corpus; on jargon-heavy, name-heavy or code corpora it often does not.
3

BM25's two knobs

Term saturation and length normalisation

k1 controls how quickly extra occurrences of a term stop helping. At low k1 the score is nearly binary — the document either contains the term or it does not. At high k1, repetition keeps adding. b controls length normalisation: b = 0 ignores document length entirely, b = 1 fully divides it out. In a corpus where long documents accumulate terms by accident, b is the knob that stops them from winning on sheer size.

Drag both sliders. The query is a single common term, so what you are watching is which documents win once the length penalty changes.

The top six documents for one query, with each document's term count (dl) shown. Change b to see the ordering move.

4

Retrieve, or just prompt?

A decision matrix, not a reflex

Retrieval adds latency, a second system to operate, a new failure surface, and a security boundary. It earns its place when at least one of the systems reasons is genuinely in play. Turn the reasons on and off and watch the verdict change: the one that never justifies retrieval on its own is "the model might not know it".

Toggle the reasons; the verdict is computed from the toggles, not the model.

💡 The test: name the property first — freshness, provenance, access, cost — and let it force the architecture. If the only reason is "the model might not know it", try the prompt and measure before standing up a retrieval stack.

Once retrieval is in, the metrics that tell you whether it worked are retrieval-side. RAGAS separates context precision (were the retrieved passages relevant?) from context recall (did they contain what was needed?), and faithfulness (did the answer stay inside them?). A single end-to-end "quality" score cannot tell you which of the three failed, which is exactly the distinction you need to fix one.

5

Retrieval is a security boundary

Filter before the model, and before the ranking you show

Once retrieval exists, a query from one user can pull a document that another user is not allowed to see. A naive retriever searches the whole index and hands the top documents to the model; the leak is already in the context by the time anyone reads the answer. The fix is to filter candidates by the caller's ACL before they enter the ranking, not after: the result list, the scores, and any snippets shown to the user are all places where the existence of a document leaks.

Two users, one query. The analyst has public access; the engineer also has the internal corpus. Toggle enforcement and watch a restricted document sit at the top of the analyst's list — then disappear.

Ranked results for the query "safety training". Red rows are documents the current user is not permitted to see.

⚠️ The leak is not just the answer. A blocked document that appears in a "related results" list, an autocomplete suggestion, or a score-ordered debug panel has already told the user it exists. Apply the ACL at index or query time, and make the filtered set the only set the ranker ever sees.

Cheat sheet

QuestionThe answer
Why retrieve, really?Freshness, provenance, per-user access, cost, auditability — systems properties, not model ignorance.
Why not just a bigger prompt?Cost, latency and context rot: sending everything is slower, pricier and often less accurate.
What is BM25?A bag-of-words ranker: rare terms count more (IDF), repeats saturate (k1), length is normalised (b).
Default BM25 parameters?k1 = 1.2, b = 0.75 (Lucene). b = 0 disables length normalisation.
Why start lexical?No training, interpretable, cheap, and hard for dense retrieval to beat without tuning.
Where does BM25 fail?Vocabulary mismatch: same meaning, no shared term. That is the embeddings gap.
Where does the ACL go?Before ranking, so restricted documents never enter the candidate set or the UI.

Further reading

6

Check your understanding

0/4 answered