Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Query transforms: rewrite the question before you search

HyDE, multi-query, step-back

Retrieval quality is bounded by the query as much as by the index. A terse question and the passage that answers it often share almost no vocabulary, which is exactly the mismatch the lexical layer cannot bridge. Three transforms rewrite the query before it is matched, and each buys recall in a different way.

HyDE (Gao et al., arXiv:2212.10496) asks the model to write a hypothetical answer to the question and embeds that instead of the question. Answers resemble answers, so the hypothetical document lands nearer the real passages than the terse query does. Multi-query expansion generates several paraphrases of the same question and unions their result lists (fused with RRF, so a document that one phrasing missed another can still catch). Step-back prompting replaces the specific question with a more general one — what class of problem is this? — and retrieves for the abstraction, which helps when the passage is written in the vocabulary of the general rule rather than the specific instance.

💡 When each helps: HyDE on short, vague or vocabulary-mismatched queries; multi-query when recall matters more than latency (each paraphrase is another retrieval call); step-back on reasoning-shaped questions whose answer depends on a principle. All three cost extra model calls and can hurt precision — a hypothetical answer is a guess, and a confident wrong guess retrieves confident wrong neighbours. Measure recall, not just top-1.

Pick a query and compare four strategies over the shared corpus. Each row lists the top three documents and marks the ones on the hand-labelled relevant set; the bar is recall@5.

raw vs HyDE vs multi-query vs step-back. Green titles are documents the labelled set counts as relevant.

2

Self-critique: retrieve, then check

Self-RAG and CRAG turn retrieval into a decision

Self-RAG (Asai et al., arXiv:2310.11511, ICLR 2024) trains the model to emit reflection tokens at generation time: should I retrieve at all, is this passage relevant, is my draft supported by it? The model learns to critique its own retrieval rather than consuming whatever the pipeline returned. CRAG (Yan et al., arXiv:2401.15884, 2024) attaches a lightweight evaluator to the retriever and grades the returned set as correct, ambiguous or incorrect; a poor grade triggers a rewrite or a fallback to web search instead of proceeding with bad evidence.

The methods differ, but the engineering lesson is shared and portable: retrieval should be a decision with a check, not a fixed stage. If nothing verifies relevance and nothing verifies that the answer stayed inside the evidence, a bad retrieval becomes a bad answer with no signal in between. That is the failure the RAGAS metrics in the last section are designed to expose.

⚠️ The trap: a self-critique step that is graded by the same model that wrote the answer is a closed loop. Judge against a held-out set of labelled passages, or at least against a different model, or the critique learns to agree with itself.
3

Local and global questions: GraphRAG

“Find the passage” versus “what themes run through this corpus”

Flat chunk retrieval answers a local question: find the passage that mentions X. It cannot answer a global question — what themes run through this corpus? — because the answer is not in any passage. It is a property of the whole set, and top-k retrieval returns unrelated passages that happen to share a word.

GraphRAG (Microsoft Research, 2024) targets that gap. An LLM pass extracts entities and relationships from the corpus, entity communities are detected, and each community is pre-summarised. A global question is then answered by reading summaries of summaries, not by hunting for a lucky chunk. The trade is real: building and refreshing the graph and its summaries is an LLM pass over the whole corpus, and the index has to be maintained as the corpus changes. So it answers a class of question flat retrieval simply cannot, and it is overkill for find-the-passage.

Toggle the two questions. The flat retriever on the left is doing exactly what it was built to do; it is just being asked the wrong kind of question.

Left: flat chunk retrieval. Right: a graph/community summary index (mock). The query is the same kind of text in tone, different in kind.

4

Multimodal and page-image retrieval

When OCR throws away the answer

PDFs, slide decks and scanned pages are where text-first pipelines lose the most. OCR turns a page into a flat stream of characters and in doing so discards layout: which number belongs to which row of a table, that a figure caption describes the chart above it, that two labels sit in the same column. A question about a value in a table then retrieves a page whose extracted text mentions none of the query words.

ColPali (Faysse et al., arXiv:2407.01449, 2024) skips OCR entirely. A vision-language model embeds page images directly — keeping the page as the retrieval unit and the layout intact — and a text query is matched against multi-vector page representations using late interaction, the ColBERT idea (Khattab & Zaharia, arXiv:2004.12832, SIGIR 2020) of scoring query tokens against document tokens and summing the best matches. The cost is a heavier encoder and much larger index storage; the win is that a question about something a reader would point to retrieves the page a reader would turn to.

💡 Pick the unit to match the document: prose documents chunk well into passages; forms, tables and slides retrieve as page images. A pipeline that forces every document through the same chunk-the-text path is making a layout decision it did not intend to make.
5

Code search: grep versus embeddings

The unfashionable answer is that lexical often wins

Code is a case where the honest answer is not the trendy one. Identifiers, error strings, file paths and API names are literal tokens. grep -r, ripgrep, or a trigram index finds them exactly, cheaply, with no index to embed and a result a human can verify by eye. Embeddings cannot reliably reproduce an exact identifier match, and they are a very expensive way to try.

Embeddings win on intent: “where do we decide whether the caller is allowed to read this record” shares no token with the function that does it, and a lexical search returns nothing. The practical answer is both, and the skill is knowing which query you are holding: literal queries to the lexical index, intent queries to the semantic one, and a reranker when the two disagree. Do not embed your way past a problem grep already solves.

// literal — exact, cheap, verifiable by eye rg "AuthTokenValidator" packages/ // intent — shares no token with the code; lexical returns nothing "where do we decide if the caller may read a record"

Six queries, each with a lexical score (does the literal token appear?) and an illustrative semantic score. The verdict column changes sides with the query — that is the point.

Blue = lexical (grep) score, pink = embedding-style semantic score. Identifiers and error strings favour grep; intent favours embeddings.

6

“RAG is dead”: an argument, not a verdict

Both sides at full strength

With million-token windows, the claim “long context replaces retrieval” is worth stating as strongly as its proponents do. The pro-window case: if the corpus fits in the window, there is no index to build or keep fresh, no embedding model to drift, no second system to operate, no ACL-filtering layer to get wrong — you paste the documents in and ask. For a single, bounded document set that fits, this is genuinely simpler, and it removes an entire class of retrieval failures (the right chunk that never got retrieved).

The pro-retrieval case: a long window is billed per token on every query. A 500k-token corpus sent whole is priced on every request, so cost scales with corpus size and query volume at once. Retrieval is also where freshness, per-user access control and citable provenance actually live — properties the model cannot supply at any window size, because it does not know who is asking or what changed an hour ago. And even inside a 1M window, attention degrades: lost in the middle shows accuracy is U-shaped rather than flat, so a passage that is technically present may not be used. A window is finite, and it is priced.

Neither side is resolved here, and pretending otherwise would be the error. The conditions decide: move the sliders and watch the verdict flip between “the window is enough” and “you need retrieval”.

Two evidence boards. Items light up when the current conditions make them bite; the banner is computed from the sliders.

7

What to actually do: measure on your corpus

The only tie-breaker

Every technique on this page is corpus-dependent. HyDE helps some query distributions and hurts others; GraphRAG earns its build cost only on global questions; grep wins on code identifiers and loses on intent; the window replaces retrieval for a small bounded set and fails on a large fresh one. The honest resolution of the debate is not a side — it is a measurement on your corpus with your questions.

So keep a small labelled set: real questions paired with the passages that answer them. Compare strategies — raw, HyDE, multi-query, step-back, hybrid, graph — on that set. The RAGAS metrics (Es et al., arXiv:2309.15217, EACL 2024) tell you which failure you have: context recall (did the passages contain what was needed?), context precision (were they relevant?), and faithfulness (did the answer stay inside them?). A retrieval change you cannot measure is a change you cannot defend — and, in the long run, cannot keep.

The contextual-retrieval numbers are the shape of the argument: on one large corpus the top-20 failure rate fell from 5.7% at baseline to 3.7% with contextual embeddings, 2.9% adding BM25, and 1.9% with reranking. Every step was a retrieval-side change measured on a labelled set.

Cheat sheet

QuestionThe answer
What does HyDE change?The query: it embeds a hypothetical answer, so the vector sits nearer real passages than a terse question does.
When does multi-query help?When recall matters more than latency — each paraphrase is another retrieval call, fused with RRF.
When does step-back help?Reasoning-shaped questions whose answer is a principle; retrieve for the abstraction, not the instance.
What do Self-RAG and CRAG add?A check: decide whether to retrieve, grade relevance, and refuse or rewrite bad evidence.
What can flat chunk retrieval not answer?A global question — “what themes run through this corpus” — because the answer is not in any passage.
What is GraphRAG for?Entity/relationship graph plus community summaries, so global questions are answered from summaries of summaries.
What does ColPali retrieve?Page images, not OCR text — layout kept intact, matched with late interaction. Heavier encoder, bigger index.
Grep or embeddings for code?Literal (identifiers, error strings) to grep; intent to embeddings; rerank when they disagree.
Is RAG dead?It is an argument. Windows are finite and billed per token; retrieval owns freshness, ACLs, citations. Measure your corpus.
What is the tie-breaker?A labelled set and the RAGAS metrics: context recall, context precision, faithfulness.

Further reading

8

Check your understanding

0/5 answered