Beyond chunks: query transforms, graph, multimodal, code — and the “RAG is dead” debate
The pipeline so far is the orthodox one: split documents, embed the chunks, retrieve a wide candidate set, rerank narrow, put the best passages in the context. This part is the edge of that map. Some questions are answered by rewriting the query before you search; some are properties of the whole corpus that no single passage contains; some live in page images that OCR destroys; some are code, where the unfashionable answer is that grep often wins. And the loudest claim in the field — that a long context makes retrieval obsolete — is treated here as an argument with two sides, not a settled fact.
Query transforms: rewrite the question before you search
HyDE, multi-query, step-back
Retrieval quality is bounded by the query as much as by the index. A terse question and the passage that answers it often share almost no vocabulary, which is exactly the mismatch the lexical layer cannot bridge. Three transforms rewrite the query before it is matched, and each buys recall in a different way.
HyDE (Gao et al., arXiv:2212.10496) asks the model to write a hypothetical answer to the question and embeds that instead of the question. Answers resemble answers, so the hypothetical document lands nearer the real passages than the terse query does. Multi-query expansion generates several paraphrases of the same question and unions their result lists (fused with RRF, so a document that one phrasing missed another can still catch). Step-back prompting replaces the specific question with a more general one — what class of problem is this? — and retrieves for the abstraction, which helps when the passage is written in the vocabulary of the general rule rather than the specific instance.
Pick a query and compare four strategies over the shared corpus. Each row lists the top three documents and marks the ones on the hand-labelled relevant set; the bar is recall@5.
raw vs HyDE vs multi-query vs step-back. Green titles are documents the labelled set counts as relevant.
Self-critique: retrieve, then check
Self-RAG and CRAG turn retrieval into a decision
Self-RAG (Asai et al., arXiv:2310.11511, ICLR 2024) trains the model to emit reflection tokens at generation time: should I retrieve at all, is this passage relevant, is my draft supported by it? The model learns to critique its own retrieval rather than consuming whatever the pipeline returned. CRAG (Yan et al., arXiv:2401.15884, 2024) attaches a lightweight evaluator to the retriever and grades the returned set as correct, ambiguous or incorrect; a poor grade triggers a rewrite or a fallback to web search instead of proceeding with bad evidence.
The methods differ, but the engineering lesson is shared and portable: retrieval should be a decision with a check, not a fixed stage. If nothing verifies relevance and nothing verifies that the answer stayed inside the evidence, a bad retrieval becomes a bad answer with no signal in between. That is the failure the RAGAS metrics in the last section are designed to expose.
Local and global questions: GraphRAG
“Find the passage” versus “what themes run through this corpus”
Flat chunk retrieval answers a local question: find the passage that mentions X. It cannot answer a global question — what themes run through this corpus? — because the answer is not in any passage. It is a property of the whole set, and top-k retrieval returns unrelated passages that happen to share a word.
GraphRAG (Microsoft Research, 2024) targets that gap. An LLM pass extracts entities and relationships from the corpus, entity communities are detected, and each community is pre-summarised. A global question is then answered by reading summaries of summaries, not by hunting for a lucky chunk. The trade is real: building and refreshing the graph and its summaries is an LLM pass over the whole corpus, and the index has to be maintained as the corpus changes. So it answers a class of question flat retrieval simply cannot, and it is overkill for find-the-passage.
Toggle the two questions. The flat retriever on the left is doing exactly what it was built to do; it is just being asked the wrong kind of question.
Left: flat chunk retrieval. Right: a graph/community summary index (mock). The query is the same kind of text in tone, different in kind.
Multimodal and page-image retrieval
When OCR throws away the answer
PDFs, slide decks and scanned pages are where text-first pipelines lose the most. OCR turns a page into a flat stream of characters and in doing so discards layout: which number belongs to which row of a table, that a figure caption describes the chart above it, that two labels sit in the same column. A question about a value in a table then retrieves a page whose extracted text mentions none of the query words.
ColPali (Faysse et al., arXiv:2407.01449, 2024) skips OCR entirely. A vision-language model embeds page images directly — keeping the page as the retrieval unit and the layout intact — and a text query is matched against multi-vector page representations using late interaction, the ColBERT idea (Khattab & Zaharia, arXiv:2004.12832, SIGIR 2020) of scoring query tokens against document tokens and summing the best matches. The cost is a heavier encoder and much larger index storage; the win is that a question about something a reader would point to retrieves the page a reader would turn to.
Code search: grep versus embeddings
The unfashionable answer is that lexical often wins
Code is a case where the honest answer is not the trendy one. Identifiers, error strings, file paths and API names are literal tokens. grep -r, ripgrep, or a trigram index finds them exactly, cheaply, with no index to embed and a result a human can verify by eye. Embeddings cannot reliably reproduce an exact identifier match, and they are a very expensive way to try.
Embeddings win on intent: “where do we decide whether the caller is allowed to read this record” shares no token with the function that does it, and a lexical search returns nothing. The practical answer is both, and the skill is knowing which query you are holding: literal queries to the lexical index, intent queries to the semantic one, and a reranker when the two disagree. Do not embed your way past a problem grep already solves.
Six queries, each with a lexical score (does the literal token appear?) and an illustrative semantic score. The verdict column changes sides with the query — that is the point.
Blue = lexical (grep) score, pink = embedding-style semantic score. Identifiers and error strings favour grep; intent favours embeddings.
“RAG is dead”: an argument, not a verdict
Both sides at full strength
With million-token windows, the claim “long context replaces retrieval” is worth stating as strongly as its proponents do. The pro-window case: if the corpus fits in the window, there is no index to build or keep fresh, no embedding model to drift, no second system to operate, no ACL-filtering layer to get wrong — you paste the documents in and ask. For a single, bounded document set that fits, this is genuinely simpler, and it removes an entire class of retrieval failures (the right chunk that never got retrieved).
The pro-retrieval case: a long window is billed per token on every query. A 500k-token corpus sent whole is priced on every request, so cost scales with corpus size and query volume at once. Retrieval is also where freshness, per-user access control and citable provenance actually live — properties the model cannot supply at any window size, because it does not know who is asking or what changed an hour ago. And even inside a 1M window, attention degrades: lost in the middle shows accuracy is U-shaped rather than flat, so a passage that is technically present may not be used. A window is finite, and it is priced.
Neither side is resolved here, and pretending otherwise would be the error. The conditions decide: move the sliders and watch the verdict flip between “the window is enough” and “you need retrieval”.
Two evidence boards. Items light up when the current conditions make them bite; the banner is computed from the sliders.
What to actually do: measure on your corpus
The only tie-breaker
Every technique on this page is corpus-dependent. HyDE helps some query distributions and hurts others; GraphRAG earns its build cost only on global questions; grep wins on code identifiers and loses on intent; the window replaces retrieval for a small bounded set and fails on a large fresh one. The honest resolution of the debate is not a side — it is a measurement on your corpus with your questions.
So keep a small labelled set: real questions paired with the passages that answer them. Compare strategies — raw, HyDE, multi-query, step-back, hybrid, graph — on that set. The RAGAS metrics (Es et al., arXiv:2309.15217, EACL 2024) tell you which failure you have: context recall (did the passages contain what was needed?), context precision (were they relevant?), and faithfulness (did the answer stay inside them?). A retrieval change you cannot measure is a change you cannot defend — and, in the long run, cannot keep.
The contextual-retrieval numbers are the shape of the argument: on one large corpus the top-20 failure rate fell from 5.7% at baseline to 3.7% with contextual embeddings, 2.9% adding BM25, and 1.9% with reranking. Every step was a retrieval-side change measured on a labelled set.
Cheat sheet
| Question | The answer |
|---|---|
| What does HyDE change? | The query: it embeds a hypothetical answer, so the vector sits nearer real passages than a terse question does. |
| When does multi-query help? | When recall matters more than latency — each paraphrase is another retrieval call, fused with RRF. |
| When does step-back help? | Reasoning-shaped questions whose answer is a principle; retrieve for the abstraction, not the instance. |
| What do Self-RAG and CRAG add? | A check: decide whether to retrieve, grade relevance, and refuse or rewrite bad evidence. |
| What can flat chunk retrieval not answer? | A global question — “what themes run through this corpus” — because the answer is not in any passage. |
| What is GraphRAG for? | Entity/relationship graph plus community summaries, so global questions are answered from summaries of summaries. |
| What does ColPali retrieve? | Page images, not OCR text — layout kept intact, matched with late interaction. Heavier encoder, bigger index. |
| Grep or embeddings for code? | Literal (identifiers, error strings) to grep; intent to embeddings; rerank when they disagree. |
| Is RAG dead? | It is an argument. Windows are finite and billed per token; retrieval owns freshness, ACLs, citations. Measure your corpus. |
| What is the tie-breaker? | A labelled set and the RAGAS metrics: context recall, context precision, faithfulness. |
Further reading
- Gao, Ma, Lin & Callan, “Precise Zero-Shot Dense Retrieval without Relevance Labels” (HyDE), 2022 — embed a hypothetical answer instead of the query.
- Asai et al., “Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection”, ICLR 2024 — reflection tokens as a retrieval decision.
- Yan et al., “Corrective Retrieval Augmented Generation” (CRAG), 2024 — grade the retrieved set, then correct or fall back.
- Microsoft Research, “GraphRAG: Unlocking LLM discovery on narrative private data”, 2024 — the global-question case and community summaries.
- Faysse et al., “ColPali: Efficient Document Retrieval with Vision Language Models”, 2024 — page-image retrieval without OCR.
- Khattab & Zaharia, “ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT”, SIGIR 2020 — the late-interaction idea ColPali reuses.
- Es et al., “RAGAS: Automated Evaluation of Retrieval Augmented Generation”, EACL 2024 — faithfulness, answer relevance, context precision, context recall.
- Liu et al., “Lost in the Middle: How Language Models Use Long Contexts”, TACL 2024 — the U-shaped curve behind the pro-retrieval evidence.