Chunking, and contextual retrieval
An embedding is computed per unit of text, and a retriever can only return what was indexed. That makes the cut the first real decision in a retrieval system: split documents too finely and a fact falls between two chunks; too coarsely and the vector is an average of everything and points at nothing. This part works through the four workhorse strategies, the boundary effect that none of them escapes, and the mitigation that has the best published numbers — contextual retrieval.
Why chunk at all
Retrieval granularity against context cost
A retriever ranks units, not documents. The unit it ranks is whatever you embedded, and the unit it returns is whatever you paste into the prompt. Those are two different pressures. Retrieval wants the unit small, so that a vector represents one idea and the nearest neighbour is actually about the query. Generation wants the unit large, so that the passage handed to the model contains the surrounding sentences the answer needs. Chunking is the negotiation between them.
The four strategies worth knowing, plus the one that changed the abstraction level, are:
- Fixed-size — slice every N characters, optionally with overlap. Trivial, uniform, and blind to meaning: it will cut a table in half mid-row.
- Recursive — split on structure (paragraphs, then sentences), greedily packing each chunk up to a size ceiling. The default in most frameworks, and a strong baseline.
- Semantic — start a new chunk when the embedding (or, cheaply, the term overlap) between adjacent sentences drifts below a threshold. Boundaries follow topic rather than punctuation.
- Parent-document — index small children for matching, but return the larger parent they sit in. Match small, return large.
- RAPTOR — recursively cluster and summarise chunks into a tree, so the same corpus can answer both a detail question (leaf) and a "what themes run through this?" question (a summary node). Multi-level abstraction rather than a single cut.
Four cuts over the same passage
Same text, different boundaries
Below, several of the corpus documents are concatenated into one passage and cut four ways by AppSim.chunk. The passage band at the top is the character axis; each row underneath is one strategy's boundaries aligned to the same axis, so you can see exactly where they disagree. Drag the size ceiling and the fixed-strategy overlap and watch the rows diverge: fixed slicing ignores the sentence ticks, recursive honours them, and semantic lets the topic decide.
Top band: the passage with sentence ticks. Four rows below: chunk boundaries for fixed, recursive, semantic and parent.
The boundary effect
A fact split across a boundary is retrievable from neither half
Here is the failure that no cleverness in the embedder can repair. If a sentence carrying the answer straddles two chunks, then neither chunk contains the sentence. The chunk on the left has the setup, the chunk on the right has the payoff, and the query — which uses the vocabulary of the whole sentence — matches neither strongly. The fact is in the index and is not retrievable. This is not a ranking error; it is an indexing invariant, and it is why overlap and boundary-aware splitting exist at all.
Slide the chunk size and pick which sentence is the target fact. The band shows the chunks and the fact's span; the sweep underneath samples seeded query variants of the fact's terms and plots the hit rate against size, with the sizes that cut the fact shaded. The troughs are exactly those sizes — the fact is in the passage and in no indexed unit.
Top: chunk boundaries and the target fact (the boxed span). Bottom: seeded retrieval hit rate against chunk size; shaded sizes are ones where a boundary cuts the fact.
Precision and recall pull in opposite directions
There is a size that maximizes neither
Small chunks are precise: the vector represents one idea, so the nearest neighbour is usually on-topic and the top result is the right passage. But small chunks lose recall, because the answer is more likely to have been split or to have left its qualifier in a neighbouring chunk. Large chunks are the reverse: the answer is almost certainly inside the retrieved block, but the block is mostly other material, so precision falls and the model has to find the needle in what you handed it.
The curve below is a seeded model of that tension, driven by the actual token length of each corpus document. Precision falls as the chunk grows; recall rises; F1 peaks somewhere you have to find rather than assume. Move the marker to read the three metrics at any size.
Precision, recall and F1 against chunk size (tokens), averaged over the corpus. The marker is the size you selected.
Contextual retrieval
Give every chunk the context it lost when it was cut
A chunk is ambiguous precisely because it was extracted. "The company's revenue grew 12%" cannot be matched by a query that mentions the company, the quarter, or the filing, because the chunk never says any of those things. Contextual retrieval (Anthropic, September 2024) fixes this cheaply: before embedding, ask a model to write one or two sentences situating the chunk inside its source document, and prepend that blurb to the chunk. The chunk now carries its own provenance, and the embedder has something to match on.
Contextual embedding alone is only the first tier. Anthropic's reported top-20 retrieval failure rate on a large corpus falls again when the contextualised chunks are combined with BM25 in a hybrid retriever, and again when a reranker reorders the result. The four tiers below are those reported numbers. Toggle a tier to run a hundred seeded queries against it and see the absolute failure rate — and notice how much larger the relative reductions sound than the absolute moves.
Reported top-20 retrieval failure rate per tier (lower is better). The selected tier is highlighted; the readout simulates 100 seeded queries at that rate.
Two cautions. First, the famous "35% → 49% → 67%" figures are relative reductions in the failure rate, not percentage points; the underlying absolute rates are 5.7%, 3.7%, 2.9% and 1.9%, so the whole journey is a 3.8-point move. Second, these are corpus-, chunking- and reranker-dependent: they are the shape of the win, not a promise about yours. Treat the tiers as a menu and measure each on your own labelled queries.
Cheat sheet
| Question | The answer |
|---|---|
| Why chunk at all? | An embedding is per unit, so the unit decides what is retrievable; small favours precision, large favours recall. |
| Which strategy first? | Recursive: split on structure and pack to a ceiling. A strong, cheap baseline before anything cleverer. |
| What does semantic chunking buy? | Boundaries follow topic drift rather than punctuation — at the cost of variable-sized chunks and a threshold to tune. |
| What is parent-document retrieval? | Match small children, return their larger parent. Decouples the match unit from the context unit. |
| What is RAPTOR? | A recursive summarisation tree, so leaf chunks answer detail questions and summary nodes answer thematic ones. |
| What is the boundary effect? | A fact split across two chunks is fully contained in neither, so it is retrievable from neither. Indexing, not ranking. |
| Does overlap fix it? | It reduces the frequency, at the cost of near-duplicate storage and duplicate results unless reranking dedupes. |
| What is contextual retrieval? | Prepend a generated context blurb to each chunk before embedding; combine with BM25 and a reranker for the best tier. |
| Do the 35/49/67% numbers mean that? | They are relative reductions. The absolute top-20 failure rate goes 5.7% → 3.7% → 2.9% → 1.9%. |
Further reading
- Anthropic, "Introducing Contextual Retrieval", 19 September 2024 — the four tiers and the absolute 5.7% → 1.9% top-20 failure rates behind step 5.
- Sarthi et al., "RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval", ICLR 2024 — the summarisation tree for multi-level abstraction.
- Günther et al., "Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models", 2024 — chunk after embedding rather than before, so each chunk vector retains document context.
- Chen et al., "Dense X Retrieval: What Retrieval Granularity Should We Use?", 2023 — propositions as retrieval units, a finer cut than the sentence.
- Jégou, Douze & Schmid, "Product Quantization for Nearest Neighbor Search", IEEE TPAMI 2011 — the compression half that Part 15 builds the index families on.
- Anthropic, "Effective context engineering for AI agents", 29 September 2025 — retrieval as selection, and why fewer, better chunks usually beat more context.