Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Why chunk at all

Retrieval granularity against context cost

A retriever ranks units, not documents. The unit it ranks is whatever you embedded, and the unit it returns is whatever you paste into the prompt. Those are two different pressures. Retrieval wants the unit small, so that a vector represents one idea and the nearest neighbour is actually about the query. Generation wants the unit large, so that the passage handed to the model contains the surrounding sentences the answer needs. Chunking is the negotiation between them.

The four strategies worth knowing, plus the one that changed the abstraction level, are:

💡 The durable idea: there is no correct chunk size, only a size that suits the questions you are asking. Make the choice visible, then measure — a chunker that looks reasonable in a design doc is a hypothesis, and retrieval metrics are the experiment.
⚠️ The trap: optimizing the chunker in isolation. Every one of these strategies changes what is indexable, and a change that helps one question class often hurts another. Chunk-size tuning without a labelled query set is guessing with extra steps.
2

Four cuts over the same passage

Same text, different boundaries

Below, several of the corpus documents are concatenated into one passage and cut four ways by AppSim.chunk. The passage band at the top is the character axis; each row underneath is one strategy's boundaries aligned to the same axis, so you can see exactly where they disagree. Drag the size ceiling and the fixed-strategy overlap and watch the rows diverge: fixed slicing ignores the sentence ticks, recursive honours them, and semantic lets the topic decide.

Top band: the passage with sentence ticks. Four rows below: chunk boundaries for fixed, recursive, semantic and parent.

💡 Read the rows, not the counts: the interesting fact is that recursive and semantic produce variable boundaries — chunk 3 may be a single sentence and chunk 4 a paragraph — while fixed is uniform by construction. Uniformity is not a virtue when the text is not uniform.
3

The boundary effect

A fact split across a boundary is retrievable from neither half

Here is the failure that no cleverness in the embedder can repair. If a sentence carrying the answer straddles two chunks, then neither chunk contains the sentence. The chunk on the left has the setup, the chunk on the right has the payoff, and the query — which uses the vocabulary of the whole sentence — matches neither strongly. The fact is in the index and is not retrievable. This is not a ranking error; it is an indexing invariant, and it is why overlap and boundary-aware splitting exist at all.

Slide the chunk size and pick which sentence is the target fact. The band shows the chunks and the fact's span; the sweep underneath samples seeded query variants of the fact's terms and plots the hit rate against size, with the sizes that cut the fact shaded. The troughs are exactly those sizes — the fact is in the passage and in no indexed unit.

Top: chunk boundaries and the target fact (the boxed span). Bottom: seeded retrieval hit rate against chunk size; shaded sizes are ones where a boundary cuts the fact.

⚠️ Overlap is a mitigation, not a fix. Overlapping windows guarantee that many splits land inside a copy of the text, at the cost of storing and searching near-duplicate chunks — and a reranker that has not been told about the duplication will happily return the same passage three times.
4

Precision and recall pull in opposite directions

There is a size that maximizes neither

Small chunks are precise: the vector represents one idea, so the nearest neighbour is usually on-topic and the top result is the right passage. But small chunks lose recall, because the answer is more likely to have been split or to have left its qualifier in a neighbouring chunk. Large chunks are the reverse: the answer is almost certainly inside the retrieved block, but the block is mostly other material, so precision falls and the model has to find the needle in what you handed it.

The curve below is a seeded model of that tension, driven by the actual token length of each corpus document. Precision falls as the chunk grows; recall rises; F1 peaks somewhere you have to find rather than assume. Move the marker to read the three metrics at any size.

Precision, recall and F1 against chunk size (tokens), averaged over the corpus. The marker is the size you selected.

💡 The move that sidesteps the tradeoff: decouple the unit you match from the unit you return. Parent-document retrieval indexes small children (precision) and returns their parent (recall), so the tension is resolved by the architecture rather than by a compromise size.
5

Contextual retrieval

Give every chunk the context it lost when it was cut

A chunk is ambiguous precisely because it was extracted. "The company's revenue grew 12%" cannot be matched by a query that mentions the company, the quarter, or the filing, because the chunk never says any of those things. Contextual retrieval (Anthropic, September 2024) fixes this cheaply: before embedding, ask a model to write one or two sentences situating the chunk inside its source document, and prepend that blurb to the chunk. The chunk now carries its own provenance, and the embedder has something to match on.

// contextual retrieval: situate the chunk, then index the situated text { "chunk": "revenue grew 12% year over year", "context": "From ACME's Q2 2023 10-K, MD&A section, on cloud segment growth.", "indexed": "<context>From ACME's Q2 2023 10-K, MD&A section…</context>\nrevenue grew 12% year over year" }

Contextual embedding alone is only the first tier. Anthropic's reported top-20 retrieval failure rate on a large corpus falls again when the contextualised chunks are combined with BM25 in a hybrid retriever, and again when a reranker reorders the result. The four tiers below are those reported numbers. Toggle a tier to run a hundred seeded queries against it and see the absolute failure rate — and notice how much larger the relative reductions sound than the absolute moves.

Reported top-20 retrieval failure rate per tier (lower is better). The selected tier is highlighted; the readout simulates 100 seeded queries at that rate.

Two cautions. First, the famous "35% → 49% → 67%" figures are relative reductions in the failure rate, not percentage points; the underlying absolute rates are 5.7%, 3.7%, 2.9% and 1.9%, so the whole journey is a 3.8-point move. Second, these are corpus-, chunking- and reranker-dependent: they are the shape of the win, not a promise about yours. Treat the tiers as a menu and measure each on your own labelled queries.

⚠️ What contextual retrieval does not do: it does not repair a boundary that cut the answer in half. It adds context around a chunk; it cannot join two chunks. That is still the job of overlap, parent-document retrieval, or a sentence-level splitter that keeps the sentence whole.

Cheat sheet

QuestionThe answer
Why chunk at all?An embedding is per unit, so the unit decides what is retrievable; small favours precision, large favours recall.
Which strategy first?Recursive: split on structure and pack to a ceiling. A strong, cheap baseline before anything cleverer.
What does semantic chunking buy?Boundaries follow topic drift rather than punctuation — at the cost of variable-sized chunks and a threshold to tune.
What is parent-document retrieval?Match small children, return their larger parent. Decouples the match unit from the context unit.
What is RAPTOR?A recursive summarisation tree, so leaf chunks answer detail questions and summary nodes answer thematic ones.
What is the boundary effect?A fact split across two chunks is fully contained in neither, so it is retrievable from neither. Indexing, not ranking.
Does overlap fix it?It reduces the frequency, at the cost of near-duplicate storage and duplicate results unless reranking dedupes.
What is contextual retrieval?Prepend a generated context blurb to each chunk before embedding; combine with BM25 and a reranker for the best tier.
Do the 35/49/67% numbers mean that?They are relative reductions. The absolute top-20 failure rate goes 5.7% → 3.7% → 2.9% → 1.9%.

Further reading

6

Check your understanding

0/5 answered