Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The corpus this volume searches

A visual index of the shared throughline

Every part in the volume works over the same sixty-document corpus — this site's own guides, summarised — and its hand-fixed coordinates in two dimensions are the "embedding" the interactive demos move through. When a part talks about a query's position, a retrieval miss, or a community in a graph, it is talking about these points. Click a legend entry below to filter the scatter and see which documents a cluster contains.

AppSim.CORPUS, coloured by guide. The same points underline every retrieval demo in the volume.

2

Glossary

Every term, linked to the part that introduces it

💡 Filter the list. Type any fragment — a term, a technique, a failure mode — and the table hides non-matching rows and reports how many remain. Matching is case-insensitive and looks at both the term and its definition.
TermWhat it meansIntroduced in
A
ACL-filtered retrievalRetrieval that applies the caller's access-control list before ranking, so a restricted passage never enters the candidate set or the UI.Part 12
AltitudeHow specific a system instruction is: too low and it over-constrains the model, too high and it says nothing actionable.Part 6
B
Bi-encoderAn embedding model that encodes query and document independently into one vector each, so search is a fast vector lookup.Part 13
BM25The bag-of-words ranker that weights rare terms by IDF, saturates repeats with k1, and normalises by length with b.Part 12
Boundary effectThe quality loss from splitting a document where a passage's meaning straddles a chunk edge.Part 14
BPE (byte-pair encoding)The subword tokenizer that merges frequent character pairs, so text becomes tokens that are neither words nor characters.Part 2
C
Cache breakpointThe marker that tells a provider where a reusable prefix ends, so the cached portion is billed cheaper.Part 11
CalibrationWhether a model's stated confidence matches its accuracy; post-RLHF models are often poorly calibrated.Part 3
ChunkingSplitting documents into retrieval units, balancing retrieval precision against the context cost of each passage.Part 14
ColBERTLate-interaction retrieval that keeps per-token vectors and scores query tokens against document tokens.Part 15
ColPaliPage-image retrieval with a vision-language model, skipping OCR so tables and layout survive.Part 17
CompactionSummarising or evicting history to fit the window; lossy, because an evicted fact is indistinguishable from one that never existed.Part 10
Constrained decodingMasking the token distribution to a grammar or schema so the output is valid by construction.Part 9
Context precisionThe fraction of retrieved passages that are relevant — the metric for a retriever that pulled in noise.Part 17
Context recallWhether the retrieved passages contained the evidence the answer needed — the metric for missing evidence.Part 17
Context rotAccuracy degradation that begins well below the advertised window, not only at the limit.Part 10
Context windowThe maximum tokens — prompt plus completion — a model can attend to in one call.Part 2
Context window as the unit of engineeringThe volume's first thesis: the smallest set of high-signal tokens that produces the outcome is what you actually design.Part 1
Contextual retrievalPrepending a short generated context to each chunk before embedding, so the chunk is not ambiguous out of context.Part 14
Cosine similarityThe dot product of two vectors divided by their norms; the standard similarity for embeddings.Part 13
CRAG (Corrective RAG)Grading the retrieved set and correcting or falling back when the grade is poor.Part 17
Cross-encoderA model that scores a query and a passage jointly; far more accurate than a bi-encoder, far too slow to retrieve with.Part 16
D
Data/instruction separationKeeping untrusted data inside delimiters and never letting it look like an instruction — the first security lesson.Part 6
E
efSearchThe HNSW query-time knob trading recall for latency; raising it does not require a rebuild.Part 15
EmbeddingA vector that represents meaning; similarity is measured by cosine distance.Part 13
Eval as the moatThe volume's second thesis: a labelled set and a deterministic check are the asset that compounds.Part 5
Eval harnessThe labelled examples plus a deterministic check that turns a quality claim into a pass rate.Part 5
F
FaithfulnessWhether the answer stayed inside the retrieved evidence rather than adding unsupported claims.Part 17
Few-shotSpecifying a task by example rather than description; the model also copies the examples' format.Part 7
Filtered ANNApproximate nearest-neighbour search that respects a metadata or ACL filter without destroying recall.Part 15
Format contagionThe tendency of a model to copy the format — including the mistakes — of its examples.Part 7
FSM maskThe finite-state machine a grammar compiles to, whose current state indexes the set of legal next tokens.Part 9
G
Golden setThe curated, trusted examples a system is measured against.Part 5
GraphRAGAn entity/relationship graph plus community summaries, built to answer global “what themes” questions.Part 17
GroundednessWhether each claim in an answer traces to a source in the context.Part 17
H
Held-out setExamples kept out of tuning so a reported score is honest.Part 5
HNSWA layered navigable small-world graph: M sets the degree, efSearch the query-time recall/latency dial.Part 15
HyDEEmbedding a hypothetical answer instead of the question, so the vector sits nearer real passages.Part 17
I
IDFInverse document frequency: the rarer a term across the corpus, the more it contributes to a BM25 score.Part 12
IVFInverted-file index: cluster vectors into nlist cells and probe the nearest nprobe at query time.Part 15
J
JIT retrievalFetching a document at the moment the model needs it, rather than preloading it into the context.Part 10
JSON SchemaThe grammar that constrains a structured output to a valid shape.Part 9
Jump-forward decodingEmitting a run of forced tokens at once when the grammar admits only one continuation.Part 9
L
Late interactionScoring query tokens against document tokens and summing the best matches — ColBERT's middle ground between bi- and cross-encoders.Part 15
LogprobThe log-probability a model assigns a token; a confidence signal, not a calibrated probability.Part 3
Lost-in-the-middleThe U-shaped accuracy curve: content at the start and end of a long context is used better than content in the middle.Part 10
M
M (HNSW degree)The number of neighbours per node in an HNSW graph; higher M means better recall and more memory.Part 15
MatryoshkaNested embeddings that can be truncated at 1/2, 1/4 or 1/8 of their width with a small recall loss.Part 13
MTEBThe embedding leaderboard; useful as a shortlist, imperfect as a predictor on your own corpus.Part 13
Multi-queryExpanding a query into several paraphrases and fusing their result lists to raise recall.Part 17
MUVERAA method that turns multi-vector ColBERT similarity into single-vector maximum-inner-product search.Part 15
N
nprobeThe number of IVF cells probed per query; the recall/latency knob.Part 15
P
Parent-document retrievalMatching on small chunks but returning the larger parent passage to the model.Part 14
Pass rateThe share of labelled cases a system gets right — the number an eval harness reports.Part 5
Prefix cacheReusing the cached computation of a byte-identical prompt prefix, billed cheaper than fresh input.Part 11
Product quantizationCompressing vectors into short codes by quantising subvectors, trading recall for memory.Part 15
Prompt versionThe identifier for a frozen prompt, so a change in score can be attributed to a change you actually made.Part 5
R
RAGRetrieval-augmented generation: retrieve passages, put them in the context, and answer with citations.Part 12
RAPTORRecursively clustering and summarising chunks into a tree, so retrieval can choose the level of abstraction.Part 14
RecallThe fraction of the true nearest neighbours an ANN index actually returns; the currency of index tuning.Part 15
RRFReciprocal rank fusion: score(d) = the sum of 1/(k + rank) across rankings, with k = 60 by default.Part 16
S
Self-RAGReflection tokens that let the model decide whether to retrieve and whether its draft is supported.Part 17
SPLADELearned sparse retrieval: a model produces an expanded, weighted sparse term vector.Part 15
Step-backReplacing a specific question with a more general one and retrieving for the abstraction.Part 17
Structured notesA durable, compact store the agent writes facts into so compaction cannot silently drop them.Part 10
T
TemperatureThe sampling divisor on logits; higher flattens the distribution, lower sharpens it, and 0 is not deterministic.Part 3
TokenThe unit a model reads and writes, produced by a tokenizer such as BPE, and not the same as a word.Part 2
Top-kSampling restricted to the k highest-probability tokens.Part 3
Top-pNucleus sampling: the smallest set of tokens whose cumulative probability exceeds p.Part 3
TPOTTime per output token after the first; with TTFT it defines a serving latency contract.Part 4
TTFTTime to first token, dominated by queueing and prefill; the responsiveness a streaming user perceives.Part 4
3

Retrieval-architecture selection table

Which index, for which job, and when not to

Read the table by finding the row whose when not to use describes you: that is usually the decisive column. Flat is the baseline every approximation is measured against; the ANN rows trade recall for latency and memory in different currencies. Hybrid plus rerank is the production default for text because lexical and dense failure modes are complementary, and the last two rows answer question shapes the others structurally cannot.

ArchitectureBest forRecall / latency / memoryOperational costWhen not to use
Flat (exact)Small collections and the recall baseline every approximation is measured against.recall 1.00 · latency highest · memory highestNone beyond storage; no tuning.Millions of vectors under a latency SLO — brute force is the thing you are trying to avoid.
IVFLarge collections where a small recall loss is acceptable and memory is not the constraint.recall 0.90–0.98 · low latency · index + data in memoryModerate: choose nlist, tune nprobe, retrain when the distribution drifts.When you cannot tolerate the true neighbour's cell not being probed, or you have no labelled set to tune nprobe against.
IVF-PQHuge collections that must fit in memory; compression matters more than the last recall points.recall 0.80–0.95 · low latency · much lower memoryBuild plus codebook tuning; quantisation error is a new, measurable regression surface.When recall is the binding constraint, or the vectors are short and dense enough that PQ error hurts.
HNSWLow-latency, high-recall serving where the graph can live in RAM.recall 0.95–0.99 · very low latency · high memory (graph)High memory; M and efSearch tuning; heavier rebuilds as the graph grows or changes.Memory-bound deployments, or indexes that change constantly and cannot absorb rebuild cost.
DiskANN / VamanaBillions of vectors that must live on SSD on a single node.recall 0.95–0.99 · low latency · index fits on SSDSSD IO is the bottleneck; heavier build; careful temperature and cache management.When RAM is ample and you need the very lowest latency, or the collection is small enough for HNSW.
Hybrid + rerankProduction text retrieval where precision matters and lexical and dense disagree usefully.recall follows the union · rerank pass adds latency · two indexes in memoryTwo indexes, an embedding model, and a cross-encoder; rerank is the dominant cost.Strict sub-10 ms budgets, or corpora small enough that BM25 alone already clears the bar.
GraphGlobal, thematic questions over a corpus — “what themes run through this?”not a chunk-level recall metric · latency from summary lookup · index = graph + summariesAn LLM pass over the whole corpus to build and refresh; extraction quality is the ceiling.Find-the-passage queries, and corpora that change faster than the graph can be rebuilt.
Page-image / ColPaliPDFs, slides, tables and forms where layout carries the answer that OCR discards.recall at the page level · heavier VLM encode latency · large multi-vector indexGPU encoding and much larger storage; a separate index beside the text one.Plain prose corpora, or strict cost and latency budgets where text extraction is good enough.

4

RAG failure modes to metrics

Name the failure, then pick the metric that moves

“The retrieval is bad” is not a diagnosis. Each failure below has a different metric and a different fix, and confusing them is how teams spend a month reranking a system whose real problem was a missing passage. Start from the symptom you actually observe, read across to the metric that makes it visible, and only then choose the fix.

Failure modeMetricFix

5

Formula and reference card

Four expressions and one chunking cheat

Every interactive demo in the volume is a version of one of these. Bytes, tokens and ranks are all approximate on purpose — accurate to a factor you can reason about is the point.

Cosine similarity

$$\cos(\mathbf{q},\mathbf{d}) \;=\; \frac{\mathbf{q}\cdot\mathbf{d}}{\lVert \mathbf{q}\rVert\,\lVert \mathbf{d}\rVert}$$

The angle between two vectors, ignoring magnitude. Meaning lives in direction; length is a nuisance the normalisation removes.

BM25

$$\text{score}(q,d) \;=\; \sum_{t\in q} \text{IDF}(t)\,\frac{f_{t,d}\,(k_1+1)}{f_{t,d}+k_1\!\left(1-b+b\,\frac{\lvert d\rvert}{\text{avgdl}}\right)}$$

Rare terms count more (IDF); repeats saturate (k1, default 1.2); long documents do not win on length alone (b, default 0.75).

Reciprocal rank fusion

$$\text{RRF}(d) \;=\; \sum_{r} \frac{1}{k+\text{rank}_r(d)}, \qquad k = 60$$

Fuses ranked lists without calibrating their scores. It only needs the ranks, which is why it is the default for lexical + dense.

Chunk-strategy cheat

fixed — cut every N characters: trivial, and it severs sentences across boundaries.

recursive — split on paragraphs and sentences, filling to a size: the sane default.

semantic — split where consecutive sentences stop resembling each other: coherent, costlier to build.

parent — match on the child, return the parent passage: precise matching, generous context.