Glossary and selection tables
This is the volume's back matter: one alphabetical index that links every term to the part that introduces it, a selection table for retrieval architectures, a RAG failure-mode to metric table you can check a real pipeline against, and a formula card for the back-of-envelope arithmetic. It is the applications counterpart to the training guide's glossary, which defines the modelling ideas, and the serving guide's glossary, which carries the deployment numbers. The companion volume, Agents in Action, continues from an agent that can search and read; where a term belongs to that volume, the link points there rather than re-teaching it.
The corpus this volume searches
A visual index of the shared throughline
Every part in the volume works over the same sixty-document corpus — this site's own guides, summarised — and its hand-fixed coordinates in two dimensions are the "embedding" the interactive demos move through. When a part talks about a query's position, a retrieval miss, or a community in a graph, it is talking about these points. Click a legend entry below to filter the scatter and see which documents a cluster contains.
AppSim.CORPUS, coloured by guide. The same points underline every retrieval demo in the volume.
Glossary
Every term, linked to the part that introduces it
| Term | What it means | Introduced in |
|---|---|---|
| A | ||
| ACL-filtered retrieval | Retrieval that applies the caller's access-control list before ranking, so a restricted passage never enters the candidate set or the UI. | Part 12 |
| Altitude | How specific a system instruction is: too low and it over-constrains the model, too high and it says nothing actionable. | Part 6 |
| B | ||
| Bi-encoder | An embedding model that encodes query and document independently into one vector each, so search is a fast vector lookup. | Part 13 |
| BM25 | The bag-of-words ranker that weights rare terms by IDF, saturates repeats with k1, and normalises by length with b. | Part 12 |
| Boundary effect | The quality loss from splitting a document where a passage's meaning straddles a chunk edge. | Part 14 |
| BPE (byte-pair encoding) | The subword tokenizer that merges frequent character pairs, so text becomes tokens that are neither words nor characters. | Part 2 |
| C | ||
| Cache breakpoint | The marker that tells a provider where a reusable prefix ends, so the cached portion is billed cheaper. | Part 11 |
| Calibration | Whether a model's stated confidence matches its accuracy; post-RLHF models are often poorly calibrated. | Part 3 |
| Chunking | Splitting documents into retrieval units, balancing retrieval precision against the context cost of each passage. | Part 14 |
| ColBERT | Late-interaction retrieval that keeps per-token vectors and scores query tokens against document tokens. | Part 15 |
| ColPali | Page-image retrieval with a vision-language model, skipping OCR so tables and layout survive. | Part 17 |
| Compaction | Summarising or evicting history to fit the window; lossy, because an evicted fact is indistinguishable from one that never existed. | Part 10 |
| Constrained decoding | Masking the token distribution to a grammar or schema so the output is valid by construction. | Part 9 |
| Context precision | The fraction of retrieved passages that are relevant — the metric for a retriever that pulled in noise. | Part 17 |
| Context recall | Whether the retrieved passages contained the evidence the answer needed — the metric for missing evidence. | Part 17 |
| Context rot | Accuracy degradation that begins well below the advertised window, not only at the limit. | Part 10 |
| Context window | The maximum tokens — prompt plus completion — a model can attend to in one call. | Part 2 |
| Context window as the unit of engineering | The volume's first thesis: the smallest set of high-signal tokens that produces the outcome is what you actually design. | Part 1 |
| Contextual retrieval | Prepending a short generated context to each chunk before embedding, so the chunk is not ambiguous out of context. | Part 14 |
| Cosine similarity | The dot product of two vectors divided by their norms; the standard similarity for embeddings. | Part 13 |
| CRAG (Corrective RAG) | Grading the retrieved set and correcting or falling back when the grade is poor. | Part 17 |
| Cross-encoder | A model that scores a query and a passage jointly; far more accurate than a bi-encoder, far too slow to retrieve with. | Part 16 |
| D | ||
| Data/instruction separation | Keeping untrusted data inside delimiters and never letting it look like an instruction — the first security lesson. | Part 6 |
| E | ||
| efSearch | The HNSW query-time knob trading recall for latency; raising it does not require a rebuild. | Part 15 |
| Embedding | A vector that represents meaning; similarity is measured by cosine distance. | Part 13 |
| Eval as the moat | The volume's second thesis: a labelled set and a deterministic check are the asset that compounds. | Part 5 |
| Eval harness | The labelled examples plus a deterministic check that turns a quality claim into a pass rate. | Part 5 |
| F | ||
| Faithfulness | Whether the answer stayed inside the retrieved evidence rather than adding unsupported claims. | Part 17 |
| Few-shot | Specifying a task by example rather than description; the model also copies the examples' format. | Part 7 |
| Filtered ANN | Approximate nearest-neighbour search that respects a metadata or ACL filter without destroying recall. | Part 15 |
| Format contagion | The tendency of a model to copy the format — including the mistakes — of its examples. | Part 7 |
| FSM mask | The finite-state machine a grammar compiles to, whose current state indexes the set of legal next tokens. | Part 9 |
| G | ||
| Golden set | The curated, trusted examples a system is measured against. | Part 5 |
| GraphRAG | An entity/relationship graph plus community summaries, built to answer global “what themes” questions. | Part 17 |
| Groundedness | Whether each claim in an answer traces to a source in the context. | Part 17 |
| H | ||
| Held-out set | Examples kept out of tuning so a reported score is honest. | Part 5 |
| HNSW | A layered navigable small-world graph: M sets the degree, efSearch the query-time recall/latency dial. | Part 15 |
| HyDE | Embedding a hypothetical answer instead of the question, so the vector sits nearer real passages. | Part 17 |
| I | ||
| IDF | Inverse document frequency: the rarer a term across the corpus, the more it contributes to a BM25 score. | Part 12 |
| IVF | Inverted-file index: cluster vectors into nlist cells and probe the nearest nprobe at query time. | Part 15 |
| J | ||
| JIT retrieval | Fetching a document at the moment the model needs it, rather than preloading it into the context. | Part 10 |
| JSON Schema | The grammar that constrains a structured output to a valid shape. | Part 9 |
| Jump-forward decoding | Emitting a run of forced tokens at once when the grammar admits only one continuation. | Part 9 |
| L | ||
| Late interaction | Scoring query tokens against document tokens and summing the best matches — ColBERT's middle ground between bi- and cross-encoders. | Part 15 |
| Logprob | The log-probability a model assigns a token; a confidence signal, not a calibrated probability. | Part 3 |
| Lost-in-the-middle | The U-shaped accuracy curve: content at the start and end of a long context is used better than content in the middle. | Part 10 |
| M | ||
| M (HNSW degree) | The number of neighbours per node in an HNSW graph; higher M means better recall and more memory. | Part 15 |
| Matryoshka | Nested embeddings that can be truncated at 1/2, 1/4 or 1/8 of their width with a small recall loss. | Part 13 |
| MTEB | The embedding leaderboard; useful as a shortlist, imperfect as a predictor on your own corpus. | Part 13 |
| Multi-query | Expanding a query into several paraphrases and fusing their result lists to raise recall. | Part 17 |
| MUVERA | A method that turns multi-vector ColBERT similarity into single-vector maximum-inner-product search. | Part 15 |
| N | ||
| nprobe | The number of IVF cells probed per query; the recall/latency knob. | Part 15 |
| P | ||
| Parent-document retrieval | Matching on small chunks but returning the larger parent passage to the model. | Part 14 |
| Pass rate | The share of labelled cases a system gets right — the number an eval harness reports. | Part 5 |
| Prefix cache | Reusing the cached computation of a byte-identical prompt prefix, billed cheaper than fresh input. | Part 11 |
| Product quantization | Compressing vectors into short codes by quantising subvectors, trading recall for memory. | Part 15 |
| Prompt version | The identifier for a frozen prompt, so a change in score can be attributed to a change you actually made. | Part 5 |
| R | ||
| RAG | Retrieval-augmented generation: retrieve passages, put them in the context, and answer with citations. | Part 12 |
| RAPTOR | Recursively clustering and summarising chunks into a tree, so retrieval can choose the level of abstraction. | Part 14 |
| Recall | The fraction of the true nearest neighbours an ANN index actually returns; the currency of index tuning. | Part 15 |
| RRF | Reciprocal rank fusion: score(d) = the sum of 1/(k + rank) across rankings, with k = 60 by default. | Part 16 |
| S | ||
| Self-RAG | Reflection tokens that let the model decide whether to retrieve and whether its draft is supported. | Part 17 |
| SPLADE | Learned sparse retrieval: a model produces an expanded, weighted sparse term vector. | Part 15 |
| Step-back | Replacing a specific question with a more general one and retrieving for the abstraction. | Part 17 |
| Structured notes | A durable, compact store the agent writes facts into so compaction cannot silently drop them. | Part 10 |
| T | ||
| Temperature | The sampling divisor on logits; higher flattens the distribution, lower sharpens it, and 0 is not deterministic. | Part 3 |
| Token | The unit a model reads and writes, produced by a tokenizer such as BPE, and not the same as a word. | Part 2 |
| Top-k | Sampling restricted to the k highest-probability tokens. | Part 3 |
| Top-p | Nucleus sampling: the smallest set of tokens whose cumulative probability exceeds p. | Part 3 |
| TPOT | Time per output token after the first; with TTFT it defines a serving latency contract. | Part 4 |
| TTFT | Time to first token, dominated by queueing and prefill; the responsiveness a streaming user perceives. | Part 4 |
Retrieval-architecture selection table
Which index, for which job, and when not to
Read the table by finding the row whose when not to use describes you: that is usually the decisive column. Flat is the baseline every approximation is measured against; the ANN rows trade recall for latency and memory in different currencies. Hybrid plus rerank is the production default for text because lexical and dense failure modes are complementary, and the last two rows answer question shapes the others structurally cannot.
| Architecture | Best for | Recall / latency / memory | Operational cost | When not to use |
|---|---|---|---|---|
| Flat (exact) | Small collections and the recall baseline every approximation is measured against. | recall 1.00 · latency highest · memory highest | None beyond storage; no tuning. | Millions of vectors under a latency SLO — brute force is the thing you are trying to avoid. |
| IVF | Large collections where a small recall loss is acceptable and memory is not the constraint. | recall 0.90–0.98 · low latency · index + data in memory | Moderate: choose nlist, tune nprobe, retrain when the distribution drifts. | When you cannot tolerate the true neighbour's cell not being probed, or you have no labelled set to tune nprobe against. |
| IVF-PQ | Huge collections that must fit in memory; compression matters more than the last recall points. | recall 0.80–0.95 · low latency · much lower memory | Build plus codebook tuning; quantisation error is a new, measurable regression surface. | When recall is the binding constraint, or the vectors are short and dense enough that PQ error hurts. |
| HNSW | Low-latency, high-recall serving where the graph can live in RAM. | recall 0.95–0.99 · very low latency · high memory (graph) | High memory; M and efSearch tuning; heavier rebuilds as the graph grows or changes. | Memory-bound deployments, or indexes that change constantly and cannot absorb rebuild cost. |
| DiskANN / Vamana | Billions of vectors that must live on SSD on a single node. | recall 0.95–0.99 · low latency · index fits on SSD | SSD IO is the bottleneck; heavier build; careful temperature and cache management. | When RAM is ample and you need the very lowest latency, or the collection is small enough for HNSW. |
| Hybrid + rerank | Production text retrieval where precision matters and lexical and dense disagree usefully. | recall follows the union · rerank pass adds latency · two indexes in memory | Two indexes, an embedding model, and a cross-encoder; rerank is the dominant cost. | Strict sub-10 ms budgets, or corpora small enough that BM25 alone already clears the bar. |
| Graph | Global, thematic questions over a corpus — “what themes run through this?” | not a chunk-level recall metric · latency from summary lookup · index = graph + summaries | An LLM pass over the whole corpus to build and refresh; extraction quality is the ceiling. | Find-the-passage queries, and corpora that change faster than the graph can be rebuilt. |
| Page-image / ColPali | PDFs, slides, tables and forms where layout carries the answer that OCR discards. | recall at the page level · heavier VLM encode latency · large multi-vector index | GPU encoding and much larger storage; a separate index beside the text one. | Plain prose corpora, or strict cost and latency budgets where text extraction is good enough. |
RAG failure modes to metrics
Name the failure, then pick the metric that moves
“The retrieval is bad” is not a diagnosis. Each failure below has a different metric and a different fix, and confusing them is how teams spend a month reranking a system whose real problem was a missing passage. Start from the symptom you actually observe, read across to the metric that makes it visible, and only then choose the fix.
| Failure mode | Metric | Fix |
|---|
Formula and reference card
Four expressions and one chunking cheat
Every interactive demo in the volume is a version of one of these. Bytes, tokens and ranks are all approximate on purpose — accurate to a factor you can reason about is the point.
Cosine similarity
The angle between two vectors, ignoring magnitude. Meaning lives in direction; length is a nuisance the normalisation removes.
BM25
Rare terms count more (IDF); repeats saturate (k1, default 1.2); long documents do not win on length alone (b, default 0.75).
Reciprocal rank fusion
Fuses ranked lists without calibrating their scores. It only needs the ranks, which is why it is the default for lexical + dense.
Chunk-strategy cheat
fixed — cut every N characters: trivial, and it severs sentences across boundaries.
recursive — split on paragraphs and sentences, filling to a size: the sane default.
semantic — split where consecutive sentences stop resembling each other: coherent, costlier to build.
parent — match on the child, return the parent passage: precise matching, generous context.