Glossary and decision tables
This is the volume's back matter: one alphabetical index that links every term to the part that introduces it, a decision table for the prompt-versus-RAG-versus-tuning question, a failure-mode-to-metric table you can check a real agent against, an injection-defense-pattern matrix, and a dated benchmark reference card. It is the applications counterpart to the companion volume Building with LLMs, Interactively and its reference page, and to the training glossary and the serving glossary, which define the modelling and deployment layers. Terms that belong to those volumes link there rather than re-teaching them; terms that belong to this volume link to the part where they first appear.
Glossary
Every term, linked to the part that introduces it
| Term | What it means | Introduced in |
|---|---|---|
| A | ||
| A/B power | The sample size a comparison needs to see an effect; an underpowered test reads noise as a win. | Part 12 |
| A2A | A protocol for agent-to-agent capability advertisement and task delegation. | Part 19 |
| Adapter serving | Serving one frozen base model with many small LoRA adapters loaded per request. | Part 14 |
| Agent | A model in a loop with tools, deciding its own next step until the task is done. | Part 1 |
| AgentDojo | A dynamic attack/defense benchmark for prompt injection in tool-using agents. | Part 15 |
| AP2 | A protocol for agent-initiated payment mandates and authorisation. | Part 19 |
| Axial coding | Grouping open codes into categories and relationships to form a failure taxonomy. | Part 8 |
| B | ||
| Backoff | Retrying a failed call after a delay, the delay growing so a struggling dependency is not hammered. | Part 17 |
| C | ||
| CaMeL | Capability-based information-flow control that tracks which values may influence which actions. | Part 16 |
| Canary | Releasing a change to a small slice of traffic before it reaches everyone. | Part 18 |
| Capability information flow | Security that decides by data provenance, not by text, which values may reach which actions. | Part 16 |
| Circuit breaker | Stopping calls to a failing dependency so the failure does not cascade. | Part 17 |
| Code execution | Letting the model write and run code as a tool, so computation happens outside the context window. | Part 4 |
| Cohen's kappa | Agreement between two annotators corrected for chance; the check before trusting a judge. | Part 10 |
| Compaction | Summarising or evicting history to fit the window; lossy, since an evicted fact is indistinguishable from one that never existed. | Part 5 |
| Complexity/value/viability/cost-of-error gate | The four questions that decide whether a task justifies an agent at all. | Part 1 |
| Computer use | Driving a GUI by screenshots and input events when no API exists. | Part 7 |
| Context editing | Deliberately removing or rewriting specific context entries rather than summarising wholesale. | Part 5 |
| Context folding | Collapsing completed sub-trajectories into a compact form as a long task proceeds. | Part 7 |
| Context precision | The fraction of retrieved passages that are relevant. | Part 11 |
| Context recall | Whether the retrieved set contained the evidence the answer needed. | Part 11 |
| Context window as a data bus (anti-pattern) | Passing data between steps through the context instead of through code, so every hop pays the tokens again. | Part 4 |
| Cost per completed task | Total spend divided by tasks that succeeded; failures are paid for too. | Part 11 |
| Coupling | How much agents must share live state; high coupling is what makes multi-agent systems fail. | Part 6 |
| Coverage matrix | Tasks crossed with failure modes, to see what the eval set does and does not exercise. | Part 9 |
| D | ||
| Deep research | The composite long-horizon case: many searches, reads and syntheses toward one report. | Part 7 |
| Degraded mode | A deliberate fallback — cached or canned — when a dependency is unavailable. | Part 17 |
| Distillation | Training a small student to reproduce a large teacher's outputs, transferring competence on a task. | Part 14 |
| Diversity collapse | Synthetic data or RL narrowing a model's outputs toward a single mode. | Part 14 |
| DPO | Direct preference optimization: tuning on preference pairs without a separate reward model. | Part 14 |
| Dual-LLM | A quarantined model reads untrusted content and may only emit structured data to the privileged model. | Part 16 |
| E | ||
| End-state eval | Grading whether the world was left correct, independent of the path taken. | Part 11 |
| Episodic memory | Memory of specific past events and turns. | Part 5 |
| Error analysis | Reading real traces to find failure modes, before choosing fixes. | Part 8 |
| Errors as data | Returning a tool's failure to the model as an ordinary result it can act on. | Part 2 |
| Evaluator-optimiser | A workflow where one call generates and another critiques, iterating to a bar. | Part 1 |
| External notes | A durable, compact store the agent writes facts into so compaction cannot silently drop them. | Part 5 |
| F | ||
| Failure taxonomy | The named, ordered list of the ways a system fails, built from traces. | Part 8 |
| Faithfulness | Whether the answer stayed inside the retrieved evidence. | Part 11 |
| Fine-tuning | Updating model weights on your data; reach for it only when prompt and RAG have saturated. | Part 13 |
| G | ||
| GAIA | A benchmark of multi-step assistant tasks needing web, files and reasoning. | Part 11 |
| GCG | Adversarial-suffix search that finds inputs making a model comply. | Part 15 |
| Golden set | The curated, trusted examples a system is measured against. | Part 9 |
| GRPO | Group-relative policy optimization: RL from verifiable rewards without a value model. | Part 14 |
| H | ||
| Harness | The loop, tools, budgets and checks around a model; it contributes as much variance as the model. | Part 7 |
| Held-out set | Examples kept out of tuning so a reported score is honest. | Part 9 |
| I | ||
| IAA (inter-annotator agreement) | How much two humans agree; the ceiling a judge can be trusted to reach. | Part 10 |
| Idempotency | Designing a tool so repeating a call has no additional effect, making retries safe. | Part 2 |
| Implicit signal | Feedback inferred from user behaviour — accept, edit, abandon — rather than an explicit rating. | Part 12 |
| Indirect injection | An instruction hidden in retrieved content rather than typed by the user. | Part 15 |
| K | ||
| KTO | Preference tuning from binary good/bad labels instead of pairs. | Part 14 |
| L | ||
| LATS | Language-agent tree search: Monte Carlo tree search over reasoning and acting. | Part 3 |
| Lethal trifecta | Private data, untrusted content and an exfiltration channel in one context. | Part 15 |
| LLM-as-judge | A model grading outputs against a rubric; fast, biased, and only as good as its validation. | Part 10 |
| Long-horizon | Tasks long enough that the harness, not the prompt, decides success. | Part 7 |
| Loop detection | Noticing and breaking a repeated identical action. | Part 7 |
| LoRA | Low-rank adapters that adapt a frozen base model cheaply. | Part 14 |
| M | ||
| MCP | The Model Context Protocol: how a model discovers and calls tools and resources. | Part 4 |
| Mem0 | A memory layer that extracts and stores durable facts across sessions. | Part 5 |
| MemGPT/Letta | An agent that manages its own memory tiers, paging context in and out. | Part 5 |
| Model context protocol | The full name of MCP; see the entry above. | Part 4 |
| Multi-agent | Several agents cooperating or delegating; powerful when read-heavy, dangerous when state is shared. | Part 6 |
| O | ||
| Open coding | Labelling raw trace observations with short descriptive codes. | Part 8 |
| Order-swap averaging | Running a pairwise comparison in both orders and averaging, to cancel position bias. | Part 10 |
| Orchestrator-worker | A lead agent that decomposes a task and delegates subtasks to workers. | Part 1 |
| Orchestrator-worker tradeoff | Delegation buys parallelism and costs coordination, duplicated context and latency. | Part 6 |
| ORPO | Odds-ratio preference optimization: SFT and preference in one stage. | Part 14 |
| OSWorld | A benchmark for real computer use across applications. | Part 11 |
| OTel GenAI | The OpenTelemetry conventions for model and agent spans. | Part 12 |
| P | ||
| PAIR | A prompt-level adversarial attack that iteratively refines a jailbreak. | Part 15 |
| Pairwise | Grading by comparing two answers instead of scoring one in isolation. | Part 10 |
| Parallel tool calls | Issuing independent tool calls in one turn rather than serially. | Part 2 |
| Parallelisation | A workflow pattern that runs independent subtasks at once and joins them. | Part 1 |
| Pareto | Ranking failure modes by frequency so you fix the few that dominate. | Part 8 |
| Pass@k | The fraction of tasks solved on at least one of k attempts; a capability measure. | Part 11 |
| pass^k | The fraction of tasks solved on all k attempts; the reliability counterpart. | Part 11 |
| Permission gate | Requiring explicit approval before a consequential action. | Part 7 |
| PII redaction | Removing or masking personal data before it is logged or sent. | Part 16 |
| Plan-and-execute | Plan the whole task up front, then execute the plan step by step. | Part 3 |
| Pointwise | Grading a single answer against a rubric, without a comparison. | Part 10 |
| Position bias | A judge preferring whichever answer is shown first. | Part 10 |
| Prefix cache | Reusing a byte-identical prompt prefix at a cheaper billed rate. | Part 17 |
| Procedural memory | Remembered how-to knowledge: skills and routines. | Part 5 |
| Progressive disclosure | Loading skill metadata always, instructions on invocation, resources on demand. | Part 4 |
| Prompt chaining | A workflow that feeds one call's output into the next as a fixed sequence. | Part 1 |
| Prompt injection | Untrusted text carrying instructions that divert an agent with tools. | Part 15 |
| Prompt version | The identifier for a frozen prompt, so a score change is attributable to a change you made. | Part 12 |
| Q | ||
| QLoRA | LoRA on a 4-bit quantised base; the common single-GPU path. | Part 14 |
| R | ||
| Rails | Input and output classifiers that flag known payloads and exfiltration channels. | Part 16 |
| ReAct | Interleaving reasoning and acting, one thought and one tool call at a time. | Part 3 |
| Red-teaming | Adversarially attacking your own system to find ways past the defenses. | Part 15 |
| Reflexion | Retrying after a failure with a self-written reflection on what went wrong. | Part 3 |
| Regression gate | Blocking a change that lowers a measured score on the labelled set. | Part 9 |
| Rejection sampling | Generating candidates and keeping only those that pass a verifier. | Part 14 |
| Result cache | Caching an application result keyed on the exact request. | Part 17 |
| Reward hacking | A policy exploiting the verifier instead of solving the task. | Part 14 |
| Risk tier | Classifying actions by reversibility and blast radius to decide how much gate they need. | Part 18 |
| RLVR | Reinforcement learning from verifiable rewards, where correctness is programmatically checkable. | Part 14 |
| Rollback | Returning to the previous known-good version. | Part 18 |
| Routing | A workflow pattern that sends an input to one of several specialised paths. | Part 1 |
| Routing cascade | Trying the cheap model first and escalating the hard cases. | Part 17 |
| Rubric | The explicit criteria a judge scores against, usually derived from the failure taxonomy. | Part 10 |
| S | ||
| Sandboxing | Confining the agent's execution to an isolated environment with bounded resources. | Part 7 |
| Self-preference bias | A judge favouring text it, or its own model family, produced. | Part 10 |
| Semantic cache | Caching by meaning rather than exact text; fast, and able to return a wrong answer. | Part 17 |
| Semantic memory | Stored facts about the world. | Part 5 |
| SKILL.md | The file that defines a skill, loaded in three stages. | Part 4 |
| Skills | Packaged instructions and resources a model loads when a task needs them. | Part 4 |
| Span | One timed operation in a trace. | Part 12 |
| Step budget | An explicit cap on the steps an agent may take before it must stop. | Part 7 |
| Streaming | Emitting output tokens as they are produced so the user sees progress. | Part 17 |
| SWE-bench Verified | A human-validated subset for resolving real GitHub issues. | Part 11 |
| Synthetic data | Training examples generated by a model; powerful, and prone to diversity collapse. | Part 14 |
| T | ||
| τ²-bench | A tool-agent-user dialogue benchmark graded on a verifiable database end state. | Part 11 |
| Temporal graph | A memory that records facts with validity times, so stale facts can be superseded. | Part 5 |
| Tool schema | The typed description of a tool's name, arguments and purpose that the model sees. | Part 2 |
| Tool-set overlap | Two tools whose purposes collide, which makes the choice unreliable. | Part 2 |
| Tool use | Letting the model call functions and read their results. | Part 2 |
| tool_result | The message carrying a tool's output back to the model. | Part 2 |
| tool_use | The model's structured request to call a tool with arguments. | Part 2 |
| Trace | The recorded tree of spans for one run. | Part 12 |
| Trajectory eval | Grading the path a run took: steps, cost, permissions, loops. | Part 11 |
| Tree of thoughts | Exploring several reasoning branches and choosing among them. | Part 3 |
| Turn detection | Deciding when a speaker has finished; the hard, variable term in voice latency. | Part 19 |
| V | ||
| Verbosity bias | A judge preferring the longer answer regardless of quality. | Part 10 |
| W | ||
| Workflow | A fixed, readable pipeline of model calls; prefer it to an agent when the steps are known. | Part 1 |
| Working memory | The current context: what the model can see right now. | Part 5 |
| X | ||
| x402 | An HTTP-native payment handshake for agents. | Part 19 |
Prompt → RAG → tune
Reach for the cheapest lever that solves the symptom
Most teams jump to the most expensive intervention — fine-tuning — for a problem that a sentence in the system prompt or a better retriever would have fixed. The order is deliberate: prompt first because it is free and instantly reversible, RAG second because facts live outside the weights, tuning last because it is the only lever that bakes a behaviour into the model and the only one that cannot be rolled back with a text edit. Read the table by finding the symptom you actually observe, then take the first fix; only escalate when that fix has been measured and found insufficient.
| Symptom | First fix | Why not the next lever yet |
|---|---|---|
| Outputs are inconsistent or the wrong shape | Sharpen the prompt: clearer instructions, delimiters, and a JSON schema via constrained decoding. | Format responds to specification and iterates in seconds. A weights change is slower to try and impossible to read. |
| The model does not know a fact, or the fact changes | Add retrieval with citations, and ground the answer in the retrieved evidence. | Facts belong outside the weights. Retrieval is fresh, auditable and updatable without a training run. |
| Retrieval returns the wrong passages | Fix retrieval first: hybrid search, reranking, better chunking. | This looks like a generation failure but is not. Prompt and tuning cannot recover evidence that was never retrieved. |
| The right evidence was retrieved and then ignored | Reorder the context, retrieve fewer and better chunks, and require citations. | The model needs to be told to stay inside the evidence; measure faithfulness, not knowledge. |
| The facts are right but the voice or format never lands | Add few-shot examples; only then consider tuning. | A pattern repeated at scale is what tuning compresses into weights. Until examples have saturated, do not. |
| Latency or cost is too high | Shorten the output, cache the prefix, and route easy traffic to a smaller model. | Output tokens dominate cost, and most traffic is easy. A routing cascade beats a bigger training bill. |
| A capability is missing (a JSON dialect, a tool-call convention) | Fine-tune with LoRA or QLoRA on labelled examples. | A narrow, well-specified behaviour is exactly what a weight update buys, and adapters keep it reversible. |
| Fine on the demo, fails in production | Build the eval set and do error analysis before changing any lever. | You cannot tune, route or roll back what you cannot measure. The labelled set and traces are the prerequisite, not the follow-up. |
Failure mode to metric
Name the failure, then pick the metric that moves
“The agent is bad” is not a diagnosis. Each failure below has a different metric and a different fix, and confusing them is how a team spends a sprint reranking a system whose real problem was a missing passage — or tuning a model when the agent was never calling the tool at all. The first five rows are the retrieval-augmented pipeline; the last three are failures that only exist once the system has tools and acts on the world, and only a trajectory eval sees them.
| Failure mode | Metric | Fix |
|---|
Injection-defense patterns
Six defenses against eight attack variants
Prompt injection is not a bug you patch; it is a class of attacks you trade off against. The matrix below is the honest picture: no single defense stops every variant, and the strongest one — capability information flow — stops seven of eight because it decides by data provenance rather than by scanning text. The variant it does not stop, a direct override from the user, is not really an injection at all: a user who authorises an action is not an attacker, and capability control governs untrusted influence, not authorised intent. The realistic posture is defense in depth plus the one structural control that actually breaks the attack: least privilege, which removes the exfiltration channel and keeps private data out of the same context as untrusted content.
Rows are the six defenses, columns the eight attack variants: override, poisoned document, encoded payload, tool-result, image exfiltration, outbound HTTP, multi-turn sleeper, homoglyph. A filled cell is a variant the defense stops; a pale cell is one it does not.
Benchmark reference card
What each benchmark measures, how it grades, and where it misleads
Scores are deliberately omitted: they change weekly, and a number without a date is a liability. What stays useful is the shape of each benchmark — what it measures, how it is graded, and its characteristic weakness as a signal. A public score is a shortlist, never a measurement of your task; leakage, scaffold, environment drift and grader design all break the transfer. Fill in the live numbers yourself, with the date you read them.
| Benchmark | Measures | Graded | Main caveat |
|---|