Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Glossary

Every term, linked to the part that introduces it

💡 Filter the list. Type any fragment — a term, a technique, a failure mode — and the table hides non-matching rows and reports how many remain. Matching is case-insensitive and looks at both the term and its definition. The letter headings hide themselves when nothing under them matches.
TermWhat it meansIntroduced in
A
A/B powerThe sample size a comparison needs to see an effect; an underpowered test reads noise as a win.Part 12
A2AA protocol for agent-to-agent capability advertisement and task delegation.Part 19
Adapter servingServing one frozen base model with many small LoRA adapters loaded per request.Part 14
AgentA model in a loop with tools, deciding its own next step until the task is done.Part 1
AgentDojoA dynamic attack/defense benchmark for prompt injection in tool-using agents.Part 15
AP2A protocol for agent-initiated payment mandates and authorisation.Part 19
Axial codingGrouping open codes into categories and relationships to form a failure taxonomy.Part 8
B
BackoffRetrying a failed call after a delay, the delay growing so a struggling dependency is not hammered.Part 17
C
CaMeLCapability-based information-flow control that tracks which values may influence which actions.Part 16
CanaryReleasing a change to a small slice of traffic before it reaches everyone.Part 18
Capability information flowSecurity that decides by data provenance, not by text, which values may reach which actions.Part 16
Circuit breakerStopping calls to a failing dependency so the failure does not cascade.Part 17
Code executionLetting the model write and run code as a tool, so computation happens outside the context window.Part 4
Cohen's kappaAgreement between two annotators corrected for chance; the check before trusting a judge.Part 10
CompactionSummarising or evicting history to fit the window; lossy, since an evicted fact is indistinguishable from one that never existed.Part 5
Complexity/value/viability/cost-of-error gateThe four questions that decide whether a task justifies an agent at all.Part 1
Computer useDriving a GUI by screenshots and input events when no API exists.Part 7
Context editingDeliberately removing or rewriting specific context entries rather than summarising wholesale.Part 5
Context foldingCollapsing completed sub-trajectories into a compact form as a long task proceeds.Part 7
Context precisionThe fraction of retrieved passages that are relevant.Part 11
Context recallWhether the retrieved set contained the evidence the answer needed.Part 11
Context window as a data bus (anti-pattern)Passing data between steps through the context instead of through code, so every hop pays the tokens again.Part 4
Cost per completed taskTotal spend divided by tasks that succeeded; failures are paid for too.Part 11
CouplingHow much agents must share live state; high coupling is what makes multi-agent systems fail.Part 6
Coverage matrixTasks crossed with failure modes, to see what the eval set does and does not exercise.Part 9
D
Deep researchThe composite long-horizon case: many searches, reads and syntheses toward one report.Part 7
Degraded modeA deliberate fallback — cached or canned — when a dependency is unavailable.Part 17
DistillationTraining a small student to reproduce a large teacher's outputs, transferring competence on a task.Part 14
Diversity collapseSynthetic data or RL narrowing a model's outputs toward a single mode.Part 14
DPODirect preference optimization: tuning on preference pairs without a separate reward model.Part 14
Dual-LLMA quarantined model reads untrusted content and may only emit structured data to the privileged model.Part 16
E
End-state evalGrading whether the world was left correct, independent of the path taken.Part 11
Episodic memoryMemory of specific past events and turns.Part 5
Error analysisReading real traces to find failure modes, before choosing fixes.Part 8
Errors as dataReturning a tool's failure to the model as an ordinary result it can act on.Part 2
Evaluator-optimiserA workflow where one call generates and another critiques, iterating to a bar.Part 1
External notesA durable, compact store the agent writes facts into so compaction cannot silently drop them.Part 5
F
Failure taxonomyThe named, ordered list of the ways a system fails, built from traces.Part 8
FaithfulnessWhether the answer stayed inside the retrieved evidence.Part 11
Fine-tuningUpdating model weights on your data; reach for it only when prompt and RAG have saturated.Part 13
G
GAIAA benchmark of multi-step assistant tasks needing web, files and reasoning.Part 11
GCGAdversarial-suffix search that finds inputs making a model comply.Part 15
Golden setThe curated, trusted examples a system is measured against.Part 9
GRPOGroup-relative policy optimization: RL from verifiable rewards without a value model.Part 14
H
HarnessThe loop, tools, budgets and checks around a model; it contributes as much variance as the model.Part 7
Held-out setExamples kept out of tuning so a reported score is honest.Part 9
I
IAA (inter-annotator agreement)How much two humans agree; the ceiling a judge can be trusted to reach.Part 10
IdempotencyDesigning a tool so repeating a call has no additional effect, making retries safe.Part 2
Implicit signalFeedback inferred from user behaviour — accept, edit, abandon — rather than an explicit rating.Part 12
Indirect injectionAn instruction hidden in retrieved content rather than typed by the user.Part 15
K
KTOPreference tuning from binary good/bad labels instead of pairs.Part 14
L
LATSLanguage-agent tree search: Monte Carlo tree search over reasoning and acting.Part 3
Lethal trifectaPrivate data, untrusted content and an exfiltration channel in one context.Part 15
LLM-as-judgeA model grading outputs against a rubric; fast, biased, and only as good as its validation.Part 10
Long-horizonTasks long enough that the harness, not the prompt, decides success.Part 7
Loop detectionNoticing and breaking a repeated identical action.Part 7
LoRALow-rank adapters that adapt a frozen base model cheaply.Part 14
M
MCPThe Model Context Protocol: how a model discovers and calls tools and resources.Part 4
Mem0A memory layer that extracts and stores durable facts across sessions.Part 5
MemGPT/LettaAn agent that manages its own memory tiers, paging context in and out.Part 5
Model context protocolThe full name of MCP; see the entry above.Part 4
Multi-agentSeveral agents cooperating or delegating; powerful when read-heavy, dangerous when state is shared.Part 6
O
Open codingLabelling raw trace observations with short descriptive codes.Part 8
Order-swap averagingRunning a pairwise comparison in both orders and averaging, to cancel position bias.Part 10
Orchestrator-workerA lead agent that decomposes a task and delegates subtasks to workers.Part 1
Orchestrator-worker tradeoffDelegation buys parallelism and costs coordination, duplicated context and latency.Part 6
ORPOOdds-ratio preference optimization: SFT and preference in one stage.Part 14
OSWorldA benchmark for real computer use across applications.Part 11
OTel GenAIThe OpenTelemetry conventions for model and agent spans.Part 12
P
PAIRA prompt-level adversarial attack that iteratively refines a jailbreak.Part 15
PairwiseGrading by comparing two answers instead of scoring one in isolation.Part 10
Parallel tool callsIssuing independent tool calls in one turn rather than serially.Part 2
ParallelisationA workflow pattern that runs independent subtasks at once and joins them.Part 1
ParetoRanking failure modes by frequency so you fix the few that dominate.Part 8
Pass@kThe fraction of tasks solved on at least one of k attempts; a capability measure.Part 11
pass^kThe fraction of tasks solved on all k attempts; the reliability counterpart.Part 11
Permission gateRequiring explicit approval before a consequential action.Part 7
PII redactionRemoving or masking personal data before it is logged or sent.Part 16
Plan-and-executePlan the whole task up front, then execute the plan step by step.Part 3
PointwiseGrading a single answer against a rubric, without a comparison.Part 10
Position biasA judge preferring whichever answer is shown first.Part 10
Prefix cacheReusing a byte-identical prompt prefix at a cheaper billed rate.Part 17
Procedural memoryRemembered how-to knowledge: skills and routines.Part 5
Progressive disclosureLoading skill metadata always, instructions on invocation, resources on demand.Part 4
Prompt chainingA workflow that feeds one call's output into the next as a fixed sequence.Part 1
Prompt injectionUntrusted text carrying instructions that divert an agent with tools.Part 15
Prompt versionThe identifier for a frozen prompt, so a score change is attributable to a change you made.Part 12
Q
QLoRALoRA on a 4-bit quantised base; the common single-GPU path.Part 14
R
RailsInput and output classifiers that flag known payloads and exfiltration channels.Part 16
ReActInterleaving reasoning and acting, one thought and one tool call at a time.Part 3
Red-teamingAdversarially attacking your own system to find ways past the defenses.Part 15
ReflexionRetrying after a failure with a self-written reflection on what went wrong.Part 3
Regression gateBlocking a change that lowers a measured score on the labelled set.Part 9
Rejection samplingGenerating candidates and keeping only those that pass a verifier.Part 14
Result cacheCaching an application result keyed on the exact request.Part 17
Reward hackingA policy exploiting the verifier instead of solving the task.Part 14
Risk tierClassifying actions by reversibility and blast radius to decide how much gate they need.Part 18
RLVRReinforcement learning from verifiable rewards, where correctness is programmatically checkable.Part 14
RollbackReturning to the previous known-good version.Part 18
RoutingA workflow pattern that sends an input to one of several specialised paths.Part 1
Routing cascadeTrying the cheap model first and escalating the hard cases.Part 17
RubricThe explicit criteria a judge scores against, usually derived from the failure taxonomy.Part 10
S
SandboxingConfining the agent's execution to an isolated environment with bounded resources.Part 7
Self-preference biasA judge favouring text it, or its own model family, produced.Part 10
Semantic cacheCaching by meaning rather than exact text; fast, and able to return a wrong answer.Part 17
Semantic memoryStored facts about the world.Part 5
SKILL.mdThe file that defines a skill, loaded in three stages.Part 4
SkillsPackaged instructions and resources a model loads when a task needs them.Part 4
SpanOne timed operation in a trace.Part 12
Step budgetAn explicit cap on the steps an agent may take before it must stop.Part 7
StreamingEmitting output tokens as they are produced so the user sees progress.Part 17
SWE-bench VerifiedA human-validated subset for resolving real GitHub issues.Part 11
Synthetic dataTraining examples generated by a model; powerful, and prone to diversity collapse.Part 14
T
τ²-benchA tool-agent-user dialogue benchmark graded on a verifiable database end state.Part 11
Temporal graphA memory that records facts with validity times, so stale facts can be superseded.Part 5
Tool schemaThe typed description of a tool's name, arguments and purpose that the model sees.Part 2
Tool-set overlapTwo tools whose purposes collide, which makes the choice unreliable.Part 2
Tool useLetting the model call functions and read their results.Part 2
tool_resultThe message carrying a tool's output back to the model.Part 2
tool_useThe model's structured request to call a tool with arguments.Part 2
TraceThe recorded tree of spans for one run.Part 12
Trajectory evalGrading the path a run took: steps, cost, permissions, loops.Part 11
Tree of thoughtsExploring several reasoning branches and choosing among them.Part 3
Turn detectionDeciding when a speaker has finished; the hard, variable term in voice latency.Part 19
V
Verbosity biasA judge preferring the longer answer regardless of quality.Part 10
W
WorkflowA fixed, readable pipeline of model calls; prefer it to an agent when the steps are known.Part 1
Working memoryThe current context: what the model can see right now.Part 5
X
x402An HTTP-native payment handshake for agents.Part 19
2

Prompt → RAG → tune

Reach for the cheapest lever that solves the symptom

Most teams jump to the most expensive intervention — fine-tuning — for a problem that a sentence in the system prompt or a better retriever would have fixed. The order is deliberate: prompt first because it is free and instantly reversible, RAG second because facts live outside the weights, tuning last because it is the only lever that bakes a behaviour into the model and the only one that cannot be rolled back with a text edit. Read the table by finding the symptom you actually observe, then take the first fix; only escalate when that fix has been measured and found insufficient.

SymptomFirst fixWhy not the next lever yet
Outputs are inconsistent or the wrong shapeSharpen the prompt: clearer instructions, delimiters, and a JSON schema via constrained decoding.Format responds to specification and iterates in seconds. A weights change is slower to try and impossible to read.
The model does not know a fact, or the fact changesAdd retrieval with citations, and ground the answer in the retrieved evidence.Facts belong outside the weights. Retrieval is fresh, auditable and updatable without a training run.
Retrieval returns the wrong passagesFix retrieval first: hybrid search, reranking, better chunking.This looks like a generation failure but is not. Prompt and tuning cannot recover evidence that was never retrieved.
The right evidence was retrieved and then ignoredReorder the context, retrieve fewer and better chunks, and require citations.The model needs to be told to stay inside the evidence; measure faithfulness, not knowledge.
The facts are right but the voice or format never landsAdd few-shot examples; only then consider tuning.A pattern repeated at scale is what tuning compresses into weights. Until examples have saturated, do not.
Latency or cost is too highShorten the output, cache the prefix, and route easy traffic to a smaller model.Output tokens dominate cost, and most traffic is easy. A routing cascade beats a bigger training bill.
A capability is missing (a JSON dialect, a tool-call convention)Fine-tune with LoRA or QLoRA on labelled examples.A narrow, well-specified behaviour is exactly what a weight update buys, and adapters keep it reversible.
Fine on the demo, fails in productionBuild the eval set and do error analysis before changing any lever.You cannot tune, route or roll back what you cannot measure. The labelled set and traces are the prerequisite, not the follow-up.
💡 The order is not a rule, it is a cost gradient. Each step down the table is more expensive, slower to iterate and harder to reverse. Take the first fix that moves the metric, and let the eval set — not intuition — tell you when to escalate.
3

Failure mode to metric

Name the failure, then pick the metric that moves

“The agent is bad” is not a diagnosis. Each failure below has a different metric and a different fix, and confusing them is how a team spends a sprint reranking a system whose real problem was a missing passage — or tuning a model when the agent was never calling the tool at all. The first five rows are the retrieval-augmented pipeline; the last three are failures that only exist once the system has tools and acts on the world, and only a trajectory eval sees them.

Failure modeMetricFix

4

Injection-defense patterns

Six defenses against eight attack variants

Prompt injection is not a bug you patch; it is a class of attacks you trade off against. The matrix below is the honest picture: no single defense stops every variant, and the strongest one — capability information flow — stops seven of eight because it decides by data provenance rather than by scanning text. The variant it does not stop, a direct override from the user, is not really an injection at all: a user who authorises an action is not an attacker, and capability control governs untrusted influence, not authorised intent. The realistic posture is defense in depth plus the one structural control that actually breaks the attack: least privilege, which removes the exfiltration channel and keeps private data out of the same context as untrusted content.

Rows are the six defenses, columns the eight attack variants: override, poisoned document, encoded payload, tool-result, image exfiltration, outbound HTTP, multi-turn sleeper, homoglyph. A filled cell is a variant the defense stops; a pale cell is one it does not.

⚠️ The least-privilege note: the single most effective control is structural, not textual. If private data and untrusted content never share a context, and there is no live exfiltration path, most of this table stops mattering. Prompt hardening and classifiers raise the bar and are worth having as depth; they are not the wall.
5

Benchmark reference card

What each benchmark measures, how it grades, and where it misleads

Scores are deliberately omitted: they change weekly, and a number without a date is a liability. What stays useful is the shape of each benchmark — what it measures, how it is graded, and its characteristic weakness as a signal. A public score is a shortlist, never a measurement of your task; leakage, scaffold, environment drift and grader design all break the transfer. Fill in the live numbers yourself, with the date you read them.

BenchmarkMeasuresGradedMain caveat

💡 Use the card, not your memory. Keep the “measures / graded / caveat” structure in front of you and fill in scores with a date. The structure survives every leaderboard turnover; the scores do not.