Few-shot: examples as specification
A description says what you want; an example shows it, and the model is far better at copying than at inferring. That is the whole power of few-shot prompting, and its whole hazard: everything you put in the example block becomes specification — the format, the label distribution, the tone, and every typo. This part covers selection and ordering, recency bias, format contagion, and the point where more examples stop helping and start costing.
Zero-shot, few-shot, and what examples specify
Show, do not only tell
Zero-shot prompting gives the model the instruction alone: "classify the intent of this message as refund, shipping, account or other." Few-shot prompting adds worked examples — labelled inputs the model can pattern-match against. The difference is not that examples teach the task from scratch; the model already understands the task. It is that examples specify the parts description is bad at: the exact output format, the treatment of ambiguous edge cases, the level of detail, and where the boundary between two similar labels falls.
An instruction says "be concise"; an example is concise, and the model will match it. That makes the example block the highest-leverage part of the prompt — and the most dangerous, because everything in it is treated as specification whether you meant it or not.
The builder below assembles a prompt from a small labelled pool. Choose how many examples to include, and whether the labels are balanced and the format consistent; the mock accuracy responds the way a real eval set would.
The assembled example block (each chip is one example, coloured by label), with the seeded mock accuracy and the prompt token cost below.
Ordering and recency
The last example is copied hardest
Examples are not a set; they are a sequence, and position carries weight. The strongest and best-documented effect is recency: the example nearest the query has the largest influence on the answer, so the format and label of the final example are the ones the model is most likely to imitate. This is the same primacy-and-recency shape that makes the middle of a long context weak, applied inside your example block.
The practical consequences are concrete. Put your canonical, cleanest example last, not first. Do not end on an edge case unless you want that edge case's format generalised. And when a task has several acceptable output formats, expect the ordering to pick one for you. The chart below holds three examples fixed and reorders them; the mock output distribution over four formats follows whichever example sits last.
Seeded mock distribution over output formats for three orderings of the same three examples (each group is one format).
Format contagion
The model copies your mistakes
If the last example is copied hardest, then everything about it is copied: not only the format you intended but the label typo, the inconsistent quoting, the stray capitalisation, and a genuinely wrong label. This is format contagion, and it is the most common way a few-shot prompt fails quietly. A single malformed example in a block of five can move the whole output distribution, because the model is doing its job — matching the pattern you gave it, noise included.
Three artefacts are worth checking for before you ship an example block: a typo in a label (POSITVE for positive), an inconsistent format (one JSON example among plain-text ones), and a wrong label on an example whose input is similar to the query. Toggle each below and watch the mock outputs copy it.
The example block at the top, three mock outputs below. The copied artefact is highlighted wherever it reappears.
Selection, balance and k-shot saturation
Which examples, how many, and when more hurts
Selection is a search problem with a small budget. Given a pool of labelled examples and room for k of them, the useful heuristics are all about coverage rather than quantity: pick examples that cover the decision boundary (the ambiguous cases), keep the labels balanced so the block does not imply a prior the task does not have, and prefer examples that pin down the output contract. A retrieval step that selects examples per query — the nearest labelled cases to this input — is usually better than a fixed block, for the same reason a reranker beats a static top-five.
How many is a curve, not a number. Accuracy rises steeply for the first few examples, flattens once the format and the boundary are pinned down, and then declines as the block bloats the context, dilutes attention and starts to overfit to the examples' incidental details. The knee is usually lower than people expect — often three to eight examples — and the decline is real enough to place with an eval.
Seeded mock accuracy against k. The curve plateaus once the format is pinned, then falls as context bloat outweighs the marginal example.
Specification, not training
Prompting examples and training data are different artefacts
It is worth separating two things that look alike. Examples in a prompt are a specification: they are read on every call, they shape this one answer, and they cost input tokens forever. Examples in a fine-tuning set are training data: they are read once, their influence is baked into weights, and they change what the model does without occupying the context at all. Same labelled inputs, opposite economics.
The decision rule follows from that. If the behaviour is a format, a label scheme or a boundary the model can be shown at inference time, put it in the prompt — it is free to change, inspectable and versionable. If you are spending a thousand tokens on examples that never change and the task is stable, that is a signal to consider tuning instead, where the same examples cost context once. The trade-off, and the reasons "just fine-tune it" is usually the wrong first move, are the subject of Agents in Action, Part 13 and Part 14.
Cheat sheet
| Question | The answer that shapes everything |
|---|---|
| Why do examples work? | The model copies patterns; an example specifies format and edge cases better than prose. |
| Which example matters most? | The last one — recency bias makes it the strongest template. |
| What should go last? | Your canonical, cleanest example. Never end on an edge case by accident. |
| What is format contagion? | Typos, inconsistent formatting and wrong labels are copied along with the intended pattern. |
| How do I choose examples? | Cover the boundary, balance the labels, pin the output contract. |
| How many examples? | Three to eight typically; find the knee with an eval, not by vibes. |
| Why does accuracy fall at large k? | Context bloat and overfitting to the examples' incidental details. |
| Few-shot vs fine-tuning? | Prompt examples are a per-call specification; a training set bakes the same labels into weights. |
Further reading
- Brown et al., "Language Models are Few-Shot Learners", NeurIPS 2020 — the GPT-3 paper that established in-context learning and the few-shot framing.
- Min et al., "Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?", EMNLP 2022 — the surprising result that label correctness matters less than the demonstration's format and label space.
- Lu et al., "Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity", ACL 2022 — ordering effects and the variance they introduce.
- Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL 2024 — the primacy-and-recency curve that explains why the last example dominates.
- Rubin, Herzig & Berant, "Learning To Retrieve Prompts for In-Context Learning", NAACL 2022 — per-query example selection instead of a fixed block.