Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Zero-shot, few-shot, and what examples specify

Show, do not only tell

Zero-shot prompting gives the model the instruction alone: "classify the intent of this message as refund, shipping, account or other." Few-shot prompting adds worked examples — labelled inputs the model can pattern-match against. The difference is not that examples teach the task from scratch; the model already understands the task. It is that examples specify the parts description is bad at: the exact output format, the treatment of ambiguous edge cases, the level of detail, and where the boundary between two similar labels falls.

An instruction says "be concise"; an example is concise, and the model will match it. That makes the example block the highest-leverage part of the prompt — and the most dangerous, because everything in it is treated as specification whether you meant it or not.

The builder below assembles a prompt from a small labelled pool. Choose how many examples to include, and whether the labels are balanced and the format consistent; the mock accuracy responds the way a real eval set would.

The assembled example block (each chip is one example, coloured by label), with the seeded mock accuracy and the prompt token cost below.

💡 The durable idea: an example is a specification with a sample attached. Curate the block the way you would curate a schema — every example earns its tokens by pinning down something the description leaves ambiguous.
2

Ordering and recency

The last example is copied hardest

Examples are not a set; they are a sequence, and position carries weight. The strongest and best-documented effect is recency: the example nearest the query has the largest influence on the answer, so the format and label of the final example are the ones the model is most likely to imitate. This is the same primacy-and-recency shape that makes the middle of a long context weak, applied inside your example block.

The practical consequences are concrete. Put your canonical, cleanest example last, not first. Do not end on an edge case unless you want that edge case's format generalised. And when a task has several acceptable output formats, expect the ordering to pick one for you. The chart below holds three examples fixed and reorders them; the mock output distribution over four formats follows whichever example sits last.

Seeded mock distribution over output formats for three orderings of the same three examples (each group is one format).

⚠️ The trap: appending a new example at the end of an existing block. That is exactly where it will be copied from, so a hasty last example silently rewrites the output format of every answer.
3

Format contagion

The model copies your mistakes

If the last example is copied hardest, then everything about it is copied: not only the format you intended but the label typo, the inconsistent quoting, the stray capitalisation, and a genuinely wrong label. This is format contagion, and it is the most common way a few-shot prompt fails quietly. A single malformed example in a block of five can move the whole output distribution, because the model is doing its job — matching the pattern you gave it, noise included.

Three artefacts are worth checking for before you ship an example block: a typo in a label (POSITVE for positive), an inconsistent format (one JSON example among plain-text ones), and a wrong label on an example whose input is similar to the query. Toggle each below and watch the mock outputs copy it.

The example block at the top, three mock outputs below. The copied artefact is highlighted wherever it reappears.

💡 Curation rule: treat the example block as production data. The same proofreading, label review and format linting you would apply to a training set applies here — because to the model, it is one.
4

Selection, balance and k-shot saturation

Which examples, how many, and when more hurts

Selection is a search problem with a small budget. Given a pool of labelled examples and room for k of them, the useful heuristics are all about coverage rather than quantity: pick examples that cover the decision boundary (the ambiguous cases), keep the labels balanced so the block does not imply a prior the task does not have, and prefer examples that pin down the output contract. A retrieval step that selects examples per query — the nearest labelled cases to this input — is usually better than a fixed block, for the same reason a reranker beats a static top-five.

How many is a curve, not a number. Accuracy rises steeply for the first few examples, flattens once the format and the boundary are pinned down, and then declines as the block bloats the context, dilutes attention and starts to overfit to the examples' incidental details. The knee is usually lower than people expect — often three to eight examples — and the decline is real enough to place with an eval.

Seeded mock accuracy against k. The curve plateaus once the format is pinned, then falls as context bloat outweighs the marginal example.

5

Specification, not training

Prompting examples and training data are different artefacts

It is worth separating two things that look alike. Examples in a prompt are a specification: they are read on every call, they shape this one answer, and they cost input tokens forever. Examples in a fine-tuning set are training data: they are read once, their influence is baked into weights, and they change what the model does without occupying the context at all. Same labelled inputs, opposite economics.

The decision rule follows from that. If the behaviour is a format, a label scheme or a boundary the model can be shown at inference time, put it in the prompt — it is free to change, inspectable and versionable. If you are spending a thousand tokens on examples that never change and the task is stable, that is a signal to consider tuning instead, where the same examples cost context once. The trade-off, and the reasons "just fine-tune it" is usually the wrong first move, are the subject of Agents in Action, Part 13 and Part 14.

💡 Carry this forward: everything in this part is a testable claim about a prompt — balance, ordering, k. Each one should move a number on the golden set from Part 5. Few-shot choice is engineering precisely when it is measured, and superstition when it is not.

Cheat sheet

QuestionThe answer that shapes everything
Why do examples work?The model copies patterns; an example specifies format and edge cases better than prose.
Which example matters most?The last one — recency bias makes it the strongest template.
What should go last?Your canonical, cleanest example. Never end on an edge case by accident.
What is format contagion?Typos, inconsistent formatting and wrong labels are copied along with the intended pattern.
How do I choose examples?Cover the boundary, balance the labels, pin the output contract.
How many examples?Three to eight typically; find the knee with an eval, not by vibes.
Why does accuracy fall at large k?Context bloat and overfitting to the examples' incidental details.
Few-shot vs fine-tuning?Prompt examples are a per-call specification; a training set bakes the same labels into weights.

Further reading

6

Check your understanding

0/4 answered