Notes
Three examples beat three paragraphs of adjectives
Describing what 'good' means is a specification problem. Showing it is easy — and the examples double as a test harness.
Most people describe what they want. "Make it concise and professional but not stuffy, with a friendly tone that isn't casual."
That sentence contains four judgement calls, and you and the model will resolve every one of them differently.
Now show it instead. Paste two outputs you would have been happy with and one you would not, and label them. The ambiguity disappears — not because the model got cleverer, but because you stopped asking it to guess what you meant by "professional".
How many
Anthropic's published guidance is three to five. Too few will not change the behaviour; too many make the model overfit to the examples you gave it, so it reproduces incidental details of your samples rather than the pattern you meant.
For code specifically, the useful breakdown is narrower — one example may be treated as a single case rather than a pattern, two establish what varies against what stays constant, and three confirm it and show an edge case. Beyond that you are spending context on examples instead of on the task.
Three things that make examples work harder
Include a counter-example. Label one preferred and one less preferred — Google's prompt guidance calls these contrastive pairs. It resolves the boundary instead of describing it. Two good examples tell the model where to aim; a pair tells it which direction is wrong.
Match the format exactly to what you want back. If you want JSON, show JSON. Anthropic's own guidance describes examples as the most reliable way to steer output format — more reliable than describing the format, which is the thing most people try first.
Use the awkward inputs, not the clean ones. Both vendors tell you to cover edge cases, and the reason is simple: examples that only show the easy path teach the easy path. Keep the outputs impeccable, though, because anything sloppy in an example gets copied faithfully.
One scope note, because saying it out loud is the point
The reasoning models now on most desks often need no examples at all. OpenAI's guidance is to write the prompt without them first and add them only when the result is wrong.
So examples are the recovery, not the opening move. Leading with five of them on a task a reasoning model would have got right unaided costs you context, costs you time, and risks constraining an answer that would have been better left open.
The bonus most people miss
The examples you use to steer the model are also the examples you can use to test it.
Change the prompt, run the same set, compare the results. That is an evaluation harness. It is small and informal, but it is the real thing: a fixed set of inputs where you know what good looks like, run against a change to see what moved.
You have built it without meaning to, and it cost you nothing. Most teams that say they have no way to evaluate their prompts are two files away from having one, and the two files are already on their disk.
The part worth remembering
Describing what "good" means is a specification problem, and it has been hard for as long as software has existed. Showing what good means is easy.
Do the easy one.
Where to go next
- Few-shot templates and counter-example patterns — the copyable version, including the contrastive pair
- Few-Shot Examples — the full lesson
- Anatomy of a Good Prompt — the six slots a prompt can fill