Few-Shot Prompting Examples: How Many to Use and Where
Few-shot prompting examples work best at three to five, in one exact shape. Learn how many to use, where to place them, and copy a reusable few-shot template.
Ask a model to classify support tickets and describe the output in a paragraph, and it mostly complies. Show it three tickets already classified in the exact shape you want, and it locks on. That's few-shot prompting examples doing the work instructions can't: teaching the pattern by demonstration instead of description. The catch is that most people either use the wrong number of examples or vary their format, and both quietly wreck the result.
The pages that rank explain the concept well and then stop short. The few-shot page on promptingguide.ai shows solid examples in about 1,200 words but doesn't prescribe how many to use for which job, and closes on a course pitch. The longer guide on prompthub.us does nail the count and the ordering across roughly 4,500 words, then never compares how the three major models actually behave, and gates the depth behind its own platform. So here's the tight version: how many, where they go, and where each model differs.
What few-shot prompting is
Few-shot prompting is showing the model several worked examples, each an input paired with its correct output, before the real input. The model learns the task from the demonstrations. No retraining, no fine-tuning, just pattern-matching from what's in the context window, which is why it's also called in-context learning.
Here's what teams reach for it to do:
- Lock an output format that a written description keeps failing to enforce
- Classify text into a fixed label set with consistent phrasing
- Teach a niche or made-up convention the model has never seen
- Show the exact tone or structure for rewrites and summaries
- Demonstrate how to handle an edge case (an empty field, an ambiguous input)
- Reduce the instruction paragraph to a couple of lines plus examples
The common thread: the examples carry information that's painful to write as rules. "Match this shape" is three examples. Writing the shape out as prose takes a paragraph the model half-follows.
How many examples, really
Three to five, for most work. That's the honest answer the thin guides dodge and the long ones bury.
The curve is steep then flat. Two examples usually establish the pattern. A third or fourth sharpens it. Past five, accuracy barely moves while token cost keeps climbing on every single call. And beyond about eight, the model can start overfitting, copying incidental details of your examples instead of the underlying rule.
Rough guide by job
Simple classification (fixed labels) -> 2 to 3 examples
Formatting / structure enforcement -> 3 to 5 examples
Nuanced tone or reasoning -> 4 to 5 examples
Edge-case handling -> add 1 example per edge case
The exception worth flagging: reasoning-heavy tasks sometimes do better with fewer, well-chosen examples plus a chain-of-thought instruction than with more examples alone. More isn't the lever people think it is.
Where the examples go, and in what shape
Two placement rules do most of the work, and both come down to how models weight tokens.
First, examples go after the instruction and before the real input. The model reads the rule, sees it demonstrated, then applies it to fresh input right where its attention is strongest. Second, put your most representative example last. Models weight recent context heavily, so the final example carries extra influence. If one example best captures the pattern, that's the one to end on.
Then the rule that matters more than count: keep every example in the identical shape. Same field order, same delimiters, same handling of blanks. Examples that vary in format teach the model to vary its output. That's not a small effect. It's the single most common reason few-shot "doesn't work" when the count looked fine.
How the three models handle few-shot
The technique carries across Claude, ChatGPT, and Gemini. The weighting doesn't. A prompt tuned on one can surprise you on another.
| Model | Few-shot behavior | Where it slips | What to do |
|---|---|---|---|
| Claude | Follows consistent examples reliably; tolerates fewer | Inconsistent formatting across examples confuses it fast | Keep 3 examples in one exact shape |
| ChatGPT (GPT-4o) | Strong, weights the last example heavily | A weak final example drags the pattern | Put the most representative example last |
| Gemini | Good, but can blend examples into the real input | Poorly delimited examples get treated as data | Wrap each example in clear delimiters |
The habit that helps all three: delimit the examples clearly and label the slot for the real input. A block like EXAMPLES: followed by NOW CLASSIFY: stops the model from reading your last demonstration as part of the input to process. Gemini needs this most, but none of the three are hurt by it.
Inconsistent examples are worse than no examples. A written instruction with zero demonstrations gives the model a clean rule to follow. Three examples in three slightly different shapes give it a rule plus evidence that the rule is negotiable, and it takes the hint. If you can't keep your examples identical in structure, cut them down to the one you can get right.
Prompt-craft patterns for few-shot
Include one hard example, not five easy ones
Five examples of the obvious case teach the obvious case. One example of the input that usually breaks, the empty field, the ambiguous label, the mixed format, teaches the model the boundary. A short set that covers the edges beats a long set of duplicates.
Show the failure output too
If a valid response is sometimes "none" or "cannot determine," show an example that returns exactly that. Otherwise the model learns every input deserves a confident answer, and it manufactures one for the cases that should come back empty. Demonstrate the honest non-answer.
Match the example format to your parser
When the output feeds a program, the examples have to be in the precise shape the parser expects, down to whether a list is empty as [] or "none". The model copies what it sees. An example showing [] teaches an empty array; one showing "none" teaches a string. Pick the one your code reads, and show only that.
Variables you'll set
| Variable | Required | What it is |
|---|---|---|
{{examples}} | Yes | The 3 to 5 worked input-output pairs, all in one identical shape |
{{input_text}} | Yes | The real input to process this run, placed after the examples |
{{label_set}} | No | The fixed set of allowed outputs, when the job is classification |
{{output_format}} | No | The exact shape each example and the answer must take, if not shown inline |
The {{examples}} slot is where reliability is won or lost. Fill it with three consistent pairs and the format locks. Fill it with five inconsistent ones and you've taught the model to improvise.
Getting started
- Write the instruction first, in one or two lines, so the examples support a rule.
- Pick three to five real input-output pairs, including at least one hard case.
- Rewrite every example into one identical shape, matching what your parser reads.
- Order them so the most representative pair comes last.
- Delimit the example block and label the slot for the real input.
- Test on inputs the examples didn't cover, not just ones that echo them.
- If accuracy stalls, fix the examples' consistency before adding more of them.
For a version where the examples live in a stable block that gets reused across runs, the Context Engineering Harness pack is built around exactly that layout.
Browse the prompt catalog →The Context Engineering Harness Prompt Pack keeps your few-shot examples in a frozen block with an {{examples}} slot and a locked output contract, so the demonstrations stay in one shape and the format holds whether you run it on Claude, ChatGPT, or Gemini. It's part of The Complete AI Prompts Bundle, a one-time lifetime license to the whole catalog plus every pack added later, worth it if you run more than one few-shot job. When the examples teach a project scaffold instead, the Golden Path Scaffolding System Prompt applies the same idea to code.
Examples are half of a reliable prompt; the frame around them is the other half. For assembling that frame so the examples don't get buried, context engineering for LLMs covers the ordering. And when the few-shot output has to be machine-readable, a structured output prompt for reliable JSON shows how the examples and the contract lock the shape together. Pair either with the Context Engineering Harness when the same examples serve every run.
Common questions
How many examples should a few-shot prompt have?
What is few-shot prompting?
Do few-shot examples work the same on Claude, ChatGPT, and Gemini?
Get the prompt packs this guide is built on
Ready-to-paste prompts with documented variables and usage guides for ChatGPT, Claude, and Gemini. One-time payment, own it forever.
More prompt guides

How to Write a System Prompt That Holds in Production
A model that keeps going off-script usually doesn't have a model problem. It has a system prompt problem. The instruction that was supposed to frame every answer is either missing, vague, or buried in…

Reusable Prompt Templates: Build Variables That Hold Up
Anyone who uses ChatGPT for real work ends up with a folder of near-identical prompts. The cold email one. The cold email one but for enterprise. The cold email one for enterprise, but shorter. Each i…

Context Engineering for LLMs: Assemble the Prompt Right
A capable model gets a question wrong. The instinct is to reword the prompt. Usually that's the wrong lever. Context engineering for LLMs starts from a different premise: the model probably could have…