Few-Shot Prompting
A prompting technique providing multiple input-output examples to guide model generation formatting.
Last reviewed: July 25, 2026
Few-shot prompting is a technique for improving an LLM’s accuracy on a specific task by including a small number of worked examples — typically 2 to 10 — directly in the prompt, demonstrating the exact input-output pattern desired before presenting the actual task. It relies on in-context learning, the ability of language models to recognize and continue a pattern shown within a single prompt, without any update to the model’s weights.
Why It Works Better Than Zero-Shot
A zero-shot prompt describes a task in words alone; a few-shot prompt shows it. For tasks with a specific, non-obvious output format — a particular JSON schema, a specific tone, an unusual classification scheme — showing examples is often far more reliable than describing the format in prose, since the model can directly pattern-match against the demonstrations rather than inferring an ambiguous textual description.
Practical Guidance
The choice, order, and diversity of examples measurably affects output quality: examples should ideally cover the range of cases the model will actually encounter, including any tricky edge cases, and should be formatted exactly as the desired output should look, since models tend to closely mimic the formatting conventions shown in the examples (spacing, capitalization, delimiter choices) as much as the underlying logic.
The Cost Tradeoff
Each additional example adds tokens to every request using that prompt, increasing both latency and per-request cost. In production systems processing high request volumes, teams often start with few-shot prompting during development to establish a reliable pattern, then invest in fine-tuning on a larger example set once the pattern is well understood — trading a one-time training cost for lower per-request token usage.
Few-Shot Learning at Scale: In-Context vs. Fine-Tuning
Few-shot prompting is sometimes framed as a lightweight alternative to fine-tuning, but the two aren’t interchangeable at scale: few-shot examples consume context window space and add token cost on every single request, while fine-tuning bakes the desired pattern into the model’s weights once, at a one-time training cost, after which every subsequent inference call is no larger or more expensive than a zero-shot prompt. This is why teams often start with few-shot prompting to validate that a pattern works and is worth investing in, then migrate to fine-tuning once request volume is high enough that the per-request token savings outweigh the upfront cost and complexity of running a fine-tuning job.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.