One of the most surprising discoveries of the GPT-3 era was in-context learning (ICL): give a large language model a few input–output examples in its prompt, and it performs the task on a new input — with no gradient updates, no fine-tuning, no change to its weights. The "learning" happens entirely during the forward pass. ICL made LLMs general-purpose tools and raised deep scientific questions about what these models compute.
The phenomenon#
English: cheese → French: fromage
English: bread → French: pain
English: water → French:The model continues with "eau". Performance generally improves from zero-shot (instruction only) to one-shot to few-shot, and larger models benefit more from examples. ICL works for classification, translation, extraction, format conversion, simple algorithms and style imitation.
What influences ICL?#
Empirical studies reveal surprising sensitivities:
- Format matters a lot: consistent input–output templates help.
- Label correctness matters less than expected (in some settings): Min et al. (2022) found that replacing demonstration labels with random labels hurt performance only modestly for many classification tasks, suggesting demonstrations often teach the label space, input distribution and format more than the exact mapping. Larger models, however, rely more on the actual mappings — Wei et al. (2023) showed large models can even follow flipped labels in context, overriding prior knowledge.
- Example order and selection: results vary with the order of examples; selecting demonstrations similar to the query (via embedding retrieval) often helps.
- Label balance and recency: models can be biased towards labels that appear frequently or last in the prompt; calibration methods correct for this.
- Number of examples: more examples help up to a point; long-context models enable "many-shot" ICL with hundreds of examples, which can approach fine-tuning performance for some tasks.
How does it work? Emerging theories#
Induction heads#
Mechanistic-interpretability studies (Olsson et al., 2022) identified induction heads: pairs of attention heads that implement "find a previous occurrence of the current token, and copy the token that followed it" — a pattern-completion mechanism: [A][B] … [A] → [B]. Induction heads form abruptly during training, coinciding with a sharp improvement in in-context learning ability, and are argued to underlie a significant part of ICL in small models.
Implicit gradient descent#
Several papers showed that transformer layers can implement steps of gradient descent on a linear regression problem defined by in-context examples (von Oswald et al., 2023; Akyürek et al., 2023), and that trained transformers' in-context predictions resemble those of learning algorithms like least squares. This frames ICL as meta-learning: pretraining on a vast variety of tasks produces a model that learns algorithms for learning.
Bayesian task inference#
Xie et al. (2022) proposed viewing ICL as implicit Bayesian inference over latent concepts: pretraining documents share latent "tasks", and demonstrations help the model infer which task is being asked for.
These views are complementary and partly supported; ICL in large models likely involves several mechanisms.
A small experiment#
from transformers import pipeline
gen = pipeline("text-generation", model="Qwen/Qwen2.5-0.5B-Instruct")
def few_shot(examples, query):
prompt = "".join(f"Review: {x}\nSentiment: {y}\n\n" for x, y in examples)
return prompt + f"Review: {query}\nSentiment:"
examples = [("The staff were kind and patient.", "positive"),
("I waited all day and nobody helped.", "negative"),
("Clean facilities and clear information.", "positive")]
q = "The forms were confusing and the office closed early."
out = gen(few_shot(examples, q), max_new_tokens=3, do_sample=False)[0]["generated_text"]
print(out.split("Sentiment:")[-1].strip())Try shuffling example order, flipping labels, or using unrelated label words ("foo"/"bar") to see how behaviour changes.
ICL vs fine-tuning#
| In-context learning | Fine-tuning | |
|---|---|---|
| Weight updates | None | Yes |
| Data needed | A few to hundreds of examples | Hundreds to thousands+ |
| Setup time | Seconds | Hours |
| Cost per query | Higher (long prompts) | Lower |
| Consistency | Sensitive to prompt details | More consistent |
| Knowledge persistence | Only within the prompt | Stored in weights |
A practical path: prototype with ICL, collect data from use, and fine-tune (or distil) when volume, consistency or cost requires it.