✨ Generative AI & LLMs · Lecture 17 of 30

In-Context Learning: How LLMs Learn from Prompts

Large language models can perform new tasks from a few examples in the prompt without weight updates. We examine what in-context learning is, what influences it, theories of how it works, and its practical limits.

One of the most surprising discoveries of the GPT-3 era was in-context learning (ICL): give a large language model a few input–output examples in its prompt, and it performs the task on a new input — with no gradient updates, no fine-tuning, no change to its weights. The "learning" happens entirely during the forward pass. ICL made LLMs general-purpose tools and raised deep scientific questions about what these models compute.

The phenomenon#

text
English: cheese        → French: fromage
English: bread         → French: pain
English: water         → French:

The model continues with "eau". Performance generally improves from zero-shot (instruction only) to one-shot to few-shot, and larger models benefit more from examples. ICL works for classification, translation, extraction, format conversion, simple algorithms and style imitation.

What influences ICL?#

Empirical studies reveal surprising sensitivities:

  • Format matters a lot: consistent input–output templates help.
  • Label correctness matters less than expected (in some settings): Min et al. (2022) found that replacing demonstration labels with random labels hurt performance only modestly for many classification tasks, suggesting demonstrations often teach the label space, input distribution and format more than the exact mapping. Larger models, however, rely more on the actual mappings — Wei et al. (2023) showed large models can even follow flipped labels in context, overriding prior knowledge.
  • Example order and selection: results vary with the order of examples; selecting demonstrations similar to the query (via embedding retrieval) often helps.
  • Label balance and recency: models can be biased towards labels that appear frequently or last in the prompt; calibration methods correct for this.
  • Number of examples: more examples help up to a point; long-context models enable "many-shot" ICL with hundreds of examples, which can approach fine-tuning performance for some tasks.

How does it work? Emerging theories#

Induction heads#

Mechanistic-interpretability studies (Olsson et al., 2022) identified induction heads: pairs of attention heads that implement "find a previous occurrence of the current token, and copy the token that followed it" — a pattern-completion mechanism: [A][B] … [A] → [B]. Induction heads form abruptly during training, coinciding with a sharp improvement in in-context learning ability, and are argued to underlie a significant part of ICL in small models.

Implicit gradient descent#

Several papers showed that transformer layers can implement steps of gradient descent on a linear regression problem defined by in-context examples (von Oswald et al., 2023; Akyürek et al., 2023), and that trained transformers' in-context predictions resemble those of learning algorithms like least squares. This frames ICL as meta-learning: pretraining on a vast variety of tasks produces a model that learns algorithms for learning.

Bayesian task inference#

Xie et al. (2022) proposed viewing ICL as implicit Bayesian inference over latent concepts: pretraining documents share latent "tasks", and demonstrations help the model infer which task is being asked for.

These views are complementary and partly supported; ICL in large models likely involves several mechanisms.

A small experiment#

python
from transformers import pipeline
gen = pipeline("text-generation", model="Qwen/Qwen2.5-0.5B-Instruct")

def few_shot(examples, query):
    prompt = "".join(f"Review: {x}\nSentiment: {y}\n\n" for x, y in examples)
    return prompt + f"Review: {query}\nSentiment:"

examples = [("The staff were kind and patient.", "positive"),
            ("I waited all day and nobody helped.", "negative"),
            ("Clean facilities and clear information.", "positive")]
q = "The forms were confusing and the office closed early."
out = gen(few_shot(examples, q), max_new_tokens=3, do_sample=False)[0]["generated_text"]
print(out.split("Sentiment:")[-1].strip())

Try shuffling example order, flipping labels, or using unrelated label words ("foo"/"bar") to see how behaviour changes.

ICL vs fine-tuning#

In-context learningFine-tuning
Weight updatesNoneYes
Data neededA few to hundreds of examplesHundreds to thousands+
Setup timeSecondsHours
Cost per queryHigher (long prompts)Lower
ConsistencySensitive to prompt detailsMore consistent
Knowledge persistenceOnly within the promptStored in weights

A practical path: prototype with ICL, collect data from use, and fine-tune (or distil) when volume, consistency or cost requires it.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Prompt Engineering: Getting Reliable Results from LLMs

Practical, evidence-based techniques for prompting LLMs — clear instructions, context, examples, output formats, decomposition, and systematic evaluation — plus prompt injection risks.

Beginner⏱ 5 min#205
✨ Generative AI & LLMs

LoRA and Parameter-Efficient Fine-Tuning (PEFT)

Full fine-tuning of billion-parameter models is expensive. PEFT methods train a tiny fraction of parameters. We derive LoRA's low-rank updates, QLoRA's 4-bit training, compare adapters and prompt tuning, and give practical recipes.

Advanced⏱ 5 min#210
✨ Generative AI & LLMs

Large Language Models: What They Are and How They Are Built

A map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.

Beginner⏱ 5 min#199