Ask a base (pretrained-only) language model "Write a poem about the monsoon" and it might continue with "Write a poem about winter. Write a poem about…" — because it has seen lists of prompts on the web. It knows a great deal but does not reliably do what you ask. Instruction tuning — supervised fine-tuning (SFT) on examples of instructions paired with good responses — transforms a base model into a helpful assistant. It is the first and most important step of post-training.
What instruction data looks like#
{"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain photosynthesis to a 10-year-old in three sentences."},
{"role": "assistant", "content": "Plants make their own food using sunlight..."}
]}Datasets cover diverse tasks: open questions, explanations, writing, summarisation, coding, maths, classification, extraction, multi-turn dialogue, refusals of harmful requests, and tool use.
Sources of instruction data#
- Reformatted NLP datasets: turn existing labelled datasets into instructions with templates. FLAN (Wei et al., 2021) and Flan 2022 (Chung et al.) fine-tuned models on hundreds to more than 1,800 tasks and showed large gains in zero-shot performance on unseen tasks — instruction tuning generalises across tasks. Also T0, Super-NaturalInstructions.
- Human-written demonstrations: annotators write high-quality responses to real user-style prompts, as in InstructGPT and Dolly. Expensive but high quality.
- Synthetic data from stronger models: Self-Instruct (Wang et al., 2023) had a model generate new instructions and responses from a small seed set; Alpaca used a strong model to generate 52K examples. Cheap and scalable, but inherits the teacher's errors and style, and may be restricted by model licence terms.
- Curated conversations from real usage (with consent and privacy protection).
Quality over quantity#
LIMA (Zhou et al., 2023) fine-tuned a 65B base model on only 1,000 carefully curated examples and obtained responses competitive with far more heavily tuned models in human evaluations. The authors proposed the superficial alignment hypothesis: most knowledge and capability come from pretraining, and instruction tuning mainly teaches the format and style of helpful interaction. Diversity and quality of examples matter more than sheer volume — though larger, high-quality datasets still help specialised skills such as maths and coding.
Chat templates and loss masking#
Conversations are serialised with special tokens marking roles, for example:
<|system|>You are a helpful assistant.<|end|>
<|user|>Explain photosynthesis to a 10-year-old.<|end|>
<|assistant|>Plants make their own food...<|end|>During SFT, the loss is usually computed only on assistant tokens — the model should learn to produce responses, not to imitate users.
import torch
def build_labels(token_ids, role_spans):
"""role_spans: list of (start, end, role). Mask everything except assistant tokens."""
labels = torch.full_like(token_ids, -100) # -100 is ignored by cross-entropy
for start, end, role in role_spans:
if role == "assistant":
labels[start:end] = token_ids[start:end]
return labels
ids = torch.arange(20)
print(build_labels(ids, [(0, 5, "system"), (5, 11, "user"), (11, 20, "assistant")]))Libraries such as Hugging Face TRL's SFTTrainer handle templating and masking; always use the same template at inference as in training.
What instruction tuning changes#
- Format following: answers questions instead of continuing text; respects requested length, style and structure (lists, JSON).
- Zero-shot generalisation to new tasks described in natural language.
- Conversation across multiple turns.
- Refusals and safety behaviours when trained with such examples.
- It can also reduce some capabilities or calibration and teach undesirable habits if data is poor (e.g. over-refusal, verbosity, confident answers to unanswerable questions).
Practical recipe for domain instruction tuning#
- Start from a strong instruction-tuned open model (usually better than tuning a base model yourself).
- Collect a few hundred to a few thousand high-quality, diverse examples from your domain, reviewed by experts — including examples where the correct answer is "I don't know" or "please contact a caseworker".
- Use parameter-efficient fine-tuning (LoRA/QLoRA) with 1–3 epochs and a small learning rate.
- Hold out an evaluation set; compare against the base model with good prompting or RAG — fine-tuning is not always necessary.
- Check for regressions in general ability and safety.