✨ Generative AI & LLMs · Lecture 12 of 30

Instruction Tuning: Teaching Language Models to Follow Directions

A base model continues text; an instruction-tuned model answers requests. We cover supervised fine-tuning data (human-written, templated and synthetic), chat formats, loss masking, what instruction tuning changes, and practical recipes.

Ask a base (pretrained-only) language model "Write a poem about the monsoon" and it might continue with "Write a poem about winter. Write a poem about…" — because it has seen lists of prompts on the web. It knows a great deal but does not reliably do what you ask. Instruction tuning — supervised fine-tuning (SFT) on examples of instructions paired with good responses — transforms a base model into a helpful assistant. It is the first and most important step of post-training.

What instruction data looks like#

json
{"messages": [
  {"role": "system", "content": "You are a helpful assistant."},
  {"role": "user", "content": "Explain photosynthesis to a 10-year-old in three sentences."},
  {"role": "assistant", "content": "Plants make their own food using sunlight..."}
]}

Datasets cover diverse tasks: open questions, explanations, writing, summarisation, coding, maths, classification, extraction, multi-turn dialogue, refusals of harmful requests, and tool use.

Sources of instruction data#

  1. Reformatted NLP datasets: turn existing labelled datasets into instructions with templates. FLAN (Wei et al., 2021) and Flan 2022 (Chung et al.) fine-tuned models on hundreds to more than 1,800 tasks and showed large gains in zero-shot performance on unseen tasks — instruction tuning generalises across tasks. Also T0, Super-NaturalInstructions.
  2. Human-written demonstrations: annotators write high-quality responses to real user-style prompts, as in InstructGPT and Dolly. Expensive but high quality.
  3. Synthetic data from stronger models: Self-Instruct (Wang et al., 2023) had a model generate new instructions and responses from a small seed set; Alpaca used a strong model to generate 52K examples. Cheap and scalable, but inherits the teacher's errors and style, and may be restricted by model licence terms.
  4. Curated conversations from real usage (with consent and privacy protection).

Quality over quantity#

LIMA (Zhou et al., 2023) fine-tuned a 65B base model on only 1,000 carefully curated examples and obtained responses competitive with far more heavily tuned models in human evaluations. The authors proposed the superficial alignment hypothesis: most knowledge and capability come from pretraining, and instruction tuning mainly teaches the format and style of helpful interaction. Diversity and quality of examples matter more than sheer volume — though larger, high-quality datasets still help specialised skills such as maths and coding.

Chat templates and loss masking#

Conversations are serialised with special tokens marking roles, for example:

text
<|system|>You are a helpful assistant.<|end|>
<|user|>Explain photosynthesis to a 10-year-old.<|end|>
<|assistant|>Plants make their own food...<|end|>

During SFT, the loss is usually computed only on assistant tokens — the model should learn to produce responses, not to imitate users.

python
import torch

def build_labels(token_ids, role_spans):
    """role_spans: list of (start, end, role). Mask everything except assistant tokens."""
    labels = torch.full_like(token_ids, -100)          # -100 is ignored by cross-entropy
    for start, end, role in role_spans:
        if role == "assistant":
            labels[start:end] = token_ids[start:end]
    return labels

ids = torch.arange(20)
print(build_labels(ids, [(0, 5, "system"), (5, 11, "user"), (11, 20, "assistant")]))

Libraries such as Hugging Face TRL's SFTTrainer handle templating and masking; always use the same template at inference as in training.

What instruction tuning changes#

  • Format following: answers questions instead of continuing text; respects requested length, style and structure (lists, JSON).
  • Zero-shot generalisation to new tasks described in natural language.
  • Conversation across multiple turns.
  • Refusals and safety behaviours when trained with such examples.
  • It can also reduce some capabilities or calibration and teach undesirable habits if data is poor (e.g. over-refusal, verbosity, confident answers to unanswerable questions).

Practical recipe for domain instruction tuning#

  1. Start from a strong instruction-tuned open model (usually better than tuning a base model yourself).
  2. Collect a few hundred to a few thousand high-quality, diverse examples from your domain, reviewed by experts — including examples where the correct answer is "I don't know" or "please contact a caseworker".
  3. Use parameter-efficient fine-tuning (LoRA/QLoRA) with 1–3 epochs and a small learning rate.
  4. Hold out an evaluation set; compare against the base model with good prompting or RAG — fine-tuning is not always necessary.
  5. Check for regressions in general ability and safety.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

LoRA and Parameter-Efficient Fine-Tuning (PEFT)

Full fine-tuning of billion-parameter models is expensive. PEFT methods train a tiny fraction of parameters. We derive LoRA's low-rank updates, QLoRA's 4-bit training, compare adapters and prompt tuning, and give practical recipes.

Advanced⏱ 5 min#210
✨ Generative AI & LLMs

Pretraining LLMs: Data Pipelines, Objectives and Infrastructure

What actually goes into pretraining an LLM? We cover data sourcing, filtering, deduplication, mixture design, tokenisation, the training objective, stability tricks, infrastructure and evaluation during pretraining.

Advanced⏱ 5 min#201
✨ Generative AI & LLMs

RLHF: Reinforcement Learning from Human Feedback

RLHF aligns language models with human preferences using a learned reward model and reinforcement learning. We derive the Bradley–Terry reward model, the KL-regularised objective optimised with PPO, and discuss reward hacking and limitations.

Advanced⏱ 5 min#203