Training a large network from scratch requires enormous data and compute. Yet most practitioners today achieve excellent results with a few thousand โ sometimes a few hundred โ labelled examples. The secret is transfer learning: start from a model pretrained on a large, general dataset, and adapt it to your task. It is the single most practical technique in modern deep learning.
Why transfer works#
Features learned on large, diverse data are general. In a CNN pretrained on ImageNet, early layers detect edges, colours and textures useful for almost any visual task; middle layers detect parts and patterns; only the last layers are specific to ImageNet's 1,000 classes. Yosinski et al. (2014) quantified this: features become progressively more task-specific with depth. In language models, pretraining on vast text teaches grammar, facts and reasoning patterns that transfer to classification, extraction and question answering.
Strategy 1: feature extraction#
Freeze the pretrained network, remove its original output layer, and train only a new head on your data:
- Very fast and cheap; works with tiny datasets.
- You can precompute features once and train a logistic regression on them.
- Limited if your domain differs substantially from the pretraining domain.
Strategy 2: fine-tuning#
Initialise from pretrained weights and continue training some or all layers on your task with a small learning rate.
- Usually better accuracy than feature extraction, especially with more data or a domain shift.
- Risk of catastrophic forgetting and overfitting on small data.
Best practices for fine-tuning#
- Replace the head with a freshly initialised layer for your classes.
- Warm up the head first: train only the new head for a few epochs (backbone frozen), then unfreeze. A random head produces large, noisy gradients that can damage pretrained features.
- Use smaller learning rates than training from scratch โ typically 10โ100ร smaller (e.g. $10^{-5}$โ$10^{-4}$ for transformers, $10^{-4}$โ$10^{-3}$ for CNNs).
- Discriminative (layer-wise) learning rates: lower rates for early, general layers; higher for later layers and the head.
- Gradual unfreezing: unfreeze layers from top to bottom over epochs (as in ULMFiT).
- Use the pretrained preprocessing: same input size, normalisation statistics and tokeniser.
- Regularise: weight decay, augmentation, early stopping.
- BatchNorm with small batches: keep BN layers in eval mode (frozen statistics).
import torch
import torch.nn as nn
from torchvision import models
model = models.resnet50(weights=models.ResNet50_Weights.IMAGENET1K_V2)
num_classes = 5 # e.g. five crop-disease categories
model.fc = nn.Linear(model.fc.in_features, num_classes)
# Stage 1: train the head only
for p in model.parameters():
p.requires_grad = False
for p in model.fc.parameters():
p.requires_grad = True
opt = torch.optim.AdamW(model.fc.parameters(), lr=1e-3)
# ... train a few epochs ...
# Stage 2: unfreeze everything with discriminative learning rates
for p in model.parameters():
p.requires_grad = True
opt = torch.optim.AdamW([
{"params": list(model.conv1.parameters()) + list(model.layer1.parameters()), "lr": 1e-5},
{"params": list(model.layer2.parameters()) + list(model.layer3.parameters()), "lr": 5e-5},
{"params": model.layer4.parameters(), "lr": 1e-4},
{"params": model.fc.parameters(), "lr": 5e-4},
], weight_decay=0.01)Choosing a strategy#
| Your data size | Similar to pretraining domain | Different domain |
|---|---|---|
| Small | Feature extraction (or fine-tune head + top layers) | Fine-tune top layers carefully; consider domain-specific pretrained models |
| Large | Fine-tune everything | Fine-tune everything (or pretrain on in-domain unlabelled data first) |
For very different domains โ satellite imagery, medical scans, microscopy, low-resource languages โ look for models pretrained on similar data, or continue self-supervised pretraining on your unlabelled in-domain data before fine-tuning (domain-adaptive pretraining).
Parameter-efficient fine-tuning#
For very large models, updating all parameters is expensive and requires storing a full copy per task. Parameter-efficient fine-tuning (PEFT) trains only a small number of added or selected parameters:
- Adapters โ small bottleneck layers inserted into each block.
- LoRA โ low-rank updates to weight matrices (covered in detail in the LLM track).
- Prompt/prefix tuning โ learned input vectors.
- BitFit โ train only biases.
These often match full fine-tuning at a tiny fraction of trainable parameters.
When transfer can hurt#
Negative transfer occurs when source and target tasks are too different or when pretrained biases conflict with the target task. Pretrained models also carry the biases and blind spots of their data: a model pretrained mostly on images from wealthy countries may perform worse on photos of households, crops or objects from elsewhere. Always evaluate on data representative of your deployment population.