๐Ÿ”— Deep Learning ยท Lecture 27 of 38

Transfer Learning and Fine-Tuning

Pretrained models let you achieve strong results with small datasets. We compare feature extraction and fine-tuning, explain discriminative learning rates and layer freezing, and discuss when transfer helps or hurts.

Training a large network from scratch requires enormous data and compute. Yet most practitioners today achieve excellent results with a few thousand โ€” sometimes a few hundred โ€” labelled examples. The secret is transfer learning: start from a model pretrained on a large, general dataset, and adapt it to your task. It is the single most practical technique in modern deep learning.

Why transfer works#

Features learned on large, diverse data are general. In a CNN pretrained on ImageNet, early layers detect edges, colours and textures useful for almost any visual task; middle layers detect parts and patterns; only the last layers are specific to ImageNet's 1,000 classes. Yosinski et al. (2014) quantified this: features become progressively more task-specific with depth. In language models, pretraining on vast text teaches grammar, facts and reasoning patterns that transfer to classification, extraction and question answering.

Strategy 1: feature extraction#

Freeze the pretrained network, remove its original output layer, and train only a new head on your data:

$$ \hat{y} = h_\psi\big(f_{\theta_{\text{frozen}}}(\mathbf{x})\big) $$
  • Very fast and cheap; works with tiny datasets.
  • You can precompute features once and train a logistic regression on them.
  • Limited if your domain differs substantially from the pretraining domain.

Strategy 2: fine-tuning#

Initialise from pretrained weights and continue training some or all layers on your task with a small learning rate.

  • Usually better accuracy than feature extraction, especially with more data or a domain shift.
  • Risk of catastrophic forgetting and overfitting on small data.

Best practices for fine-tuning#

  1. Replace the head with a freshly initialised layer for your classes.
  2. Warm up the head first: train only the new head for a few epochs (backbone frozen), then unfreeze. A random head produces large, noisy gradients that can damage pretrained features.
  3. Use smaller learning rates than training from scratch โ€” typically 10โ€“100ร— smaller (e.g. $10^{-5}$โ€“$10^{-4}$ for transformers, $10^{-4}$โ€“$10^{-3}$ for CNNs).
  4. Discriminative (layer-wise) learning rates: lower rates for early, general layers; higher for later layers and the head.
  5. Gradual unfreezing: unfreeze layers from top to bottom over epochs (as in ULMFiT).
  6. Use the pretrained preprocessing: same input size, normalisation statistics and tokeniser.
  7. Regularise: weight decay, augmentation, early stopping.
  8. BatchNorm with small batches: keep BN layers in eval mode (frozen statistics).
python
import torch
import torch.nn as nn
from torchvision import models

model = models.resnet50(weights=models.ResNet50_Weights.IMAGENET1K_V2)
num_classes = 5                                     # e.g. five crop-disease categories
model.fc = nn.Linear(model.fc.in_features, num_classes)

# Stage 1: train the head only
for p in model.parameters():
    p.requires_grad = False
for p in model.fc.parameters():
    p.requires_grad = True
opt = torch.optim.AdamW(model.fc.parameters(), lr=1e-3)
# ... train a few epochs ...

# Stage 2: unfreeze everything with discriminative learning rates
for p in model.parameters():
    p.requires_grad = True
opt = torch.optim.AdamW([
    {"params": list(model.conv1.parameters()) + list(model.layer1.parameters()), "lr": 1e-5},
    {"params": list(model.layer2.parameters()) + list(model.layer3.parameters()), "lr": 5e-5},
    {"params": model.layer4.parameters(), "lr": 1e-4},
    {"params": model.fc.parameters(), "lr": 5e-4},
], weight_decay=0.01)

Choosing a strategy#

Your data sizeSimilar to pretraining domainDifferent domain
SmallFeature extraction (or fine-tune head + top layers)Fine-tune top layers carefully; consider domain-specific pretrained models
LargeFine-tune everythingFine-tune everything (or pretrain on in-domain unlabelled data first)

For very different domains โ€” satellite imagery, medical scans, microscopy, low-resource languages โ€” look for models pretrained on similar data, or continue self-supervised pretraining on your unlabelled in-domain data before fine-tuning (domain-adaptive pretraining).

Parameter-efficient fine-tuning#

For very large models, updating all parameters is expensive and requires storing a full copy per task. Parameter-efficient fine-tuning (PEFT) trains only a small number of added or selected parameters:

  • Adapters โ€” small bottleneck layers inserted into each block.
  • LoRA โ€” low-rank updates to weight matrices (covered in detail in the LLM track).
  • Prompt/prefix tuning โ€” learned input vectors.
  • BitFit โ€” train only biases.

These often match full fine-tuning at a tiny fraction of trainable parameters.

When transfer can hurt#

Negative transfer occurs when source and target tasks are too different or when pretrained biases conflict with the target task. Pretrained models also carry the biases and blind spots of their data: a model pretrained mostly on images from wealthy countries may perform worse on photos of households, crops or objects from elsewhere. Always evaluate on data representative of your deployment population.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Embeddings: Turning Discrete Things into Meaningful Vectors

Words, users, products and categories become dense vectors whose geometry encodes meaning. We explain embedding layers, how embeddings are learned, how to measure similarity, and how they power search and recommendation.

Beginnerโฑ 5 min#122
๐Ÿ”— Deep Learning

PyTorch Fundamentals: Tensors, Autograd, Modules and the Training Loop

A practical tour of PyTorch โ€” tensors and devices, autograd, nn.Module, Dataset and DataLoader, optimisers, and a complete, correct training and evaluation loop you can reuse in every project.

Beginnerโฑ 5 min#124
๐Ÿ”— Deep Learning

Autoencoders: Compression, Denoising and Representation Learning

An autoencoder learns to reconstruct its input through a bottleneck, discovering compact representations without labels. We cover undercomplete, denoising, sparse and convolutional autoencoders and their uses.

Intermediateโฑ 5 min#121