Fine-tuning all weights of a 7-billion-parameter model requires storing gradients and optimiser states for 7 billion parameters — well over 100 GB of GPU memory — and saving a full copy of the model for every task. Parameter-Efficient Fine-Tuning (PEFT) methods freeze the pretrained model and train a small number of new or selected parameters, often matching full fine-tuning quality. LoRA is the most widely used, and with QLoRA a large model can be fine-tuned on a single consumer GPU.
The intuition#
Aghajanyan et al. (2020) found that pretrained models have a low intrinsic dimension for fine-tuning: good solutions for downstream tasks can be found in a surprisingly low-dimensional subspace of parameter changes. If the update $\Delta\mathbf{W}$ needed for a task is approximately low rank, we can parameterise it cheaply.
LoRA: Low-Rank Adaptation#
Hu et al. (2021) keep a pretrained weight matrix $\mathbf{W}_0 \in \mathbb{R}^{d \times k}$ frozen and add a trainable low-rank update:
- $\mathbf{A}$ is initialised randomly and $\mathbf{B}$ at zero, so training starts exactly from the pretrained model.
- $\alpha$ is a scaling hyperparameter; the factor $\alpha/r$ keeps update magnitudes comparable across ranks.
- Trainable parameters per matrix: $r(d + k)$ instead of $dk$. For $d = k = 4096$ and $r = 8$: 65,536 instead of 16.8 million — about 0.4%.
The forward pass becomes $\mathbf{h} = \mathbf{W}_0\mathbf{x} + \frac{\alpha}{r}\mathbf{B}\mathbf{A}\mathbf{x}$.
Advantages:
- Large memory savings: no optimiser state for frozen weights.
- No inference latency: after training, merge $\mathbf{W} = \mathbf{W}_0 + \frac{\alpha}{r}\mathbf{B}\mathbf{A}$.
- Tiny, swappable adapters: store one base model and many task-specific adapters of a few megabytes; switch at serving time (multi-LoRA serving).
import torch
import torch.nn as nn
class LoRALinear(nn.Module):
def __init__(self, base: nn.Linear, r=8, alpha=16, dropout=0.05):
super().__init__()
self.base = base
for p in self.base.parameters():
p.requires_grad = False # freeze pretrained weights
self.A = nn.Parameter(torch.randn(r, base.in_features) * 0.01)
self.B = nn.Parameter(torch.zeros(base.out_features, r)) # zero init -> no change at start
self.scale, self.drop = alpha / r, nn.Dropout(dropout)
def forward(self, x):
return self.base(x) + self.scale * (self.drop(x) @ self.A.T @ self.B.T)
def merge(self):
self.base.weight.data += self.scale * self.B @ self.A # fold into the base weight
layer = LoRALinear(nn.Linear(4096, 4096), r=8)
trainable = sum(p.numel() for p in layer.parameters() if p.requires_grad)
total = sum(p.numel() for p in layer.parameters())
print(f"trainable {trainable:,} of {total:,} ({100 * trainable / total:.2f}%)")Where to apply LoRA#
The original paper applied LoRA to attention query and value projections. Later practice (e.g. the QLoRA paper) found applying it to all linear layers (attention and MLP projections) improves quality. Typical settings: rank $r$ = 8–64, $\alpha$ = $r$ to $2r$, dropout 0.05, learning rate around $1\times10^{-4}$ to $2\times10^{-4}$ (higher than full fine-tuning).
QLoRA: fine-tuning quantised models#
Dettmers et al. (2023) combined LoRA with a frozen base model quantised to 4 bits:
- NF4 (4-bit NormalFloat): a data type whose quantisation levels match normally distributed weights.
- Double quantisation: quantise the quantisation constants too.
- Paged optimisers: move optimiser memory spikes to CPU memory when needed.
Gradients flow through the dequantised 4-bit weights into bf16 LoRA adapters. QLoRA made it possible to fine-tune a 65B-parameter model on a single 48 GB GPU while matching 16-bit fine-tuning quality on their benchmarks.
# Practical QLoRA with Hugging Face (sketch)
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
import torch
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct", quantization_config=bnb,
device_map="auto")
model = prepare_model_for_kbit_training(model)
config = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"])
model = get_peft_model(model, config)
model.print_trainable_parameters() # typically well under 1% of parameters
# Then train with TRL's SFTTrainer on your instruction data.Other PEFT methods#
| Method | What is trained | Notes |
|---|---|---|
| Adapters (Houlsby et al., 2019) | Small bottleneck MLPs inserted in each layer | Adds some inference latency unless fused |
| Prefix / prompt tuning | Learned "virtual token" vectors prepended to inputs or keys/values | Very few parameters; weaker for small models |
| (IA)³ | Learned vectors that rescale activations | Extremely few parameters |
| BitFit | Bias terms only | Simple baseline |
| LoRA variants | DoRA (decomposes magnitude and direction), rsLoRA, LoRA+ | Often small gains over LoRA |
Limitations and cautions#
- LoRA can learn less new knowledge than full fine-tuning for very different domains (and also forgets less — "LoRA learns less and forgets less", Biderman et al., 2024).
- Rank and target modules matter; tune them.
- Fine-tuning — even with LoRA — can weaken safety behaviours; re-evaluate safety.
- Adapters inherit the base model's licence and limitations.