✨ Generative AI & LLMs · Lecture 20 of 30

LoRA and Parameter-Efficient Fine-Tuning (PEFT)

Full fine-tuning of billion-parameter models is expensive. PEFT methods train a tiny fraction of parameters. We derive LoRA's low-rank updates, QLoRA's 4-bit training, compare adapters and prompt tuning, and give practical recipes.

Fine-tuning all weights of a 7-billion-parameter model requires storing gradients and optimiser states for 7 billion parameters — well over 100 GB of GPU memory — and saving a full copy of the model for every task. Parameter-Efficient Fine-Tuning (PEFT) methods freeze the pretrained model and train a small number of new or selected parameters, often matching full fine-tuning quality. LoRA is the most widely used, and with QLoRA a large model can be fine-tuned on a single consumer GPU.

The intuition#

Aghajanyan et al. (2020) found that pretrained models have a low intrinsic dimension for fine-tuning: good solutions for downstream tasks can be found in a surprisingly low-dimensional subspace of parameter changes. If the update $\Delta\mathbf{W}$ needed for a task is approximately low rank, we can parameterise it cheaply.

LoRA: Low-Rank Adaptation#

Hu et al. (2021) keep a pretrained weight matrix $\mathbf{W}_0 \in \mathbb{R}^{d \times k}$ frozen and add a trainable low-rank update:

$$ \mathbf{W} = \mathbf{W}_0 + \Delta\mathbf{W} = \mathbf{W}_0 + \frac{\alpha}{r}\,\mathbf{B}\mathbf{A}, \qquad \mathbf{B} \in \mathbb{R}^{d \times r},\; \mathbf{A} \in \mathbb{R}^{r \times k},\; r \ll \min(d, k) $$
  • $\mathbf{A}$ is initialised randomly and $\mathbf{B}$ at zero, so training starts exactly from the pretrained model.
  • $\alpha$ is a scaling hyperparameter; the factor $\alpha/r$ keeps update magnitudes comparable across ranks.
  • Trainable parameters per matrix: $r(d + k)$ instead of $dk$. For $d = k = 4096$ and $r = 8$: 65,536 instead of 16.8 million — about 0.4%.

The forward pass becomes $\mathbf{h} = \mathbf{W}_0\mathbf{x} + \frac{\alpha}{r}\mathbf{B}\mathbf{A}\mathbf{x}$.

Advantages:

  • Large memory savings: no optimiser state for frozen weights.
  • No inference latency: after training, merge $\mathbf{W} = \mathbf{W}_0 + \frac{\alpha}{r}\mathbf{B}\mathbf{A}$.
  • Tiny, swappable adapters: store one base model and many task-specific adapters of a few megabytes; switch at serving time (multi-LoRA serving).
python
import torch
import torch.nn as nn

class LoRALinear(nn.Module):
    def __init__(self, base: nn.Linear, r=8, alpha=16, dropout=0.05):
        super().__init__()
        self.base = base
        for p in self.base.parameters():
            p.requires_grad = False                            # freeze pretrained weights
        self.A = nn.Parameter(torch.randn(r, base.in_features) * 0.01)
        self.B = nn.Parameter(torch.zeros(base.out_features, r))   # zero init -> no change at start
        self.scale, self.drop = alpha / r, nn.Dropout(dropout)
    def forward(self, x):
        return self.base(x) + self.scale * (self.drop(x) @ self.A.T @ self.B.T)
    def merge(self):
        self.base.weight.data += self.scale * self.B @ self.A   # fold into the base weight

layer = LoRALinear(nn.Linear(4096, 4096), r=8)
trainable = sum(p.numel() for p in layer.parameters() if p.requires_grad)
total = sum(p.numel() for p in layer.parameters())
print(f"trainable {trainable:,} of {total:,} ({100 * trainable / total:.2f}%)")

Where to apply LoRA#

The original paper applied LoRA to attention query and value projections. Later practice (e.g. the QLoRA paper) found applying it to all linear layers (attention and MLP projections) improves quality. Typical settings: rank $r$ = 8–64, $\alpha$ = $r$ to $2r$, dropout 0.05, learning rate around $1\times10^{-4}$ to $2\times10^{-4}$ (higher than full fine-tuning).

QLoRA: fine-tuning quantised models#

Dettmers et al. (2023) combined LoRA with a frozen base model quantised to 4 bits:

  • NF4 (4-bit NormalFloat): a data type whose quantisation levels match normally distributed weights.
  • Double quantisation: quantise the quantisation constants too.
  • Paged optimisers: move optimiser memory spikes to CPU memory when needed.

Gradients flow through the dequantised 4-bit weights into bf16 LoRA adapters. QLoRA made it possible to fine-tune a 65B-parameter model on a single 48 GB GPU while matching 16-bit fine-tuning quality on their benchmarks.

python
# Practical QLoRA with Hugging Face (sketch)
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
import torch

bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct", quantization_config=bnb,
                                             device_map="auto")
model = prepare_model_for_kbit_training(model)
config = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
                    target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"])
model = get_peft_model(model, config)
model.print_trainable_parameters()          # typically well under 1% of parameters
# Then train with TRL's SFTTrainer on your instruction data.

Other PEFT methods#

MethodWhat is trainedNotes
Adapters (Houlsby et al., 2019)Small bottleneck MLPs inserted in each layerAdds some inference latency unless fused
Prefix / prompt tuningLearned "virtual token" vectors prepended to inputs or keys/valuesVery few parameters; weaker for small models
(IA)³Learned vectors that rescale activationsExtremely few parameters
BitFitBias terms onlySimple baseline
LoRA variantsDoRA (decomposes magnitude and direction), rsLoRA, LoRA+Often small gains over LoRA

Limitations and cautions#

  • LoRA can learn less new knowledge than full fine-tuning for very different domains (and also forgets less — "LoRA learns less and forgets less", Biderman et al., 2024).
  • Rank and target modules matter; tune them.
  • Fine-tuning — even with LoRA — can weaken safety behaviours; re-evaluate safety.
  • Adapters inherit the base model's licence and limitations.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

In-Context Learning: How LLMs Learn from Prompts

Large language models can perform new tasks from a few examples in the prompt without weight updates. We examine what in-context learning is, what influences it, theories of how it works, and its practical limits.

Intermediate⏱ 5 min#207
✨ Generative AI & LLMs

Prompt Engineering: Getting Reliable Results from LLMs

Practical, evidence-based techniques for prompting LLMs — clear instructions, context, examples, output formats, decomposition, and systematic evaluation — plus prompt injection risks.

Beginner⏱ 5 min#205
✨ Generative AI & LLMs

AI Agents and Tool Use: LLMs That Act

Agents let LLMs call tools, observe results and pursue multi-step goals. We cover function calling, the ReAct loop, planning and memory, multi-agent patterns, evaluation, and — critically — safety, permissions and human oversight.

Intermediate⏱ 5 min#217