✨ Generative AI & LLMs · Lecture 21 of 30

Quantising Large Language Models for Efficient Inference

Quantisation shrinks LLM weights to 8, 4 or fewer bits so models run on smaller GPUs, laptops and phones. We cover outlier features, weight-only vs activation quantisation, GPTQ, AWQ, formats like GGUF, and how to evaluate quality loss.

A 70-billion-parameter model stored in 16-bit precision needs about 140 GB of memory just for its weights — more than any single consumer GPU. Quantised to 4 bits, it needs about 35 GB. For smaller models, quantisation turns "needs a data-centre GPU" into "runs on a laptop". Because LLM inference is usually limited by memory bandwidth (moving weights from memory to compute units for each generated token), fewer bits also means faster generation. This lecture builds on the general quantisation lecture in the Deep Learning track and focuses on what is special about LLMs.

Memory arithmetic#

PrecisionBytes per parameter7B model70B model
FP32428 GB280 GB
FP16 / BF16214 GB140 GB
INT817 GB70 GB
4-bit0.5~3.5 GB~35 GB

(Plus memory for the KV cache and activations, and small overheads for scales.)

Why LLMs are hard to quantise: outlier features#

Dettmers et al. (2022) discovered that once transformers exceed a few billion parameters, a small number of hidden dimensions develop extremely large activation values ("emergent outlier features"). With per-tensor INT8 quantisation, these outliers force a large scale, wiping out precision for all other values and severely degrading the model.

Solutions:

  • LLM.int8(): perform matrix multiplication for outlier dimensions in 16-bit and the rest in INT8 (mixed-precision decomposition).
  • SmoothQuant (Xiao et al., 2023): mathematically migrate quantisation difficulty from activations to weights by per-channel scaling ($\mathbf{Y} = (\mathbf{X}\,\text{diag}(\mathbf{s})^{-1})(\text{diag}(\mathbf{s})\mathbf{W})$), enabling INT8 weights and activations.
  • Rotation-based methods (e.g. QuaRot, SpinQuant) apply orthogonal rotations that spread outliers across dimensions before quantising.

Weight-only vs weight-and-activation quantisation#

  • Weight-only (e.g. W4A16: 4-bit weights, 16-bit activations): weights are dequantised on the fly inside the matrix-multiplication kernel. Great for memory and for bandwidth-bound single-user generation. The most common approach for local and small-batch inference.
  • Weight + activation (W8A8, FP8): enables low-precision arithmetic on tensor cores — better for high-throughput, large-batch serving.
  • KV-cache quantisation reduces memory for long contexts.

Post-training methods for 4-bit weights#

Naive round-to-nearest at 4 bits loses noticeable quality. Better methods use a small calibration set of text:

  • GPTQ (Frantar et al., 2022): quantise weights column by column and update the remaining unquantised weights to compensate for the error, using approximate second-order (Hessian) information from calibration activations — an efficient descendant of Optimal Brain Surgeon. It quantises large models in hours on a single GPU.
  • AWQ (Lin et al., 2023): observes that a small fraction of weight channels matters most — those multiplied by large activations — and protects them by scaling before quantisation, without backpropagation or reconstruction.
  • Group-wise scales: one scale (and zero point) per group of 64–128 weights handles varying ranges within a row.

The objective these methods approximately solve is layer-wise output reconstruction:

$$ \hat{\mathbf{W}} = \arg\min_{\hat{\mathbf{W}} \in \mathcal{Q}}\left\|\mathbf{W}\mathbf{X} - \hat{\mathbf{W}}\mathbf{X}\right\|_2^2 $$

where $\mathbf{X}$ are calibration inputs to the layer and $\mathcal{Q}$ is the set of quantised matrices.

python
import numpy as np

def quantize_groupwise(w, bits=4, group=128):
    """Symmetric group-wise quantisation of a 1-D weight row."""
    qmax = 2 ** (bits - 1) - 1
    w = w.reshape(-1, group)
    scale = np.abs(w).max(axis=1, keepdims=True) / qmax
    q = np.clip(np.round(w / scale), -qmax - 1, qmax)
    return (q * scale).reshape(-1)

rng = np.random.default_rng(0)
row = rng.normal(0, 0.02, 4096); row[[10, 500]] = 0.8          # two outlier weights
for g in [4096, 128, 32]:
    err = np.abs(quantize_groupwise(row, 4, g) - row).mean()
    print(f"group size {g:>4}: mean abs error {err:.5f}")

Smaller groups isolate outliers and reduce error at the cost of storing more scales.

Formats and runtimes#

  • GGUF with llama.cpp: popular for CPU and Apple-silicon inference, with many quantisation variants (e.g. 4-bit "K-quants" with mixed precision across layers).
  • GPTQ/AWQ checkpoints served with GPU inference engines (e.g. vLLM, TensorRT-LLM, text-generation-inference).
  • bitsandbytes for 8-bit and 4-bit loading in PyTorch (also used by QLoRA).
  • MLX on Apple silicon; ONNX Runtime and ExecuTorch for mobile.

Evaluating quantised models#

Quantisation error is not uniform across tasks. Measure:

  • Perplexity on held-out text (a sensitive, cheap indicator);
  • Task benchmarks relevant to you — reasoning, maths and code often degrade more than general chat;
  • Multilingual performance — lower-resource languages can degrade more;
  • Long-context behaviour, especially with KV-cache quantisation;
  • Latency, throughput and memory on the target hardware.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

LoRA and Parameter-Efficient Fine-Tuning (PEFT)

Full fine-tuning of billion-parameter models is expensive. PEFT methods train a tiny fraction of parameters. We derive LoRA's low-rank updates, QLoRA's 4-bit training, compare adapters and prompt tuning, and give practical recipes.

Advanced⏱ 5 min#210
✨ Generative AI & LLMs

Mixture of Experts: Scaling Parameters Without Scaling Compute

Mixture-of-experts layers route each token to a few of many expert networks, so models gain parameters without proportional compute. We cover gating, top-k routing, load balancing, capacity, training and serving challenges.

Advanced⏱ 5 min#212
✨ Generative AI & LLMs

Vector Databases and Approximate Nearest Neighbour Search

Embedding-based applications need fast similarity search over millions of vectors. We explain exact vs approximate search, HNSW graphs, IVF and product quantisation, filtering, and how to choose and operate a vector store.

Intermediate⏱ 6 min#209