A 70-billion-parameter model stored in 16-bit precision needs about 140 GB of memory just for its weights — more than any single consumer GPU. Quantised to 4 bits, it needs about 35 GB. For smaller models, quantisation turns "needs a data-centre GPU" into "runs on a laptop". Because LLM inference is usually limited by memory bandwidth (moving weights from memory to compute units for each generated token), fewer bits also means faster generation. This lecture builds on the general quantisation lecture in the Deep Learning track and focuses on what is special about LLMs.
Memory arithmetic#
| Precision | Bytes per parameter | 7B model | 70B model |
|---|---|---|---|
| FP32 | 4 | 28 GB | 280 GB |
| FP16 / BF16 | 2 | 14 GB | 140 GB |
| INT8 | 1 | 7 GB | 70 GB |
| 4-bit | 0.5 | ~3.5 GB | ~35 GB |
(Plus memory for the KV cache and activations, and small overheads for scales.)
Why LLMs are hard to quantise: outlier features#
Dettmers et al. (2022) discovered that once transformers exceed a few billion parameters, a small number of hidden dimensions develop extremely large activation values ("emergent outlier features"). With per-tensor INT8 quantisation, these outliers force a large scale, wiping out precision for all other values and severely degrading the model.
Solutions:
- LLM.int8(): perform matrix multiplication for outlier dimensions in 16-bit and the rest in INT8 (mixed-precision decomposition).
- SmoothQuant (Xiao et al., 2023): mathematically migrate quantisation difficulty from activations to weights by per-channel scaling ($\mathbf{Y} = (\mathbf{X}\,\text{diag}(\mathbf{s})^{-1})(\text{diag}(\mathbf{s})\mathbf{W})$), enabling INT8 weights and activations.
- Rotation-based methods (e.g. QuaRot, SpinQuant) apply orthogonal rotations that spread outliers across dimensions before quantising.
Weight-only vs weight-and-activation quantisation#
- Weight-only (e.g. W4A16: 4-bit weights, 16-bit activations): weights are dequantised on the fly inside the matrix-multiplication kernel. Great for memory and for bandwidth-bound single-user generation. The most common approach for local and small-batch inference.
- Weight + activation (W8A8, FP8): enables low-precision arithmetic on tensor cores — better for high-throughput, large-batch serving.
- KV-cache quantisation reduces memory for long contexts.
Post-training methods for 4-bit weights#
Naive round-to-nearest at 4 bits loses noticeable quality. Better methods use a small calibration set of text:
- GPTQ (Frantar et al., 2022): quantise weights column by column and update the remaining unquantised weights to compensate for the error, using approximate second-order (Hessian) information from calibration activations — an efficient descendant of Optimal Brain Surgeon. It quantises large models in hours on a single GPU.
- AWQ (Lin et al., 2023): observes that a small fraction of weight channels matters most — those multiplied by large activations — and protects them by scaling before quantisation, without backpropagation or reconstruction.
- Group-wise scales: one scale (and zero point) per group of 64–128 weights handles varying ranges within a row.
The objective these methods approximately solve is layer-wise output reconstruction:
where $\mathbf{X}$ are calibration inputs to the layer and $\mathcal{Q}$ is the set of quantised matrices.
import numpy as np
def quantize_groupwise(w, bits=4, group=128):
"""Symmetric group-wise quantisation of a 1-D weight row."""
qmax = 2 ** (bits - 1) - 1
w = w.reshape(-1, group)
scale = np.abs(w).max(axis=1, keepdims=True) / qmax
q = np.clip(np.round(w / scale), -qmax - 1, qmax)
return (q * scale).reshape(-1)
rng = np.random.default_rng(0)
row = rng.normal(0, 0.02, 4096); row[[10, 500]] = 0.8 # two outlier weights
for g in [4096, 128, 32]:
err = np.abs(quantize_groupwise(row, 4, g) - row).mean()
print(f"group size {g:>4}: mean abs error {err:.5f}")Smaller groups isolate outliers and reduce error at the cost of storing more scales.
Formats and runtimes#
- GGUF with llama.cpp: popular for CPU and Apple-silicon inference, with many quantisation variants (e.g. 4-bit "K-quants" with mixed precision across layers).
- GPTQ/AWQ checkpoints served with GPU inference engines (e.g. vLLM, TensorRT-LLM, text-generation-inference).
- bitsandbytes for 8-bit and 4-bit loading in PyTorch (also used by QLoRA).
- MLX on Apple silicon; ONNX Runtime and ExecuTorch for mobile.
Evaluating quantised models#
Quantisation error is not uniform across tasks. Measure:
- Perplexity on held-out text (a sensitive, cheap indicator);
- Task benchmarks relevant to you — reasoning, maths and code often degrade more than general chat;
- Multilingual performance — lower-resource languages can degrade more;
- Long-context behaviour, especially with KV-cache quantisation;
- Latency, throughput and memory on the target hardware.