Deep learning became practical because of GPUs — processors originally built for video games. Knowing what your hardware can and cannot do helps you choose batch sizes, precision and model sizes, estimate training time and cost, and diagnose slow training. This lecture gives students the hardware literacy needed for practical ML.
Why GPUs?#
A CPU has a few powerful cores optimised for sequential, branching logic and low latency. A GPU has thousands of simpler cores optimised for parallel arithmetic on large arrays. Neural network training is dominated by matrix multiplications and convolutions — massively parallel operations — so GPUs can be orders of magnitude faster.
Modern NVIDIA GPUs include Tensor Cores, specialised units for mixed-precision matrix multiply–accumulate, delivering far higher throughput in FP16/BF16/FP8/INT8 than in FP32. Other accelerators include Google TPUs (systolic arrays for matrix operations), AMD GPUs (ROCm software stack), Apple silicon (unified memory, MPS/MLX), and mobile NPUs.
The two numbers that matter most#
1. Memory capacity (GB of VRAM)#
Determines what fits: model weights, gradients, optimiser states, activations and batch. Out-of-memory errors are the most common limit students hit.
Rough training memory for a model with $P$ parameters using Adam in mixed precision: about 16–18 bytes per parameter (weights, gradients, FP32 master weights and two Adam moments) plus activations. Inference in FP16 needs about 2 bytes per parameter plus activations/KV cache.
2. Memory bandwidth (GB/s) and compute (FLOP/s)#
Determine speed. An operation is:
- compute-bound if it performs many operations per byte moved (large matrix multiplications);
- memory-bound if it moves a lot of data per operation (element-wise ops, normalisation, LLM decoding).
The ratio of FLOPs to bytes (arithmetic intensity) tells which applies — the "roofline model". This explains why LLM token generation depends on bandwidth, while training large batches depends on compute.
Estimating training time#
For transformers, training compute is about $6ND$ FLOPs (parameters × tokens). Divide by achievable throughput (peak × utilisation, typically 30–60%):
def training_days(params, tokens, n_gpus, peak_tflops, mfu=0.4):
flops = 6 * params * tokens
per_second = n_gpus * peak_tflops * 1e12 * mfu
return flops / per_second / 86_400
# A 125M-parameter model on 2.5B tokens with one GPU at 100 TFLOP/s (bf16) peak
print(round(training_days(125e6, 2.5e9, 1, 100), 2), "days")
# A 7B model on 1T tokens with 64 GPUs at 800 TFLOP/s peak
print(round(training_days(7e9, 1e12, 64, 800), 1), "days")Practical tips to use hardware well#
- Keep the GPU busy: profile data loading; use multiple DataLoader workers, pinned memory and prefetching. A GPU at 30% utilisation is usually waiting for data.
- Use mixed precision (BF16/FP16) — often 2–3× faster and half the activation memory.
- Maximise batch size within memory; use gradient accumulation for larger effective batches.
- Activation checkpointing trades compute for memory.
- Compile models (
torch.compile) and use fused kernels (FlashAttention). - Avoid CPU–GPU synchronisation in the training loop (
.item(), printing tensors each step). - Monitor:
nvidia-smifor utilisation and memory; the PyTorch profiler for bottlenecks.
nvidia-smi --query-gpu=name,utilization.gpu,memory.used,memory.total --format=csv -l 5Choosing hardware as a student or small organisation#
| Option | Pros | Cons |
|---|---|---|
| Laptop CPU / Apple silicon | Free, always available | Slow for deep learning; fine for classical ML and small models |
| Free cloud notebooks (e.g. Colab, Kaggle) | Free GPUs for learning | Session limits, variable availability |
| Consumer GPU workstation | Full control; cost-effective for continuous use | Upfront cost, power, maintenance |
| Cloud GPU instances | Scale on demand; latest hardware | Can become expensive; must manage data security and shut down idle machines |
| University/HPC clusters | Powerful, often free for students | Queues, scheduler learning curve |
Energy and environmental cost#
Training and serving large models consume significant electricity; the carbon footprint depends heavily on the energy mix of the data centre. Report compute and energy where possible (tools like CodeCarbon estimate emissions), choose efficient models, and avoid unnecessary large-scale runs — themes discussed further in the Ethics track.