⚙️ MLOps & Engineering · Lecture 11 of 15

GPUs and Hardware for Machine Learning

Understanding hardware helps you train faster and cheaper. We explain why GPUs suit deep learning, the roles of memory capacity and bandwidth, precision and tensor cores, estimating requirements, and choosing between local, cloud and free resources.

Deep learning became practical because of GPUs — processors originally built for video games. Knowing what your hardware can and cannot do helps you choose batch sizes, precision and model sizes, estimate training time and cost, and diagnose slow training. This lecture gives students the hardware literacy needed for practical ML.

Why GPUs?#

A CPU has a few powerful cores optimised for sequential, branching logic and low latency. A GPU has thousands of simpler cores optimised for parallel arithmetic on large arrays. Neural network training is dominated by matrix multiplications and convolutions — massively parallel operations — so GPUs can be orders of magnitude faster.

Modern NVIDIA GPUs include Tensor Cores, specialised units for mixed-precision matrix multiply–accumulate, delivering far higher throughput in FP16/BF16/FP8/INT8 than in FP32. Other accelerators include Google TPUs (systolic arrays for matrix operations), AMD GPUs (ROCm software stack), Apple silicon (unified memory, MPS/MLX), and mobile NPUs.

The two numbers that matter most#

1. Memory capacity (GB of VRAM)#

Determines what fits: model weights, gradients, optimiser states, activations and batch. Out-of-memory errors are the most common limit students hit.

Rough training memory for a model with $P$ parameters using Adam in mixed precision: about 16–18 bytes per parameter (weights, gradients, FP32 master weights and two Adam moments) plus activations. Inference in FP16 needs about 2 bytes per parameter plus activations/KV cache.

2. Memory bandwidth (GB/s) and compute (FLOP/s)#

Determine speed. An operation is:

  • compute-bound if it performs many operations per byte moved (large matrix multiplications);
  • memory-bound if it moves a lot of data per operation (element-wise ops, normalisation, LLM decoding).

The ratio of FLOPs to bytes (arithmetic intensity) tells which applies — the "roofline model". This explains why LLM token generation depends on bandwidth, while training large batches depends on compute.

Estimating training time#

For transformers, training compute is about $6ND$ FLOPs (parameters × tokens). Divide by achievable throughput (peak × utilisation, typically 30–60%):

python
def training_days(params, tokens, n_gpus, peak_tflops, mfu=0.4):
    flops = 6 * params * tokens
    per_second = n_gpus * peak_tflops * 1e12 * mfu
    return flops / per_second / 86_400

# A 125M-parameter model on 2.5B tokens with one GPU at 100 TFLOP/s (bf16) peak
print(round(training_days(125e6, 2.5e9, 1, 100), 2), "days")
# A 7B model on 1T tokens with 64 GPUs at 800 TFLOP/s peak
print(round(training_days(7e9, 1e12, 64, 800), 1), "days")

Practical tips to use hardware well#

  1. Keep the GPU busy: profile data loading; use multiple DataLoader workers, pinned memory and prefetching. A GPU at 30% utilisation is usually waiting for data.
  2. Use mixed precision (BF16/FP16) — often 2–3× faster and half the activation memory.
  3. Maximise batch size within memory; use gradient accumulation for larger effective batches.
  4. Activation checkpointing trades compute for memory.
  5. Compile models (torch.compile) and use fused kernels (FlashAttention).
  6. Avoid CPU–GPU synchronisation in the training loop (.item(), printing tensors each step).
  7. Monitor: nvidia-smi for utilisation and memory; the PyTorch profiler for bottlenecks.
bash
nvidia-smi --query-gpu=name,utilization.gpu,memory.used,memory.total --format=csv -l 5

Choosing hardware as a student or small organisation#

OptionProsCons
Laptop CPU / Apple siliconFree, always availableSlow for deep learning; fine for classical ML and small models
Free cloud notebooks (e.g. Colab, Kaggle)Free GPUs for learningSession limits, variable availability
Consumer GPU workstationFull control; cost-effective for continuous useUpfront cost, power, maintenance
Cloud GPU instancesScale on demand; latest hardwareCan become expensive; must manage data security and shut down idle machines
University/HPC clustersPowerful, often free for studentsQueues, scheduler learning curve

Energy and environmental cost#

Training and serving large models consume significant electricity; the carbon footprint depends heavily on the energy mix of the data centre. Report compute and energy where possible (tools like CodeCarbon estimate emissions), choose efficient models, and avoid unnecessary large-scale runs — themes discussed further in the Ethics track.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

Docker and Containers for Machine Learning

"It works on my machine" is not a deployment strategy. We explain containers, write efficient Dockerfiles for ML training and serving, handle GPUs, and cover image size, security and orchestration basics.

Beginner⏱ 5 min#247
⚙️ MLOps & Engineering

Edge AI and TinyML: Running Models on Phones and Microcontrollers

Running models on-device brings privacy, offline operation, low latency and low cost. We cover the edge hardware spectrum, the optimisation pipeline, TensorFlow Lite, ONNX Runtime and TinyML on microcontrollers, and field-deployment lessons.

Intermediate⏱ 5 min#251
⚙️ MLOps & Engineering

Data Labelling and Annotation: Building High-Quality Datasets

Labels are the foundation of supervised learning, yet labelling is often rushed. We cover annotation guidelines, workflows and tools, measuring agreement, handling label noise, model-assisted labelling, and fair treatment of annotators.

Beginner⏱ 5 min#253