🔗

Deep Learning

Neural networks, backpropagation, optimisers, normalisation, CNNs, RNNs and training at scale.

  1. 01From Biological Neurons to Artificial Neural NetworksWe open the Deep Learning track by tracing the path from biological neurons to artificial ones, defining a neural network precisely, and explaining why depth and learned representations changed AI.Beginner4 min
  2. 02The Perceptron: The First Learning MachineRosenblatt's perceptron learned to classify by correcting its mistakes. We derive its learning rule, prove the convergence theorem, reveal its XOR limitation, and see how it foreshadowed modern networks.Beginner5 min
  3. 03Multilayer Perceptrons and the Universal Approximation TheoremWith one hidden layer, a network can approximate any continuous function — so why go deep? We state the universal approximation theorem, build intuition with bumps, and explain the efficiency advantages of depth.Intermediate5 min
  4. 04Activation Functions: Sigmoid, Tanh, ReLU, GELU, SwiGLU and SoftmaxThe choice of non-linearity shapes how gradients flow and how networks learn. We compare the classical and modern activations, their derivatives and failure modes, and which to use where.Beginner5 min
  5. 05Loss Functions in Deep Learning: What Are We Really Optimising?The loss function defines what "good" means to a network. We survey regression, classification, ranking and representation-learning losses, their probabilistic meaning, and common implementation mistakes.Beginner5 min
  6. 06Backpropagation Derived Step by StepBackpropagation computes every gradient in a network at about the cost of one forward pass. We derive it for a two-layer network by hand, generalise to any depth, implement it in NumPy and verify it numerically.Intermediate6 min
  7. 07Computational Graphs and Automatic DifferentiationFrameworks compute gradients of arbitrary programs automatically. We compare symbolic, numerical and automatic differentiation, contrast forward and reverse mode, and build a tiny reverse-mode autodiff engine.Intermediate6 min
  8. 08Optimisers I: SGD, Momentum and Nesterov AccelerationPlain SGD zig-zags through ravines and crawls across plateaus. Momentum accumulates velocity to fix both. We derive heavy-ball and Nesterov momentum, analyse their effect on ill-conditioned problems, and give tuning advice.Intermediate5 min
  9. 09Optimisers II: AdaGrad, RMSProp, Adam and AdamWAdaptive optimisers give each parameter its own learning rate. We derive AdaGrad, RMSProp and Adam including bias correction, explain why AdamW decouples weight decay, and survey newer optimisers.Intermediate6 min
  10. 10Learning Rate Schedules, Warm-up and the LR Range TestThe learning rate is the most important hyperparameter, and it should change during training. We compare step, exponential, cosine and one-cycle schedules, explain why warm-up stabilises large models, and find good rates quickly.Intermediate5 min
  11. 11Weight Initialisation: Xavier, He and Why It MattersBad initial weights make signals explode or vanish before training even starts. We derive variance-preserving initialisation for tanh (Xavier) and ReLU (He) networks and discuss modern practice for deep and residual models.Intermediate4 min
  12. 12Vanishing and Exploding Gradients — Causes and CuresGradients are products of many Jacobians, so they can shrink or grow exponentially with depth. We analyse why, how to diagnose it, and the arsenal of fixes from ReLU and initialisation to residuals, normalisation, clipping and gating.Intermediate4 min
  13. 13Batch Normalisation: Faster, More Stable TrainingBatchNorm normalises each feature using mini-batch statistics, then rescales it with learned parameters. We derive the forward pass, explain training-versus-inference behaviour, debate why it works, and list its pitfalls.Intermediate5 min
  14. 14Beyond BatchNorm: Layer, Group, Instance and RMS NormalisationNormalisation layers differ only in which axes they average over — yet that choice decides where they work. We compare LayerNorm, GroupNorm, InstanceNorm and RMSNorm, and the pre-norm versus post-norm debate in transformers.Intermediate4 min
  15. 15Dropout: Regularisation by Random DeletionRandomly switching off neurons during training prevents co-adaptation and approximates an ensemble of exponentially many networks. We cover inverted dropout, where to apply it, its variants, and Monte Carlo dropout for uncertainty.Beginner5 min
  16. 16Regularisation in Deep Learning: Weight Decay, Early Stopping, Augmentation and MoreDeep networks can memorise anything, yet generalise well when regularised properly. We survey the toolkit — weight decay, early stopping, data augmentation, mixup and cutmix, label smoothing — and how to combine them.Intermediate5 min
  17. 17Convolutional Neural Networks: The Core IdeasConvolutions exploit the structure of images through local connectivity, weight sharing and translation equivariance. We define the convolution operation, count parameters, and build a CNN that learns hierarchical features.Beginner5 min
  18. 18Padding, Stride, Pooling and Receptive FieldsThe geometry of convolutional layers determines output sizes, computational cost and what each unit can see. We derive the output-size formula, compare pooling types, and compute receptive fields.Beginner5 min
  19. 19Recurrent Neural Networks: Modelling SequencesSequences need memory. RNNs carry a hidden state through time with shared weights. We define the vanilla RNN, unroll it, derive backpropagation through time, and see why long dependencies are hard.Intermediate5 min
  20. 20Long Short-Term Memory (LSTM): Gated Memory ExplainedLSTMs add a protected cell state and three gates that decide what to forget, write and reveal. We walk through the equations, explain why they preserve gradients, and apply them to sequence tasks.Intermediate5 min
  21. 21Gated Recurrent Units (GRU) and Choosing a Recurrent CellThe GRU simplifies the LSTM to two gates and one state while keeping long-term memory. We derive its equations, compare it with LSTMs empirically, and give practical guidance for recurrent models.Intermediate5 min
  22. 22Sequence-to-Sequence Models and the Encoder–Decoder FrameworkTranslation maps a sequence to another of different length. We build the encoder–decoder architecture, train it with teacher forcing, decode with greedy and beam search, and expose the bottleneck that motivated attention.Intermediate5 min
  23. 23The Attention Mechanism: Learning Where to LookAttention lets a model compute a weighted focus over all input positions for each output. We derive Bahdanau and Luong attention, generalise to queries, keys and values, and see why it became the foundation of transformers.Intermediate5 min
  24. 24Residual Connections: Why Very Deep Networks Became TrainableDeeper plain networks can train worse than shallower ones. Residual connections fix this by learning corrections to the identity. We explain the degradation problem, the gradient highway, and variants from ResNet to transformers.Intermediate5 min
  25. 25Autoencoders: Compression, Denoising and Representation LearningAn autoencoder learns to reconstruct its input through a bottleneck, discovering compact representations without labels. We cover undercomplete, denoising, sparse and convolutional autoencoders and their uses.Intermediate5 min
  26. 26Embeddings: Turning Discrete Things into Meaningful VectorsWords, users, products and categories become dense vectors whose geometry encodes meaning. We explain embedding layers, how embeddings are learned, how to measure similarity, and how they power search and recommendation.Beginner5 min
  27. 27Transfer Learning and Fine-TuningPretrained models let you achieve strong results with small datasets. We compare feature extraction and fine-tuning, explain discriminative learning rates and layer freezing, and discuss when transfer helps or hurts.Intermediate5 min
  28. 28PyTorch Fundamentals: Tensors, Autograd, Modules and the Training LoopA practical tour of PyTorch — tensors and devices, autograd, nn.Module, Dataset and DataLoader, optimisers, and a complete, correct training and evaluation loop you can reuse in every project.Beginner5 min
  29. 29TensorFlow and Keras FundamentalsKeras offers a high-level, productive API for deep learning. We build models with the Sequential and Functional APIs, train with fit and callbacks, write a custom training step, and export for deployment.Beginner4 min
  30. 30Debugging Neural Network Training: A Systematic RecipeNeural networks fail silently — they train, but badly. We present a systematic recipe for finding bugs, from data inspection and overfitting a single batch to monitoring activations, gradients and learning curves.Intermediate6 min
  31. 31Mixed-Precision Training and GPU EfficiencyTraining in 16-bit arithmetic roughly halves memory and can multiply throughput. We explain float16 and bfloat16, loss scaling, automatic mixed precision, and other practical techniques to make GPUs work harder.Advanced5 min
  32. 32Distributed Training: Data, Model, Pipeline and Sharded ParallelismLarge models and datasets need many accelerators. We explain data parallelism with all-reduce, sharded data parallelism (ZeRO/FSDP), tensor and pipeline model parallelism, and how they combine at scale.Advanced6 min
  33. 33Graph Neural Networks: Learning on Relational DataMolecules, social networks, road maps and knowledge graphs are graphs. GNNs learn from them by passing messages between neighbours. We derive message passing, GCN and GAT layers, and survey node, edge and graph-level tasks.Advanced6 min
  34. 34Knowledge Distillation: Teaching Small Models with Large OnesA large "teacher" model's soft predictions contain rich information that can train a much smaller "student". We derive the distillation loss with temperature, discuss dark knowledge, and survey feature and LLM distillation.Intermediate5 min
  35. 35Model Compression: Pruning and QuantisationNeural networks are highly redundant. We remove unnecessary weights with pruning, represent the rest with fewer bits via quantisation, and discuss the lottery ticket hypothesis and deployment on edge devices.Advanced6 min
  36. 36Neural Architecture Search and Automated Machine LearningCan algorithms design better networks than humans? We review search spaces, reinforcement-learning and evolutionary search, differentiable NAS, weight sharing and hardware-aware search — and the lessons of the NAS era.Advanced5 min
  37. 37Double Descent and the Generalisation Mystery of Deep LearningOver-parameterised networks can fit random labels yet generalise on real data, and test error can fall again beyond the interpolation threshold. We explore double descent, benign overfitting and implicit regularisation.Advanced6 min
  38. 38Loss Landscapes, Saddle Points and Flat MinimaWhat does the surface that SGD descends actually look like? We study critical points in high dimensions, visualise loss landscapes, discuss sharp versus flat minima, mode connectivity and why architecture shapes trainability.Advanced6 min