Welcome to the Deep Learning track. Over the next lectures we will build โ from first principles โ the technology behind modern image recognition, speech recognition, machine translation and generative AI. Let us begin where the field began: with the neuron.
The biological inspiration#
A biological neuron receives signals through thousands of dendrites, integrates them in the cell body, and โ if the combined input exceeds a threshold โ fires an electrical spike down its axon to other neurons via synapses. Learning in the brain involves changing synaptic strengths; Donald Hebb (1949) summarised it as "cells that fire together wire together". The human brain has roughly 86 billion neurons and on the order of $10^{14}$ synapses.
The artificial neuron#
An artificial neuron computes a weighted sum of its inputs plus a bias, then applies a non-linear activation function $\phi$:
With $\phi$ = step function, this is McCulloch and Pitts' 1943 threshold unit; with $\phi$ = sigmoid it is logistic regression. The weights $\mathbf{w}$ and bias $b$ are learned.
Networks of neurons#
A feedforward neural network (multilayer perceptron) arranges neurons in layers. Each layer computes
and the final layer produces the output (logits for classification, values for regression). The whole network is a composition of functions:
"Deep" learning simply means many layers.
Why non-linearity is essential#
Without activation functions, each layer is a linear map, and a composition of linear maps is linear: $\mathbf{W}^{(2)}\mathbf{W}^{(1)}\mathbf{x} = \mathbf{W}'\mathbf{x}$. A hundred linear layers are no more expressive than one. Non-linear activations let networks represent curved decision boundaries and complex functions.
Representation learning: the key idea#
Classical ML relies on hand-engineered features. Deep networks learn features automatically, layer by layer. In an image classifier:
- early layers detect edges and colour blobs;
- middle layers combine them into textures and parts (eyes, wheels);
- late layers represent whole objects.
Each layer transforms the data into a representation in which the next task is easier. Remember the phrase from the ML track: a neural network is "linear regression on learned features" โ the last layer is a linear model, and everything before it learns the features.
Why deep learning took off after 2010#
The ideas are decades old. Three ingredients arrived together:
- Data โ large labelled datasets like ImageNet (over a million images) and web-scale text.
- Compute โ GPUs made the massive matrix multiplications fast and affordable.
- Algorithms and tricks โ ReLU activations, good initialisation, dropout, batch normalisation, residual connections, the Adam optimiser.
Together they allowed networks with many layers to be trained reliably, and performance scaled with data and model size.
A first network in PyTorch#
import torch
import torch.nn as nn
model = nn.Sequential(
nn.Linear(2, 16), nn.ReLU(),
nn.Linear(16, 16), nn.ReLU(),
nn.Linear(16, 1), # a logit for binary classification
)
print(model)
print("parameters:", sum(p.numel() for p in model.parameters()))
x = torch.randn(4, 2) # a batch of 4 examples with 2 features
print(torch.sigmoid(model(x)).squeeze()) # predicted probabilities (untrained)This tiny network has 337 parameters. In the coming lectures we will learn how to train it โ computing gradients by backpropagation and updating weights by gradient descent.
What deep learning is good (and not so good) at#
Excels at: perception (images, audio), language, large-scale pattern recognition, generating content, learning from raw high-dimensional data.
Struggles with: small datasets without pretrained models, guarantees and verification, out-of-distribution robustness, explaining its decisions, and tasks where tabular gradient boosting is simpler and stronger.