Every year a student asks me: "Sir, the libraries do everything. Why must I learn mathematics?" It is a fair question, and it deserves an honest answer. You can train a model with five lines of code. But the moment your model does not converge, overfits, produces NaNs, or behaves unfairly, those five lines offer no help. Mathematics is the language in which the reasons are written.
The four pillars#
Machine learning rests on four mathematical pillars:
| Pillar | What it gives ML | Where you will meet it |
|---|---|---|
| Linear algebra | Language for data and models: vectors, matrices, transformations | Every model; embeddings; PCA; attention |
| Calculus | How outputs change with inputs; gradients | Training by gradient descent; backpropagation |
| Probability & statistics | Modelling uncertainty; learning from samples | Loss functions, Bayesian methods, evaluation |
| Optimisation | Finding the best parameters | Every training algorithm |
Information theory — entropy, cross-entropy, KL divergence — ties probability to learning and appears in almost every loss function.
One equation, all four pillars#
Consider logistic regression trained with gradient descent:
- $\mathbf{w}^\top \mathbf{x}_i$ is a dot product — linear algebra.
- $\sigma$ turns a score into a probability — probability.
- The term in the sum is the gradient of the cross-entropy loss — calculus and information theory.
- The update rule is gradient descent — optimisation.
Every idea in this course is some elaboration of this pattern.
What mathematics lets you do#
- Debug. Exploding losses? You will recognise exploding gradients from the chain rule. Model predicts only one class? You will reason about class imbalance and decision thresholds.
- Choose. Should you use MSE or cross-entropy? L1 or L2 regularisation? Mathematics explains the consequences.
- Read research. Papers are written in mathematics. Without it you are limited to blog summaries.
- Invent. New methods come from understanding why old ones work.
How much do you need?#
You do not need to be a mathematician. You need fluency in a specific toolkit:
- Linear algebra: vectors, matrices, matrix multiplication, transpose, inverse, rank, eigenvalues, SVD, norms, projections.
- Calculus: derivatives, partial derivatives, the chain rule, gradients, Jacobians, a little Taylor expansion.
- Probability: random variables, distributions (Bernoulli, categorical, Gaussian), expectation, variance, Bayes' rule, independence, maximum likelihood.
- Optimisation: convexity, gradient descent and its variants, constrained optimisation basics.
Learn mathematics with code#
The best way for engineers to learn mathematics is to compute it. Verify every identity numerically:
import numpy as np
rng = np.random.default_rng(0)
A = rng.normal(size=(3, 3))
B = rng.normal(size=(3, 3))
# Identity: (AB)^T = B^T A^T
print(np.allclose((A @ B).T, B.T @ A.T)) # True
# Numerical derivative vs analytic derivative of f(x) = x^3
f = lambda x: x**3
x, h = 2.0, 1e-5
numeric = (f(x + h) - f(x - h)) / (2 * h)
print(numeric, 3 * x**2) # ~12.0, 12.0Notation used in this course#
- Scalars: lowercase italics $x, \eta$.
- Vectors: bold lowercase $\mathbf{x}$ (column vectors by default).
- Matrices: bold uppercase $\mathbf{X}, \mathbf{W}$.
- $\mathbf{x}^\top$: transpose. $\|\mathbf{x}\|$: Euclidean norm.
- $\nabla_{\mathbf{w}} L$: gradient of $L$ with respect to $\mathbf{w}$.
- $\mathbb{E}[X]$: expectation. $p(x)$: probability density or mass.
- Dataset: $\mathcal{D} = \{(\mathbf{x}_i, y_i)\}_{i=1}^n$ with $n$ examples of dimension $d$.