Training a model means adjusting parameters to reduce a loss. To know which way to adjust, we need to know how the loss changes when each parameter changes. That is exactly what derivatives measure. In this lecture we build the calculus you need for the rest of the course — especially the chain rule, which is the mathematical heart of backpropagation.
The derivative#
For $f: \mathbb{R} \to \mathbb{R}$, the derivative at $x$ is
It is the slope of the tangent line — the best linear approximation of $f$ near $x$:
Rules to know by heart: power rule $(x^n)' = nx^{n-1}$; $(e^x)' = e^x$; $(\ln x)' = 1/x$; product rule $(uv)' = u'v + uv'$; quotient rule; and the chain rule below.
Derivatives of ML's favourite functions:
| Function | Derivative |
|---|---|
| Sigmoid $\sigma(x) = \frac{1}{1 + e^{-x}}$ | $\sigma(x)(1 - \sigma(x))$ |
| $\tanh(x)$ | $1 - \tanh^2(x)$ |
| ReLU $\max(0, x)$ | $1$ if $x > 0$, $0$ if $x < 0$ |
| Softplus $\ln(1 + e^x)$ | $\sigma(x)$ |
Partial derivatives and the gradient#
For $f: \mathbb{R}^d \to \mathbb{R}$, the partial derivative $\partial f/\partial x_i$ measures change along coordinate $i$ holding the others fixed. Stacking them gives the gradient:
Why the gradient points uphill#
The directional derivative in unit direction $\mathbf{u}$ is $D_{\mathbf{u}}f = \nabla f^\top\mathbf{u} = \|\nabla f\|\cos\theta$. It is maximised when $\mathbf{u}$ points along $\nabla f$. Therefore:
- $\nabla f$ points in the direction of steepest ascent;
- $-\nabla f$ points in the direction of steepest descent;
- $\nabla f$ is perpendicular to the level sets (contours) of $f$.
This one fact justifies gradient descent: $\mathbf{w} \leftarrow \mathbf{w} - \eta\nabla L(\mathbf{w})$.
The chain rule#
If $y = f(u)$ and $u = g(x)$, then
For a long composition $L = f_4(f_3(f_2(f_1(x))))$ — which is what a neural network is — the derivative is a product of local derivatives:
Multivariable chain rule. If $L$ depends on $x$ through several intermediate variables $u_1, \dots, u_k$, we sum over paths:
Worked example: logistic loss#
Let $z = wx + b$, $p = \sigma(z)$, $L = -[y\ln p + (1 - y)\ln(1 - p)]$. By the chain rule:
Multiplying, most terms cancel:
This beautifully simple result — prediction minus target, times input — is why sigmoid pairs naturally with cross-entropy.
Taylor series#
Near a point $\mathbf{x}_0$:
where $\mathbf{H}$ is the Hessian matrix of second derivatives. The first-order term justifies gradient descent; the second-order term underlies Newton's method and explains why curvature limits the step size.
Checking gradients numerically#
import numpy as np
def loss(w, x, y):
p = 1 / (1 + np.exp(-(w @ x)))
return -(y * np.log(p) + (1 - y) * np.log(1 - p))
def analytic_grad(w, x, y):
p = 1 / (1 + np.exp(-(w @ x)))
return (p - y) * x
def numeric_grad(f, w, eps=1e-6):
g = np.zeros_like(w)
for i in range(len(w)):
e = np.zeros_like(w); e[i] = eps
g[i] = (f(w + e) - f(w - e)) / (2 * eps) # central difference
return g
rng = np.random.default_rng(1)
w, x, y = rng.normal(size=3), rng.normal(size=3), 1.0
a = analytic_grad(w, x, y)
n = numeric_grad(lambda v: loss(v, x, y), w)
print(a, n, "rel. error:", np.linalg.norm(a - n) / np.linalg.norm(a + n))