∑ Mathematics for ML · Lecture 7 of 25

Derivatives, Gradients and the Chain Rule

Learning is the art of nudging parameters in the right direction. We review derivatives, partial derivatives, gradients, directional derivatives, Taylor expansions and the chain rule that makes backpropagation possible.

Training a model means adjusting parameters to reduce a loss. To know which way to adjust, we need to know how the loss changes when each parameter changes. That is exactly what derivatives measure. In this lecture we build the calculus you need for the rest of the course — especially the chain rule, which is the mathematical heart of backpropagation.

The derivative#

For $f: \mathbb{R} \to \mathbb{R}$, the derivative at $x$ is

$$ f'(x) = \lim_{h \to 0} \frac{f(x + h) - f(x)}{h} $$

It is the slope of the tangent line — the best linear approximation of $f$ near $x$:

$$ f(x + h) \approx f(x) + f'(x)\,h $$

Rules to know by heart: power rule $(x^n)' = nx^{n-1}$; $(e^x)' = e^x$; $(\ln x)' = 1/x$; product rule $(uv)' = u'v + uv'$; quotient rule; and the chain rule below.

Derivatives of ML's favourite functions:

FunctionDerivative
Sigmoid $\sigma(x) = \frac{1}{1 + e^{-x}}$$\sigma(x)(1 - \sigma(x))$
$\tanh(x)$$1 - \tanh^2(x)$
ReLU $\max(0, x)$$1$ if $x > 0$, $0$ if $x < 0$
Softplus $\ln(1 + e^x)$$\sigma(x)$

Partial derivatives and the gradient#

For $f: \mathbb{R}^d \to \mathbb{R}$, the partial derivative $\partial f/\partial x_i$ measures change along coordinate $i$ holding the others fixed. Stacking them gives the gradient:

$$ \nabla f(\mathbf{x}) = \begin{bmatrix} \frac{\partial f}{\partial x_1} & \cdots & \frac{\partial f}{\partial x_d} \end{bmatrix}^\top $$

Why the gradient points uphill#

The directional derivative in unit direction $\mathbf{u}$ is $D_{\mathbf{u}}f = \nabla f^\top\mathbf{u} = \|\nabla f\|\cos\theta$. It is maximised when $\mathbf{u}$ points along $\nabla f$. Therefore:

  • $\nabla f$ points in the direction of steepest ascent;
  • $-\nabla f$ points in the direction of steepest descent;
  • $\nabla f$ is perpendicular to the level sets (contours) of $f$.

This one fact justifies gradient descent: $\mathbf{w} \leftarrow \mathbf{w} - \eta\nabla L(\mathbf{w})$.

The chain rule#

If $y = f(u)$ and $u = g(x)$, then

$$ \frac{dy}{dx} = \frac{dy}{du}\cdot\frac{du}{dx} $$

For a long composition $L = f_4(f_3(f_2(f_1(x))))$ — which is what a neural network is — the derivative is a product of local derivatives:

$$ \frac{dL}{dx} = f_4'(\cdot)\; f_3'(\cdot)\; f_2'(\cdot)\; f_1'(x) $$

Multivariable chain rule. If $L$ depends on $x$ through several intermediate variables $u_1, \dots, u_k$, we sum over paths:

$$ \frac{\partial L}{\partial x} = \sum_{j=1}^{k} \frac{\partial L}{\partial u_j}\,\frac{\partial u_j}{\partial x} $$

Worked example: logistic loss#

Let $z = wx + b$, $p = \sigma(z)$, $L = -[y\ln p + (1 - y)\ln(1 - p)]$. By the chain rule:

$$ \frac{\partial L}{\partial p} = -\frac{y}{p} + \frac{1 - y}{1 - p}, \qquad \frac{\partial p}{\partial z} = p(1 - p), \qquad \frac{\partial z}{\partial w} = x $$

Multiplying, most terms cancel:

$$ \frac{\partial L}{\partial w} = (p - y)\,x, \qquad \frac{\partial L}{\partial b} = p - y $$

This beautifully simple result — prediction minus target, times input — is why sigmoid pairs naturally with cross-entropy.

Taylor series#

Near a point $\mathbf{x}_0$:

$$ f(\mathbf{x}_0 + \boldsymbol{\delta}) \approx f(\mathbf{x}_0) + \nabla f(\mathbf{x}_0)^\top\boldsymbol{\delta} + \tfrac{1}{2}\boldsymbol{\delta}^\top\mathbf{H}(\mathbf{x}_0)\boldsymbol{\delta} $$

where $\mathbf{H}$ is the Hessian matrix of second derivatives. The first-order term justifies gradient descent; the second-order term underlies Newton's method and explains why curvature limits the step size.

Checking gradients numerically#

python
import numpy as np

def loss(w, x, y):
    p = 1 / (1 + np.exp(-(w @ x)))
    return -(y * np.log(p) + (1 - y) * np.log(1 - p))

def analytic_grad(w, x, y):
    p = 1 / (1 + np.exp(-(w @ x)))
    return (p - y) * x

def numeric_grad(f, w, eps=1e-6):
    g = np.zeros_like(w)
    for i in range(len(w)):
        e = np.zeros_like(w); e[i] = eps
        g[i] = (f(w + e) - f(w - e)) / (2 * eps)      # central difference
    return g

rng = np.random.default_rng(1)
w, x, y = rng.normal(size=3), rng.normal(size=3), 1.0
a = analytic_grad(w, x, y)
n = numeric_grad(lambda v: loss(v, x, y), w)
print(a, n, "rel. error:", np.linalg.norm(a - n) / np.linalg.norm(a + n))
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Matrix Calculus Essentials: Jacobians, Hessians and Vector Derivatives

Deep learning differentiates vectors with respect to matrices. We learn the Jacobian, Hessian, key vector-derivative identities, and the shape-checking discipline that makes backprop derivations painless.

Intermediate⏱ 5 min#032
∑ Mathematics for ML

Why Mathematics Matters for Machine Learning

You can call library functions without mathematics — until something breaks. We explain which branches of mathematics ML uses, why, and how to learn them efficiently.

Beginner⏱ 4 min#025
∑ Mathematics for ML

Norms, Inner Products, Projections and Distances

How long is a vector, and how similar are two vectors? We study L1, L2 and L∞ norms, dot products, cosine similarity, orthogonal projections and distance metrics used across ML.

Beginner⏱ 5 min#030