∑ Mathematics for ML · Lecture 1 of 25

Why Mathematics Matters for Machine Learning

You can call library functions without mathematics — until something breaks. We explain which branches of mathematics ML uses, why, and how to learn them efficiently.

Every year a student asks me: "Sir, the libraries do everything. Why must I learn mathematics?" It is a fair question, and it deserves an honest answer. You can train a model with five lines of code. But the moment your model does not converge, overfits, produces NaNs, or behaves unfairly, those five lines offer no help. Mathematics is the language in which the reasons are written.

The four pillars#

Machine learning rests on four mathematical pillars:

PillarWhat it gives MLWhere you will meet it
Linear algebraLanguage for data and models: vectors, matrices, transformationsEvery model; embeddings; PCA; attention
CalculusHow outputs change with inputs; gradientsTraining by gradient descent; backpropagation
Probability & statisticsModelling uncertainty; learning from samplesLoss functions, Bayesian methods, evaluation
OptimisationFinding the best parametersEvery training algorithm

Information theory — entropy, cross-entropy, KL divergence — ties probability to learning and appears in almost every loss function.

One equation, all four pillars#

Consider logistic regression trained with gradient descent:

$$ \mathbf{w} \leftarrow \mathbf{w} - \eta \, \frac{1}{n} \sum_{i=1}^{n} \big(\sigma(\mathbf{w}^\top \mathbf{x}_i) - y_i\big)\, \mathbf{x}_i $$
  • $\mathbf{w}^\top \mathbf{x}_i$ is a dot product — linear algebra.
  • $\sigma$ turns a score into a probability — probability.
  • The term in the sum is the gradient of the cross-entropy loss — calculus and information theory.
  • The update rule is gradient descent — optimisation.

Every idea in this course is some elaboration of this pattern.

What mathematics lets you do#

  1. Debug. Exploding losses? You will recognise exploding gradients from the chain rule. Model predicts only one class? You will reason about class imbalance and decision thresholds.
  2. Choose. Should you use MSE or cross-entropy? L1 or L2 regularisation? Mathematics explains the consequences.
  3. Read research. Papers are written in mathematics. Without it you are limited to blog summaries.
  4. Invent. New methods come from understanding why old ones work.

How much do you need?#

You do not need to be a mathematician. You need fluency in a specific toolkit:

  • Linear algebra: vectors, matrices, matrix multiplication, transpose, inverse, rank, eigenvalues, SVD, norms, projections.
  • Calculus: derivatives, partial derivatives, the chain rule, gradients, Jacobians, a little Taylor expansion.
  • Probability: random variables, distributions (Bernoulli, categorical, Gaussian), expectation, variance, Bayes' rule, independence, maximum likelihood.
  • Optimisation: convexity, gradient descent and its variants, constrained optimisation basics.

Learn mathematics with code#

The best way for engineers to learn mathematics is to compute it. Verify every identity numerically:

python
import numpy as np

rng = np.random.default_rng(0)
A = rng.normal(size=(3, 3))
B = rng.normal(size=(3, 3))

# Identity: (AB)^T = B^T A^T
print(np.allclose((A @ B).T, B.T @ A.T))          # True

# Numerical derivative vs analytic derivative of f(x) = x^3
f = lambda x: x**3
x, h = 2.0, 1e-5
numeric = (f(x + h) - f(x - h)) / (2 * h)
print(numeric, 3 * x**2)                           # ~12.0, 12.0

Notation used in this course#

  • Scalars: lowercase italics $x, \eta$.
  • Vectors: bold lowercase $\mathbf{x}$ (column vectors by default).
  • Matrices: bold uppercase $\mathbf{X}, \mathbf{W}$.
  • $\mathbf{x}^\top$: transpose. $\|\mathbf{x}\|$: Euclidean norm.
  • $\nabla_{\mathbf{w}} L$: gradient of $L$ with respect to $\mathbf{w}$.
  • $\mathbb{E}[X]$: expectation. $p(x)$: probability density or mass.
  • Dataset: $\mathcal{D} = \{(\mathbf{x}_i, y_i)\}_{i=1}^n$ with $n$ examples of dimension $d$.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Vectors and Vector Spaces: The Language of Data

Every data point, word and image becomes a vector. We define vector spaces, linear combinations, span, independence, basis and dimension — and see why embeddings live in them.

Beginner⏱ 5 min#026
∑ Mathematics for ML

Matrices and Linear Transformations

A matrix is not just a table of numbers — it is a function that transforms space. We cover matrix multiplication four ways, rank, inverses, determinants and the fundamental subspaces.

Beginner⏱ 5 min#027
∑ Mathematics for ML

Eigenvalues and Eigenvectors: The Natural Axes of a Transformation

Eigenvectors are directions a matrix merely stretches. We derive the characteristic equation, diagonalisation and the spectral theorem, and connect them to PCA, PageRank, Markov chains and training stability.

Intermediate⏱ 5 min#028