Mathematics for ML
Linear algebra, calculus, probability, statistics, information theory and optimisation.
- 01Why Mathematics Matters for Machine LearningYou can call library functions without mathematics โ until something breaks. We explain which branches of mathematics ML uses, why, and how to learn them efficiently.
- 02Vectors and Vector Spaces: The Language of DataEvery data point, word and image becomes a vector. We define vector spaces, linear combinations, span, independence, basis and dimension โ and see why embeddings live in them.
- 03Matrices and Linear TransformationsA matrix is not just a table of numbers โ it is a function that transforms space. We cover matrix multiplication four ways, rank, inverses, determinants and the fundamental subspaces.
- 04Eigenvalues and Eigenvectors: The Natural Axes of a TransformationEigenvectors are directions a matrix merely stretches. We derive the characteristic equation, diagonalisation and the spectral theorem, and connect them to PCA, PageRank, Markov chains and training stability.
- 05Singular Value Decomposition: The Swiss Army Knife of Linear AlgebraEvery matrix โ any shape, any rank โ factors as rotation, scaling, rotation. We derive the SVD, prove the EckartโYoung low-rank theorem, and apply it to compression, PCA, recommenders and least squares.
- 06Norms, Inner Products, Projections and DistancesHow long is a vector, and how similar are two vectors? We study L1, L2 and Lโ norms, dot products, cosine similarity, orthogonal projections and distance metrics used across ML.
- 07Derivatives, Gradients and the Chain RuleLearning is the art of nudging parameters in the right direction. We review derivatives, partial derivatives, gradients, directional derivatives, Taylor expansions and the chain rule that makes backpropagation possible.
- 08Matrix Calculus Essentials: Jacobians, Hessians and Vector DerivativesDeep learning differentiates vectors with respect to matrices. We learn the Jacobian, Hessian, key vector-derivative identities, and the shape-checking discipline that makes backprop derivations painless.
- 09Probability Fundamentals: Sample Spaces, Axioms and Conditional ProbabilityMachine learning is reasoning under uncertainty. We build probability from Kolmogorov's axioms, then master conditional probability, the product and sum rules, and independence.
- 10Random Variables and Probability DistributionsA random variable turns outcomes into numbers. We study discrete and continuous distributions โ Bernoulli, categorical, binomial, Poisson, uniform, exponential, Beta โ and when ML uses each.
- 11Bayes' Theorem: Updating Beliefs with EvidenceBayes' theorem is the mathematical rule for learning from evidence. We derive it, work through the famous medical-test example, and see how it underlies Naive Bayes, Bayesian inference and spam filters.
- 12Expectation, Variance, Covariance and CorrelationSummaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.
- 13The Gaussian Distribution: Why It Is EverywhereThe bell curve appears in noise models, weight initialisation, VAEs and diffusion models. We study univariate and multivariate Gaussians, the central limit theorem, and the closure properties that make Gaussians so convenient.
- 14Maximum Likelihood Estimation: How Models Learn from DataMost loss functions in ML are negative log-likelihoods in disguise. We define MLE, derive estimators for Bernoulli and Gaussian models, and prove that MSE and cross-entropy arise from maximum likelihood.
- 15MAP Estimation, Priors and the Bayesian View of RegularisationAdding a prior to maximum likelihood gives MAP estimation โ and reveals that L2 and L1 regularisation are Gaussian and Laplace priors. We also meet conjugate priors and full Bayesian inference.
- 16Information Theory for ML: Entropy, Cross-Entropy and KL DivergenceShannon's theory of information explains our loss functions. We derive entropy, cross-entropy, KL divergence and mutual information, and show why minimising cross-entropy is maximum likelihood.
- 17Convex Optimisation Basics: Why Some Problems Are EasyIn a convex problem every local minimum is global. We define convex sets and functions, learn practical tests for convexity, and see which ML models are convex and which are not.
- 18Gradient Descent: Theory, Step Sizes and ConvergenceThe simplest optimisation algorithm trains the largest models in the world. We analyse gradient descent, the role of the learning rate, stochastic gradients, and the convergence rates you should know.
- 19Lagrange Multipliers and Constrained Optimisation (KKT Conditions)Many ML problems impose constraints: margins, budgets, probabilities that sum to one. We derive Lagrange multipliers, the KKT conditions and duality โ the mathematics behind support vector machines.
- 20Statistical Hypothesis Testing for ML: Is Model A Really Better?A 1% accuracy gain may be noise. We cover confidence intervals, p-values, paired tests, McNemar's test, the bootstrap and multiple-comparison pitfalls so your experimental claims hold up.
- 21Sampling and Monte Carlo MethodsWhen integrals are intractable, we estimate them by sampling. We cover Monte Carlo estimation, inverse-transform and rejection sampling, importance sampling and Markov chain Monte Carlo.
- 22Markov Chains: Memoryless Processes and Stationary DistributionsMarkov chains model sequences where the future depends only on the present. We study transition matrices, stationary distributions, ergodicity and mixing, with applications from PageRank to MCMC and RL.
- 23Numerical Stability: Floating Point, Log-Sum-Exp and Avoiding NaNsMathematically correct code can still produce NaN. We study floating-point arithmetic, overflow and underflow, catastrophic cancellation, the log-sum-exp trick, stable softmax and mixed-precision pitfalls.
- 24The Curse of DimensionalityHigh-dimensional spaces behave strangely: volume hides in corners, distances concentrate and data becomes sparse. We quantify the curse, explain why ML still works, and survey the remedies.
- 25Tensors and Tensor Operations: Broadcasting, Reshaping and EinsumDeep learning code manipulates multi-dimensional arrays. We master tensor shapes, indexing, broadcasting, reshaping versus transposing, reductions and einsum โ the skills that prevent most deep-learning bugs.