โˆ‘

Mathematics for ML

Linear algebra, calculus, probability, statistics, information theory and optimisation.

  1. 01Why Mathematics Matters for Machine LearningYou can call library functions without mathematics โ€” until something breaks. We explain which branches of mathematics ML uses, why, and how to learn them efficiently.Beginner4 min
  2. 02Vectors and Vector Spaces: The Language of DataEvery data point, word and image becomes a vector. We define vector spaces, linear combinations, span, independence, basis and dimension โ€” and see why embeddings live in them.Beginner5 min
  3. 03Matrices and Linear TransformationsA matrix is not just a table of numbers โ€” it is a function that transforms space. We cover matrix multiplication four ways, rank, inverses, determinants and the fundamental subspaces.Beginner5 min
  4. 04Eigenvalues and Eigenvectors: The Natural Axes of a TransformationEigenvectors are directions a matrix merely stretches. We derive the characteristic equation, diagonalisation and the spectral theorem, and connect them to PCA, PageRank, Markov chains and training stability.Intermediate5 min
  5. 05Singular Value Decomposition: The Swiss Army Knife of Linear AlgebraEvery matrix โ€” any shape, any rank โ€” factors as rotation, scaling, rotation. We derive the SVD, prove the Eckartโ€“Young low-rank theorem, and apply it to compression, PCA, recommenders and least squares.Intermediate4 min
  6. 06Norms, Inner Products, Projections and DistancesHow long is a vector, and how similar are two vectors? We study L1, L2 and Lโˆž norms, dot products, cosine similarity, orthogonal projections and distance metrics used across ML.Beginner5 min
  7. 07Derivatives, Gradients and the Chain RuleLearning is the art of nudging parameters in the right direction. We review derivatives, partial derivatives, gradients, directional derivatives, Taylor expansions and the chain rule that makes backpropagation possible.Beginner5 min
  8. 08Matrix Calculus Essentials: Jacobians, Hessians and Vector DerivativesDeep learning differentiates vectors with respect to matrices. We learn the Jacobian, Hessian, key vector-derivative identities, and the shape-checking discipline that makes backprop derivations painless.Intermediate5 min
  9. 09Probability Fundamentals: Sample Spaces, Axioms and Conditional ProbabilityMachine learning is reasoning under uncertainty. We build probability from Kolmogorov's axioms, then master conditional probability, the product and sum rules, and independence.Beginner5 min
  10. 10Random Variables and Probability DistributionsA random variable turns outcomes into numbers. We study discrete and continuous distributions โ€” Bernoulli, categorical, binomial, Poisson, uniform, exponential, Beta โ€” and when ML uses each.Beginner5 min
  11. 11Bayes' Theorem: Updating Beliefs with EvidenceBayes' theorem is the mathematical rule for learning from evidence. We derive it, work through the famous medical-test example, and see how it underlies Naive Bayes, Bayesian inference and spam filters.Beginner5 min
  12. 12Expectation, Variance, Covariance and CorrelationSummaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.Beginner5 min
  13. 13The Gaussian Distribution: Why It Is EverywhereThe bell curve appears in noise models, weight initialisation, VAEs and diffusion models. We study univariate and multivariate Gaussians, the central limit theorem, and the closure properties that make Gaussians so convenient.Intermediate5 min
  14. 14Maximum Likelihood Estimation: How Models Learn from DataMost loss functions in ML are negative log-likelihoods in disguise. We define MLE, derive estimators for Bernoulli and Gaussian models, and prove that MSE and cross-entropy arise from maximum likelihood.Intermediate5 min
  15. 15MAP Estimation, Priors and the Bayesian View of RegularisationAdding a prior to maximum likelihood gives MAP estimation โ€” and reveals that L2 and L1 regularisation are Gaussian and Laplace priors. We also meet conjugate priors and full Bayesian inference.Intermediate5 min
  16. 16Information Theory for ML: Entropy, Cross-Entropy and KL DivergenceShannon's theory of information explains our loss functions. We derive entropy, cross-entropy, KL divergence and mutual information, and show why minimising cross-entropy is maximum likelihood.Intermediate5 min
  17. 17Convex Optimisation Basics: Why Some Problems Are EasyIn a convex problem every local minimum is global. We define convex sets and functions, learn practical tests for convexity, and see which ML models are convex and which are not.Intermediate5 min
  18. 18Gradient Descent: Theory, Step Sizes and ConvergenceThe simplest optimisation algorithm trains the largest models in the world. We analyse gradient descent, the role of the learning rate, stochastic gradients, and the convergence rates you should know.Intermediate5 min
  19. 19Lagrange Multipliers and Constrained Optimisation (KKT Conditions)Many ML problems impose constraints: margins, budgets, probabilities that sum to one. We derive Lagrange multipliers, the KKT conditions and duality โ€” the mathematics behind support vector machines.Advanced5 min
  20. 20Statistical Hypothesis Testing for ML: Is Model A Really Better?A 1% accuracy gain may be noise. We cover confidence intervals, p-values, paired tests, McNemar's test, the bootstrap and multiple-comparison pitfalls so your experimental claims hold up.Intermediate6 min
  21. 21Sampling and Monte Carlo MethodsWhen integrals are intractable, we estimate them by sampling. We cover Monte Carlo estimation, inverse-transform and rejection sampling, importance sampling and Markov chain Monte Carlo.Intermediate5 min
  22. 22Markov Chains: Memoryless Processes and Stationary DistributionsMarkov chains model sequences where the future depends only on the present. We study transition matrices, stationary distributions, ergodicity and mixing, with applications from PageRank to MCMC and RL.Intermediate5 min
  23. 23Numerical Stability: Floating Point, Log-Sum-Exp and Avoiding NaNsMathematically correct code can still produce NaN. We study floating-point arithmetic, overflow and underflow, catastrophic cancellation, the log-sum-exp trick, stable softmax and mixed-precision pitfalls.Intermediate5 min
  24. 24The Curse of DimensionalityHigh-dimensional spaces behave strangely: volume hides in corners, distances concentrate and data becomes sparse. We quantify the curse, explain why ML still works, and survey the remedies.Intermediate5 min
  25. 25Tensors and Tensor Operations: Broadcasting, Reshaping and EinsumDeep learning code manipulates multi-dimensional arrays. We master tensor shapes, indexing, broadcasting, reshaping versus transposing, reductions and einsum โ€” the skills that prevent most deep-learning bugs.Beginner6 min