Machine Learning
Regression, classification, trees, ensembles, clustering, evaluation and learning theory.
- 01What Is Machine Learning? The Learning Problem FormalisedTom Mitchell's definition, the formal learning problem, empirical risk minimisation and the central goal of generalisation — the conceptual foundation for every model in this track.
- 02Types of Machine Learning: Supervised, Unsupervised, Self-Supervised and ReinforcementLearning problems differ by the kind of feedback available. We map the major paradigms, their typical tasks and algorithms, and the hybrid settings — semi-supervised, weak and transfer learning — that dominate practice.
- 03The Machine Learning Workflow: From Problem to Deployed ModelSuccessful ML projects follow a disciplined process — problem framing, data collection, exploration, baselines, iteration, evaluation and deployment. We walk through it with a complete scikit-learn example.
- 04Linear Regression from First PrinciplesThe most important model in statistics and ML. We derive least squares geometrically and analytically, solve it with the normal equations and gradient descent, and interpret coefficients carefully.
- 05Polynomial Regression and Basis Functions: Non-Linearity with Linear ModelsLinear models can fit curves if we transform the inputs. We study polynomial, spline and radial basis features, watch overfitting happen as degree grows, and connect it to model selection.
- 06Regularisation: Ridge, Lasso and Elastic NetPenalising large weights tames overfitting. We derive ridge regression's closed form, explain why lasso yields sparse models, combine them in elastic net, and tune the penalty by cross-validation.
- 07Logistic Regression: Probabilistic Classification Done RightDespite its name, logistic regression is a classifier — and one of the most reliable. We derive it from log-odds, train it by maximum likelihood, interpret its coefficients and understand its decision boundary.
- 08Softmax Regression and Multiclass Classification StrategiesHow do we classify into more than two classes? We generalise logistic regression to softmax regression, derive its gradient, and compare it with one-vs-rest and one-vs-one strategies.
- 09The Bias–Variance Trade-off: Derivation and IntuitionWhy do simple models underfit and complex models overfit? We derive the bias–variance decomposition of expected squared error, visualise it, and discuss how modern deep learning complicates the classical picture.
- 10Overfitting and Underfitting: Diagnosis with Learning CurvesBefore fixing a model you must diagnose it. We learn to read learning curves and validation curves, recognise high bias and high variance, and choose the right remedy instead of guessing.
- 11Train, Validation and Test Splits — and Cross-Validation Done RightTrustworthy evaluation starts with correct data splits. We cover hold-out validation, k-fold, stratified, group and time-series cross-validation, nested CV for tuning, and the leakage traps that invalidate results.
- 12Evaluation Metrics for Classification: Accuracy, Precision, Recall and F1Accuracy can be dangerously misleading. We build the confusion matrix, define precision, recall, specificity, F-scores and balanced accuracy, and learn to choose metrics from the costs of errors.
- 13ROC Curves, AUC and Precision–Recall CurvesA classifier's score supports many thresholds. ROC and precision–recall curves summarise all of them. We construct both, interpret AUC probabilistically, and learn when each curve tells the truth.
- 14Regression Metrics: MSE, RMSE, MAE, R² and BeyondHow good is a numeric prediction? We compare MSE, RMSE, MAE, R², MAPE and quantile loss, explain how each responds to outliers and scale, and match metrics to real decisions.
- 15k-Nearest Neighbours: Learning by SimilarityThe simplest learning algorithm stores the data and asks the neighbours. We analyse k-NN's bias–variance behaviour, distance choices, scaling, efficient search structures and its surprising theoretical guarantees.
- 16Naive Bayes Classifiers: Simple, Fast and Surprisingly StrongNaive Bayes applies Bayes' theorem with a bold independence assumption. We derive Gaussian, multinomial and Bernoulli variants, explain why it works despite being "naive", and build a text classifier.
- 17Decision Trees: Splitting Criteria, Pruning and InterpretabilityDecision trees learn a flowchart of if–then questions. We derive Gini impurity and information gain, build a tree greedily, control overfitting with pruning, and see why trees are the building blocks of the best tabular models.
- 18Bagging and the Bootstrap: Variance Reduction by AveragingAveraging many noisy models trained on resampled data produces one stable model. We study the bootstrap, derive why bagging reduces variance, and use out-of-bag error as a free validation estimate.
- 19Random Forests: Decorrelated Trees and Robust PredictionsRandom forests add feature randomness to bagged trees, breaking their correlation. We explain the algorithm, its hyperparameters, OOB estimates, feature importance pitfalls and when forests are the right tool.
- 20Boosting I: AdaBoost and the Power of Weak LearnersCan many weak rules of thumb combine into a strong classifier? AdaBoost answered yes. We walk through the reweighting algorithm, derive it as exponential-loss minimisation, and discuss its margins and sensitivity to noise.
- 21Boosting II: Gradient Boosting MachinesGradient boosting performs gradient descent in function space, fitting each new tree to the negative gradient of the loss. We derive the algorithm, explain shrinkage and subsampling, and tune it properly.
- 22XGBoost, LightGBM and CatBoost: Modern Gradient Boosting LibrariesThree libraries turned gradient boosting into an industrial tool. We compare their key innovations — second-order optimisation, histogram splits, leaf-wise growth, GOSS, ordered target statistics — and show how to use each well.
- 23Support Vector Machines: Maximum-Margin ClassificationAmong all separating hyperplanes, SVMs pick the one with the widest margin. We derive the hard- and soft-margin formulations, the hinge loss, support vectors and the role of the C parameter.
- 24Kernel Methods and the Kernel TrickKernels let linear algorithms learn non-linear functions by computing inner products in high- or infinite-dimensional feature spaces — without ever visiting them. We study feature maps, Mercer's theorem, common kernels and their limits.
- 25k-Means Clustering: Algorithm, Objective and Pitfallsk-means partitions data into k groups by alternating assignment and update steps. We derive it as coordinate descent, cover k-means++ initialisation, choosing k, and the assumptions that make it fail.
- 26Hierarchical Clustering and DendrogramsHierarchical clustering builds a whole tree of nested clusters instead of a single partition. We compare linkage criteria, read dendrograms, and learn when a hierarchy is more useful than k flat clusters.
- 27DBSCAN and Density-Based ClusteringDBSCAN defines clusters as dense regions separated by sparse ones. It finds arbitrarily shaped clusters, labels outliers as noise and needs no k. We study core points, parameter selection and HDBSCAN.
- 28Gaussian Mixture Models and the EM AlgorithmGaussian mixtures model data as a blend of Gaussian components with soft cluster memberships. We derive the Expectation–Maximisation algorithm, prove it increases likelihood, and choose the number of components with BIC.
- 29Principal Component Analysis (PCA): Theory and PracticePCA finds the orthogonal directions of maximum variance. We derive it two ways — maximum variance and minimum reconstruction error — compute it via SVD, choose the number of components, and discuss its limits.
- 30t-SNE and UMAP: Visualising High-Dimensional DataNon-linear embeddings reveal cluster structure that PCA hides. We explain how t-SNE and UMAP work, what their hyperparameters do, and — critically — how not to misread their plots.
- 31Linear Discriminant Analysis: Supervised Dimensionality ReductionUnlike PCA, LDA uses labels to find projections that separate classes. We derive Fisher's criterion, the generative Gaussian view, and compare LDA with PCA, QDA and logistic regression.
- 32Feature Engineering: Turning Raw Data into SignalBetter features beat better algorithms. We survey the craft — transformations, interactions, aggregations, date and text features, domain-driven ratios — and how to engineer features without leaking the target.
- 33Feature Scaling and Normalisation: Standardisation, Min–Max and Robust ScalingMany algorithms silently assume features share a scale. We explain which models need scaling and why, compare standardisation, min–max, robust and quantile scaling, and show how to apply them without leakage.
- 34Handling Missing Data: Mechanisms, Imputation and IndicatorsMissing values are rarely random. We classify missingness as MCAR, MAR or MNAR, compare deletion and imputation strategies from simple to iterative, and show why a missingness indicator is often a feature in itself.
- 35Encoding Categorical Variables: One-Hot, Ordinal, Target and BeyondModels need numbers, but many features are categories. We compare one-hot, ordinal, frequency, target and hashing encoders, handle high cardinality and unseen categories, and avoid target-encoding leakage.
- 36Learning from Imbalanced DataWhen one class is rare, naive models ignore it. We cover the right metrics, class weighting, over- and under-sampling, SMOTE, threshold moving and calibration — and when each is appropriate.
- 37Feature Selection: Filter, Wrapper and Embedded MethodsMore features are not always better. We compare filter methods (correlation, mutual information), wrappers (RFE, sequential selection) and embedded methods (lasso, tree importance), and learn to select without leaking.
- 38Hyperparameter Tuning: Grid, Random, Bayesian and Early-Stopping MethodsHyperparameters control how models learn. We compare grid search, random search, Bayesian optimisation and successive halving, explain why random search beats grid search, and tune efficiently with Optuna.
- 39Anomaly Detection: Finding the UnusualFraud, faults and intrusions are rare and varied. We cover statistical, distance-, density- and tree-based detectors, autoencoders for complex data, and how to evaluate detectors when labels are scarce.
- 40Ensemble Learning III: Voting, Stacking and BlendingDifferent models make different mistakes. We combine them with hard and soft voting, weighted averaging, and stacking with a meta-learner trained on out-of-fold predictions — and discuss when the extra complexity pays off.
- 41Recommender Systems I: Content-Based and Collaborative FilteringRecommendation engines drive what billions of people watch, read and buy. We compare content-based and collaborative approaches, build user- and item-based neighbourhood models, and confront cold start and popularity bias.
- 42Recommender Systems II: Matrix Factorisation and Latent FactorsMatrix factorisation represents users and items as vectors in a shared latent space. We derive the regularised objective, train it with SGD and ALS, add biases and implicit feedback, and connect it to modern embedding models.
- 43Time Series Forecasting: Stationarity, ARIMA and EvaluationForecasting demand, rainfall or arrivals needs models that respect time. We decompose series into trend and seasonality, test for stationarity, build ARIMA and SARIMA models, and evaluate with rolling-origin backtesting.
- 44Semi-Supervised Learning: Learning from Few Labels and Many Unlabelled ExamplesLabels are expensive; unlabelled data is cheap. We study the assumptions that make unlabelled data useful and the main techniques — self-training, label propagation, consistency regularisation and FixMatch.
- 45Active Learning: Letting the Model Choose What to LabelIf labelling is expensive, label the most informative examples first. We cover pool-based active learning, uncertainty and diversity sampling, query-by-committee, and the practical pitfalls of real annotation loops.
- 46Learning Theory: PAC Learning, VC Dimension and Generalisation BoundsWhy should a model that fits training data work on new data? Learning theory answers precisely. We develop PAC learning, finite-class bounds, the VC dimension, and discuss what these bounds do and do not explain about deep learning.
- 47Association Rule Mining: Apriori, FP-Growth and Market-Basket AnalysisWhich items appear together? Association rule mining discovers patterns like "bread and butter imply milk". We define support, confidence and lift, derive the Apriori algorithm, and interpret rules responsibly.