What Is Artificial Intelligence? A Rigorous Introduction
We open the course by asking the hardest question first — what do we actually mean by "intelligence" in a machine? Four classical definitions, one working definition, and a map of the field.
Rigorous, intuitive lectures on Machine Learning, Deep Learning and AI — from the mathematics of gradient descent to transformers, diffusion models and reinforcement learning. Written by Janin A Apurba.
Search, logic, knowledge, uncertainty and the ideas that started the field.
24 lectures →02Linear algebra, calculus, probability, statistics, information theory and optimisation.
25 lectures →03Regression, classification, trees, ensembles, clustering, evaluation and learning theory.
47 lectures →04Neural networks, backpropagation, optimisers, normalisation, CNNs, RNNs and training at scale.
38 lectures →05From pixels to perception: classification, detection, segmentation, ViTs and 3D vision.
27 lectures →06Language models, embeddings, attention, BERT, GPT, speech and multilingual NLP.
29 lectures →07VAEs, GANs, diffusion, large language models, RAG, fine-tuning and AI agents.
30 lectures →08MDPs, dynamic programming, Q-learning, policy gradients, PPO and AlphaZero.
21 lectures →09Pipelines, reproducibility, serving, monitoring and shipping ML systems to production.
15 lectures →10Fairness, explainability, privacy, safety, regulation, humanitarian AI and your career.
17 lectures →We open the course by asking the hardest question first — what do we actually mean by "intelligence" in a machine? Four classical definitions, one working definition, and a map of the field.
Backpropagation computes every gradient in a network at about the cost of one forward pass. We derive it for a two-layer network by hand, generalise to any depth, implement it in NumPy and verify it numerically.
The 2017 Transformer replaced recurrence with attention and became the foundation of modern AI. We walk through embeddings, positional encoding, multi-head self-attention, feed-forward layers, residuals, normalisation, masking and the encoder–decoder design.
Diffusion models gradually add noise to data and train a network to reverse the process. We derive the DDPM forward and reverse processes, the simple noise-prediction loss, sampling, the score-based view, and faster samplers like DDIM.
TD methods for control learn action values and improve the policy on the fly. We derive on-policy SARSA and off-policy Q-learning, implement both on the cliff-walking problem, and explain why they learn different paths.
RAG connects an LLM to a searchable knowledge base so answers are current, specific and citable. We build the full pipeline — ingestion, chunking, embeddings, retrieval, re-ranking, prompting with citations — and evaluate and harden it.
We open the course by asking the hardest question first — what do we actually mean by "intelligence" in a machine? Four classical definitions, one working definition, and a map of the field.
Seventy years of ambition, disappointment and breakthroughs. Understanding the history of AI teaches you why today's methods look the way they do — and why humility is a scientific virtue.
Turing replaced "Can machines think?" with a game. We examine the imitation game, the Chinese Room argument, and what modern AI evaluation has learned from seventy years of debate.
The agent is the central abstraction of AI. We define agents formally, characterise environments along six dimensions, and study the architectures from simple reflex agents to learning agents.
Many AI problems reduce to finding a path in a huge graph. We formalise search problems and analyse the classic blind strategies for completeness, optimality, time and space.
Knowledge about where the goal lies transforms search. We derive A*, prove its optimality with admissible heuristics, and learn how to invent good heuristics by relaxing problems.
When an opponent is trying to defeat you, search must account for their choices. We derive minimax, prove alpha–beta pruning correct, and see how real game engines evaluate positions.
Timetabling, map colouring, Sudoku and circuit layout share a structure. CSPs exploit that structure with backtracking, variable-ordering heuristics and constraint propagation such as AC-3.
Logic gives an agent a language for knowledge and a mechanical way to draw conclusions. We cover syntax, truth tables, entailment, resolution and the SAT problem that powers modern solvers.
First-order logic lets us talk about objects and relations with quantifiers. We study its syntax and semantics, unification, generalised modus ponens, and resolution-based theorem proving.
How should an intelligent system store what it knows? We compare semantic networks, frames, description logics and modern knowledge graphs, and discuss the trade-off between expressiveness and tractability.
Expert systems were AI's first commercial success. We build a small rule engine, examine certainty factors from MYCIN, and ask why rule-based systems remain useful — and where they fail.
A Bayesian network encodes a joint probability distribution compactly using conditional independence. We learn the semantics, d-separation, exact inference by enumeration and variable elimination, and approximate sampling.
When the world changes over time and we only see noisy observations, Hidden Markov Models let us infer what is really happening. We derive the forward algorithm, Viterbi decoding and Baum–Welch learning.
Planning is search with structured, factored states. We represent actions with preconditions and effects, compare forward and backward planning, and see how domain-independent heuristics are derived automatically.
For decades AI was split between those who manipulate symbols and those who train networks. We compare the two paradigms honestly and examine how modern research tries to combine their strengths.
When only the final configuration matters, we can abandon paths and move through the space of complete solutions. Local search is simple, memory-light and the conceptual ancestor of gradient descent.
Nature optimises through selection, crossover and mutation. We implement a genetic algorithm from scratch, discuss the schema theorem and survey evolution strategies and neuroevolution.
Is 29°C "hot"? Fuzzy logic replaces true/false with degrees of membership. We build a Mamdani fuzzy controller step by step: fuzzification, rule evaluation, aggregation and defuzzification.
When the game tree is too vast and positions too hard to evaluate, MCTS builds an asymmetric tree guided by random simulations and the UCB1 bandit formula. We implement it and connect it to AlphaZero.
Rational agents must act under uncertainty. We develop utility theory from axioms, the principle of maximum expected utility, decision networks and the value of perfect information.
Ants find shortest paths and birds flock without a leader. We study how simple local rules produce intelligent collective behaviour and implement PSO and ACO for optimisation problems.
What would it mean for AI to be "general"? We define narrow and general intelligence, examine how to measure generality, and discuss the arguments about superintelligence with scientific care.
A guided map of today's AI ecosystem — the paradigms, the tools, the research frontiers and how the tracks of this course fit together — so you always know where you are in the journey.
You can call library functions without mathematics — until something breaks. We explain which branches of mathematics ML uses, why, and how to learn them efficiently.
Every data point, word and image becomes a vector. We define vector spaces, linear combinations, span, independence, basis and dimension — and see why embeddings live in them.
A matrix is not just a table of numbers — it is a function that transforms space. We cover matrix multiplication four ways, rank, inverses, determinants and the fundamental subspaces.
Eigenvectors are directions a matrix merely stretches. We derive the characteristic equation, diagonalisation and the spectral theorem, and connect them to PCA, PageRank, Markov chains and training stability.
Every matrix — any shape, any rank — factors as rotation, scaling, rotation. We derive the SVD, prove the Eckart–Young low-rank theorem, and apply it to compression, PCA, recommenders and least squares.
How long is a vector, and how similar are two vectors? We study L1, L2 and L∞ norms, dot products, cosine similarity, orthogonal projections and distance metrics used across ML.
Learning is the art of nudging parameters in the right direction. We review derivatives, partial derivatives, gradients, directional derivatives, Taylor expansions and the chain rule that makes backpropagation possible.
Deep learning differentiates vectors with respect to matrices. We learn the Jacobian, Hessian, key vector-derivative identities, and the shape-checking discipline that makes backprop derivations painless.
Machine learning is reasoning under uncertainty. We build probability from Kolmogorov's axioms, then master conditional probability, the product and sum rules, and independence.
A random variable turns outcomes into numbers. We study discrete and continuous distributions — Bernoulli, categorical, binomial, Poisson, uniform, exponential, Beta — and when ML uses each.
Bayes' theorem is the mathematical rule for learning from evidence. We derive it, work through the famous medical-test example, and see how it underlies Naive Bayes, Bayesian inference and spam filters.
Summaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.
The bell curve appears in noise models, weight initialisation, VAEs and diffusion models. We study univariate and multivariate Gaussians, the central limit theorem, and the closure properties that make Gaussians so convenient.
Most loss functions in ML are negative log-likelihoods in disguise. We define MLE, derive estimators for Bernoulli and Gaussian models, and prove that MSE and cross-entropy arise from maximum likelihood.
Adding a prior to maximum likelihood gives MAP estimation — and reveals that L2 and L1 regularisation are Gaussian and Laplace priors. We also meet conjugate priors and full Bayesian inference.
Shannon's theory of information explains our loss functions. We derive entropy, cross-entropy, KL divergence and mutual information, and show why minimising cross-entropy is maximum likelihood.
In a convex problem every local minimum is global. We define convex sets and functions, learn practical tests for convexity, and see which ML models are convex and which are not.
The simplest optimisation algorithm trains the largest models in the world. We analyse gradient descent, the role of the learning rate, stochastic gradients, and the convergence rates you should know.
Many ML problems impose constraints: margins, budgets, probabilities that sum to one. We derive Lagrange multipliers, the KKT conditions and duality — the mathematics behind support vector machines.
A 1% accuracy gain may be noise. We cover confidence intervals, p-values, paired tests, McNemar's test, the bootstrap and multiple-comparison pitfalls so your experimental claims hold up.
When integrals are intractable, we estimate them by sampling. We cover Monte Carlo estimation, inverse-transform and rejection sampling, importance sampling and Markov chain Monte Carlo.
Markov chains model sequences where the future depends only on the present. We study transition matrices, stationary distributions, ergodicity and mixing, with applications from PageRank to MCMC and RL.
Mathematically correct code can still produce NaN. We study floating-point arithmetic, overflow and underflow, catastrophic cancellation, the log-sum-exp trick, stable softmax and mixed-precision pitfalls.
High-dimensional spaces behave strangely: volume hides in corners, distances concentrate and data becomes sparse. We quantify the curse, explain why ML still works, and survey the remedies.
Deep learning code manipulates multi-dimensional arrays. We master tensor shapes, indexing, broadcasting, reshaping versus transposing, reductions and einsum — the skills that prevent most deep-learning bugs.
Tom Mitchell's definition, the formal learning problem, empirical risk minimisation and the central goal of generalisation — the conceptual foundation for every model in this track.
Learning problems differ by the kind of feedback available. We map the major paradigms, their typical tasks and algorithms, and the hybrid settings — semi-supervised, weak and transfer learning — that dominate practice.
Successful ML projects follow a disciplined process — problem framing, data collection, exploration, baselines, iteration, evaluation and deployment. We walk through it with a complete scikit-learn example.
The most important model in statistics and ML. We derive least squares geometrically and analytically, solve it with the normal equations and gradient descent, and interpret coefficients carefully.
Linear models can fit curves if we transform the inputs. We study polynomial, spline and radial basis features, watch overfitting happen as degree grows, and connect it to model selection.
Penalising large weights tames overfitting. We derive ridge regression's closed form, explain why lasso yields sparse models, combine them in elastic net, and tune the penalty by cross-validation.
Despite its name, logistic regression is a classifier — and one of the most reliable. We derive it from log-odds, train it by maximum likelihood, interpret its coefficients and understand its decision boundary.
How do we classify into more than two classes? We generalise logistic regression to softmax regression, derive its gradient, and compare it with one-vs-rest and one-vs-one strategies.
Why do simple models underfit and complex models overfit? We derive the bias–variance decomposition of expected squared error, visualise it, and discuss how modern deep learning complicates the classical picture.
Before fixing a model you must diagnose it. We learn to read learning curves and validation curves, recognise high bias and high variance, and choose the right remedy instead of guessing.
Trustworthy evaluation starts with correct data splits. We cover hold-out validation, k-fold, stratified, group and time-series cross-validation, nested CV for tuning, and the leakage traps that invalidate results.
Accuracy can be dangerously misleading. We build the confusion matrix, define precision, recall, specificity, F-scores and balanced accuracy, and learn to choose metrics from the costs of errors.
A classifier's score supports many thresholds. ROC and precision–recall curves summarise all of them. We construct both, interpret AUC probabilistically, and learn when each curve tells the truth.
How good is a numeric prediction? We compare MSE, RMSE, MAE, R², MAPE and quantile loss, explain how each responds to outliers and scale, and match metrics to real decisions.
The simplest learning algorithm stores the data and asks the neighbours. We analyse k-NN's bias–variance behaviour, distance choices, scaling, efficient search structures and its surprising theoretical guarantees.
Naive Bayes applies Bayes' theorem with a bold independence assumption. We derive Gaussian, multinomial and Bernoulli variants, explain why it works despite being "naive", and build a text classifier.
Decision trees learn a flowchart of if–then questions. We derive Gini impurity and information gain, build a tree greedily, control overfitting with pruning, and see why trees are the building blocks of the best tabular models.
Averaging many noisy models trained on resampled data produces one stable model. We study the bootstrap, derive why bagging reduces variance, and use out-of-bag error as a free validation estimate.
Random forests add feature randomness to bagged trees, breaking their correlation. We explain the algorithm, its hyperparameters, OOB estimates, feature importance pitfalls and when forests are the right tool.
Can many weak rules of thumb combine into a strong classifier? AdaBoost answered yes. We walk through the reweighting algorithm, derive it as exponential-loss minimisation, and discuss its margins and sensitivity to noise.
Gradient boosting performs gradient descent in function space, fitting each new tree to the negative gradient of the loss. We derive the algorithm, explain shrinkage and subsampling, and tune it properly.
Three libraries turned gradient boosting into an industrial tool. We compare their key innovations — second-order optimisation, histogram splits, leaf-wise growth, GOSS, ordered target statistics — and show how to use each well.
Among all separating hyperplanes, SVMs pick the one with the widest margin. We derive the hard- and soft-margin formulations, the hinge loss, support vectors and the role of the C parameter.
Kernels let linear algorithms learn non-linear functions by computing inner products in high- or infinite-dimensional feature spaces — without ever visiting them. We study feature maps, Mercer's theorem, common kernels and their limits.
k-means partitions data into k groups by alternating assignment and update steps. We derive it as coordinate descent, cover k-means++ initialisation, choosing k, and the assumptions that make it fail.
Hierarchical clustering builds a whole tree of nested clusters instead of a single partition. We compare linkage criteria, read dendrograms, and learn when a hierarchy is more useful than k flat clusters.
DBSCAN defines clusters as dense regions separated by sparse ones. It finds arbitrarily shaped clusters, labels outliers as noise and needs no k. We study core points, parameter selection and HDBSCAN.
Gaussian mixtures model data as a blend of Gaussian components with soft cluster memberships. We derive the Expectation–Maximisation algorithm, prove it increases likelihood, and choose the number of components with BIC.
PCA finds the orthogonal directions of maximum variance. We derive it two ways — maximum variance and minimum reconstruction error — compute it via SVD, choose the number of components, and discuss its limits.
Non-linear embeddings reveal cluster structure that PCA hides. We explain how t-SNE and UMAP work, what their hyperparameters do, and — critically — how not to misread their plots.
Unlike PCA, LDA uses labels to find projections that separate classes. We derive Fisher's criterion, the generative Gaussian view, and compare LDA with PCA, QDA and logistic regression.
Better features beat better algorithms. We survey the craft — transformations, interactions, aggregations, date and text features, domain-driven ratios — and how to engineer features without leaking the target.
Many algorithms silently assume features share a scale. We explain which models need scaling and why, compare standardisation, min–max, robust and quantile scaling, and show how to apply them without leakage.
Missing values are rarely random. We classify missingness as MCAR, MAR or MNAR, compare deletion and imputation strategies from simple to iterative, and show why a missingness indicator is often a feature in itself.
Models need numbers, but many features are categories. We compare one-hot, ordinal, frequency, target and hashing encoders, handle high cardinality and unseen categories, and avoid target-encoding leakage.
When one class is rare, naive models ignore it. We cover the right metrics, class weighting, over- and under-sampling, SMOTE, threshold moving and calibration — and when each is appropriate.
More features are not always better. We compare filter methods (correlation, mutual information), wrappers (RFE, sequential selection) and embedded methods (lasso, tree importance), and learn to select without leaking.
Hyperparameters control how models learn. We compare grid search, random search, Bayesian optimisation and successive halving, explain why random search beats grid search, and tune efficiently with Optuna.
Fraud, faults and intrusions are rare and varied. We cover statistical, distance-, density- and tree-based detectors, autoencoders for complex data, and how to evaluate detectors when labels are scarce.
Different models make different mistakes. We combine them with hard and soft voting, weighted averaging, and stacking with a meta-learner trained on out-of-fold predictions — and discuss when the extra complexity pays off.
Recommendation engines drive what billions of people watch, read and buy. We compare content-based and collaborative approaches, build user- and item-based neighbourhood models, and confront cold start and popularity bias.
Matrix factorisation represents users and items as vectors in a shared latent space. We derive the regularised objective, train it with SGD and ALS, add biases and implicit feedback, and connect it to modern embedding models.
Forecasting demand, rainfall or arrivals needs models that respect time. We decompose series into trend and seasonality, test for stationarity, build ARIMA and SARIMA models, and evaluate with rolling-origin backtesting.
Labels are expensive; unlabelled data is cheap. We study the assumptions that make unlabelled data useful and the main techniques — self-training, label propagation, consistency regularisation and FixMatch.
If labelling is expensive, label the most informative examples first. We cover pool-based active learning, uncertainty and diversity sampling, query-by-committee, and the practical pitfalls of real annotation loops.
Why should a model that fits training data work on new data? Learning theory answers precisely. We develop PAC learning, finite-class bounds, the VC dimension, and discuss what these bounds do and do not explain about deep learning.
Which items appear together? Association rule mining discovers patterns like "bread and butter imply milk". We define support, confidence and lift, derive the Apriori algorithm, and interpret rules responsibly.
We open the Deep Learning track by tracing the path from biological neurons to artificial ones, defining a neural network precisely, and explaining why depth and learned representations changed AI.
Rosenblatt's perceptron learned to classify by correcting its mistakes. We derive its learning rule, prove the convergence theorem, reveal its XOR limitation, and see how it foreshadowed modern networks.
With one hidden layer, a network can approximate any continuous function — so why go deep? We state the universal approximation theorem, build intuition with bumps, and explain the efficiency advantages of depth.
The choice of non-linearity shapes how gradients flow and how networks learn. We compare the classical and modern activations, their derivatives and failure modes, and which to use where.
The loss function defines what "good" means to a network. We survey regression, classification, ranking and representation-learning losses, their probabilistic meaning, and common implementation mistakes.
Backpropagation computes every gradient in a network at about the cost of one forward pass. We derive it for a two-layer network by hand, generalise to any depth, implement it in NumPy and verify it numerically.
Frameworks compute gradients of arbitrary programs automatically. We compare symbolic, numerical and automatic differentiation, contrast forward and reverse mode, and build a tiny reverse-mode autodiff engine.
Plain SGD zig-zags through ravines and crawls across plateaus. Momentum accumulates velocity to fix both. We derive heavy-ball and Nesterov momentum, analyse their effect on ill-conditioned problems, and give tuning advice.
Adaptive optimisers give each parameter its own learning rate. We derive AdaGrad, RMSProp and Adam including bias correction, explain why AdamW decouples weight decay, and survey newer optimisers.
The learning rate is the most important hyperparameter, and it should change during training. We compare step, exponential, cosine and one-cycle schedules, explain why warm-up stabilises large models, and find good rates quickly.
Bad initial weights make signals explode or vanish before training even starts. We derive variance-preserving initialisation for tanh (Xavier) and ReLU (He) networks and discuss modern practice for deep and residual models.
Gradients are products of many Jacobians, so they can shrink or grow exponentially with depth. We analyse why, how to diagnose it, and the arsenal of fixes from ReLU and initialisation to residuals, normalisation, clipping and gating.
BatchNorm normalises each feature using mini-batch statistics, then rescales it with learned parameters. We derive the forward pass, explain training-versus-inference behaviour, debate why it works, and list its pitfalls.
Normalisation layers differ only in which axes they average over — yet that choice decides where they work. We compare LayerNorm, GroupNorm, InstanceNorm and RMSNorm, and the pre-norm versus post-norm debate in transformers.
Randomly switching off neurons during training prevents co-adaptation and approximates an ensemble of exponentially many networks. We cover inverted dropout, where to apply it, its variants, and Monte Carlo dropout for uncertainty.
Deep networks can memorise anything, yet generalise well when regularised properly. We survey the toolkit — weight decay, early stopping, data augmentation, mixup and cutmix, label smoothing — and how to combine them.
Convolutions exploit the structure of images through local connectivity, weight sharing and translation equivariance. We define the convolution operation, count parameters, and build a CNN that learns hierarchical features.
The geometry of convolutional layers determines output sizes, computational cost and what each unit can see. We derive the output-size formula, compare pooling types, and compute receptive fields.
Sequences need memory. RNNs carry a hidden state through time with shared weights. We define the vanilla RNN, unroll it, derive backpropagation through time, and see why long dependencies are hard.
LSTMs add a protected cell state and three gates that decide what to forget, write and reveal. We walk through the equations, explain why they preserve gradients, and apply them to sequence tasks.
The GRU simplifies the LSTM to two gates and one state while keeping long-term memory. We derive its equations, compare it with LSTMs empirically, and give practical guidance for recurrent models.
Translation maps a sequence to another of different length. We build the encoder–decoder architecture, train it with teacher forcing, decode with greedy and beam search, and expose the bottleneck that motivated attention.
Attention lets a model compute a weighted focus over all input positions for each output. We derive Bahdanau and Luong attention, generalise to queries, keys and values, and see why it became the foundation of transformers.
Deeper plain networks can train worse than shallower ones. Residual connections fix this by learning corrections to the identity. We explain the degradation problem, the gradient highway, and variants from ResNet to transformers.
An autoencoder learns to reconstruct its input through a bottleneck, discovering compact representations without labels. We cover undercomplete, denoising, sparse and convolutional autoencoders and their uses.
Words, users, products and categories become dense vectors whose geometry encodes meaning. We explain embedding layers, how embeddings are learned, how to measure similarity, and how they power search and recommendation.
Pretrained models let you achieve strong results with small datasets. We compare feature extraction and fine-tuning, explain discriminative learning rates and layer freezing, and discuss when transfer helps or hurts.
A practical tour of PyTorch — tensors and devices, autograd, nn.Module, Dataset and DataLoader, optimisers, and a complete, correct training and evaluation loop you can reuse in every project.
Keras offers a high-level, productive API for deep learning. We build models with the Sequential and Functional APIs, train with fit and callbacks, write a custom training step, and export for deployment.
Neural networks fail silently — they train, but badly. We present a systematic recipe for finding bugs, from data inspection and overfitting a single batch to monitoring activations, gradients and learning curves.
Training in 16-bit arithmetic roughly halves memory and can multiply throughput. We explain float16 and bfloat16, loss scaling, automatic mixed precision, and other practical techniques to make GPUs work harder.
Large models and datasets need many accelerators. We explain data parallelism with all-reduce, sharded data parallelism (ZeRO/FSDP), tensor and pipeline model parallelism, and how they combine at scale.
Molecules, social networks, road maps and knowledge graphs are graphs. GNNs learn from them by passing messages between neighbours. We derive message passing, GCN and GAT layers, and survey node, edge and graph-level tasks.
A large "teacher" model's soft predictions contain rich information that can train a much smaller "student". We derive the distillation loss with temperature, discuss dark knowledge, and survey feature and LLM distillation.
Neural networks are highly redundant. We remove unnecessary weights with pruning, represent the rest with fewer bits via quantisation, and discuss the lottery ticket hypothesis and deployment on edge devices.
Can algorithms design better networks than humans? We review search spaces, reinforcement-learning and evolutionary search, differentiable NAS, weight sharing and hardware-aware search — and the lessons of the NAS era.
Over-parameterised networks can fit random labels yet generalise on real data, and test error can fall again beyond the interpolation threshold. We explore double descent, benign overfitting and implicit regularisation.
What does the surface that SGD descends actually look like? We study critical points in high dimensions, visualise loss landscapes, discuss sharp versus flat minima, mode connectivity and why architecture shapes trainability.
We open the Computer Vision track by asking how a machine can see. We cover how images are represented, why vision is hard, the landscape of vision tasks, and how deep learning transformed the field.
Before deep learning, images were processed with hand-designed filters. We study histograms, blurring, sharpening, gradients, the Sobel and Canny edge detectors, and morphological operations — the vocabulary CNNs later learned.
Before CNNs, vision relied on carefully engineered features. We study corner detection, SIFT keypoints and descriptors, HOG for pedestrian detection and the bag-of-visual-words model — ideas still used in geometry and robotics.
Two architectures bookend the rise of CNNs. LeNet-5 read handwritten digits in the 1990s; AlexNet won ImageNet in 2012 and launched the deep learning era. We dissect both and the innovations that made AlexNet work.
In 2014 two architectures pushed CNNs deeper in opposite styles: VGG with uniform stacks of 3×3 convolutions, GoogLeNet with parallel multi-scale Inception modules and 1×1 bottlenecks. We compare their designs and lessons.
ResNet's residual blocks enabled 152-layer networks and became the default vision backbone. We study basic and bottleneck blocks, the full ResNet-50 layout, training recipe, and descendants such as ResNeXt and ConvNeXt.
How should a network grow when you have more compute — deeper, wider or higher resolution? EfficientNet's compound scaling answers "all three, in balance". We cover MBConv blocks, the B0–B7 family and EfficientNetV2.
Phones, drones and microcontrollers need vision models that are small and fast. We derive the cost savings of depthwise separable convolutions and study MobileNet V1–V3, ShuffleNet and design principles for efficient inference.
A practical walkthrough of a real image classifier — collecting and splitting data, preprocessing, choosing a pretrained backbone, training, evaluating per class, inspecting errors and exporting the model.
Augmentation multiplies your data by encoding known invariances. We survey geometric, photometric and occlusion augmentations, automated policies like RandAugment, mixing methods, and augmentation for detection and segmentation.
Detection asks what objects are in an image and where. We define bounding boxes, IoU and mAP, then trace the two-stage R-CNN family from selective search to region proposal networks and feature pyramids.
"You Only Look Once" reframed detection as a single regression problem, enabling real-time performance. We study the original YOLO grid formulation, its evolution, anchor-free heads, and practical training with modern tools.
One-stage detectors face an extreme imbalance between background and objects. We study SSD's multi-scale default boxes, then derive RetinaNet's focal loss, which let one-stage detectors match two-stage accuracy.
Segmentation labels every pixel. We cover fully convolutional networks, the encoder–decoder U-Net with skip connections, DeepLab's atrous convolutions, loss functions like Dice, and evaluation with IoU.
Instance segmentation separates each individual object with its own mask. We study Mask R-CNN's mask branch and RoIAlign, compare instance, semantic and panoptic segmentation, and survey query-based models like Mask2Former.
Transformers conquered language, then vision. We dissect ViT's patch embeddings, class token and positional encodings, compare inductive biases with CNNs, and survey DeiT, Swin and hierarchical designs.
Labels are expensive, images are abundant. Self-supervised methods learn visual representations from unlabelled images through contrastive learning, self-distillation or masked reconstruction. We compare the main families and how to use them.
CLIP learns a shared embedding space for images and text from hundreds of millions of image–caption pairs. We explain its contrastive training, zero-shot classification with prompts, retrieval, limitations and its role in generative models.
Pose estimation locates body joints in images and video. We cover keypoint heatmap regression, top-down versus bottom-up approaches, part affinity fields, evaluation with OKS, 3-D pose, and applications from health to sport.
Face recognition maps faces to embeddings where the same person is close. We study the pipeline, triplet and angular-margin losses, verification versus identification, evaluation — and the serious ethical questions this technology raises.
Video adds time to vision. We study optical flow, two-stream networks, 3-D convolutions, (2+1)-D factorisation, video transformers and tracking, plus the computational tricks that make video models practical.
Images are 2-D projections of a 3-D world. We cover camera geometry, stereo and monocular depth, structure from motion, point-cloud networks like PointNet, and neural scene representations such as NeRF and Gaussian splatting.
Much of the world's information is locked in scanned forms, receipts and handwritten records. We cover text detection, recognition with CRNN and CTC, layout-aware models, multilingual challenges and end-to-end document understanding.
Deep learning can detect disease in X-rays, retinal scans and pathology slides. We survey modalities and tasks, discuss data and labelling challenges, shortcut learning, rigorous clinical validation, and deployment responsibilities.
Vision is following language towards general-purpose foundation models. We study the Segment Anything Model's promptable design and data engine, open-vocabulary detection, and how foundation models change vision workflows.
Which pixels made the model decide? We study gradient saliency, Grad-CAM, integrated gradients and occlusion, show how they reveal shortcuts, and discuss sanity checks that expose unreliable explanations.
Imperceptible perturbations can make a network confidently wrong. We derive FGSM and PGD attacks, explain why adversarial examples exist, cover physical and black-box attacks, and evaluate defences including adversarial training.
We open the NLP track with the question of why human language is so hard for machines — ambiguity, context, compositionality — and map the tasks, eras and methods of natural language processing.
Raw text must be converted into units a model can process. We cover normalisation, word and sentence tokenisation, stop words, stemming versus lemmatisation, and how preprocessing needs differ for classical and neural models.
The simplest way to turn documents into vectors is to count words. We build bag-of-words and TF-IDF representations, derive the IDF formula, use cosine similarity for retrieval, and train strong linear text classifiers.
A language model assigns probabilities to sequences of words. We derive n-gram models from the chain rule and Markov assumption, fix zero probabilities with smoothing, generate text, and evaluate with perplexity.
Word2Vec learns dense word vectors by predicting context words. We derive the skip-gram and CBOW objectives, negative sampling, and explore analogies, similarity and the limitations of static embeddings.
GloVe learns embeddings from a global co-occurrence matrix with a weighted least-squares objective; FastText represents words as bags of character n-grams, handling rare and unseen words. We compare them with word2vec and discuss evaluation.
Modern language models split text into subword units. We derive byte-pair encoding step by step, compare WordPiece and Unigram LM tokenisation, discuss byte-level BPE, and examine how tokenisation affects multilingual fairness and cost.
Text classification is the most widely deployed NLP task. We compare TF-IDF baselines, CNN and RNN classifiers, and fine-tuned transformers, and cover label design, imbalance, multilingual data and evaluation.
Sentiment analysis detects opinions and emotions in text. We compare lexicon-based and learned approaches, tackle negation and sarcasm, extend to aspect-based sentiment and emotions, and discuss evaluation and misuse.
NER locates and classifies entity mentions in text. We cover BIO tagging, feature-based and neural approaches, transformer token classification with subword alignment, entity-level evaluation, and domain adaptation.
POS tagging assigns grammatical categories to words. We compare HMM taggers with discriminative Conditional Random Fields, derive the CRF likelihood and Viterbi decoding, and see why CRF layers still appear on top of neural encoders.
Machine translation is one of NLP's oldest and most impactful tasks. We trace its evolution to neural systems, cover training data, subword vocabularies, back-translation, evaluation with BLEU and COMET, and the challenges of low-resource languages.
The 2017 Transformer replaced recurrence with attention and became the foundation of modern AI. We walk through embeddings, positional encoding, multi-head self-attention, feed-forward layers, residuals, normalisation, masking and the encoder–decoder design.
A deeper look at self-attention — what attention heads learn, the geometry of queries and keys, computational complexity, causal masking, KV caching, and efficient variants such as multi-query and grouped-query attention.
Attention is order-blind, so transformers need positional information. We compare absolute sinusoidal and learned encodings with relative methods — rotary embeddings (RoPE) and ALiBi — and discuss extending context length.
BERT showed that a bidirectional transformer pretrained on unlabelled text could be fine-tuned to beat task-specific models across NLP. We cover masked language modelling, input format, fine-tuning patterns, and successors such as RoBERTa, DeBERTa and multilingual encoders.
GPT models are decoder-only transformers trained to predict the next token. We trace GPT-1 through GPT-3's few-shot learning to instruction-tuned assistants, and explain why next-token prediction at scale produces broad capabilities.
T5 casts every NLP task as text in, text out; BART pretrains as a denoising autoencoder. We cover span corruption, the text-to-text framework, the lessons of T5's systematic study, and when encoder–decoders are the right choice.
How to adapt a pretrained language model to your task reliably — choosing a model, preparing data, hyperparameters, handling small data and instability, continued pretraining, and evaluating properly.
Question answering systems return answers, not documents. We cover extractive reading comprehension with span prediction, open-domain retriever–reader pipelines, generative QA, evaluation metrics and the problem of unanswerable questions.
Summarisation condenses documents while preserving key information. We compare extractive methods (TextRank) with abstractive neural models, evaluate with ROUGE and factual-consistency checks, and discuss long documents and hallucination.
How do we know if a language system is good? We survey intrinsic and extrinsic evaluation, overlap metrics, embedding-based metrics, learned metrics, LLM judges, human evaluation, benchmarks and their pitfalls.
Search is the most used NLP application. We cover indexing, BM25, dense bi-encoder retrieval, cross-encoder re-ranking, hybrid search, approximate nearest neighbours, and evaluation with recall@k, MRR and nDCG.
Topic models discover themes in large document collections without labels. We derive Latent Dirichlet Allocation's generative story, compare it with NMF and embedding-based BERTopic, and discuss evaluating and interpreting topics.
Most of the world's languages have little digital data. We examine why this matters, how multilingual models enable cross-lingual transfer, the specific challenges of languages like Bangla, and practical strategies for building NLP in low-resource settings.
We compare rule-based, task-oriented and open-domain dialogue systems; cover intent detection, slot filling and dialogue state tracking; and show how LLM-based assistants with retrieval and tools are designed, evaluated and deployed safely.
Speech recognition converts audio into text. We cover audio features and spectrograms, the classical HMM–GMM pipeline, end-to-end neural models with CTC and attention, self-supervised wav2vec 2.0, Whisper, and evaluation with word error rate.
Text-to-speech systems turn written text into natural-sounding speech. We cover the TTS pipeline, text normalisation and phonemes, acoustic models like Tacotron and FastSpeech, neural vocoders, end-to-end and zero-shot voice models, evaluation, and voice-cloning ethics.
Self-attention's quadratic cost limits context length. We survey sparse and local attention, low-rank and kernel-based linear attention, IO-aware FlashAttention, and alternatives such as state-space models — with their trade-offs.
We open the Generative AI track by contrasting models that draw boundaries with models that learn the data distribution itself, and map the families of deep generative models — autoregressive, VAEs, GANs, flows and diffusion.
VAEs turn autoencoders into generative models by learning a smooth, probabilistic latent space. We derive the evidence lower bound, the reparameterisation trick and the KL term, and discuss blurriness, posterior collapse and β-VAE.
GANs train a generator to fool a discriminator in a two-player game. We derive the minimax objective and its optimum, study training dynamics, mode collapse and the non-saturating loss, and build a DCGAN.
A tour of influential GAN architectures — conditional GANs, paired and unpaired image-to-image translation, progressive growing and StyleGAN's style-based generator — plus the deepfake concerns they raised.
Normalising flows transform a simple distribution into a complex one through invertible mappings, giving exact likelihoods and fast sampling. We derive the change-of-variables formula, coupling layers, RealNVP and Glow, and continuous flows.
Diffusion models gradually add noise to data and train a network to reverse the process. We derive the DDPM forward and reverse processes, the simple noise-prediction loss, sampling, the score-based view, and faster samplers like DDIM.
Running diffusion in a compressed latent space made high-resolution text-to-image generation affordable. We dissect latent diffusion — autoencoder, U-Net with cross-attention, text encoder — and practical techniques like img2img, inpainting, ControlNet and fine-tuning.
Conditional diffusion models often ignore their prompt. Guidance amplifies the condition. We derive classifier guidance from Bayes' rule, then classifier-free guidance, and analyse the fidelity–diversity trade-off controlled by the guidance scale.
A map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.
Language-model loss follows smooth power laws in model size, data and compute. We examine the Kaplan and Chinchilla scaling laws, compute-optimal training, the debate on emergent abilities, and what scaling means for the field.
What actually goes into pretraining an LLM? We cover data sourcing, filtering, deduplication, mixture design, tokenisation, the training objective, stability tricks, infrastructure and evaluation during pretraining.
A base model continues text; an instruction-tuned model answers requests. We cover supervised fine-tuning data (human-written, templated and synthetic), chat formats, loss masking, what instruction tuning changes, and practical recipes.
RLHF aligns language models with human preferences using a learned reward model and reinforcement learning. We derive the Bradley–Terry reward model, the KL-regularised objective optimised with PPO, and discuss reward hacking and limitations.
DPO aligns language models to preferences with a simple classification-style loss — no reward model, no RL loop. We derive DPO from the KL-regularised RLHF objective, implement it, and survey variants and practical considerations.
Practical, evidence-based techniques for prompting LLMs — clear instructions, context, examples, output formats, decomposition, and systematic evaluation — plus prompt injection risks.
Asking models to show intermediate steps dramatically improves reasoning. We cover chain-of-thought prompting, self-consistency, tree search, program-aided reasoning, reasoning models trained with RL, test-time compute and the faithfulness question.
Large language models can perform new tasks from a few examples in the prompt without weight updates. We examine what in-context learning is, what influences it, theories of how it works, and its practical limits.
RAG connects an LLM to a searchable knowledge base so answers are current, specific and citable. We build the full pipeline — ingestion, chunking, embeddings, retrieval, re-ranking, prompting with citations — and evaluate and harden it.
Embedding-based applications need fast similarity search over millions of vectors. We explain exact vs approximate search, HNSW graphs, IVF and product quantisation, filtering, and how to choose and operate a vector store.
Full fine-tuning of billion-parameter models is expensive. PEFT methods train a tiny fraction of parameters. We derive LoRA's low-rank updates, QLoRA's 4-bit training, compare adapters and prompt tuning, and give practical recipes.
Quantisation shrinks LLM weights to 8, 4 or fewer bits so models run on smaller GPUs, laptops and phones. We cover outlier features, weight-only vs activation quantisation, GPTQ, AWQ, formats like GGUF, and how to evaluate quality loss.
Mixture-of-experts layers route each token to a few of many expert networks, so models gain parameters without proportional compute. We cover gating, top-k routing, load balancing, capacity, training and serving challenges.
Serving LLMs efficiently is a systems problem. We analyse prefill vs decode phases, the KV cache and PagedAttention, continuous batching, speculative decoding, and the latency and throughput metrics that matter.
A language model outputs probabilities; a decoding strategy turns them into text. We compare greedy and beam search with temperature, top-k, nucleus and min-p sampling, repetition penalties and constrained decoding, and when to use each.
LLMs sometimes produce fluent but false or unsupported content. We classify hallucinations, explain why they arise from training and decoding, and survey detection methods and mitigation — from retrieval and citations to calibrated abstention.
How good is an LLM — and for what? We survey capability benchmarks, human-preference arenas, safety evaluations, contamination and saturation problems, and how to build task-specific evaluation suites for your own applications.
Agents let LLMs call tools, observe results and pursue multi-step goals. We cover function calling, the ReAct loop, planning and memory, multi-agent patterns, evaluation, and — critically — safety, permissions and human oversight.
Multimodal models understand and generate across text, images, audio and video. We cover fusion strategies, vision–language model architectures (LLaVA-style), training stages, capabilities like document understanding and VQA, and known failure modes.
A systems view of modern text-to-image and text-to-video models — architectures, training data, controllability, evaluation, and the provenance and safety measures that responsible use requires.
A practical capstone for the Generative AI track — scoping a use case, choosing models, designing prompts, RAG and tools, evaluation, guardrails, cost and latency, deployment, monitoring and governance.
Reinforcement learning studies agents that learn to act from rewards. We define the agent–environment loop, rewards, returns, policies and value functions, contrast RL with supervised learning, and map the field.
MDPs formalise sequential decision making under uncertainty. We define states, actions, transition probabilities, rewards and discounting, discuss the Markov property, episodic vs continuing tasks, and partial observability.
Value functions satisfy recursive consistency conditions. We derive the Bellman expectation and optimality equations for V and Q, interpret backup diagrams, and solve a small MDP exactly with linear algebra.
When the MDP model is known, dynamic programming computes optimal policies exactly. We implement iterative policy evaluation, policy improvement, policy iteration and value iteration on a gridworld, and discuss their limits.
Without a model, an agent can estimate values by averaging actual returns from complete episodes. We cover first-visit and every-visit MC prediction, MC control with ε-greedy policies, and off-policy learning with importance sampling.
TD learning combines Monte Carlo sampling with dynamic-programming bootstrapping, updating after every step from the TD error. We derive TD(0), compare bias and variance with MC, and introduce n-step returns and TD(λ).
TD methods for control learn action values and improve the policy on the fly. We derive on-policy SARSA and off-policy Q-learning, implement both on the cliff-walking problem, and explain why they learn different paths.
Bandits are RL without states — the purest form of the exploration problem. We compare ε-greedy, optimistic initialisation, UCB and Thompson sampling, define regret, and look at contextual bandits for real decisions.
Tabular methods cannot scale to large or continuous state spaces. We replace tables with parameterised functions, derive semi-gradient TD, discuss linear features and tile coding, and understand the deadly triad that makes deep RL unstable.
In 2013–2015 DeepMind's DQN learned to play dozens of Atari games from raw pixels with one algorithm. We dissect the Q-network, experience replay, target networks and preprocessing, and implement DQN for CartPole.
A series of improvements fixed DQN's weaknesses: overestimation, inefficient replay, poor value decomposition and myopic returns. We study each idea and how Rainbow combined six of them.
Instead of learning values and acting greedily, policy gradient methods optimise the policy directly. We derive the policy gradient theorem and REINFORCE, reduce variance with baselines, and implement it on CartPole.
Actor–critic methods pair a policy (actor) with a learned value function (critic) to cut variance and learn online. We derive the one-step actor–critic, A2C/A3C, entropy regularisation and Generalised Advantage Estimation.
Large policy updates can destroy performance. TRPO constrains each update with a KL trust region; PPO achieves similar stability with a simple clipped objective. We derive both and implement PPO's core update.
Robots need continuous actions — torques, velocities, steering angles. We study off-policy actor–critic methods for continuous control: DDPG's deterministic policy gradient, TD3's fixes, and SAC's maximum-entropy framework.
Model-based agents learn a model of the environment and use it to plan or generate imagined experience. We cover Dyna, model-predictive control, model errors and ensembles, MBPO, and latent world models like Dreamer and MuZero.
DeepMind's Go programs combined deep neural networks with Monte Carlo tree search and self-play. We trace AlphaGo's supervised and RL training, AlphaZero's tabula-rasa self-play, MuZero's learned model, and the lessons for AI.
When many learning agents share an environment, each faces a moving target. We introduce Markov games, Nash equilibria, independent learners, centralised training with decentralised execution, self-play and emergent behaviour.
When rewards are hard to specify but demonstrations are available, agents can learn by imitation. We cover behavioural cloning and its compounding errors, DAgger, inverse RL, maximum-entropy IRL and adversarial imitation (GAIL).
In many domains, trial-and-error exploration is unsafe or impossible, but logged data exists. Offline RL learns policies from fixed datasets. We explain the distributional shift problem and methods such as BCQ, CQL and IQL, plus off-policy evaluation.
Games are forgiving; the real world is not. We examine the challenges of deploying RL — sample efficiency, safety, reward design, partial observability — and the techniques that make it work: simulation, domain randomisation, safe RL and human oversight.
Most ML projects fail not because of the model but because of everything around it. We define MLOps, walk through the ML lifecycle, describe maturity levels, and examine the hidden technical debt of machine learning systems.
Models are only as good as their data — and data changes. We cover versioning datasets, validating schemas and distributions, building reproducible data pipelines, and orchestration tools.
ML development involves hundreds of runs with different data, code and hyperparameters. We cover what to track, how to use MLflow for runs, metrics and artefacts, the model registry, and good experiment hygiene.
Can someone else — or you, in six months — get the same result? We examine sources of non-reproducibility, from seeds and GPUs to data and environments, and practical steps for reproducible research and production.
A model creates value only when its predictions reach people and systems. We compare batch, online and streaming serving, build a FastAPI prediction service with validation, and cover latency, scaling and safe rollout.
"It works on my machine" is not a deployment strategy. We explain containers, write efficient Dockerfiles for ML training and serving, handle GPUs, and cover image size, security and orchestration basics.
Deployed models degrade silently as the world changes. We define data drift, concept drift and prediction drift, measure drift with PSI and statistical tests, monitor without labels, and design alerts and retraining triggers.
Continuous integration and delivery bring software-engineering discipline to ML. We cover the testing pyramid for ML — code, data and model tests — quality gates, continuous training, and a practical GitHub Actions workflow.
Feature stores manage features as shared, versioned assets available offline for training and online for serving. We explain training–serving skew, point-in-time correctness, online vs offline stores, and when a feature store is worth it.
Running models on-device brings privacy, offline operation, low latency and low cost. We cover the edge hardware spectrum, the optimisation pipeline, TensorFlow Lite, ONNX Runtime and TinyML on microcontrollers, and field-deployment lessons.
Understanding hardware helps you train faster and cheaper. We explain why GPUs suit deep learning, the roles of memory capacity and bandwidth, precision and tensor cores, estimating requirements, and choosing between local, cloud and free resources.
Labels are the foundation of supervised learning, yet labelling is often rushed. We cover annotation guidelines, workflows and tools, measuring agreement, handling label noise, model-assisted labelling, and fair treatment of annotators.
Offline metrics do not guarantee real-world impact. Online experiments measure what a model actually changes. We design randomised A/B tests for ML, compute sample sizes, avoid common pitfalls, and discuss ethics of experimenting with people.
Notebooks are great for exploration and terrible for production. We cover a clean project layout, configuration, modular code, typing and testing, logging, packaging, and a workflow for moving from exploration to maintainable software.
Documentation is how models and datasets are understood, audited and used responsibly. We cover model cards, datasheets for datasets, system cards, what to include, and how documentation supports accountability and regulation.
AI systems make or shape decisions about people at scale. We examine real harms, the major principles of responsible AI, why good intentions are not enough, and how ethics becomes concrete engineering practice.
What does it mean for a model to be fair, and where does unfairness come from? We examine sources of bias, the failure of "fairness through unawareness", proxy variables, and practical strategies across the ML pipeline.
We formalise group fairness — demographic parity, equal opportunity, equalised odds, predictive parity and calibration — compute them with Fairlearn, and prove why several cannot hold simultaneously except in special cases.
Why did the model decide that? We distinguish interpretable models from post-hoc explanations, derive Shapley values and SHAP, explain LIME, cover global vs local explanations and counterfactuals, and discuss the limits of explanations.
Models can leak the data they were trained on. We examine re-identification and model attacks, why anonymisation often fails, the mathematics of differential privacy, DP-SGD, and practical privacy-by-design.
Federated learning trains a shared model across many devices or institutions while raw data stays local. We derive FedAvg, discuss non-IID data, communication costs, secure aggregation, privacy limits and real applications.
As AI systems grow more capable, ensuring they pursue intended goals becomes critical. We cover specification gaming, reward hacking, goal misgeneralisation, current alignment techniques, interpretability, evaluations and governance of frontier models.
AI systems introduce new attack surfaces. We survey threats across the ML lifecycle — data poisoning and backdoors, evasion, model extraction, privacy attacks, prompt injection and supply-chain risks — and practical defences.
Governments and organisations are creating rules for AI. We survey risk-based regulation with the EU AI Act, data-protection law, international principles and standards, and how organisations build AI governance in practice.
AI can help humanitarian organisations anticipate crises, map needs and serve people in their own languages — but the stakes and risks are exceptionally high. We survey applications, principles, pitfalls and how students can contribute responsibly.
Generative AI makes realistic fake images, audio and video cheap. We examine how deepfakes are made, the harms they cause, why detection is an arms race, provenance and watermarking solutions, and the role of policy and media literacy.
Training and running AI models consumes electricity, emits carbon, uses water and requires hardware. We explain how to estimate these costs, what drives them, and practical steps toward efficient, sustainable AI.
Will AI take our jobs? We look at evidence on automation and augmentation, task-based analysis of exposure, early productivity studies of generative AI, distributional effects, and how individuals, organisations and policy can respond.
Generative AI is transforming how students learn and teachers teach. We discuss benefits like personalised tutoring, risks to learning and integrity, why AI-text detectors fail, and practical guidelines for students and educators.
A practical guide to AI careers — the main roles, the skills each needs, a staged learning roadmap, how to build a portfolio that stands out, finding opportunities, and growing throughout your career.
Research papers are how new ideas enter the field. We present a three-pass reading method, a guide to each section, questions for critical reading, common red flags in ML papers, and tools for keeping up.
A practical guide to undergraduate and master's research in ML — choosing a question, reviewing literature, designing rigorous experiments, avoiding common pitfalls, writing clearly, and conducting research ethically.