Generative AI & LLMs
VAEs, GANs, diffusion, large language models, RAG, fine-tuning and AI agents.
- 01Generative vs Discriminative Models: Learning to CreateWe open the Generative AI track by contrasting models that draw boundaries with models that learn the data distribution itself, and map the families of deep generative models — autoregressive, VAEs, GANs, flows and diffusion.
- 02Variational Autoencoders: Probabilistic Latent SpacesVAEs turn autoencoders into generative models by learning a smooth, probabilistic latent space. We derive the evidence lower bound, the reparameterisation trick and the KL term, and discuss blurriness, posterior collapse and β-VAE.
- 03Generative Adversarial Networks: The Generator–Discriminator GameGANs train a generator to fool a discriminator in a two-player game. We derive the minimax objective and its optimum, study training dynamics, mode collapse and the non-saturating loss, and build a DCGAN.
- 04GAN Variants: Conditional GANs, Pix2Pix, CycleGAN and StyleGANA tour of influential GAN architectures — conditional GANs, paired and unpaired image-to-image translation, progressive growing and StyleGAN's style-based generator — plus the deepfake concerns they raised.
- 05Normalising Flows: Exact Likelihood with Invertible NetworksNormalising flows transform a simple distribution into a complex one through invertible mappings, giving exact likelihoods and fast sampling. We derive the change-of-variables formula, coupling layers, RealNVP and Glow, and continuous flows.
- 06Diffusion Models: Generating by Learning to DenoiseDiffusion models gradually add noise to data and train a network to reverse the process. We derive the DDPM forward and reverse processes, the simple noise-prediction loss, sampling, the score-based view, and faster samplers like DDIM.
- 07Latent Diffusion and Stable Diffusion: Text-to-Image at ScaleRunning diffusion in a compressed latent space made high-resolution text-to-image generation affordable. We dissect latent diffusion — autoencoder, U-Net with cross-attention, text encoder — and practical techniques like img2img, inpainting, ControlNet and fine-tuning.
- 08Guidance in Diffusion Models: Classifier and Classifier-Free GuidanceConditional diffusion models often ignore their prompt. Guidance amplifies the condition. We derive classifier guidance from Bayes' rule, then classifier-free guidance, and analyse the fidelity–diversity trade-off controlled by the guidance scale.
- 09Large Language Models: What They Are and How They Are BuiltA map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.
- 10Scaling Laws: How Performance Grows with Compute, Data and ParametersLanguage-model loss follows smooth power laws in model size, data and compute. We examine the Kaplan and Chinchilla scaling laws, compute-optimal training, the debate on emergent abilities, and what scaling means for the field.
- 11Pretraining LLMs: Data Pipelines, Objectives and InfrastructureWhat actually goes into pretraining an LLM? We cover data sourcing, filtering, deduplication, mixture design, tokenisation, the training objective, stability tricks, infrastructure and evaluation during pretraining.
- 12Instruction Tuning: Teaching Language Models to Follow DirectionsA base model continues text; an instruction-tuned model answers requests. We cover supervised fine-tuning data (human-written, templated and synthetic), chat formats, loss masking, what instruction tuning changes, and practical recipes.
- 13RLHF: Reinforcement Learning from Human FeedbackRLHF aligns language models with human preferences using a learned reward model and reinforcement learning. We derive the Bradley–Terry reward model, the KL-regularised objective optimised with PPO, and discuss reward hacking and limitations.
- 14Direct Preference Optimisation (DPO) and BeyondDPO aligns language models to preferences with a simple classification-style loss — no reward model, no RL loop. We derive DPO from the KL-regularised RLHF objective, implement it, and survey variants and practical considerations.
- 15Prompt Engineering: Getting Reliable Results from LLMsPractical, evidence-based techniques for prompting LLMs — clear instructions, context, examples, output formats, decomposition, and systematic evaluation — plus prompt injection risks.
- 16Chain-of-Thought and Reasoning in Language ModelsAsking models to show intermediate steps dramatically improves reasoning. We cover chain-of-thought prompting, self-consistency, tree search, program-aided reasoning, reasoning models trained with RL, test-time compute and the faithfulness question.
- 17In-Context Learning: How LLMs Learn from PromptsLarge language models can perform new tasks from a few examples in the prompt without weight updates. We examine what in-context learning is, what influences it, theories of how it works, and its practical limits.
- 18Retrieval-Augmented Generation (RAG): Grounding LLMs in Your DocumentsRAG connects an LLM to a searchable knowledge base so answers are current, specific and citable. We build the full pipeline — ingestion, chunking, embeddings, retrieval, re-ranking, prompting with citations — and evaluate and harden it.
- 19Vector Databases and Approximate Nearest Neighbour SearchEmbedding-based applications need fast similarity search over millions of vectors. We explain exact vs approximate search, HNSW graphs, IVF and product quantisation, filtering, and how to choose and operate a vector store.
- 20LoRA and Parameter-Efficient Fine-Tuning (PEFT)Full fine-tuning of billion-parameter models is expensive. PEFT methods train a tiny fraction of parameters. We derive LoRA's low-rank updates, QLoRA's 4-bit training, compare adapters and prompt tuning, and give practical recipes.
- 21Quantising Large Language Models for Efficient InferenceQuantisation shrinks LLM weights to 8, 4 or fewer bits so models run on smaller GPUs, laptops and phones. We cover outlier features, weight-only vs activation quantisation, GPTQ, AWQ, formats like GGUF, and how to evaluate quality loss.
- 22Mixture of Experts: Scaling Parameters Without Scaling ComputeMixture-of-experts layers route each token to a few of many expert networks, so models gain parameters without proportional compute. We cover gating, top-k routing, load balancing, capacity, training and serving challenges.
- 23LLM Inference: KV Caching, Batching and Speculative DecodingServing LLMs efficiently is a systems problem. We analyse prefill vs decode phases, the KV cache and PagedAttention, continuous batching, speculative decoding, and the latency and throughput metrics that matter.
- 24Decoding Strategies: Greedy, Beam Search, Temperature, Top-k and Top-pA language model outputs probabilities; a decoding strategy turns them into text. We compare greedy and beam search with temperature, top-k, nucleus and min-p sampling, repetition penalties and constrained decoding, and when to use each.
- 25Hallucination in LLMs: Causes, Detection and MitigationLLMs sometimes produce fluent but false or unsupported content. We classify hallucinations, explain why they arise from training and decoding, and survey detection methods and mitigation — from retrieval and citations to calibrated abstention.
- 26Evaluating Large Language Models: Benchmarks, Arenas and Custom EvalsHow good is an LLM — and for what? We survey capability benchmarks, human-preference arenas, safety evaluations, contamination and saturation problems, and how to build task-specific evaluation suites for your own applications.
- 27AI Agents and Tool Use: LLMs That ActAgents let LLMs call tools, observe results and pursue multi-step goals. We cover function calling, the ReAct loop, planning and memory, multi-agent patterns, evaluation, and — critically — safety, permissions and human oversight.
- 28Multimodal Models: Vision–Language and BeyondMultimodal models understand and generate across text, images, audio and video. We cover fusion strategies, vision–language model architectures (LLaVA-style), training stages, capabilities like document understanding and VQA, and known failure modes.
- 29Text-to-Image and Text-to-Video Generation: Systems, Control and ProvenanceA systems view of modern text-to-image and text-to-video models — architectures, training data, controllability, evaluation, and the provenance and safety measures that responsible use requires.
- 30Building an LLM Application End to End: From Idea to ProductionA practical capstone for the Generative AI track — scoping a use case, choosing models, designing prompts, RAG and tools, evaluation, guardrails, cost and latency, deployment, monitoring and governance.