✨

Generative AI & LLMs

VAEs, GANs, diffusion, large language models, RAG, fine-tuning and AI agents.

  1. 01Generative vs Discriminative Models: Learning to CreateWe open the Generative AI track by contrasting models that draw boundaries with models that learn the data distribution itself, and map the families of deep generative models — autoregressive, VAEs, GANs, flows and diffusion.Beginner5 min
  2. 02Variational Autoencoders: Probabilistic Latent SpacesVAEs turn autoencoders into generative models by learning a smooth, probabilistic latent space. We derive the evidence lower bound, the reparameterisation trick and the KL term, and discuss blurriness, posterior collapse and β-VAE.Advanced5 min
  3. 03Generative Adversarial Networks: The Generator–Discriminator GameGANs train a generator to fool a discriminator in a two-player game. We derive the minimax objective and its optimum, study training dynamics, mode collapse and the non-saturating loss, and build a DCGAN.Advanced5 min
  4. 04GAN Variants: Conditional GANs, Pix2Pix, CycleGAN and StyleGANA tour of influential GAN architectures — conditional GANs, paired and unpaired image-to-image translation, progressive growing and StyleGAN's style-based generator — plus the deepfake concerns they raised.Advanced5 min
  5. 05Normalising Flows: Exact Likelihood with Invertible NetworksNormalising flows transform a simple distribution into a complex one through invertible mappings, giving exact likelihoods and fast sampling. We derive the change-of-variables formula, coupling layers, RealNVP and Glow, and continuous flows.Advanced5 min
  6. 06Diffusion Models: Generating by Learning to DenoiseDiffusion models gradually add noise to data and train a network to reverse the process. We derive the DDPM forward and reverse processes, the simple noise-prediction loss, sampling, the score-based view, and faster samplers like DDIM.Advanced6 min
  7. 07Latent Diffusion and Stable Diffusion: Text-to-Image at ScaleRunning diffusion in a compressed latent space made high-resolution text-to-image generation affordable. We dissect latent diffusion — autoencoder, U-Net with cross-attention, text encoder — and practical techniques like img2img, inpainting, ControlNet and fine-tuning.Advanced5 min
  8. 08Guidance in Diffusion Models: Classifier and Classifier-Free GuidanceConditional diffusion models often ignore their prompt. Guidance amplifies the condition. We derive classifier guidance from Bayes' rule, then classifier-free guidance, and analyse the fidelity–diversity trade-off controlled by the guidance scale.Advanced5 min
  9. 09Large Language Models: What They Are and How They Are BuiltA map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.Beginner5 min
  10. 10Scaling Laws: How Performance Grows with Compute, Data and ParametersLanguage-model loss follows smooth power laws in model size, data and compute. We examine the Kaplan and Chinchilla scaling laws, compute-optimal training, the debate on emergent abilities, and what scaling means for the field.Advanced5 min
  11. 11Pretraining LLMs: Data Pipelines, Objectives and InfrastructureWhat actually goes into pretraining an LLM? We cover data sourcing, filtering, deduplication, mixture design, tokenisation, the training objective, stability tricks, infrastructure and evaluation during pretraining.Advanced5 min
  12. 12Instruction Tuning: Teaching Language Models to Follow DirectionsA base model continues text; an instruction-tuned model answers requests. We cover supervised fine-tuning data (human-written, templated and synthetic), chat formats, loss masking, what instruction tuning changes, and practical recipes.Intermediate5 min
  13. 13RLHF: Reinforcement Learning from Human FeedbackRLHF aligns language models with human preferences using a learned reward model and reinforcement learning. We derive the Bradley–Terry reward model, the KL-regularised objective optimised with PPO, and discuss reward hacking and limitations.Advanced5 min
  14. 14Direct Preference Optimisation (DPO) and BeyondDPO aligns language models to preferences with a simple classification-style loss — no reward model, no RL loop. We derive DPO from the KL-regularised RLHF objective, implement it, and survey variants and practical considerations.Advanced5 min
  15. 15Prompt Engineering: Getting Reliable Results from LLMsPractical, evidence-based techniques for prompting LLMs — clear instructions, context, examples, output formats, decomposition, and systematic evaluation — plus prompt injection risks.Beginner5 min
  16. 16Chain-of-Thought and Reasoning in Language ModelsAsking models to show intermediate steps dramatically improves reasoning. We cover chain-of-thought prompting, self-consistency, tree search, program-aided reasoning, reasoning models trained with RL, test-time compute and the faithfulness question.Intermediate5 min
  17. 17In-Context Learning: How LLMs Learn from PromptsLarge language models can perform new tasks from a few examples in the prompt without weight updates. We examine what in-context learning is, what influences it, theories of how it works, and its practical limits.Intermediate5 min
  18. 18Retrieval-Augmented Generation (RAG): Grounding LLMs in Your DocumentsRAG connects an LLM to a searchable knowledge base so answers are current, specific and citable. We build the full pipeline — ingestion, chunking, embeddings, retrieval, re-ranking, prompting with citations — and evaluate and harden it.Intermediate6 min
  19. 19Vector Databases and Approximate Nearest Neighbour SearchEmbedding-based applications need fast similarity search over millions of vectors. We explain exact vs approximate search, HNSW graphs, IVF and product quantisation, filtering, and how to choose and operate a vector store.Intermediate6 min
  20. 20LoRA and Parameter-Efficient Fine-Tuning (PEFT)Full fine-tuning of billion-parameter models is expensive. PEFT methods train a tiny fraction of parameters. We derive LoRA's low-rank updates, QLoRA's 4-bit training, compare adapters and prompt tuning, and give practical recipes.Advanced5 min
  21. 21Quantising Large Language Models for Efficient InferenceQuantisation shrinks LLM weights to 8, 4 or fewer bits so models run on smaller GPUs, laptops and phones. We cover outlier features, weight-only vs activation quantisation, GPTQ, AWQ, formats like GGUF, and how to evaluate quality loss.Advanced5 min
  22. 22Mixture of Experts: Scaling Parameters Without Scaling ComputeMixture-of-experts layers route each token to a few of many expert networks, so models gain parameters without proportional compute. We cover gating, top-k routing, load balancing, capacity, training and serving challenges.Advanced5 min
  23. 23LLM Inference: KV Caching, Batching and Speculative DecodingServing LLMs efficiently is a systems problem. We analyse prefill vs decode phases, the KV cache and PagedAttention, continuous batching, speculative decoding, and the latency and throughput metrics that matter.Advanced6 min
  24. 24Decoding Strategies: Greedy, Beam Search, Temperature, Top-k and Top-pA language model outputs probabilities; a decoding strategy turns them into text. We compare greedy and beam search with temperature, top-k, nucleus and min-p sampling, repetition penalties and constrained decoding, and when to use each.Intermediate5 min
  25. 25Hallucination in LLMs: Causes, Detection and MitigationLLMs sometimes produce fluent but false or unsupported content. We classify hallucinations, explain why they arise from training and decoding, and survey detection methods and mitigation — from retrieval and citations to calibrated abstention.Intermediate6 min
  26. 26Evaluating Large Language Models: Benchmarks, Arenas and Custom EvalsHow good is an LLM — and for what? We survey capability benchmarks, human-preference arenas, safety evaluations, contamination and saturation problems, and how to build task-specific evaluation suites for your own applications.Intermediate5 min
  27. 27AI Agents and Tool Use: LLMs That ActAgents let LLMs call tools, observe results and pursue multi-step goals. We cover function calling, the ReAct loop, planning and memory, multi-agent patterns, evaluation, and — critically — safety, permissions and human oversight.Intermediate5 min
  28. 28Multimodal Models: Vision–Language and BeyondMultimodal models understand and generate across text, images, audio and video. We cover fusion strategies, vision–language model architectures (LLaVA-style), training stages, capabilities like document understanding and VQA, and known failure modes.Intermediate5 min
  29. 29Text-to-Image and Text-to-Video Generation: Systems, Control and ProvenanceA systems view of modern text-to-image and text-to-video models — architectures, training data, controllability, evaluation, and the provenance and safety measures that responsible use requires.Intermediate5 min
  30. 30Building an LLM Application End to End: From Idea to ProductionA practical capstone for the Generative AI track — scoping a use case, choosing models, designing prompts, RAG and tools, evaluation, guardrails, cost and latency, deployment, monitoring and governance.Intermediate5 min