Reinforcement Learning
MDPs, dynamic programming, Q-learning, policy gradients, PPO and AlphaZero.
- 01Introduction to Reinforcement Learning: Learning by InteractionReinforcement learning studies agents that learn to act from rewards. We define the agent–environment loop, rewards, returns, policies and value functions, contrast RL with supervised learning, and map the field.
- 02Markov Decision Processes: The Mathematical Framework of RLMDPs formalise sequential decision making under uncertainty. We define states, actions, transition probabilities, rewards and discounting, discuss the Markov property, episodic vs continuing tasks, and partial observability.
- 03The Bellman Equations: Recursive Structure of ValueValue functions satisfy recursive consistency conditions. We derive the Bellman expectation and optimality equations for V and Q, interpret backup diagrams, and solve a small MDP exactly with linear algebra.
- 04Dynamic Programming: Policy Evaluation, Policy Iteration and Value IterationWhen the MDP model is known, dynamic programming computes optimal policies exactly. We implement iterative policy evaluation, policy improvement, policy iteration and value iteration on a gridworld, and discuss their limits.
- 05Monte Carlo Methods in Reinforcement LearningWithout a model, an agent can estimate values by averaging actual returns from complete episodes. We cover first-visit and every-visit MC prediction, MC control with ε-greedy policies, and off-policy learning with importance sampling.
- 06Temporal-Difference Learning: Learning from GuessesTD learning combines Monte Carlo sampling with dynamic-programming bootstrapping, updating after every step from the TD error. We derive TD(0), compare bias and variance with MC, and introduce n-step returns and TD(λ).
- 07Q-Learning and SARSA: Model-Free ControlTD methods for control learn action values and improve the policy on the fly. We derive on-policy SARSA and off-policy Q-learning, implement both on the cliff-walking problem, and explain why they learn different paths.
- 08Multi-Armed Bandits: The Exploration–Exploitation DilemmaBandits are RL without states — the purest form of the exploration problem. We compare ε-greedy, optimistic initialisation, UCB and Thompson sampling, define regret, and look at contextual bandits for real decisions.
- 09Function Approximation in RL: From Tables to Neural NetworksTabular methods cannot scale to large or continuous state spaces. We replace tables with parameterised functions, derive semi-gradient TD, discuss linear features and tile coding, and understand the deadly triad that makes deep RL unstable.
- 10Deep Q-Networks (DQN): Human-Level Atari from PixelsIn 2013–2015 DeepMind's DQN learned to play dozens of Atari games from raw pixels with one algorithm. We dissect the Q-network, experience replay, target networks and preprocessing, and implement DQN for CartPole.
- 11Improving DQN: Double, Dueling, Prioritised Replay and RainbowA series of improvements fixed DQN's weaknesses: overestimation, inefficient replay, poor value decomposition and myopic returns. We study each idea and how Rainbow combined six of them.
- 12Policy Gradient Methods: REINFORCE and the Policy Gradient TheoremInstead of learning values and acting greedily, policy gradient methods optimise the policy directly. We derive the policy gradient theorem and REINFORCE, reduce variance with baselines, and implement it on CartPole.
- 13Actor–Critic Methods: A2C, A3C and Advantage EstimationActor–critic methods pair a policy (actor) with a learned value function (critic) to cut variance and learn online. We derive the one-step actor–critic, A2C/A3C, entropy regularisation and Generalised Advantage Estimation.
- 14PPO and TRPO: Stable Policy Optimisation with Trust RegionsLarge policy updates can destroy performance. TRPO constrains each update with a KL trust region; PPO achieves similar stability with a simple clipped objective. We derive both and implement PPO's core update.
- 15Continuous Control: DDPG, TD3 and Soft Actor–CriticRobots need continuous actions — torques, velocities, steering angles. We study off-policy actor–critic methods for continuous control: DDPG's deterministic policy gradient, TD3's fixes, and SAC's maximum-entropy framework.
- 16Model-Based Reinforcement Learning: Learning and Planning with World ModelsModel-based agents learn a model of the environment and use it to plan or generate imagined experience. We cover Dyna, model-predictive control, model errors and ensembles, MBPO, and latent world models like Dreamer and MuZero.
- 17AlphaGo, AlphaZero and MuZero: Search Meets Deep LearningDeepMind's Go programs combined deep neural networks with Monte Carlo tree search and self-play. We trace AlphaGo's supervised and RL training, AlphaZero's tabula-rasa self-play, MuZero's learned model, and the lessons for AI.
- 18Multi-Agent Reinforcement Learning: Cooperation, Competition and EquilibriaWhen many learning agents share an environment, each faces a moving target. We introduce Markov games, Nash equilibria, independent learners, centralised training with decentralised execution, self-play and emergent behaviour.
- 19Imitation Learning and Inverse Reinforcement LearningWhen rewards are hard to specify but demonstrations are available, agents can learn by imitation. We cover behavioural cloning and its compounding errors, DAgger, inverse RL, maximum-entropy IRL and adversarial imitation (GAIL).
- 20Offline Reinforcement Learning: Learning from Logged DataIn many domains, trial-and-error exploration is unsafe or impossible, but logged data exists. Offline RL learns policies from fixed datasets. We explain the distributional shift problem and methods such as BCQ, CQL and IQL, plus off-policy evaluation.
- 21Reinforcement Learning in the Real World: Robotics, Sim-to-Real and SafetyGames are forgiving; the real world is not. We examine the challenges of deploying RL — sample efficiency, safety, reward design, partial observability — and the techniques that make it work: simulation, domain randomisation, safe RL and human oversight.