🎮

Reinforcement Learning

MDPs, dynamic programming, Q-learning, policy gradients, PPO and AlphaZero.

  1. 01Introduction to Reinforcement Learning: Learning by InteractionReinforcement learning studies agents that learn to act from rewards. We define the agent–environment loop, rewards, returns, policies and value functions, contrast RL with supervised learning, and map the field.Beginner5 min
  2. 02Markov Decision Processes: The Mathematical Framework of RLMDPs formalise sequential decision making under uncertainty. We define states, actions, transition probabilities, rewards and discounting, discuss the Markov property, episodic vs continuing tasks, and partial observability.Intermediate5 min
  3. 03The Bellman Equations: Recursive Structure of ValueValue functions satisfy recursive consistency conditions. We derive the Bellman expectation and optimality equations for V and Q, interpret backup diagrams, and solve a small MDP exactly with linear algebra.Intermediate4 min
  4. 04Dynamic Programming: Policy Evaluation, Policy Iteration and Value IterationWhen the MDP model is known, dynamic programming computes optimal policies exactly. We implement iterative policy evaluation, policy improvement, policy iteration and value iteration on a gridworld, and discuss their limits.Intermediate5 min
  5. 05Monte Carlo Methods in Reinforcement LearningWithout a model, an agent can estimate values by averaging actual returns from complete episodes. We cover first-visit and every-visit MC prediction, MC control with ε-greedy policies, and off-policy learning with importance sampling.Intermediate5 min
  6. 06Temporal-Difference Learning: Learning from GuessesTD learning combines Monte Carlo sampling with dynamic-programming bootstrapping, updating after every step from the TD error. We derive TD(0), compare bias and variance with MC, and introduce n-step returns and TD(λ).Intermediate5 min
  7. 07Q-Learning and SARSA: Model-Free ControlTD methods for control learn action values and improve the policy on the fly. We derive on-policy SARSA and off-policy Q-learning, implement both on the cliff-walking problem, and explain why they learn different paths.Intermediate5 min
  8. 08Multi-Armed Bandits: The Exploration–Exploitation DilemmaBandits are RL without states — the purest form of the exploration problem. We compare ε-greedy, optimistic initialisation, UCB and Thompson sampling, define regret, and look at contextual bandits for real decisions.Intermediate5 min
  9. 09Function Approximation in RL: From Tables to Neural NetworksTabular methods cannot scale to large or continuous state spaces. We replace tables with parameterised functions, derive semi-gradient TD, discuss linear features and tile coding, and understand the deadly triad that makes deep RL unstable.Advanced4 min
  10. 10Deep Q-Networks (DQN): Human-Level Atari from PixelsIn 2013–2015 DeepMind's DQN learned to play dozens of Atari games from raw pixels with one algorithm. We dissect the Q-network, experience replay, target networks and preprocessing, and implement DQN for CartPole.Advanced5 min
  11. 11Improving DQN: Double, Dueling, Prioritised Replay and RainbowA series of improvements fixed DQN's weaknesses: overestimation, inefficient replay, poor value decomposition and myopic returns. We study each idea and how Rainbow combined six of them.Advanced5 min
  12. 12Policy Gradient Methods: REINFORCE and the Policy Gradient TheoremInstead of learning values and acting greedily, policy gradient methods optimise the policy directly. We derive the policy gradient theorem and REINFORCE, reduce variance with baselines, and implement it on CartPole.Advanced5 min
  13. 13Actor–Critic Methods: A2C, A3C and Advantage EstimationActor–critic methods pair a policy (actor) with a learned value function (critic) to cut variance and learn online. We derive the one-step actor–critic, A2C/A3C, entropy regularisation and Generalised Advantage Estimation.Advanced4 min
  14. 14PPO and TRPO: Stable Policy Optimisation with Trust RegionsLarge policy updates can destroy performance. TRPO constrains each update with a KL trust region; PPO achieves similar stability with a simple clipped objective. We derive both and implement PPO's core update.Advanced5 min
  15. 15Continuous Control: DDPG, TD3 and Soft Actor–CriticRobots need continuous actions — torques, velocities, steering angles. We study off-policy actor–critic methods for continuous control: DDPG's deterministic policy gradient, TD3's fixes, and SAC's maximum-entropy framework.Advanced5 min
  16. 16Model-Based Reinforcement Learning: Learning and Planning with World ModelsModel-based agents learn a model of the environment and use it to plan or generate imagined experience. We cover Dyna, model-predictive control, model errors and ensembles, MBPO, and latent world models like Dreamer and MuZero.Advanced5 min
  17. 17AlphaGo, AlphaZero and MuZero: Search Meets Deep LearningDeepMind's Go programs combined deep neural networks with Monte Carlo tree search and self-play. We trace AlphaGo's supervised and RL training, AlphaZero's tabula-rasa self-play, MuZero's learned model, and the lessons for AI.Advanced6 min
  18. 18Multi-Agent Reinforcement Learning: Cooperation, Competition and EquilibriaWhen many learning agents share an environment, each faces a moving target. We introduce Markov games, Nash equilibria, independent learners, centralised training with decentralised execution, self-play and emergent behaviour.Advanced5 min
  19. 19Imitation Learning and Inverse Reinforcement LearningWhen rewards are hard to specify but demonstrations are available, agents can learn by imitation. We cover behavioural cloning and its compounding errors, DAgger, inverse RL, maximum-entropy IRL and adversarial imitation (GAIL).Advanced5 min
  20. 20Offline Reinforcement Learning: Learning from Logged DataIn many domains, trial-and-error exploration is unsafe or impossible, but logged data exists. Offline RL learns policies from fixed datasets. We explain the distributional shift problem and methods such as BCQ, CQL and IQL, plus off-policy evaluation.Advanced5 min
  21. 21Reinforcement Learning in the Real World: Robotics, Sim-to-Real and SafetyGames are forgiving; the real world is not. We examine the challenges of deploying RL — sample efficiency, safety, reward design, partial observability — and the techniques that make it work: simulation, domain randomisation, safe RL and human oversight.Advanced5 min