A child learning to ride a bicycle receives no labelled dataset of "correct pedal pressures". They try, wobble, fall, adjust — and gradually improve through interaction and feedback. Reinforcement learning (RL) formalises this kind of learning. It is how AlphaGo mastered Go, how robots learn to walk in simulation, how data-centre cooling has been optimised, and how language models are aligned with human preferences.
The agent–environment loop#
At each time step $t$:
- The agent observes a state $S_t$.
- It chooses an action $A_t$ according to its policy $\pi(a \mid s)$.
- The environment returns a reward $R_{t+1}$ and a new state $S_{t+1}$.
┌──────── action A_t ────────┐
Agent Environment
└── state S_{t+1}, reward R_{t+1} ──┘The agent's goal is to maximise cumulative reward — not the immediate reward.
Rewards and returns#
The reward hypothesis (Sutton & Barto): goals can be expressed as maximising the expected cumulative sum of a scalar reward. The return from time $t$ is the discounted sum of future rewards:
The discount factor $\gamma \in [0, 1]$ makes near rewards count more than distant ones, keeps infinite sums finite, and reflects uncertainty about the future. $\gamma$ close to 1 means far-sighted; close to 0 means myopic.
Policies and value functions#
- A policy $\pi$ maps states to actions (deterministic) or to distributions over actions (stochastic).
- The state-value function $V^\pi(s) = \mathbb{E}_\pi[G_t \mid S_t = s]$: how good it is to be in $s$ when following $\pi$.
- The action-value function $Q^\pi(s, a) = \mathbb{E}_\pi[G_t \mid S_t = s, A_t = a]$: how good it is to take $a$ in $s$ and then follow $\pi$.
An optimal policy $\pi^*$ achieves the highest value in every state. If we know the optimal $Q^*$, acting optimally is easy: choose $\arg\max_a Q^*(s, a)$.
How RL differs from supervised learning#
| Aspect | Supervised learning | Reinforcement learning |
|---|---|---|
| Feedback | Correct label for each input | Scalar reward, evaluative not instructive |
| Timing | Immediate | Often delayed (credit assignment problem) |
| Data | Fixed i.i.d. dataset | Generated by the agent's own actions (non-i.i.d.) |
| Key challenge | Generalisation | Exploration, credit assignment, stability |
The exploration–exploitation dilemma#
Should the agent exploit what it knows works, or explore actions that might be better? Exploit too much and it never discovers the best strategy; explore too much and it wastes reward. Every RL algorithm must balance the two — for example with ε-greedy action selection (random action with probability ε).
A first environment#
# pip install gymnasium
import gymnasium as gym
env = gym.make("CartPole-v1") # balance a pole on a moving cart
obs, info = env.reset(seed=0)
total, done = 0.0, False
while not done:
action = env.action_space.sample() # a random policy
obs, reward, terminated, truncated, info = env.step(action)
total += reward
done = terminated or truncated
print("random policy return:", total) # typically around 20; an optimal policy reaches 500
env.close()Gymnasium (the maintained successor of OpenAI Gym) provides standard environments for learning RL.
A map of RL methods#
- Model-based vs model-free: does the agent learn/use a model of the environment's dynamics?
- Value-based (learn $Q$, act greedily): Q-learning, DQN.
- Policy-based (optimise $\pi$ directly): REINFORCE, policy gradients.
- Actor–critic (both): A2C, PPO, SAC.
- On-policy vs off-policy: learn about the policy being executed, or about a different (e.g. greedy) policy from any data.
- Online vs offline: interact with the environment, or learn from a fixed logged dataset.
Applications and cautions#
Games (Atari, Go, StarCraft), robotics, recommendation and advertising, resource allocation, chip design, controlling fusion plasma in simulation-trained controllers, and aligning LLMs (RLHF). In the real world, RL faces sample inefficiency, safety during exploration, and reward misspecification — so many successful applications rely on simulators, offline data or tight human oversight.