🎮 Reinforcement Learning · Lecture 1 of 21

Introduction to Reinforcement Learning: Learning by Interaction

Reinforcement learning studies agents that learn to act from rewards. We define the agent–environment loop, rewards, returns, policies and value functions, contrast RL with supervised learning, and map the field.

A child learning to ride a bicycle receives no labelled dataset of "correct pedal pressures". They try, wobble, fall, adjust — and gradually improve through interaction and feedback. Reinforcement learning (RL) formalises this kind of learning. It is how AlphaGo mastered Go, how robots learn to walk in simulation, how data-centre cooling has been optimised, and how language models are aligned with human preferences.

The agent–environment loop#

At each time step $t$:

  1. The agent observes a state $S_t$.
  2. It chooses an action $A_t$ according to its policy $\pi(a \mid s)$.
  3. The environment returns a reward $R_{t+1}$ and a new state $S_{t+1}$.
text
        ┌──────── action A_t ────────┐
   Agent                              Environment
        └── state S_{t+1}, reward R_{t+1} ──┘

The agent's goal is to maximise cumulative reward — not the immediate reward.

Rewards and returns#

The reward hypothesis (Sutton & Barto): goals can be expressed as maximising the expected cumulative sum of a scalar reward. The return from time $t$ is the discounted sum of future rewards:

$$ G_t = R_{t+1} + \gamma R_{t+2} + \gamma^2 R_{t+3} + \dots = \sum_{k=0}^{\infty}\gamma^k R_{t+k+1} $$

The discount factor $\gamma \in [0, 1]$ makes near rewards count more than distant ones, keeps infinite sums finite, and reflects uncertainty about the future. $\gamma$ close to 1 means far-sighted; close to 0 means myopic.

Policies and value functions#

  • A policy $\pi$ maps states to actions (deterministic) or to distributions over actions (stochastic).
  • The state-value function $V^\pi(s) = \mathbb{E}_\pi[G_t \mid S_t = s]$: how good it is to be in $s$ when following $\pi$.
  • The action-value function $Q^\pi(s, a) = \mathbb{E}_\pi[G_t \mid S_t = s, A_t = a]$: how good it is to take $a$ in $s$ and then follow $\pi$.

An optimal policy $\pi^*$ achieves the highest value in every state. If we know the optimal $Q^*$, acting optimally is easy: choose $\arg\max_a Q^*(s, a)$.

How RL differs from supervised learning#

AspectSupervised learningReinforcement learning
FeedbackCorrect label for each inputScalar reward, evaluative not instructive
TimingImmediateOften delayed (credit assignment problem)
DataFixed i.i.d. datasetGenerated by the agent's own actions (non-i.i.d.)
Key challengeGeneralisationExploration, credit assignment, stability

The exploration–exploitation dilemma#

Should the agent exploit what it knows works, or explore actions that might be better? Exploit too much and it never discovers the best strategy; explore too much and it wastes reward. Every RL algorithm must balance the two — for example with ε-greedy action selection (random action with probability ε).

A first environment#

python
# pip install gymnasium
import gymnasium as gym

env = gym.make("CartPole-v1")             # balance a pole on a moving cart
obs, info = env.reset(seed=0)
total, done = 0.0, False
while not done:
    action = env.action_space.sample()    # a random policy
    obs, reward, terminated, truncated, info = env.step(action)
    total += reward
    done = terminated or truncated
print("random policy return:", total)     # typically around 20; an optimal policy reaches 500
env.close()

Gymnasium (the maintained successor of OpenAI Gym) provides standard environments for learning RL.

A map of RL methods#

  • Model-based vs model-free: does the agent learn/use a model of the environment's dynamics?
  • Value-based (learn $Q$, act greedily): Q-learning, DQN.
  • Policy-based (optimise $\pi$ directly): REINFORCE, policy gradients.
  • Actor–critic (both): A2C, PPO, SAC.
  • On-policy vs off-policy: learn about the policy being executed, or about a different (e.g. greedy) policy from any data.
  • Online vs offline: interact with the environment, or learn from a fixed logged dataset.

Applications and cautions#

Games (Atari, Go, StarCraft), robotics, recommendation and advertising, resource allocation, chip design, controlling fusion plasma in simulation-trained controllers, and aligning LLMs (RLHF). In the real world, RL faces sample inefficiency, safety during exploration, and reward misspecification — so many successful applications rely on simulators, offline data or tight human oversight.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

🎮 Reinforcement Learning

Markov Decision Processes: The Mathematical Framework of RL

MDPs formalise sequential decision making under uncertainty. We define states, actions, transition probabilities, rewards and discounting, discuss the Markov property, episodic vs continuing tasks, and partial observability.

Intermediate⏱ 5 min#222
🎮 Reinforcement Learning

The Bellman Equations: Recursive Structure of Value

Value functions satisfy recursive consistency conditions. We derive the Bellman expectation and optimality equations for V and Q, interpret backup diagrams, and solve a small MDP exactly with linear algebra.

Intermediate⏱ 4 min#223
🎮 Reinforcement Learning

Multi-Armed Bandits: The Exploration–Exploitation Dilemma

Bandits are RL without states — the purest form of the exploration problem. We compare ε-greedy, optimistic initialisation, UCB and Thompson sampling, define regret, and look at contextual bandits for real decisions.

Intermediate⏱ 5 min#228