🎮 Reinforcement Learning · Lecture 19 of 21

Imitation Learning and Inverse Reinforcement Learning

When rewards are hard to specify but demonstrations are available, agents can learn by imitation. We cover behavioural cloning and its compounding errors, DAgger, inverse RL, maximum-entropy IRL and adversarial imitation (GAIL).

Writing a reward function for "drive safely and courteously" or "fold laundry neatly" is extremely difficult. Showing someone how to do it is easy. Imitation learning trains agents from expert demonstrations rather than hand-designed rewards. It underpins autonomous-driving research, robot manipulation, game agents — and, in a sense, supervised fine-tuning of language models on human-written responses.

Behavioural cloning (BC)#

The simplest approach: treat demonstrations as a supervised dataset of (state, action) pairs and train a policy by classification or regression:

$$ \min_\theta\;\mathbb{E}_{(s, a) \sim \mathcal{D}_{\text{expert}}}\big[-\log\pi_\theta(a \mid s)\big] $$

It is simple, needs no environment interaction, and works well when demonstrations are plentiful and cover the situations the agent will face. ALVINN (1989) steered a vehicle with a neural network trained on human driving; modern robot learning uses BC with expressive policies (e.g. diffusion policies or transformers) and large demonstration datasets.

The compounding-error problem#

BC violates the i.i.d. assumption. The expert rarely makes mistakes, so demonstrations contain few examples of recovering from errors. When the learned policy makes a small mistake, it drifts into states unlike the training data, where it makes bigger mistakes, drifting further — errors compound. Ross and Bagnell showed that the expected cost of BC can grow quadratically with the task horizon $T$ (as $O(\epsilon T^2)$ for per-step error rate $\epsilon$), whereas an ideal learner's cost grows linearly.

DAgger: dataset aggregation#

Ross, Gordon and Bagnell (2011) proposed DAgger:

  1. Train a policy on expert data.
  2. Run the learner's policy to collect the states it actually visits.
  3. Ask the expert to label the correct action for those states.
  4. Aggregate into the dataset and retrain; repeat.

DAgger trains on the learner's own state distribution, fixing compounding errors and achieving linear-in-horizon error bounds. The cost: an expert available to label on demand (possible in simulation or with a scripted planner; harder with humans).

python
import numpy as np
from sklearn.neighbors import KNeighborsClassifier

def dagger(env, expert, n_iters=5, rollouts=10):
    X, y = [], []
    policy = None
    for it in range(n_iters):
        for _ in range(rollouts):
            s, _ = env.reset(); done = False
            while not done:
                X.append(s); y.append(expert(s))                     # expert labels every visited state
                a = expert(s) if policy is None else int(policy.predict([s])[0])   # learner acts
                s, _, term, trunc, _ = env.step(a); done = term or trunc
        policy = KNeighborsClassifier(5).fit(np.array(X), np.array(y))  # retrain on aggregated data
    return policy

Inverse reinforcement learning (IRL)#

Instead of copying actions, infer the reward function that the expert appears to optimise, then find a policy for that reward with RL (Ng & Russell, 2000; Abbeel & Ng, 2004). Motivations:

  • A reward is a compact, transferable description of the task — it generalises to new situations and dynamics better than copied actions.
  • Understanding why the expert acts, not just what they do.

The problem is ill-posed: many rewards explain the same behaviour (a zero reward makes every policy optimal). Maximum-entropy IRL (Ziebart et al., 2008) resolves ambiguity by assuming the expert chooses trajectories with probability

$$ P(\tau) \propto \exp\big(R_\psi(\tau)\big) $$

and fits $\psi$ by maximum likelihood — preferring the least committed explanation consistent with demonstrations. It was used, for example, to model taxi drivers' route choices.

Adversarial imitation: GAIL#

Generative Adversarial Imitation Learning (Ho & Ermon, 2016) links IRL with GANs. A discriminator $D(s, a)$ learns to distinguish expert state–action pairs from the agent's; the agent (trained with RL, e.g. TRPO/PPO) receives reward for fooling the discriminator, such as $-\log(1 - D(s, a))$. GAIL matches the expert's state–action distribution directly, avoiding compounding errors, with far fewer demonstrations than BC — though it requires environment interaction.

Comparing approaches#

MethodNeeds environment interactionNeeds interactive expertHandles compounding errorsRecovers reward
Behavioural cloningNoNoNoNo
DAggerYesYesYesNo
IRL (e.g. MaxEnt)Yes (RL inner loop)NoYesYes
GAILYesNoYesImplicitly

Connections and cautions#

  • RLHF is related: rather than demonstrations, it learns a reward from preferences — another way of inferring human intent.
  • Offline RL (next lecture) learns from logged data that may include non-expert behaviour, using rewards.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

🎮 Reinforcement Learning

Multi-Agent Reinforcement Learning: Cooperation, Competition and Equilibria

When many learning agents share an environment, each faces a moving target. We introduce Markov games, Nash equilibria, independent learners, centralised training with decentralised execution, self-play and emergent behaviour.

Advanced⏱ 5 min#238
🎮 Reinforcement Learning

Offline Reinforcement Learning: Learning from Logged Data

In many domains, trial-and-error exploration is unsafe or impossible, but logged data exists. Offline RL learns policies from fixed datasets. We explain the distributional shift problem and methods such as BCQ, CQL and IQL, plus off-policy evaluation.

Advanced⏱ 5 min#240
🎮 Reinforcement Learning

AlphaGo, AlphaZero and MuZero: Search Meets Deep Learning

DeepMind's Go programs combined deep neural networks with Monte Carlo tree search and self-play. We trace AlphaGo's supervised and RL training, AlphaZero's tabula-rasa self-play, MuZero's learned model, and the lessons for AI.

Advanced⏱ 6 min#237