Writing a reward function for "drive safely and courteously" or "fold laundry neatly" is extremely difficult. Showing someone how to do it is easy. Imitation learning trains agents from expert demonstrations rather than hand-designed rewards. It underpins autonomous-driving research, robot manipulation, game agents — and, in a sense, supervised fine-tuning of language models on human-written responses.
Behavioural cloning (BC)#
The simplest approach: treat demonstrations as a supervised dataset of (state, action) pairs and train a policy by classification or regression:
It is simple, needs no environment interaction, and works well when demonstrations are plentiful and cover the situations the agent will face. ALVINN (1989) steered a vehicle with a neural network trained on human driving; modern robot learning uses BC with expressive policies (e.g. diffusion policies or transformers) and large demonstration datasets.
The compounding-error problem#
BC violates the i.i.d. assumption. The expert rarely makes mistakes, so demonstrations contain few examples of recovering from errors. When the learned policy makes a small mistake, it drifts into states unlike the training data, where it makes bigger mistakes, drifting further — errors compound. Ross and Bagnell showed that the expected cost of BC can grow quadratically with the task horizon $T$ (as $O(\epsilon T^2)$ for per-step error rate $\epsilon$), whereas an ideal learner's cost grows linearly.
DAgger: dataset aggregation#
Ross, Gordon and Bagnell (2011) proposed DAgger:
- Train a policy on expert data.
- Run the learner's policy to collect the states it actually visits.
- Ask the expert to label the correct action for those states.
- Aggregate into the dataset and retrain; repeat.
DAgger trains on the learner's own state distribution, fixing compounding errors and achieving linear-in-horizon error bounds. The cost: an expert available to label on demand (possible in simulation or with a scripted planner; harder with humans).
import numpy as np
from sklearn.neighbors import KNeighborsClassifier
def dagger(env, expert, n_iters=5, rollouts=10):
X, y = [], []
policy = None
for it in range(n_iters):
for _ in range(rollouts):
s, _ = env.reset(); done = False
while not done:
X.append(s); y.append(expert(s)) # expert labels every visited state
a = expert(s) if policy is None else int(policy.predict([s])[0]) # learner acts
s, _, term, trunc, _ = env.step(a); done = term or trunc
policy = KNeighborsClassifier(5).fit(np.array(X), np.array(y)) # retrain on aggregated data
return policyInverse reinforcement learning (IRL)#
Instead of copying actions, infer the reward function that the expert appears to optimise, then find a policy for that reward with RL (Ng & Russell, 2000; Abbeel & Ng, 2004). Motivations:
- A reward is a compact, transferable description of the task — it generalises to new situations and dynamics better than copied actions.
- Understanding why the expert acts, not just what they do.
The problem is ill-posed: many rewards explain the same behaviour (a zero reward makes every policy optimal). Maximum-entropy IRL (Ziebart et al., 2008) resolves ambiguity by assuming the expert chooses trajectories with probability
and fits $\psi$ by maximum likelihood — preferring the least committed explanation consistent with demonstrations. It was used, for example, to model taxi drivers' route choices.
Adversarial imitation: GAIL#
Generative Adversarial Imitation Learning (Ho & Ermon, 2016) links IRL with GANs. A discriminator $D(s, a)$ learns to distinguish expert state–action pairs from the agent's; the agent (trained with RL, e.g. TRPO/PPO) receives reward for fooling the discriminator, such as $-\log(1 - D(s, a))$. GAIL matches the expert's state–action distribution directly, avoiding compounding errors, with far fewer demonstrations than BC — though it requires environment interaction.
Comparing approaches#
| Method | Needs environment interaction | Needs interactive expert | Handles compounding errors | Recovers reward |
|---|---|---|---|---|
| Behavioural cloning | No | No | No | No |
| DAgger | Yes | Yes | Yes | No |
| IRL (e.g. MaxEnt) | Yes (RL inner loop) | No | Yes | Yes |
| GAIL | Yes | No | Yes | Implicitly |
Connections and cautions#
- RLHF is related: rather than demonstrations, it learns a reward from preferences — another way of inferring human intent.
- Offline RL (next lecture) learns from logged data that may include non-expert behaviour, using rewards.