REINFORCE waits until the end of an episode and suffers from high variance. Actor–critic methods fix both problems by learning a critic — a value function — that evaluates the actor's actions after every step. The actor learns from the critic's feedback; the critic learns from the rewards. This combination of policy gradients with TD learning is the backbone of most modern deep RL algorithms, including PPO and SAC.
The idea#
- Actor: policy $\pi_\theta(a \mid s)$, updated in the direction suggested by the critic.
- Critic: value function $V_w(s)$ (or $Q_w(s, a)$), updated by TD learning.
Replace the Monte Carlo return in the policy gradient with a bootstrapped estimate of the advantage, using the TD error:
$\delta_t$ is an estimate of $A(s_t, a_t) = Q(s_t, a_t) - V(s_t)$: positive if the action turned out better than expected.
One-step actor–critic updates:
Updates happen every step, with far lower variance than REINFORCE — at the cost of bias from the critic's errors.
A2C and A3C#
A3C (Asynchronous Advantage Actor–Critic, Mnih et al., 2016) ran many actor–learners in parallel on CPU threads, each interacting with its own copy of the environment and asynchronously updating shared parameters. Parallel, diverse experience decorrelated the data without a replay buffer, enabling on-policy deep RL to work stably.
A2C is the synchronous version: $N$ parallel environments step together; their batch of transitions produces one gradient update. It is simpler, uses GPUs better, and performs as well or better.
The combined loss for a batch:
The entropy bonus encourages the policy to stay stochastic longer, preventing premature convergence to a deterministic, possibly suboptimal policy.
Generalised Advantage Estimation (GAE)#
How should we estimate advantages? One-step TD errors are low-variance but biased; Monte Carlo returns are unbiased but high-variance. Schulman et al. (2016) proposed GAE, an exponentially weighted average of multi-step advantage estimates — the λ-return idea applied to advantages:
$\lambda = 0$ gives the one-step TD advantage; $\lambda = 1$ gives the Monte Carlo advantage. Values around 0.95 work well in practice. GAE is computed efficiently backwards over a rollout:
import torch
def compute_gae(rewards, values, dones, last_value, gamma=0.99, lam=0.95):
"""rewards, values, dones: tensors of shape (T,). Returns advantages and value targets."""
T = len(rewards)
adv = torch.zeros(T)
gae, next_value = 0.0, last_value
for t in reversed(range(T)):
nonterminal = 1.0 - dones[t]
delta = rewards[t] + gamma * next_value * nonterminal - values[t]
gae = delta + gamma * lam * nonterminal * gae
adv[t] = gae
next_value = values[t]
return adv, adv + values # advantages and returns (critic targets)
r = torch.tensor([1.0, 1.0, 1.0, 0.0]); v = torch.tensor([2.5, 2.0, 1.2, 0.5]); d = torch.tensor([0, 0, 0, 1.0])
print(compute_gae(r, v, d, last_value=0.0))Architecture choices#
Actor and critic can be separate networks or share a trunk with two heads. Sharing saves computation and helps representation learning, but the value and policy losses can interfere; separate networks are often more stable for continuous control.
Why actor–critic dominates#
- Works with discrete and continuous actions.
- Learns online from partial episodes.
- Controllable bias–variance trade-off (via GAE's λ and n-step returns).
- Scales with parallel environments.
Its main remaining issue is step-size sensitivity: a large policy update can wreck performance, and on-policy data is discarded after use. The next lecture shows how PPO and TRPO constrain updates.