🎮 Reinforcement Learning · Lecture 13 of 21

Actor–Critic Methods: A2C, A3C and Advantage Estimation

Actor–critic methods pair a policy (actor) with a learned value function (critic) to cut variance and learn online. We derive the one-step actor–critic, A2C/A3C, entropy regularisation and Generalised Advantage Estimation.

REINFORCE waits until the end of an episode and suffers from high variance. Actor–critic methods fix both problems by learning a critic — a value function — that evaluates the actor's actions after every step. The actor learns from the critic's feedback; the critic learns from the rewards. This combination of policy gradients with TD learning is the backbone of most modern deep RL algorithms, including PPO and SAC.

The idea#

  • Actor: policy $\pi_\theta(a \mid s)$, updated in the direction suggested by the critic.
  • Critic: value function $V_w(s)$ (or $Q_w(s, a)$), updated by TD learning.

Replace the Monte Carlo return in the policy gradient with a bootstrapped estimate of the advantage, using the TD error:

$$ \delta_t = r_{t+1} + \gamma V_w(s_{t+1}) - V_w(s_t) $$

$\delta_t$ is an estimate of $A(s_t, a_t) = Q(s_t, a_t) - V(s_t)$: positive if the action turned out better than expected.

One-step actor–critic updates:

$$ w \leftarrow w + \alpha_w\,\delta_t\,\nabla_wV_w(s_t), \qquad \theta \leftarrow \theta + \alpha_\theta\,\delta_t\,\nabla_\theta\log\pi_\theta(a_t \mid s_t) $$

Updates happen every step, with far lower variance than REINFORCE — at the cost of bias from the critic's errors.

A2C and A3C#

A3C (Asynchronous Advantage Actor–Critic, Mnih et al., 2016) ran many actor–learners in parallel on CPU threads, each interacting with its own copy of the environment and asynchronously updating shared parameters. Parallel, diverse experience decorrelated the data without a replay buffer, enabling on-policy deep RL to work stably.

A2C is the synchronous version: $N$ parallel environments step together; their batch of transitions produces one gradient update. It is simpler, uses GPUs better, and performs as well or better.

The combined loss for a batch:

$$ \mathcal{L} = \underbrace{-\sum_t\log\pi_\theta(a_t \mid s_t)\,\hat{A}_t}_{\text{policy}} + c_v\underbrace{\sum_t\big(\hat{R}_t - V_w(s_t)\big)^2}_{\text{value}} - c_e\underbrace{\sum_t\mathcal{H}\big(\pi_\theta(\cdot \mid s_t)\big)}_{\text{entropy bonus}} $$

The entropy bonus encourages the policy to stay stochastic longer, preventing premature convergence to a deterministic, possibly suboptimal policy.

Generalised Advantage Estimation (GAE)#

How should we estimate advantages? One-step TD errors are low-variance but biased; Monte Carlo returns are unbiased but high-variance. Schulman et al. (2016) proposed GAE, an exponentially weighted average of multi-step advantage estimates — the λ-return idea applied to advantages:

$$ \hat{A}_t^{\text{GAE}(\gamma, \lambda)} = \sum_{l=0}^{\infty}(\gamma\lambda)^l\,\delta_{t+l} $$

$\lambda = 0$ gives the one-step TD advantage; $\lambda = 1$ gives the Monte Carlo advantage. Values around 0.95 work well in practice. GAE is computed efficiently backwards over a rollout:

python
import torch

def compute_gae(rewards, values, dones, last_value, gamma=0.99, lam=0.95):
    """rewards, values, dones: tensors of shape (T,). Returns advantages and value targets."""
    T = len(rewards)
    adv = torch.zeros(T)
    gae, next_value = 0.0, last_value
    for t in reversed(range(T)):
        nonterminal = 1.0 - dones[t]
        delta = rewards[t] + gamma * next_value * nonterminal - values[t]
        gae = delta + gamma * lam * nonterminal * gae
        adv[t] = gae
        next_value = values[t]
    return adv, adv + values                        # advantages and returns (critic targets)

r = torch.tensor([1.0, 1.0, 1.0, 0.0]); v = torch.tensor([2.5, 2.0, 1.2, 0.5]); d = torch.tensor([0, 0, 0, 1.0])
print(compute_gae(r, v, d, last_value=0.0))

Architecture choices#

Actor and critic can be separate networks or share a trunk with two heads. Sharing saves computation and helps representation learning, but the value and policy losses can interfere; separate networks are often more stable for continuous control.

Why actor–critic dominates#

  • Works with discrete and continuous actions.
  • Learns online from partial episodes.
  • Controllable bias–variance trade-off (via GAE's λ and n-step returns).
  • Scales with parallel environments.

Its main remaining issue is step-size sensitivity: a large policy update can wreck performance, and on-policy data is discarded after use. The next lecture shows how PPO and TRPO constrain updates.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

🎮 Reinforcement Learning

Policy Gradient Methods: REINFORCE and the Policy Gradient Theorem

Instead of learning values and acting greedily, policy gradient methods optimise the policy directly. We derive the policy gradient theorem and REINFORCE, reduce variance with baselines, and implement it on CartPole.

Advanced⏱ 5 min#232
🎮 Reinforcement Learning

PPO and TRPO: Stable Policy Optimisation with Trust Regions

Large policy updates can destroy performance. TRPO constrains each update with a KL trust region; PPO achieves similar stability with a simple clipped objective. We derive both and implement PPO's core update.

Advanced⏱ 5 min#234
🎮 Reinforcement Learning

Improving DQN: Double, Dueling, Prioritised Replay and Rainbow

A series of improvements fixed DQN's weaknesses: overestimation, inefficient replay, poor value decomposition and myopic returns. We study each idea and how Rainbow combined six of them.

Advanced⏱ 5 min#231