A robot arm does not choose among 18 joystick buttons; it outputs continuous torques for each joint. Q-learning's $\max_a Q(s, a)$ becomes an optimisation problem at every step, and on-policy methods like PPO discard data after each update — costly when experience comes from real hardware or slow simulators. A family of off-policy actor–critic methods addresses continuous control with good sample efficiency: DDPG, TD3 and SAC.
Deterministic policy gradients#
Silver et al. (2014) showed that for a deterministic policy $a = \mu_\theta(s)$, the policy gradient takes a simple form:
Move the action in the direction that increases the critic's Q-value — backpropagate through the critic into the actor.
DDPG#
Deep Deterministic Policy Gradient (Lillicrap et al., 2016) combined this with DQN's stabilisers:
- An actor $\mu_\theta(s)$ and a critic $Q_\phi(s, a)$ (the action is an input to the critic).
- Experience replay (off-policy).
- Target networks for both actor and critic, updated softly: $\phi' \leftarrow \tau\phi + (1 - \tau)\phi'$.
- Exploration noise added to actions (Gaussian or Ornstein–Uhlenbeck noise), since the policy is deterministic.
Critic target:
DDPG learned many continuous control tasks from state vectors and even pixels, but it is brittle: sensitive to hyperparameters and prone to Q-value overestimation, which the actor exploits.
TD3: Twin Delayed DDPG#
Fujimoto, van Hoof and Meger (2018) identified overestimation as DDPG's main failure and proposed three fixes:
- Clipped double Q-learning: learn two critics and use the minimum for targets:
- Target policy smoothing: add clipped noise to the target action, $\tilde{a}' = \mu_{\theta'}(s') + \text{clip}(\epsilon, -c, c)$, so the critic cannot exploit sharp peaks in Q.
- Delayed policy updates: update the actor (and targets) less often than the critics (e.g. every 2 critic steps), letting value estimates settle.
TD3 was substantially more reliable than DDPG on standard MuJoCo benchmarks.
Soft Actor–Critic (SAC)#
Haarnoja et al. (2018) framed control as maximum-entropy RL: maximise reward and policy entropy,
The temperature $\alpha$ trades off reward and randomness. Benefits: built-in exploration, robustness (the policy keeps multiple good behaviours), and stable learning. SAC uses:
- a stochastic Gaussian policy with a tanh squashing to bound actions, trained via the reparameterisation trick;
- twin critics with the clipped double-Q minimum (as in TD3), and a soft value target including the entropy term:
- automatic temperature tuning: adjust $\alpha$ so that policy entropy stays near a target (e.g. $-\dim(\mathcal{A})$).
SAC became a default choice for continuous control thanks to its sample efficiency and robustness to hyperparameters, and it has been used to train real robots (e.g. learning to walk on hardware within hours).
import torch, torch.nn as nn
class GaussianActor(nn.Module):
def __init__(self, obs_dim, act_dim, hidden=256, act_limit=1.0):
super().__init__()
self.net = nn.Sequential(nn.Linear(obs_dim, hidden), nn.ReLU(), nn.Linear(hidden, hidden), nn.ReLU())
self.mu, self.log_std = nn.Linear(hidden, act_dim), nn.Linear(hidden, act_dim)
self.act_limit = act_limit
def forward(self, s):
h = self.net(s)
mu, log_std = self.mu(h), self.log_std(h).clamp(-20, 2)
std = log_std.exp()
u = mu + std * torch.randn_like(mu) # reparameterised sample
a = torch.tanh(u) # squash to (-1, 1)
logp = torch.distributions.Normal(mu, std).log_prob(u).sum(-1)
logp -= torch.log(1 - a.pow(2) + 1e-6).sum(-1) # tanh change-of-variables correction
return self.act_limit * a, logp
actor = GaussianActor(obs_dim=17, act_dim=6)
a, logp = actor(torch.randn(4, 17))
print(a.shape, logp.shape)
# In practice:
# from stable_baselines3 import SAC
# SAC("MlpPolicy", "Pendulum-v1", verbose=0).learn(50_000)Choosing an algorithm#
| Situation | Suggested |
|---|---|
| Continuous actions, sample efficiency matters | SAC (or TD3) |
| Massive parallel simulation, wall-clock matters | PPO |
| Discrete actions | DQN family or PPO |
| Need simple, robust baseline | PPO or SAC via Stable-Baselines3 |
Sim-to-real#
Most continuous-control policies are trained in simulation and transferred to real robots. The reality gap (differences in friction, delays, sensor noise) is bridged with domain randomisation (randomising physics parameters during training), system identification and fine-tuning on hardware — a topic of the final lecture of this track.