🎮 Reinforcement Learning · Lecture 15 of 21

Continuous Control: DDPG, TD3 and Soft Actor–Critic

Robots need continuous actions — torques, velocities, steering angles. We study off-policy actor–critic methods for continuous control: DDPG's deterministic policy gradient, TD3's fixes, and SAC's maximum-entropy framework.

A robot arm does not choose among 18 joystick buttons; it outputs continuous torques for each joint. Q-learning's $\max_a Q(s, a)$ becomes an optimisation problem at every step, and on-policy methods like PPO discard data after each update — costly when experience comes from real hardware or slow simulators. A family of off-policy actor–critic methods addresses continuous control with good sample efficiency: DDPG, TD3 and SAC.

Deterministic policy gradients#

Silver et al. (2014) showed that for a deterministic policy $a = \mu_\theta(s)$, the policy gradient takes a simple form:

$$ \nabla_\theta J = \mathbb{E}_{s}\Big[\nabla_aQ(s, a)\big|_{a = \mu_\theta(s)}\,\nabla_\theta\mu_\theta(s)\Big] $$

Move the action in the direction that increases the critic's Q-value — backpropagate through the critic into the actor.

DDPG#

Deep Deterministic Policy Gradient (Lillicrap et al., 2016) combined this with DQN's stabilisers:

  • An actor $\mu_\theta(s)$ and a critic $Q_\phi(s, a)$ (the action is an input to the critic).
  • Experience replay (off-policy).
  • Target networks for both actor and critic, updated softly: $\phi' \leftarrow \tau\phi + (1 - \tau)\phi'$.
  • Exploration noise added to actions (Gaussian or Ornstein–Uhlenbeck noise), since the policy is deterministic.

Critic target:

$$ y = r + \gamma\,Q_{\phi'}\big(s', \mu_{\theta'}(s')\big) $$

DDPG learned many continuous control tasks from state vectors and even pixels, but it is brittle: sensitive to hyperparameters and prone to Q-value overestimation, which the actor exploits.

TD3: Twin Delayed DDPG#

Fujimoto, van Hoof and Meger (2018) identified overestimation as DDPG's main failure and proposed three fixes:

  1. Clipped double Q-learning: learn two critics and use the minimum for targets:
$$ y = r + \gamma\min_{i=1,2}Q_{\phi'_i}\big(s', \tilde{a}'\big) $$
  1. Target policy smoothing: add clipped noise to the target action, $\tilde{a}' = \mu_{\theta'}(s') + \text{clip}(\epsilon, -c, c)$, so the critic cannot exploit sharp peaks in Q.
  2. Delayed policy updates: update the actor (and targets) less often than the critics (e.g. every 2 critic steps), letting value estimates settle.

TD3 was substantially more reliable than DDPG on standard MuJoCo benchmarks.

Soft Actor–Critic (SAC)#

Haarnoja et al. (2018) framed control as maximum-entropy RL: maximise reward and policy entropy,

$$ J(\pi) = \sum_t\mathbb{E}\Big[r(s_t, a_t) + \alpha\,\mathcal{H}\big(\pi(\cdot \mid s_t)\big)\Big] $$

The temperature $\alpha$ trades off reward and randomness. Benefits: built-in exploration, robustness (the policy keeps multiple good behaviours), and stable learning. SAC uses:

  • a stochastic Gaussian policy with a tanh squashing to bound actions, trained via the reparameterisation trick;
  • twin critics with the clipped double-Q minimum (as in TD3), and a soft value target including the entropy term:
$$ y = r + \gamma\Big(\min_iQ_{\phi'_i}(s', a') - \alpha\log\pi_\theta(a' \mid s')\Big), \quad a' \sim \pi_\theta(\cdot \mid s') $$
  • automatic temperature tuning: adjust $\alpha$ so that policy entropy stays near a target (e.g. $-\dim(\mathcal{A})$).

SAC became a default choice for continuous control thanks to its sample efficiency and robustness to hyperparameters, and it has been used to train real robots (e.g. learning to walk on hardware within hours).

python
import torch, torch.nn as nn

class GaussianActor(nn.Module):
    def __init__(self, obs_dim, act_dim, hidden=256, act_limit=1.0):
        super().__init__()
        self.net = nn.Sequential(nn.Linear(obs_dim, hidden), nn.ReLU(), nn.Linear(hidden, hidden), nn.ReLU())
        self.mu, self.log_std = nn.Linear(hidden, act_dim), nn.Linear(hidden, act_dim)
        self.act_limit = act_limit
    def forward(self, s):
        h = self.net(s)
        mu, log_std = self.mu(h), self.log_std(h).clamp(-20, 2)
        std = log_std.exp()
        u = mu + std * torch.randn_like(mu)                        # reparameterised sample
        a = torch.tanh(u)                                          # squash to (-1, 1)
        logp = torch.distributions.Normal(mu, std).log_prob(u).sum(-1)
        logp -= torch.log(1 - a.pow(2) + 1e-6).sum(-1)             # tanh change-of-variables correction
        return self.act_limit * a, logp

actor = GaussianActor(obs_dim=17, act_dim=6)
a, logp = actor(torch.randn(4, 17))
print(a.shape, logp.shape)

# In practice:
# from stable_baselines3 import SAC
# SAC("MlpPolicy", "Pendulum-v1", verbose=0).learn(50_000)

Choosing an algorithm#

SituationSuggested
Continuous actions, sample efficiency mattersSAC (or TD3)
Massive parallel simulation, wall-clock mattersPPO
Discrete actionsDQN family or PPO
Need simple, robust baselinePPO or SAC via Stable-Baselines3

Sim-to-real#

Most continuous-control policies are trained in simulation and transferred to real robots. The reality gap (differences in friction, delays, sensor noise) is bridged with domain randomisation (randomising physics parameters during training), system identification and fine-tuning on hardware — a topic of the final lecture of this track.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

🎮 Reinforcement Learning

Reinforcement Learning in the Real World: Robotics, Sim-to-Real and Safety

Games are forgiving; the real world is not. We examine the challenges of deploying RL — sample efficiency, safety, reward design, partial observability — and the techniques that make it work: simulation, domain randomisation, safe RL and human oversight.

Advanced⏱ 5 min#241
🎮 Reinforcement Learning

PPO and TRPO: Stable Policy Optimisation with Trust Regions

Large policy updates can destroy performance. TRPO constrains each update with a KL trust region; PPO achieves similar stability with a simple clipped objective. We derive both and implement PPO's core update.

Advanced⏱ 5 min#234
🎮 Reinforcement Learning

Model-Based Reinforcement Learning: Learning and Planning with World Models

Model-based agents learn a model of the environment and use it to plan or generate imagined experience. We cover Dyna, model-predictive control, model errors and ensembles, MBPO, and latent world models like Dreamer and MuZero.

Advanced⏱ 5 min#236