🎮 Reinforcement Learning · Lecture 21 of 21

Reinforcement Learning in the Real World: Robotics, Sim-to-Real and Safety

Games are forgiving; the real world is not. We examine the challenges of deploying RL — sample efficiency, safety, reward design, partial observability — and the techniques that make it work: simulation, domain randomisation, safe RL and human oversight.

Reinforcement learning's most famous successes happened in games and simulators, where an agent can fail millions of times at no cost. The physical and social world is different: robots break, experiments are slow, rewards are unclear and mistakes can hurt people. Yet RL has achieved remarkable real-world results — dexterous robot hands, agile legged locomotion, data-centre cooling, and aligning language models. This closing lecture of the track examines what it takes to make RL work outside the simulator.

Challenges of real-world RL#

Dulac-Arnold et al. (2019) catalogued the main challenges:

  1. Sample efficiency: real interactions are slow and expensive.
  2. Safety: exploration must not cause damage or harm.
  3. Reward specification: rewards are hard to define and easy to game.
  4. Partial observability and noise: sensors are imperfect; delays exist.
  5. High-dimensional, continuous state and action spaces.
  6. Non-stationarity: the world changes (wear, seasons, user behaviour).
  7. Constraints: physical limits, regulations, budgets.
  8. Explainability and the need for operators to trust the system.
  9. Learning from limited logged data (offline RL).

Simulation and sim-to-real transfer#

The dominant strategy is to train in simulation and transfer to reality. The difference between simulator and world — the reality gap — causes policies to fail on hardware. Techniques:

  • Domain randomisation: randomise physical parameters (masses, friction, motor strength, delays), visual appearance (textures, lighting, camera position) and noise during training, so the real world looks like just another variation. OpenAI's Dactyl (2018–2019) used massive domain randomisation to train a robot hand to manipulate a cube and solve a Rubik's Cube.
  • System identification: measure real parameters and calibrate the simulator.
  • Domain adaptation: adapt the policy or its perception with a small amount of real data.
  • Teacher–student training: a "teacher" policy with privileged simulator information (exact terrain, friction) trains a "student" that uses only real sensors — used for robust quadruped locomotion over rough terrain.
  • Massively parallel GPU simulation (e.g. Isaac Gym-style simulators) lets legged robots learn to walk in minutes of wall-clock time.
python
import numpy as np

class RandomizedPendulumParams:
    """Sample physics parameters each episode for domain randomisation."""
    def __init__(self, rng=None):
        self.rng = rng or np.random.default_rng()
    def sample(self):
        return {
            "mass": self.rng.uniform(0.8, 1.2),           # ±20% around nominal
            "length": self.rng.uniform(0.9, 1.1),
            "friction": self.rng.uniform(0.0, 0.1),
            "action_delay_steps": int(self.rng.integers(0, 3)),
            "obs_noise_std": self.rng.uniform(0.0, 0.02),
        }

params = RandomizedPendulumParams(np.random.default_rng(0))
print([params.sample()["mass"].round(3) for _ in range(5)])
# At every env.reset(), apply params.sample() to the simulator before the episode.

Safe reinforcement learning#

Safety must be designed in:

  • Constrained MDPs: maximise reward subject to expected cost limits, $\mathbb{E}[\sum_t\gamma^tc_t] \le d$, solved with Lagrangian methods or constrained policy optimisation (CPO).
  • Safety layers and shields: a verified filter overrides actions that would violate constraints (e.g. keep a drone within a geofence).
  • Conservative exploration: explore only within a known-safe region; expand it cautiously.
  • Learning from demonstrations and offline data to avoid dangerous random exploration.
  • Human oversight: emergency stops, approval of new policies, gradual rollout.

Reward design in practice#

  • Start from the true objective and measurable outcomes; avoid rewarding proxies that can be gamed.
  • Combine sparse task rewards with carefully designed shaping (potential-based shaping preserves optimal policies — Ng et al., 1999).
  • Add explicit penalties for unsafe or undesirable side effects, and test for specification gaming in simulation before deployment.
  • When a reward is too hard to write, learn it from demonstrations (IRL) or preferences (RLHF).

Success stories#

  • Robotics: dexterous in-hand manipulation (Dactyl), robust quadruped and humanoid locomotion trained in simulation, robot learning from large demonstration datasets combined with RL fine-tuning.
  • Industrial control: DeepMind reported substantial energy savings in data-centre cooling using RL-based recommendations and later autonomous control with safety constraints; RL has been explored for controlling plasma shapes in a tokamak fusion reactor (trained in simulation).
  • Recommendation and operations: contextual bandits and RL for personalisation, inventory and logistics.
  • Language models: RLHF and RL with verifiable rewards — perhaps RL's most widespread deployment today.

Responsible deployment#

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

🎮 Reinforcement Learning

Continuous Control: DDPG, TD3 and Soft Actor–Critic

Robots need continuous actions — torques, velocities, steering angles. We study off-policy actor–critic methods for continuous control: DDPG's deterministic policy gradient, TD3's fixes, and SAC's maximum-entropy framework.

Advanced⏱ 5 min#235
🎮 Reinforcement Learning

Offline Reinforcement Learning: Learning from Logged Data

In many domains, trial-and-error exploration is unsafe or impossible, but logged data exists. Offline RL learns policies from fixed datasets. We explain the distributional shift problem and methods such as BCQ, CQL and IQL, plus off-policy evaluation.

Advanced⏱ 5 min#240
🎮 Reinforcement Learning

Imitation Learning and Inverse Reinforcement Learning

When rewards are hard to specify but demonstrations are available, agents can learn by imitation. We cover behavioural cloning and its compounding errors, DAgger, inverse RL, maximum-entropy IRL and adversarial imitation (GAIL).

Advanced⏱ 5 min#239