Reinforcement learning's most famous successes happened in games and simulators, where an agent can fail millions of times at no cost. The physical and social world is different: robots break, experiments are slow, rewards are unclear and mistakes can hurt people. Yet RL has achieved remarkable real-world results — dexterous robot hands, agile legged locomotion, data-centre cooling, and aligning language models. This closing lecture of the track examines what it takes to make RL work outside the simulator.
Challenges of real-world RL#
Dulac-Arnold et al. (2019) catalogued the main challenges:
- Sample efficiency: real interactions are slow and expensive.
- Safety: exploration must not cause damage or harm.
- Reward specification: rewards are hard to define and easy to game.
- Partial observability and noise: sensors are imperfect; delays exist.
- High-dimensional, continuous state and action spaces.
- Non-stationarity: the world changes (wear, seasons, user behaviour).
- Constraints: physical limits, regulations, budgets.
- Explainability and the need for operators to trust the system.
- Learning from limited logged data (offline RL).
Simulation and sim-to-real transfer#
The dominant strategy is to train in simulation and transfer to reality. The difference between simulator and world — the reality gap — causes policies to fail on hardware. Techniques:
- Domain randomisation: randomise physical parameters (masses, friction, motor strength, delays), visual appearance (textures, lighting, camera position) and noise during training, so the real world looks like just another variation. OpenAI's Dactyl (2018–2019) used massive domain randomisation to train a robot hand to manipulate a cube and solve a Rubik's Cube.
- System identification: measure real parameters and calibrate the simulator.
- Domain adaptation: adapt the policy or its perception with a small amount of real data.
- Teacher–student training: a "teacher" policy with privileged simulator information (exact terrain, friction) trains a "student" that uses only real sensors — used for robust quadruped locomotion over rough terrain.
- Massively parallel GPU simulation (e.g. Isaac Gym-style simulators) lets legged robots learn to walk in minutes of wall-clock time.
import numpy as np
class RandomizedPendulumParams:
"""Sample physics parameters each episode for domain randomisation."""
def __init__(self, rng=None):
self.rng = rng or np.random.default_rng()
def sample(self):
return {
"mass": self.rng.uniform(0.8, 1.2), # ±20% around nominal
"length": self.rng.uniform(0.9, 1.1),
"friction": self.rng.uniform(0.0, 0.1),
"action_delay_steps": int(self.rng.integers(0, 3)),
"obs_noise_std": self.rng.uniform(0.0, 0.02),
}
params = RandomizedPendulumParams(np.random.default_rng(0))
print([params.sample()["mass"].round(3) for _ in range(5)])
# At every env.reset(), apply params.sample() to the simulator before the episode.Safe reinforcement learning#
Safety must be designed in:
- Constrained MDPs: maximise reward subject to expected cost limits, $\mathbb{E}[\sum_t\gamma^tc_t] \le d$, solved with Lagrangian methods or constrained policy optimisation (CPO).
- Safety layers and shields: a verified filter overrides actions that would violate constraints (e.g. keep a drone within a geofence).
- Conservative exploration: explore only within a known-safe region; expand it cautiously.
- Learning from demonstrations and offline data to avoid dangerous random exploration.
- Human oversight: emergency stops, approval of new policies, gradual rollout.
Reward design in practice#
- Start from the true objective and measurable outcomes; avoid rewarding proxies that can be gamed.
- Combine sparse task rewards with carefully designed shaping (potential-based shaping preserves optimal policies — Ng et al., 1999).
- Add explicit penalties for unsafe or undesirable side effects, and test for specification gaming in simulation before deployment.
- When a reward is too hard to write, learn it from demonstrations (IRL) or preferences (RLHF).
Success stories#
- Robotics: dexterous in-hand manipulation (Dactyl), robust quadruped and humanoid locomotion trained in simulation, robot learning from large demonstration datasets combined with RL fine-tuning.
- Industrial control: DeepMind reported substantial energy savings in data-centre cooling using RL-based recommendations and later autonomous control with safety constraints; RL has been explored for controlling plasma shapes in a tokamak fusion reactor (trained in simulation).
- Recommendation and operations: contextual bandits and RL for personalisation, inventory and logistics.
- Language models: RLHF and RL with verifiable rewards — perhaps RL's most widespread deployment today.