A hospital cannot let an RL agent experiment with treatments on patients to see what happens. A public agency cannot randomly vary how it allocates assistance just to explore. Yet both have years of logged data โ decisions taken and outcomes observed. Offline RL (also called batch RL) aims to learn good policies purely from such fixed datasets, without any new interaction. It promises to bring RL to healthcare, education, logistics and recommendation โ but it is fundamentally harder than it first appears.
The setting#
Given a dataset $\mathcal{D} = \{(s_i, a_i, r_i, s'_i)\}$ collected by some behaviour policy $\pi_\beta$ (humans, an old system, a mix), learn a policy $\pi$ that performs well when deployed โ without collecting more data.
Why not just run off-policy RL on the dataset?#
Q-learning is off-policy, so in principle it can learn from any data. In practice, naive offline Q-learning or DQN on a fixed dataset fails badly. The reason is distributional shift combined with the max in the Bellman target:
The max considers all actions, including actions that never appear in the data for state $s'$. Their Q-values are pure extrapolation by the function approximator โ often wildly overestimated. The policy then prefers exactly these out-of-distribution (OOD) actions, and because no new data is collected, the errors are never corrected. In online RL, trying an overestimated action reveals the truth; offline, the agent can never check.
Approach 1: constrain the policy to the data#
Keep the learned policy close to the behaviour policy:
- BCQ (Batch-Constrained Q-learning, Fujimoto et al., 2019): a generative model proposes actions similar to those in the data; the policy chooses only among them.
- TD3+BC (Fujimoto & Gu, 2021): add a behavioural-cloning term to the TD3 actor loss, $\max_\pi\big[\lambda Q(s, \pi(s)) - (\pi(s) - a)^2\big]$ โ remarkably simple and strong.
- BRAC / BEAR: penalise divergence (KL, MMD) from the behaviour policy.
Approach 2: pessimistic values#
Make OOD actions look bad rather than good:
- Conservative Q-Learning (CQL) (Kumar et al., 2020) adds a regulariser that pushes down Q-values on actions sampled from the learned policy (or all actions via log-sum-exp) and pushes up Q-values on dataset actions:
The learned Q-function lower-bounds the true value of the policy โ a principled form of pessimism.
Approach 3: avoid querying OOD actions at all#
- Implicit Q-Learning (IQL) (Kostrikov et al., 2022) never evaluates actions outside the dataset. It fits a value function with expectile regression on dataset Q-values (approximating a max over in-distribution actions), and extracts a policy by advantage-weighted regression: behavioural cloning weighted by $\exp(\beta A(s, a))$ โ imitate the good actions in the data more strongly.
import torch
def expectile_loss(diff, tau=0.7):
"""Asymmetric L2: weight positive errors by tau, negative by (1 - tau)."""
weight = torch.where(diff > 0, tau, 1 - tau)
return (weight * diff.pow(2)).mean()
def awr_policy_loss(log_prob_data_actions, advantages, beta=3.0, max_weight=100.0):
"""Advantage-weighted regression: clone dataset actions, weighted by exp(beta * A)."""
w = torch.exp(beta * advantages).clamp(max=max_weight)
return -(w.detach() * log_prob_data_actions).mean()
diff = torch.randn(256) # Q(s, a) - V(s) on dataset pairs
print(expectile_loss(diff).item(), awr_policy_loss(torch.randn(256), torch.randn(256) * 0.1).item())Off-policy evaluation (OPE)#
Before deploying a learned policy, estimate its value from logged data:
- Importance sampling estimators (high variance for long horizons);
- Direct methods (fit a model or Q-function and evaluate);
- Doubly robust estimators combining both.
OPE is hard and estimates can be badly wrong; in high-stakes settings, follow OPE with carefully monitored small-scale pilots.
Data matters most#
Offline RL cannot find good actions that the data never tried. Dataset coverage and quality determine what is achievable:
- Expert-only data โ behavioural cloning may do as well.
- Diverse, "medium" data with varied behaviour โ offline RL can outperform the behaviour policy by stitching together good parts of different trajectories.
- Confounded data (decisions based on information not recorded) โ estimates can be biased, a problem shared with causal inference.