Traffic intersections with many autonomous vehicles, teams of warehouse robots, trading agents in markets, players in a strategy game, LLM agents negotiating on behalf of users โ many real problems involve multiple agents acting simultaneously. Multi-agent reinforcement learning (MARL) studies how agents learn when their outcomes depend on each other's behaviour. It brings together RL and game theory, and introduces challenges absent in single-agent RL.
Markov games#
A Markov game (stochastic game) generalises the MDP to $N$ agents:
- a shared state $s$;
- each agent $i$ chooses an action $a^i$; the joint action $\mathbf{a} = (a^1, \dots, a^N)$ determines the transition $P(s' \mid s, \mathbf{a})$;
- each agent receives its own reward $r^i(s, \mathbf{a})$.
Settings:
| Setting | Rewards | Example |
|---|---|---|
| Fully cooperative | Shared reward | Robot team moving a heavy object |
| Fully competitive (zero-sum) | $r^1 = -r^2$ | Chess, Go, poker duels |
| Mixed / general-sum | Different, partly aligned | Traffic, markets, negotiation, social dilemmas |
With partial observability (each agent sees only local observations) we get Dec-POMDPs.
Solution concepts#
What is "optimal" when outcomes depend on others? Game theory offers equilibria:
- A Nash equilibrium is a joint policy where no agent can improve its expected return by changing its own policy unilaterally.
- In two-player zero-sum games, Nash equilibria correspond to minimax strategies โ well defined and computable in principle.
- In general-sum games, there may be many equilibria, some much better for everyone than others (e.g. mutual cooperation vs mutual defection in the prisoner's dilemma).
Social dilemmas โ situations where individually rational behaviour leads to collectively poor outcomes (overusing a shared resource) โ are a major research theme, relevant to understanding cooperation among AI agents and humans.
Challenges#
- Non-stationarity: from each agent's perspective, the environment changes as other agents learn โ violating the stationary MDP assumption and destabilising learning.
- Credit assignment: with a shared team reward, which agent's action caused success?
- Scalability: the joint action space grows exponentially with the number of agents.
- Partial observability and communication.
- Equilibrium selection and coordination: agents must converge on compatible conventions (drive on the left or the right?).
Approaches#
Independent learners#
Each agent runs its own RL algorithm (e.g. independent Q-learning or independent PPO), treating others as part of the environment. Simple and scalable; surprisingly competitive in some benchmarks, but without convergence guarantees due to non-stationarity.
Centralised training, decentralised execution (CTDE)#
During training (often in simulation), use global information; at execution, each agent acts on local observations only.
- MADDPG (Lowe et al., 2017): each agent's critic sees all agents' observations and actions; actors use local observations.
- Value decomposition for cooperative tasks: VDN sums per-agent Q-values; QMIX combines them with a monotonic mixing network, so each agent's greedy local action is consistent with the joint greedy action.
- MAPPO: PPO with a centralised value function โ a strong, simple baseline in cooperative benchmarks.
Self-play and populations#
For competitive games, agents train against copies of themselves (AlphaZero). To avoid cycles and overfitting to one opponent, train against populations or leagues of past and diverse agents โ the approach behind DeepMind's AlphaStar (StarCraft II, reaching Grandmaster level) and OpenAI Five (Dota 2).
Game-theoretic methods#
Counterfactual regret minimisation (CFR) converges to equilibria in imperfect-information games; combined with search and learning it produced superhuman poker agents (Libratus, and Pluribus for six-player no-limit Texas hold'em).
A tiny example: independent learners in a coordination game#
import numpy as np
# Two agents choose A or B. Payoff 1 each if they match, 0 otherwise (a coordination game).
payoff = np.array([[1, 0], [0, 1]])
rng = np.random.default_rng(0)
Q1, Q2, alpha, eps = np.zeros(2), np.zeros(2), 0.1, 0.1
for t in range(2000):
a1 = rng.integers(2) if rng.random() < eps else int(np.argmax(Q1))
a2 = rng.integers(2) if rng.random() < eps else int(np.argmax(Q2))
r = payoff[a1, a2]
Q1[a1] += alpha * (r - Q1[a1]); Q2[a2] += alpha * (r - Q2[a2])
print("agent 1 prefers", "AB"[int(np.argmax(Q1))], "| agent 2 prefers", "AB"[int(np.argmax(Q2))], Q1.round(2), Q2.round(2))Run it with different seeds: the agents usually settle on a shared convention โ sometimes A, sometimes B โ illustrating equilibrium selection.
Emergent behaviour#
MARL can produce surprising emergent strategies. In OpenAI's hide-and-seek environment (Baker et al., 2020), teams discovered a sequence of strategies and counter-strategies โ building shelters, using ramps, "box surfing" โ none explicitly rewarded. Emergent communication, tool use and social conventions are active research areas.