๐ŸŽฎ Reinforcement Learning ยท Lecture 18 of 21

Multi-Agent Reinforcement Learning: Cooperation, Competition and Equilibria

When many learning agents share an environment, each faces a moving target. We introduce Markov games, Nash equilibria, independent learners, centralised training with decentralised execution, self-play and emergent behaviour.

Traffic intersections with many autonomous vehicles, teams of warehouse robots, trading agents in markets, players in a strategy game, LLM agents negotiating on behalf of users โ€” many real problems involve multiple agents acting simultaneously. Multi-agent reinforcement learning (MARL) studies how agents learn when their outcomes depend on each other's behaviour. It brings together RL and game theory, and introduces challenges absent in single-agent RL.

Markov games#

A Markov game (stochastic game) generalises the MDP to $N$ agents:

  • a shared state $s$;
  • each agent $i$ chooses an action $a^i$; the joint action $\mathbf{a} = (a^1, \dots, a^N)$ determines the transition $P(s' \mid s, \mathbf{a})$;
  • each agent receives its own reward $r^i(s, \mathbf{a})$.

Settings:

SettingRewardsExample
Fully cooperativeShared rewardRobot team moving a heavy object
Fully competitive (zero-sum)$r^1 = -r^2$Chess, Go, poker duels
Mixed / general-sumDifferent, partly alignedTraffic, markets, negotiation, social dilemmas

With partial observability (each agent sees only local observations) we get Dec-POMDPs.

Solution concepts#

What is "optimal" when outcomes depend on others? Game theory offers equilibria:

  • A Nash equilibrium is a joint policy where no agent can improve its expected return by changing its own policy unilaterally.
  • In two-player zero-sum games, Nash equilibria correspond to minimax strategies โ€” well defined and computable in principle.
  • In general-sum games, there may be many equilibria, some much better for everyone than others (e.g. mutual cooperation vs mutual defection in the prisoner's dilemma).

Social dilemmas โ€” situations where individually rational behaviour leads to collectively poor outcomes (overusing a shared resource) โ€” are a major research theme, relevant to understanding cooperation among AI agents and humans.

Challenges#

  1. Non-stationarity: from each agent's perspective, the environment changes as other agents learn โ€” violating the stationary MDP assumption and destabilising learning.
  2. Credit assignment: with a shared team reward, which agent's action caused success?
  3. Scalability: the joint action space grows exponentially with the number of agents.
  4. Partial observability and communication.
  5. Equilibrium selection and coordination: agents must converge on compatible conventions (drive on the left or the right?).

Approaches#

Independent learners#

Each agent runs its own RL algorithm (e.g. independent Q-learning or independent PPO), treating others as part of the environment. Simple and scalable; surprisingly competitive in some benchmarks, but without convergence guarantees due to non-stationarity.

Centralised training, decentralised execution (CTDE)#

During training (often in simulation), use global information; at execution, each agent acts on local observations only.

  • MADDPG (Lowe et al., 2017): each agent's critic sees all agents' observations and actions; actors use local observations.
  • Value decomposition for cooperative tasks: VDN sums per-agent Q-values; QMIX combines them with a monotonic mixing network, so each agent's greedy local action is consistent with the joint greedy action.
  • MAPPO: PPO with a centralised value function โ€” a strong, simple baseline in cooperative benchmarks.

Self-play and populations#

For competitive games, agents train against copies of themselves (AlphaZero). To avoid cycles and overfitting to one opponent, train against populations or leagues of past and diverse agents โ€” the approach behind DeepMind's AlphaStar (StarCraft II, reaching Grandmaster level) and OpenAI Five (Dota 2).

Game-theoretic methods#

Counterfactual regret minimisation (CFR) converges to equilibria in imperfect-information games; combined with search and learning it produced superhuman poker agents (Libratus, and Pluribus for six-player no-limit Texas hold'em).

A tiny example: independent learners in a coordination game#

python
import numpy as np

# Two agents choose A or B. Payoff 1 each if they match, 0 otherwise (a coordination game).
payoff = np.array([[1, 0], [0, 1]])
rng = np.random.default_rng(0)
Q1, Q2, alpha, eps = np.zeros(2), np.zeros(2), 0.1, 0.1

for t in range(2000):
    a1 = rng.integers(2) if rng.random() < eps else int(np.argmax(Q1))
    a2 = rng.integers(2) if rng.random() < eps else int(np.argmax(Q2))
    r = payoff[a1, a2]
    Q1[a1] += alpha * (r - Q1[a1]); Q2[a2] += alpha * (r - Q2[a2])
print("agent 1 prefers", "AB"[int(np.argmax(Q1))], "| agent 2 prefers", "AB"[int(np.argmax(Q2))], Q1.round(2), Q2.round(2))

Run it with different seeds: the agents usually settle on a shared convention โ€” sometimes A, sometimes B โ€” illustrating equilibrium selection.

Emergent behaviour#

MARL can produce surprising emergent strategies. In OpenAI's hide-and-seek environment (Baker et al., 2020), teams discovered a sequence of strategies and counter-strategies โ€” building shelters, using ramps, "box surfing" โ€” none explicitly rewarded. Emergent communication, tool use and social conventions are active research areas.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐ŸŽฎ Reinforcement Learning

AlphaGo, AlphaZero and MuZero: Search Meets Deep Learning

DeepMind's Go programs combined deep neural networks with Monte Carlo tree search and self-play. We trace AlphaGo's supervised and RL training, AlphaZero's tabula-rasa self-play, MuZero's learned model, and the lessons for AI.

Advancedโฑ 6 min#237
๐ŸŽฎ Reinforcement Learning

Imitation Learning and Inverse Reinforcement Learning

When rewards are hard to specify but demonstrations are available, agents can learn by imitation. We cover behavioural cloning and its compounding errors, DAgger, inverse RL, maximum-entropy IRL and adversarial imitation (GAIL).

Advancedโฑ 5 min#239
๐ŸŽฎ Reinforcement Learning

Model-Based Reinforcement Learning: Learning and Planning with World Models

Model-based agents learn a model of the environment and use it to plan or generate imagined experience. We cover Dyna, model-predictive control, model errors and ensembles, MBPO, and latent world models like Dreamer and MuZero.

Advancedโฑ 5 min#236