Probability tells an agent what it should believe. Decision theory tells it what it should do. Combining the two gives the principle at the heart of rational agency: choose the action with maximum expected utility (MEU). Today we build this theory carefully, because it underlies reinforcement learning, Bayesian optimisation and responsible deployment of any ML system that makes decisions.
Maximum expected utility#
Suppose action $a$ leads to outcome $s'$ with probability $P(s' \mid a, e)$ given evidence $e$, and $U(s')$ measures how desirable $s'$ is. The expected utility is
and a rational agent chooses
Why utilities? The von Neumann–Morgenstern axioms#
Why should preferences be representable by a number, and why should we maximise its expectation? Von Neumann and Morgenstern (1944) showed that if an agent's preferences over lotteries (probability distributions over outcomes) satisfy a few reasonable axioms, then a utility function exists and the agent behaves as a maximiser of expected utility. The axioms are:
- Orderability — any two lotteries can be compared.
- Transitivity — if $A \succ B$ and $B \succ C$ then $A \succ C$.
- Continuity — if $A \succ B \succ C$, some mixture of $A$ and $C$ is equivalent to $B$.
- Substitutability — indifferent lotteries can be swapped inside more complex lotteries.
- Monotonicity — a higher chance of a preferred outcome is preferred.
- Decomposability — compound lotteries reduce to simple ones.
Utility of money and risk attitudes#
Utility is not the same as money. Most people prefer a guaranteed 1 million to a 50% chance of 3 million, even though the expected money of the gamble is higher. This is captured by a concave utility function, such as $U(x) = \log x$, which models risk aversion. The difference between the expected monetary value of a lottery and its certainty equivalent is the risk premium — the basis of the insurance industry.
- Concave $U$: risk-averse.
- Linear $U$: risk-neutral.
- Convex $U$: risk-seeking.
Human irrationality#
Kahneman and Tversky's experiments showed humans systematically deviate from expected-utility theory: we weigh losses more than gains (loss aversion), are influenced by how options are framed, and overweight small probabilities. Their prospect theory describes these behaviours. AI systems that interact with humans must account for them, and designers must beware of building systems that exploit them.
Decision networks#
A decision network (influence diagram) extends a Bayesian network with:
- chance nodes (ovals) — random variables;
- decision nodes (rectangles) — choices;
- utility nodes (diamonds) — the utility as a function of parents.
To evaluate: for each value of the decision node, set it, compute posterior probabilities of the utility node's parents, compute expected utility, and pick the best.
The value of information#
Should a doctor order an expensive test before treating? Should an oil company buy a seismic survey? Decision theory answers precisely. The value of perfect information (VPI) about variable $E_j$ is
— the expected utility of deciding after learning $E_j$, minus the expected utility of deciding now. Key properties:
- VPI is never negative (in expectation, information cannot hurt a rational agent).
- VPI is zero if the information would not change the optimal decision.
- VPI is not additive — two tests may be redundant.
A myopic information-gathering agent requests the observation with the highest VPI minus cost, as long as it is positive.
def expected_utility(action, p_state, utility):
return sum(p * utility[action][s] for s, p in p_state.items())
# Should we deploy a new model? States: model is good / bad.
p = {"good": 0.6, "bad": 0.4}
U = {"deploy": {"good": 100, "bad": -150}, "keep_old": {"good": 0, "bad": 0}}
eu_now = max(expected_utility(a, p, U) for a in U)
# Perfect information: we would learn the true state first, then choose the best action.
eu_informed = sum(p[s] * max(U[a][s] for a in U) for s in p)
print("EU now:", eu_now, " EU with info:", eu_informed, " VPI:", eu_informed - eu_now)Here deploying now has expected utility $0.6 \times 100 - 0.4 \times 150 = 0$, the same as keeping the old model. With perfect information we deploy only when good: $0.6 \times 100 = 60$. So a perfect evaluation is worth up to 60 utility units — a quantitative argument for investing in a thorough offline and A/B evaluation before deployment.