A classifier that says "this email is spam" is less useful than one that says "this email is spam with probability 0.97". Probability is the language that lets models express — and reason about — uncertainty. Today we lay its foundations with care, because sloppy probabilistic reasoning is one of the most common sources of error in data science.
Sample spaces and events#
- The sample space $\Omega$ is the set of all possible outcomes of an experiment. Rolling a die: $\Omega = \{1, 2, 3, 4, 5, 6\}$.
- An event $A \subseteq \Omega$ is a set of outcomes. "Even number": $A = \{2, 4, 6\}$.
- Events combine with set operations: $A \cup B$ (A or B), $A \cap B$ (A and B), $A^c$ (not A).
Kolmogorov's axioms#
A probability measure $P$ assigns numbers to events such that:
- $P(A) \ge 0$ for every event $A$;
- $P(\Omega) = 1$;
- For disjoint events $A_1, A_2, \dots$: $P(\bigcup_i A_i) = \sum_i P(A_i)$.
Everything else follows. For example:
- $P(A^c) = 1 - P(A)$;
- $P(\emptyset) = 0$;
- Inclusion–exclusion: $P(A \cup B) = P(A) + P(B) - P(A \cap B)$;
- If $A \subseteq B$ then $P(A) \le P(B)$.
Conditional probability#
The probability of $A$ given that $B$ occurred is
Conditioning restricts the sample space to $B$ and renormalises. Rearranging gives the product rule:
and extending it to many events gives the chain rule of probability:
The law of total probability#
If $B_1, \dots, B_k$ partition $\Omega$ (disjoint and covering everything):
This is the sum rule or marginalisation. It lets us compute a probability by considering all the ways it can happen.
Independence#
Events $A$ and $B$ are independent if $P(A \cap B) = P(A)P(B)$, equivalently $P(A \mid B) = P(A)$ — knowing $B$ tells you nothing about $A$.
Conditional independence: $A \perp B \mid C$ if $P(A, B \mid C) = P(A \mid C)\,P(B \mid C)$. Conditional independence is the assumption that makes Naive Bayes, HMMs and Bayesian networks tractable.
Simulation: the best way to build intuition#
import numpy as np
rng = np.random.default_rng(42)
N = 1_000_000
# The birthday problem: probability that at least two of 23 people share a birthday
bdays = rng.integers(0, 365, size=(100_000, 23))
shared = np.array([len(set(row)) < 23 for row in bdays])
print("birthday (simulated):", shared.mean()) # ~0.507
# Conditional probability by simulation: two dice, P(sum = 8 | first die is even)
d1, d2 = rng.integers(1, 7, N), rng.integers(1, 7, N)
cond = d1 % 2 == 0
print("P(sum=8 | d1 even):", ((d1 + d2 == 8) & cond).sum() / cond.sum()) # 3/18 = 0.1667The birthday result surprises almost everyone: with only 23 people, a shared birthday is more likely than not. Our intuitions about probability are unreliable — which is precisely why we formalise and simulate.
Probability in ML, in one paragraph#
A supervised model estimates a conditional distribution $p(y \mid \mathbf{x})$. Training maximises the probability the model assigns to the observed data (maximum likelihood). Generative models estimate $p(\mathbf{x})$ itself. Evaluation estimates probabilities of errors from finite samples. Uncertainty estimates tell us when to trust predictions. Every one of these rests on today's axioms.