∑ Mathematics for ML · Lecture 9 of 25

Probability Fundamentals: Sample Spaces, Axioms and Conditional Probability

Machine learning is reasoning under uncertainty. We build probability from Kolmogorov's axioms, then master conditional probability, the product and sum rules, and independence.

A classifier that says "this email is spam" is less useful than one that says "this email is spam with probability 0.97". Probability is the language that lets models express — and reason about — uncertainty. Today we lay its foundations with care, because sloppy probabilistic reasoning is one of the most common sources of error in data science.

Sample spaces and events#

  • The sample space $\Omega$ is the set of all possible outcomes of an experiment. Rolling a die: $\Omega = \{1, 2, 3, 4, 5, 6\}$.
  • An event $A \subseteq \Omega$ is a set of outcomes. "Even number": $A = \{2, 4, 6\}$.
  • Events combine with set operations: $A \cup B$ (A or B), $A \cap B$ (A and B), $A^c$ (not A).

Kolmogorov's axioms#

A probability measure $P$ assigns numbers to events such that:

  1. $P(A) \ge 0$ for every event $A$;
  2. $P(\Omega) = 1$;
  3. For disjoint events $A_1, A_2, \dots$: $P(\bigcup_i A_i) = \sum_i P(A_i)$.

Everything else follows. For example:

  • $P(A^c) = 1 - P(A)$;
  • $P(\emptyset) = 0$;
  • Inclusion–exclusion: $P(A \cup B) = P(A) + P(B) - P(A \cap B)$;
  • If $A \subseteq B$ then $P(A) \le P(B)$.

Conditional probability#

The probability of $A$ given that $B$ occurred is

$$ P(A \mid B) = \frac{P(A \cap B)}{P(B)}, \qquad P(B) > 0 $$

Conditioning restricts the sample space to $B$ and renormalises. Rearranging gives the product rule:

$$ P(A \cap B) = P(A \mid B)\,P(B) = P(B \mid A)\,P(A) $$

and extending it to many events gives the chain rule of probability:

$$ P(A_1, A_2, \dots, A_n) = P(A_1)\,P(A_2 \mid A_1)\,P(A_3 \mid A_1, A_2) \cdots P(A_n \mid A_{1:n-1}) $$

The law of total probability#

If $B_1, \dots, B_k$ partition $\Omega$ (disjoint and covering everything):

$$ P(A) = \sum_{i=1}^{k} P(A \mid B_i)\,P(B_i) $$

This is the sum rule or marginalisation. It lets us compute a probability by considering all the ways it can happen.

Independence#

Events $A$ and $B$ are independent if $P(A \cap B) = P(A)P(B)$, equivalently $P(A \mid B) = P(A)$ — knowing $B$ tells you nothing about $A$.

Conditional independence: $A \perp B \mid C$ if $P(A, B \mid C) = P(A \mid C)\,P(B \mid C)$. Conditional independence is the assumption that makes Naive Bayes, HMMs and Bayesian networks tractable.

Simulation: the best way to build intuition#

python
import numpy as np

rng = np.random.default_rng(42)
N = 1_000_000

# The birthday problem: probability that at least two of 23 people share a birthday
bdays = rng.integers(0, 365, size=(100_000, 23))
shared = np.array([len(set(row)) < 23 for row in bdays])
print("birthday (simulated):", shared.mean())           # ~0.507

# Conditional probability by simulation: two dice, P(sum = 8 | first die is even)
d1, d2 = rng.integers(1, 7, N), rng.integers(1, 7, N)
cond = d1 % 2 == 0
print("P(sum=8 | d1 even):", ((d1 + d2 == 8) & cond).sum() / cond.sum())   # 3/18 = 0.1667

The birthday result surprises almost everyone: with only 23 people, a shared birthday is more likely than not. Our intuitions about probability are unreliable — which is precisely why we formalise and simulate.

Probability in ML, in one paragraph#

A supervised model estimates a conditional distribution $p(y \mid \mathbf{x})$. Training maximises the probability the model assigns to the observed data (maximum likelihood). Generative models estimate $p(\mathbf{x})$ itself. Evaluation estimates probabilities of errors from finite samples. Uncertainty estimates tell us when to trust predictions. Every one of these rests on today's axioms.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Random Variables and Probability Distributions

A random variable turns outcomes into numbers. We study discrete and continuous distributions — Bernoulli, categorical, binomial, Poisson, uniform, exponential, Beta — and when ML uses each.

Beginner⏱ 5 min#034
∑ Mathematics for ML

Bayes' Theorem: Updating Beliefs with Evidence

Bayes' theorem is the mathematical rule for learning from evidence. We derive it, work through the famous medical-test example, and see how it underlies Naive Bayes, Bayesian inference and spam filters.

Beginner⏱ 5 min#035
∑ Mathematics for ML

Expectation, Variance, Covariance and Correlation

Summaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.

Beginner⏱ 5 min#036