If a test for a rare disease is 99% accurate and you test positive, how worried should you be? Most people — including, in published studies, many doctors — answer "99%". The correct answer is often much lower. The tool that gets this right is Bayes' theorem, and it is the mathematical formalisation of learning from evidence.
Derivation#
From the product rule, $P(A \cap B) = P(A \mid B)P(B) = P(B \mid A)P(A)$. Divide by $P(B)$:
In the language of learning, with hypothesis $H$ and evidence (data) $D$:
- Prior $P(H)$: belief before seeing data.
- Likelihood $P(D \mid H)$: how probable the data is if $H$ were true.
- Evidence $P(D) = \sum_{H'} P(D \mid H')P(H')$: a normalising constant.
- Posterior $P(H \mid D)$: updated belief.
In words: posterior ∝ likelihood × prior.
The medical test example#
A disease affects 1 in 1,000 people. A test has sensitivity 99% ($P(+ \mid \text{disease}) = 0.99$) and specificity 95% ($P(- \mid \text{healthy}) = 0.95$, so a 5% false-positive rate). You test positive. What is $P(\text{disease} \mid +)$?
Less than 2%! Among 1,000 people, about 1 is sick and tests positive, but about 50 healthy people also test positive. A positive result raises the probability from 0.1% to 1.9% — a nineteen-fold increase — but the disease remains unlikely.
A natural-frequencies table makes it vivid (per 100,000 people):
| Test + | Test − | Total | |
|---|---|---|---|
| Disease | 99 | 1 | 100 |
| Healthy | 4,995 | 94,905 | 99,900 |
| Total | 5,094 | 94,906 | 100,000 |
Sequential updating#
Today's posterior becomes tomorrow's prior. If the patient takes a second, independent test and is again positive:
Evidence accumulates. This is how Bayesian agents learn continuously — the same logic as the filtering algorithms in HMMs.
Odds form#
Bayes' theorem is often cleaner in odds:
For our test, the likelihood ratio of a positive result is $0.99/0.05 = 19.8$. Prior odds of 1:999 become posterior odds of about 19.8:999 ≈ 1:50. Taking logarithms turns multiplication into addition: each independent piece of evidence adds its log-likelihood ratio. Logistic regression's scores are exactly such log-odds.
Naive Bayes: Bayes' theorem as a classifier#
To classify an email with words $w_1, \dots, w_n$ as spam or ham, apply Bayes' theorem with the naive assumption that words are conditionally independent given the class:
import math
from collections import Counter
train = [("win money now", "spam"), ("cheap money offer", "spam"),
("meeting at noon", "ham"), ("lecture notes attached", "ham"),
("win a free lecture", "spam"), ("noon meeting moved", "ham")]
counts = {"spam": Counter(), "ham": Counter()}
docs = Counter()
for text, label in train:
docs[label] += 1
counts[label].update(text.split())
vocab = set(w for c in counts.values() for w in c)
def log_posterior(text, label, alpha=1.0): # Laplace smoothing
total = sum(counts[label].values())
lp = math.log(docs[label] / sum(docs.values()))
for w in text.split():
lp += math.log((counts[label][w] + alpha) / (total + alpha * len(vocab)))
return lp
msg = "free money meeting"
scores = {c: log_posterior(msg, c) for c in counts}
print(max(scores, key=scores.get), scores)We work with log probabilities to avoid underflow and use Laplace smoothing so that an unseen word does not force a probability of zero.
Bayesian inference beyond classification#
Bayes' theorem applies to model parameters too:
The posterior over parameters expresses what we know and how uncertain we remain. Predictions average over it. This is the foundation of Bayesian linear regression, Gaussian processes, Bayesian optimisation for hyperparameter tuning and Thompson sampling for bandits.