∑ Mathematics for ML · Lecture 11 of 25

Bayes' Theorem: Updating Beliefs with Evidence

Bayes' theorem is the mathematical rule for learning from evidence. We derive it, work through the famous medical-test example, and see how it underlies Naive Bayes, Bayesian inference and spam filters.

If a test for a rare disease is 99% accurate and you test positive, how worried should you be? Most people — including, in published studies, many doctors — answer "99%". The correct answer is often much lower. The tool that gets this right is Bayes' theorem, and it is the mathematical formalisation of learning from evidence.

Derivation#

From the product rule, $P(A \cap B) = P(A \mid B)P(B) = P(B \mid A)P(A)$. Divide by $P(B)$:

$$ P(A \mid B) = \frac{P(B \mid A)\,P(A)}{P(B)} $$

In the language of learning, with hypothesis $H$ and evidence (data) $D$:

$$ \underbrace{P(H \mid D)}_{\text{posterior}} = \frac{\overbrace{P(D \mid H)}^{\text{likelihood}}\;\overbrace{P(H)}^{\text{prior}}}{\underbrace{P(D)}_{\text{evidence}}} $$
  • Prior $P(H)$: belief before seeing data.
  • Likelihood $P(D \mid H)$: how probable the data is if $H$ were true.
  • Evidence $P(D) = \sum_{H'} P(D \mid H')P(H')$: a normalising constant.
  • Posterior $P(H \mid D)$: updated belief.

In words: posterior ∝ likelihood × prior.

The medical test example#

A disease affects 1 in 1,000 people. A test has sensitivity 99% ($P(+ \mid \text{disease}) = 0.99$) and specificity 95% ($P(- \mid \text{healthy}) = 0.95$, so a 5% false-positive rate). You test positive. What is $P(\text{disease} \mid +)$?

$$ P(\text{disease} \mid +) = \frac{0.99 \times 0.001}{0.99 \times 0.001 + 0.05 \times 0.999} = \frac{0.00099}{0.05094} \approx 0.019 $$

Less than 2%! Among 1,000 people, about 1 is sick and tests positive, but about 50 healthy people also test positive. A positive result raises the probability from 0.1% to 1.9% — a nineteen-fold increase — but the disease remains unlikely.

A natural-frequencies table makes it vivid (per 100,000 people):

Test +Test −Total
Disease991100
Healthy4,99594,90599,900
Total5,09494,906100,000

Sequential updating#

Today's posterior becomes tomorrow's prior. If the patient takes a second, independent test and is again positive:

$$ P(\text{disease} \mid +, +) = \frac{0.99 \times 0.019}{0.99 \times 0.019 + 0.05 \times 0.981} \approx 0.28 $$

Evidence accumulates. This is how Bayesian agents learn continuously — the same logic as the filtering algorithms in HMMs.

Odds form#

Bayes' theorem is often cleaner in odds:

$$ \underbrace{\frac{P(H \mid D)}{P(\neg H \mid D)}}_{\text{posterior odds}} = \underbrace{\frac{P(D \mid H)}{P(D \mid \neg H)}}_{\text{likelihood ratio}} \times \underbrace{\frac{P(H)}{P(\neg H)}}_{\text{prior odds}} $$

For our test, the likelihood ratio of a positive result is $0.99/0.05 = 19.8$. Prior odds of 1:999 become posterior odds of about 19.8:999 ≈ 1:50. Taking logarithms turns multiplication into addition: each independent piece of evidence adds its log-likelihood ratio. Logistic regression's scores are exactly such log-odds.

Naive Bayes: Bayes' theorem as a classifier#

To classify an email with words $w_1, \dots, w_n$ as spam or ham, apply Bayes' theorem with the naive assumption that words are conditionally independent given the class:

$$ P(\text{spam} \mid w_{1:n}) \propto P(\text{spam})\prod_{i=1}^{n} P(w_i \mid \text{spam}) $$
python
import math
from collections import Counter

train = [("win money now", "spam"), ("cheap money offer", "spam"),
         ("meeting at noon", "ham"), ("lecture notes attached", "ham"),
         ("win a free lecture", "spam"), ("noon meeting moved", "ham")]

counts = {"spam": Counter(), "ham": Counter()}
docs = Counter()
for text, label in train:
    docs[label] += 1
    counts[label].update(text.split())
vocab = set(w for c in counts.values() for w in c)

def log_posterior(text, label, alpha=1.0):               # Laplace smoothing
    total = sum(counts[label].values())
    lp = math.log(docs[label] / sum(docs.values()))
    for w in text.split():
        lp += math.log((counts[label][w] + alpha) / (total + alpha * len(vocab)))
    return lp

msg = "free money meeting"
scores = {c: log_posterior(msg, c) for c in counts}
print(max(scores, key=scores.get), scores)

We work with log probabilities to avoid underflow and use Laplace smoothing so that an unseen word does not force a probability of zero.

Bayesian inference beyond classification#

Bayes' theorem applies to model parameters too:

$$ p(\boldsymbol{\theta} \mid \mathcal{D}) \propto p(\mathcal{D} \mid \boldsymbol{\theta})\,p(\boldsymbol{\theta}) $$

The posterior over parameters expresses what we know and how uncertain we remain. Predictions average over it. This is the foundation of Bayesian linear regression, Gaussian processes, Bayesian optimisation for hyperparameter tuning and Thompson sampling for bandits.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Random Variables and Probability Distributions

A random variable turns outcomes into numbers. We study discrete and continuous distributions — Bernoulli, categorical, binomial, Poisson, uniform, exponential, Beta — and when ML uses each.

Beginner⏱ 5 min#034
∑ Mathematics for ML

Expectation, Variance, Covariance and Correlation

Summaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.

Beginner⏱ 5 min#036
∑ Mathematics for ML

Probability Fundamentals: Sample Spaces, Axioms and Conditional Probability

Machine learning is reasoning under uncertainty. We build probability from Kolmogorov's axioms, then master conditional probability, the product and sum rules, and independence.

Beginner⏱ 5 min#033