Maximum likelihood trusts the data completely. With three coin flips all heads, MLE concludes the coin can never land tails. A sensible person would say: "Probably biased towards heads, but I've seen only three flips." That common-sense caution is a prior, and combining it with the likelihood gives Maximum A Posteriori (MAP) estimation. Along the way we will discover that regularisation is secretly Bayesian.
From MLE to MAP#
Bayes' theorem for parameters:
MAP chooses the mode of the posterior:
The negative log-prior acts as a regulariser.
Regularisation is a prior#
Gaussian prior → L2 (ridge, weight decay)#
Place an independent Gaussian prior on each weight: $w_j \sim \mathcal{N}(0, \tau^2)$. Then
So MAP with a Gaussian prior equals L2-regularised maximum likelihood with $\lambda = \sigma^2/\tau^2$ in linear regression. A small prior variance $\tau^2$ (strong belief that weights are near zero) means strong regularisation.
Laplace prior → L1 (lasso)#
With $w_j \sim \text{Laplace}(0, b)$: $-\ln p(\mathbf{w}) = \frac{1}{b}\|\mathbf{w}\|_1 + \text{const}$ — the lasso penalty. The Laplace density has a sharp peak at zero, encouraging many weights to be exactly zero.
Conjugate priors#
A prior is conjugate to a likelihood if the posterior is in the same family as the prior. Then Bayesian updating is simple arithmetic.
Beta–Bernoulli. Prior $\theta \sim \text{Beta}(\alpha, \beta)$. After $k$ heads and $n - k$ tails:
The prior parameters behave like pseudo-counts: $\alpha - 1$ imaginary heads and $\beta - 1$ imaginary tails. The MAP estimate is
With $\alpha = \beta = 2$ and three heads out of three, $\hat\theta_{\text{MAP}} = 4/5 = 0.8$ — confident but not certain. With $\alpha = \beta = 1$ (uniform prior) MAP equals MLE. Laplace smoothing in Naive Bayes is exactly this idea.
Other conjugate pairs: Gamma–Poisson, Dirichlet–Categorical, and Gaussian–Gaussian (for the mean with known variance).
import numpy as np
from scipy import stats
k, n = 3, 3 # three heads in three flips
for a, b in [(1, 1), (2, 2), (10, 10)]:
post = stats.beta(a + k, b + n - k)
map_est = (k + a - 1) / (n + a + b - 2)
lo, hi = post.ppf([0.025, 0.975])
print(f"prior Beta({a},{b}): MAP={map_est:.2f}, posterior mean={post.mean():.2f}, "
f"95% credible interval=({lo:.2f}, {hi:.2f})")MAP versus full Bayesian inference#
MAP returns a single point — the posterior mode. Full Bayesian inference keeps the whole posterior and makes predictions by averaging over it:
This posterior predictive distribution accounts for parameter uncertainty — predictions become less confident far from the training data. The integral is usually intractable, so we approximate it using:
- Conjugate closed forms (Bayesian linear regression, Gaussian processes);
- MCMC sampling (Metropolis–Hastings, Hamiltonian Monte Carlo);
- Variational inference — optimise a simple distribution to approximate the posterior (the basis of VAEs);
- Practical deep-learning approximations — deep ensembles, MC dropout, Laplace approximations.
| MLE | MAP | Full Bayes | |
|---|---|---|---|
| Uses prior | No | Yes | Yes |
| Output | Point estimate | Point estimate | Distribution |
| Overfitting on small data | High | Reduced | Reduced |
| Uncertainty estimates | No | No | Yes |
| Cost | Low | Low | High |
Caveats of MAP#
- The mode can be unrepresentative of the posterior (e.g. a sharp spike with little mass).
- MAP is not invariant to reparameterisation: the mode of $p(\theta)$ does not map to the mode of $p(g(\theta))$, because densities pick up a Jacobian factor.
- As $n \to \infty$ the likelihood dominates the prior and MAP → MLE. Priors matter most when data is scarce — exactly when you need them.