∑ Mathematics for ML · Lecture 15 of 25

MAP Estimation, Priors and the Bayesian View of Regularisation

Adding a prior to maximum likelihood gives MAP estimation — and reveals that L2 and L1 regularisation are Gaussian and Laplace priors. We also meet conjugate priors and full Bayesian inference.

Maximum likelihood trusts the data completely. With three coin flips all heads, MLE concludes the coin can never land tails. A sensible person would say: "Probably biased towards heads, but I've seen only three flips." That common-sense caution is a prior, and combining it with the likelihood gives Maximum A Posteriori (MAP) estimation. Along the way we will discover that regularisation is secretly Bayesian.

From MLE to MAP#

Bayes' theorem for parameters:

$$ p(\boldsymbol{\theta} \mid \mathcal{D}) = \frac{p(\mathcal{D} \mid \boldsymbol{\theta})\,p(\boldsymbol{\theta})}{p(\mathcal{D})} $$

MAP chooses the mode of the posterior:

$$ \hat{\boldsymbol{\theta}}_{\text{MAP}} = \arg\max_{\boldsymbol{\theta}}\; p(\mathcal{D} \mid \boldsymbol{\theta})\,p(\boldsymbol{\theta}) = \arg\min_{\boldsymbol{\theta}}\; \Big[\underbrace{-\ln p(\mathcal{D} \mid \boldsymbol{\theta})}_{\text{data fit (NLL)}} \; \underbrace{- \ln p(\boldsymbol{\theta})}_{\text{penalty}}\Big] $$

The negative log-prior acts as a regulariser.

Regularisation is a prior#

Gaussian prior → L2 (ridge, weight decay)#

Place an independent Gaussian prior on each weight: $w_j \sim \mathcal{N}(0, \tau^2)$. Then

$$ -\ln p(\mathbf{w}) = \frac{1}{2\tau^2}\|\mathbf{w}\|_2^2 + \text{const} $$

So MAP with a Gaussian prior equals L2-regularised maximum likelihood with $\lambda = \sigma^2/\tau^2$ in linear regression. A small prior variance $\tau^2$ (strong belief that weights are near zero) means strong regularisation.

Laplace prior → L1 (lasso)#

With $w_j \sim \text{Laplace}(0, b)$: $-\ln p(\mathbf{w}) = \frac{1}{b}\|\mathbf{w}\|_1 + \text{const}$ — the lasso penalty. The Laplace density has a sharp peak at zero, encouraging many weights to be exactly zero.

Conjugate priors#

A prior is conjugate to a likelihood if the posterior is in the same family as the prior. Then Bayesian updating is simple arithmetic.

Beta–Bernoulli. Prior $\theta \sim \text{Beta}(\alpha, \beta)$. After $k$ heads and $n - k$ tails:

$$ \theta \mid \mathcal{D} \sim \text{Beta}(\alpha + k,\; \beta + n - k) $$

The prior parameters behave like pseudo-counts: $\alpha - 1$ imaginary heads and $\beta - 1$ imaginary tails. The MAP estimate is

$$ \hat\theta_{\text{MAP}} = \frac{k + \alpha - 1}{n + \alpha + \beta - 2} $$

With $\alpha = \beta = 2$ and three heads out of three, $\hat\theta_{\text{MAP}} = 4/5 = 0.8$ — confident but not certain. With $\alpha = \beta = 1$ (uniform prior) MAP equals MLE. Laplace smoothing in Naive Bayes is exactly this idea.

Other conjugate pairs: Gamma–Poisson, Dirichlet–Categorical, and Gaussian–Gaussian (for the mean with known variance).

python
import numpy as np
from scipy import stats

k, n = 3, 3                     # three heads in three flips
for a, b in [(1, 1), (2, 2), (10, 10)]:
    post = stats.beta(a + k, b + n - k)
    map_est = (k + a - 1) / (n + a + b - 2)
    lo, hi = post.ppf([0.025, 0.975])
    print(f"prior Beta({a},{b}): MAP={map_est:.2f}, posterior mean={post.mean():.2f}, "
          f"95% credible interval=({lo:.2f}, {hi:.2f})")

MAP versus full Bayesian inference#

MAP returns a single point — the posterior mode. Full Bayesian inference keeps the whole posterior and makes predictions by averaging over it:

$$ p(y^* \mid \mathbf{x}^*, \mathcal{D}) = \int p(y^* \mid \mathbf{x}^*, \boldsymbol{\theta})\,p(\boldsymbol{\theta} \mid \mathcal{D})\,d\boldsymbol{\theta} $$

This posterior predictive distribution accounts for parameter uncertainty — predictions become less confident far from the training data. The integral is usually intractable, so we approximate it using:

  • Conjugate closed forms (Bayesian linear regression, Gaussian processes);
  • MCMC sampling (Metropolis–Hastings, Hamiltonian Monte Carlo);
  • Variational inference — optimise a simple distribution to approximate the posterior (the basis of VAEs);
  • Practical deep-learning approximations — deep ensembles, MC dropout, Laplace approximations.
MLEMAPFull Bayes
Uses priorNoYesYes
OutputPoint estimatePoint estimateDistribution
Overfitting on small dataHighReducedReduced
Uncertainty estimatesNoNoYes
CostLowLowHigh

Caveats of MAP#

  • The mode can be unrepresentative of the posterior (e.g. a sharp spike with little mass).
  • MAP is not invariant to reparameterisation: the mode of $p(\theta)$ does not map to the mode of $p(g(\theta))$, because densities pick up a Jacobian factor.
  • As $n \to \infty$ the likelihood dominates the prior and MAP → MLE. Priors matter most when data is scarce — exactly when you need them.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Maximum Likelihood Estimation: How Models Learn from Data

Most loss functions in ML are negative log-likelihoods in disguise. We define MLE, derive estimators for Bernoulli and Gaussian models, and prove that MSE and cross-entropy arise from maximum likelihood.

Intermediate⏱ 5 min#038
∑ Mathematics for ML

Expectation, Variance, Covariance and Correlation

Summaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.

Beginner⏱ 5 min#036
∑ Mathematics for ML

Statistical Hypothesis Testing for ML: Is Model A Really Better?

A 1% accuracy gain may be noise. We cover confidence intervals, p-values, paired tests, McNemar's test, the bootstrap and multiple-comparison pitfalls so your experimental claims hold up.

Intermediate⏱ 6 min#044