∑ Mathematics for ML · Lecture 14 of 25

Maximum Likelihood Estimation: How Models Learn from Data

Most loss functions in ML are negative log-likelihoods in disguise. We define MLE, derive estimators for Bernoulli and Gaussian models, and prove that MSE and cross-entropy arise from maximum likelihood.

Why do we train regression models with mean squared error and classifiers with cross-entropy? Are these arbitrary choices? No — both arise from a single principle: maximum likelihood estimation (MLE). Once you understand MLE, loss functions stop being a list to memorise and become consequences of your assumptions about the data.

The likelihood function#

Suppose data $\mathcal{D} = \{x_1, \dots, x_n\}$ are drawn independently from a model $p(x \mid \boldsymbol{\theta})$ with unknown parameters $\boldsymbol{\theta}$. The likelihood is the probability of the observed data viewed as a function of the parameters:

$$ \mathcal{L}(\boldsymbol{\theta}) = p(\mathcal{D} \mid \boldsymbol{\theta}) = \prod_{i=1}^{n} p(x_i \mid \boldsymbol{\theta}) $$

The maximum likelihood estimate is

$$ \hat{\boldsymbol{\theta}}_{\text{MLE}} = \arg\max_{\boldsymbol{\theta}} \prod_{i=1}^{n} p(x_i \mid \boldsymbol{\theta}) = \arg\min_{\boldsymbol{\theta}} \left[ -\sum_{i=1}^{n} \ln p(x_i \mid \boldsymbol{\theta}) \right] $$

We take logarithms because (1) products of many small numbers underflow, (2) sums are easier to differentiate, and (3) the log is monotonic, so the maximiser is unchanged. The quantity in brackets is the negative log-likelihood (NLL) — the loss function.

Example 1: coin flips (Bernoulli)#

Observing $k$ heads in $n$ flips, with $p(x \mid \theta) = \theta^x(1-\theta)^{1-x}$:

$$ \ln\mathcal{L}(\theta) = k\ln\theta + (n - k)\ln(1 - \theta) $$

Setting the derivative to zero: $\frac{k}{\theta} - \frac{n - k}{1 - \theta} = 0 \Rightarrow \hat{\theta} = \frac{k}{n}$. The MLE is the observed frequency — reassuringly intuitive.

Example 2: Gaussian#

For $x_i \sim \mathcal{N}(\mu, \sigma^2)$:

$$ \ln\mathcal{L} = -\frac{n}{2}\ln(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_i(x_i - \mu)^2 $$

giving $\hat\mu = \frac{1}{n}\sum_i x_i$ and $\hat\sigma^2 = \frac{1}{n}\sum_i(x_i - \hat\mu)^2$. Note the MLE variance divides by $n$, not $n - 1$: it is slightly biased, underestimating variance for small samples.

From MLE to loss functions#

Regression → Mean Squared Error#

Model the target as $y = f_{\boldsymbol{\theta}}(\mathbf{x}) + \epsilon$ with $\epsilon \sim \mathcal{N}(0, \sigma^2)$. Then

$$ -\ln p(y \mid \mathbf{x}, \boldsymbol{\theta}) = \frac{1}{2\sigma^2}\big(y - f_{\boldsymbol{\theta}}(\mathbf{x})\big)^2 + \text{const} $$

Summing over data, minimising NLL = minimising squared error. MSE is the loss that corresponds to assuming Gaussian noise. Assume Laplace noise instead and you get mean absolute error.

Classification → Cross-Entropy#

Model $y \in \{0, 1\}$ as Bernoulli with probability $\hat{p} = \sigma(f_{\boldsymbol{\theta}}(\mathbf{x}))$:

$$ -\ln p(y \mid \mathbf{x}, \boldsymbol{\theta}) = -\big[y\ln\hat{p} + (1 - y)\ln(1 - \hat{p})\big] $$

That is binary cross-entropy. For $K$ classes with a categorical model and softmax outputs, the NLL is categorical cross-entropy $-\ln\hat{p}_{y}$. Language models are trained by exactly this loss on next-token prediction.

Assumed output distributionNegative log-likelihood = loss
Gaussian (fixed variance)Mean squared error
LaplaceMean absolute error
BernoulliBinary cross-entropy
CategoricalSoftmax cross-entropy
PoissonPoisson deviance

MLE with gradient descent#

When no closed form exists (logistic regression, neural networks), we minimise the NLL numerically:

python
import numpy as np

rng = np.random.default_rng(0)
n = 2000
X = np.column_stack([np.ones(n), rng.normal(size=(n, 2))])
true_w = np.array([-0.5, 2.0, -1.0])
y = rng.random(n) < 1 / (1 + np.exp(-X @ true_w))        # Bernoulli labels

w = np.zeros(3)
for step in range(2000):
    p = 1 / (1 + np.exp(-X @ w))
    nll = -np.mean(y * np.log(p + 1e-12) + (1 - y) * np.log(1 - p + 1e-12))
    grad = X.T @ (p - y) / n
    w -= 0.5 * grad
print("estimated:", w.round(2), " true:", true_w, " final NLL:", round(nll, 4))

The estimate recovers the true parameters closely — MLE is consistent.

Properties of MLE#

Under regularity conditions, as $n \to \infty$ the MLE is:

  • Consistent — converges to the true parameter.
  • Asymptotically normal — its sampling distribution approaches a Gaussian with covariance given by the inverse Fisher information.
  • Asymptotically efficient — achieves the lowest possible variance (the Cramér–Rao bound).
  • Invariant — the MLE of $g(\theta)$ is $g(\hat\theta)$.

The weakness of MLE: overfitting#

With little data, MLE trusts the data completely. Flip a coin three times, get three heads, and MLE says $\hat\theta = 1$: tails are impossible. With a flexible model, MLE fits noise. The remedies are regularisation and, equivalently, priors — which leads to MAP estimation in the next lecture.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

MAP Estimation, Priors and the Bayesian View of Regularisation

Adding a prior to maximum likelihood gives MAP estimation — and reveals that L2 and L1 regularisation are Gaussian and Laplace priors. We also meet conjugate priors and full Bayesian inference.

Intermediate⏱ 5 min#039
∑ Mathematics for ML

Expectation, Variance, Covariance and Correlation

Summaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.

Beginner⏱ 5 min#036
∑ Mathematics for ML

Statistical Hypothesis Testing for ML: Is Model A Really Better?

A 1% accuracy gain may be noise. We cover confidence intervals, p-values, paired tests, McNemar's test, the bootstrap and multiple-comparison pitfalls so your experimental claims hold up.

Intermediate⏱ 6 min#044