Why do we train regression models with mean squared error and classifiers with cross-entropy? Are these arbitrary choices? No — both arise from a single principle: maximum likelihood estimation (MLE). Once you understand MLE, loss functions stop being a list to memorise and become consequences of your assumptions about the data.
The likelihood function#
Suppose data $\mathcal{D} = \{x_1, \dots, x_n\}$ are drawn independently from a model $p(x \mid \boldsymbol{\theta})$ with unknown parameters $\boldsymbol{\theta}$. The likelihood is the probability of the observed data viewed as a function of the parameters:
The maximum likelihood estimate is
We take logarithms because (1) products of many small numbers underflow, (2) sums are easier to differentiate, and (3) the log is monotonic, so the maximiser is unchanged. The quantity in brackets is the negative log-likelihood (NLL) — the loss function.
Example 1: coin flips (Bernoulli)#
Observing $k$ heads in $n$ flips, with $p(x \mid \theta) = \theta^x(1-\theta)^{1-x}$:
Setting the derivative to zero: $\frac{k}{\theta} - \frac{n - k}{1 - \theta} = 0 \Rightarrow \hat{\theta} = \frac{k}{n}$. The MLE is the observed frequency — reassuringly intuitive.
Example 2: Gaussian#
For $x_i \sim \mathcal{N}(\mu, \sigma^2)$:
giving $\hat\mu = \frac{1}{n}\sum_i x_i$ and $\hat\sigma^2 = \frac{1}{n}\sum_i(x_i - \hat\mu)^2$. Note the MLE variance divides by $n$, not $n - 1$: it is slightly biased, underestimating variance for small samples.
From MLE to loss functions#
Regression → Mean Squared Error#
Model the target as $y = f_{\boldsymbol{\theta}}(\mathbf{x}) + \epsilon$ with $\epsilon \sim \mathcal{N}(0, \sigma^2)$. Then
Summing over data, minimising NLL = minimising squared error. MSE is the loss that corresponds to assuming Gaussian noise. Assume Laplace noise instead and you get mean absolute error.
Classification → Cross-Entropy#
Model $y \in \{0, 1\}$ as Bernoulli with probability $\hat{p} = \sigma(f_{\boldsymbol{\theta}}(\mathbf{x}))$:
That is binary cross-entropy. For $K$ classes with a categorical model and softmax outputs, the NLL is categorical cross-entropy $-\ln\hat{p}_{y}$. Language models are trained by exactly this loss on next-token prediction.
| Assumed output distribution | Negative log-likelihood = loss |
|---|---|
| Gaussian (fixed variance) | Mean squared error |
| Laplace | Mean absolute error |
| Bernoulli | Binary cross-entropy |
| Categorical | Softmax cross-entropy |
| Poisson | Poisson deviance |
MLE with gradient descent#
When no closed form exists (logistic regression, neural networks), we minimise the NLL numerically:
import numpy as np
rng = np.random.default_rng(0)
n = 2000
X = np.column_stack([np.ones(n), rng.normal(size=(n, 2))])
true_w = np.array([-0.5, 2.0, -1.0])
y = rng.random(n) < 1 / (1 + np.exp(-X @ true_w)) # Bernoulli labels
w = np.zeros(3)
for step in range(2000):
p = 1 / (1 + np.exp(-X @ w))
nll = -np.mean(y * np.log(p + 1e-12) + (1 - y) * np.log(1 - p + 1e-12))
grad = X.T @ (p - y) / n
w -= 0.5 * grad
print("estimated:", w.round(2), " true:", true_w, " final NLL:", round(nll, 4))The estimate recovers the true parameters closely — MLE is consistent.
Properties of MLE#
Under regularity conditions, as $n \to \infty$ the MLE is:
- Consistent — converges to the true parameter.
- Asymptotically normal — its sampling distribution approaches a Gaussian with covariance given by the inverse Fisher information.
- Asymptotically efficient — achieves the lowest possible variance (the Cramér–Rao bound).
- Invariant — the MLE of $g(\theta)$ is $g(\hat\theta)$.
The weakness of MLE: overfitting#
With little data, MLE trusts the data completely. Flip a coin three times, get three heads, and MLE says $\hat\theta = 1$: tails are impossible. With a flexible model, MLE fits noise. The remedies are regularisation and, equivalently, priors — which leads to MAP estimation in the next lecture.