A neural network does exactly one thing during training: it reduces its loss. If the loss does not capture what you actually care about, the network will faithfully optimise the wrong thing. Choosing a loss is therefore a modelling decision, not a detail. Today we survey the losses you will use most often and connect each to its underlying assumptions.
Regression losses#
Mean squared error (MSE / L2):
Negative log-likelihood of Gaussian noise; predicts the conditional mean; sensitive to outliers.
Mean absolute error (MAE / L1): $\frac{1}{n}\sum_i|y_i - \hat{y}_i|$. Laplace noise; predicts the median; robust but has a constant gradient magnitude, which can make fine convergence slow.
Huber (smooth L1):
Quadratic near zero, linear for large errors โ the best of both. Used in object-detection box regression and in DQN.
Quantile (pinball) loss for prediction intervals; Gaussian NLL with a predicted variance for heteroscedastic uncertainty.
Classification losses#
Binary cross-entropy (BCE): for logit $z$ and label $y \in \{0,1\}$:
Categorical cross-entropy: for logits $\mathbf{z}$ and class $c$:
Both are negative log-likelihoods and have the clean gradient "probability minus target".
Refinements#
- Class weights โ up-weight rare classes.
- Label smoothing โ replace one-hot targets with $(1 - \epsilon)$ on the true class and $\epsilon/(K-1)$ elsewhere. It discourages over-confident logits and often improves calibration and generalisation (widely used in image classification and translation).
- Focal loss (Lin et al., 2017):
where $p_t$ is the predicted probability of the true class. The factor $(1 - p_t)^\gamma$ down-weights easy, well-classified examples, focusing training on hard ones. It was introduced for dense object detection, where easy background examples vastly outnumber objects.
Margin and ranking losses#
- Hinge loss $\max(0, 1 - yz)$ โ SVM-style classification.
- Pairwise ranking loss โ for a relevant item $i$ and irrelevant $j$: $\max(0, m - s_i + s_j)$ or the logistic version $\log(1 + e^{-(s_i - s_j)})$ (as in BPR for recommenders).
- Triplet loss โ for an anchor $a$, positive $p$ (same identity) and negative $n$:
Used in face recognition (FaceNet) to learn embeddings where same-identity faces are close.
Contrastive losses for representation learning#
InfoNCE (used in SimCLR, CLIP and dense retrieval): given a query $\mathbf{q}$, one positive key $\mathbf{k}^+$ and many negatives, with similarity $s$ and temperature $\tau$:
It is a softmax cross-entropy where the "class" is "which key is the positive". Other examples in the batch serve as negatives, so larger batches give harder, more informative contrasts.
Generative-model losses#
- Reconstruction + KL (VAE ELBO).
- Adversarial losses (GANs).
- Denoising MSE on predicted noise (diffusion models).
- Next-token cross-entropy (language models).
We will meet each in the Generative AI track.
Composite losses#
Real systems combine terms: detection = classification + box regression; multi-task models = weighted sum of task losses; regularised objectives add weight decay or auxiliary losses. Balancing weights matters โ a term with a larger scale can dominate gradients. Normalise terms or tune their weights on validation data.
import torch
import torch.nn as nn
import torch.nn.functional as F
logits = torch.tensor([[2.0, 0.5, -1.0], [0.1, 0.2, 3.0]])
target = torch.tensor([0, 2])
print("CE:", F.cross_entropy(logits, target).item())
print("CE + label smoothing 0.1:", F.cross_entropy(logits, target, label_smoothing=0.1).item())
# WRONG: softmax applied twice
print("double-softmax bug:", F.cross_entropy(F.softmax(logits, 1), target).item())
def focal_loss(logits, target, gamma=2.0):
logp = F.log_softmax(logits, dim=1).gather(1, target[:, None]).squeeze(1)
return (-(1 - logp.exp()) ** gamma * logp).mean()
print("focal:", focal_loss(logits, target).item())
pred, y = torch.tensor([1.0, 2.0, 10.0]), torch.tensor([1.2, 1.8, 2.0])
print("MSE:", F.mse_loss(pred, y).item(), " Huber:", F.huber_loss(pred, y, delta=1.0).item())