๐Ÿ”— Deep Learning ยท Lecture 5 of 38

Loss Functions in Deep Learning: What Are We Really Optimising?

The loss function defines what "good" means to a network. We survey regression, classification, ranking and representation-learning losses, their probabilistic meaning, and common implementation mistakes.

A neural network does exactly one thing during training: it reduces its loss. If the loss does not capture what you actually care about, the network will faithfully optimise the wrong thing. Choosing a loss is therefore a modelling decision, not a detail. Today we survey the losses you will use most often and connect each to its underlying assumptions.

Regression losses#

Mean squared error (MSE / L2):

$$ L = \frac{1}{n}\sum_i(y_i - \hat{y}_i)^2 $$

Negative log-likelihood of Gaussian noise; predicts the conditional mean; sensitive to outliers.

Mean absolute error (MAE / L1): $\frac{1}{n}\sum_i|y_i - \hat{y}_i|$. Laplace noise; predicts the median; robust but has a constant gradient magnitude, which can make fine convergence slow.

Huber (smooth L1):

$$ L_\delta(r) = \begin{cases} \frac{1}{2}r^2 & |r| \le \delta \\ \delta\left(|r| - \frac{1}{2}\delta\right) & |r| > \delta \end{cases} $$

Quadratic near zero, linear for large errors โ€” the best of both. Used in object-detection box regression and in DQN.

Quantile (pinball) loss for prediction intervals; Gaussian NLL with a predicted variance for heteroscedastic uncertainty.

Classification losses#

Binary cross-entropy (BCE): for logit $z$ and label $y \in \{0,1\}$:

$$ L = -\big[y\log\sigma(z) + (1 - y)\log(1 - \sigma(z))\big] $$

Categorical cross-entropy: for logits $\mathbf{z}$ and class $c$:

$$ L = -\log\frac{e^{z_c}}{\sum_k e^{z_k}} = -z_c + \log\sum_k e^{z_k} $$

Both are negative log-likelihoods and have the clean gradient "probability minus target".

Refinements#

  • Class weights โ€” up-weight rare classes.
  • Label smoothing โ€” replace one-hot targets with $(1 - \epsilon)$ on the true class and $\epsilon/(K-1)$ elsewhere. It discourages over-confident logits and often improves calibration and generalisation (widely used in image classification and translation).
  • Focal loss (Lin et al., 2017):
$$ L = -(1 - p_t)^\gamma\log p_t $$

where $p_t$ is the predicted probability of the true class. The factor $(1 - p_t)^\gamma$ down-weights easy, well-classified examples, focusing training on hard ones. It was introduced for dense object detection, where easy background examples vastly outnumber objects.

Margin and ranking losses#

  • Hinge loss $\max(0, 1 - yz)$ โ€” SVM-style classification.
  • Pairwise ranking loss โ€” for a relevant item $i$ and irrelevant $j$: $\max(0, m - s_i + s_j)$ or the logistic version $\log(1 + e^{-(s_i - s_j)})$ (as in BPR for recommenders).
  • Triplet loss โ€” for an anchor $a$, positive $p$ (same identity) and negative $n$:
$$ L = \max\big(0,\; \|f(a) - f(p)\|^2 - \|f(a) - f(n)\|^2 + m\big) $$

Used in face recognition (FaceNet) to learn embeddings where same-identity faces are close.

Contrastive losses for representation learning#

InfoNCE (used in SimCLR, CLIP and dense retrieval): given a query $\mathbf{q}$, one positive key $\mathbf{k}^+$ and many negatives, with similarity $s$ and temperature $\tau$:

$$ L = -\log\frac{\exp(s(\mathbf{q}, \mathbf{k}^+)/\tau)}{\sum_{j}\exp(s(\mathbf{q}, \mathbf{k}_j)/\tau)} $$

It is a softmax cross-entropy where the "class" is "which key is the positive". Other examples in the batch serve as negatives, so larger batches give harder, more informative contrasts.

Generative-model losses#

  • Reconstruction + KL (VAE ELBO).
  • Adversarial losses (GANs).
  • Denoising MSE on predicted noise (diffusion models).
  • Next-token cross-entropy (language models).

We will meet each in the Generative AI track.

Composite losses#

Real systems combine terms: detection = classification + box regression; multi-task models = weighted sum of task losses; regularised objectives add weight decay or auxiliary losses. Balancing weights matters โ€” a term with a larger scale can dominate gradients. Normalise terms or tune their weights on validation data.

python
import torch
import torch.nn as nn
import torch.nn.functional as F

logits = torch.tensor([[2.0, 0.5, -1.0], [0.1, 0.2, 3.0]])
target = torch.tensor([0, 2])

print("CE:", F.cross_entropy(logits, target).item())
print("CE + label smoothing 0.1:", F.cross_entropy(logits, target, label_smoothing=0.1).item())
# WRONG: softmax applied twice
print("double-softmax bug:", F.cross_entropy(F.softmax(logits, 1), target).item())

def focal_loss(logits, target, gamma=2.0):
    logp = F.log_softmax(logits, dim=1).gather(1, target[:, None]).squeeze(1)
    return (-(1 - logp.exp()) ** gamma * logp).mean()
print("focal:", focal_loss(logits, target).item())

pred, y = torch.tensor([1.0, 2.0, 10.0]), torch.tensor([1.2, 1.8, 2.0])
print("MSE:", F.mse_loss(pred, y).item(), " Huber:", F.huber_loss(pred, y, delta=1.0).item())
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Activation Functions: Sigmoid, Tanh, ReLU, GELU, SwiGLU and Softmax

The choice of non-linearity shapes how gradients flow and how networks learn. We compare the classical and modern activations, their derivatives and failure modes, and which to use where.

Beginnerโฑ 5 min#100
๐Ÿ”— Deep Learning

Backpropagation Derived Step by Step

Backpropagation computes every gradient in a network at about the cost of one forward pass. We derive it for a two-layer network by hand, generalise to any depth, implement it in NumPy and verify it numerically.

Intermediateโฑ 6 min#102
๐Ÿ”— Deep Learning

Multilayer Perceptrons and the Universal Approximation Theorem

With one hidden layer, a network can approximate any continuous function โ€” so why go deep? We state the universal approximation theorem, build intuition with bumps, and explain the efficiency advantages of depth.

Intermediateโฑ 5 min#099