In 2013 Szegedy and colleagues discovered something unsettling: adding a tiny, carefully crafted perturbation โ invisible to humans โ to an image could make a state-of-the-art network misclassify it with high confidence. A panda becomes a "gibbon"; a stop sign becomes a "speed limit". These adversarial examples reveal that networks do not see the world the way we do, and they raise security concerns for any system where an attacker might manipulate inputs.
Threat model#
An adversarial example $\mathbf{x}' = \mathbf{x} + \boldsymbol{\delta}$ changes the model's prediction while the perturbation is small under some norm:
Commonly $\ell_\infty$ with $\epsilon = 8/255$ (each pixel changes by at most about 3%), or $\ell_2$. Attacks can be untargeted (any wrong class) or targeted (a chosen class).
Knowledge levels:
- White-box โ the attacker knows the model and its gradients.
- Black-box โ only queries (or nothing at all) are available.
FGSM: the fast gradient sign method#
Goodfellow, Shlens and Szegedy (2015) took one step in the direction that increases the loss the most under an $\ell_\infty$ constraint:
PGD: projected gradient descent#
Madry et al. (2018) iterate smaller steps and project back into the allowed ball after each:
starting from a random point in the ball. PGD is a strong, standard first-order attack.
import torch
import torch.nn.functional as F
def fgsm(model, x, y, eps=8 / 255):
x = x.clone().requires_grad_(True)
F.cross_entropy(model(x), y).backward()
return (x + eps * x.grad.sign()).clamp(0, 1).detach()
def pgd(model, x, y, eps=8 / 255, alpha=2 / 255, steps=10):
x_adv = (x + torch.empty_like(x).uniform_(-eps, eps)).clamp(0, 1)
for _ in range(steps):
x_adv.requires_grad_(True)
grad, = torch.autograd.grad(F.cross_entropy(model(x_adv), y), x_adv)
x_adv = x_adv.detach() + alpha * grad.sign()
x_adv = torch.min(torch.max(x_adv, x - eps), x + eps).clamp(0, 1) # project to the eps-ball
return x_adv.detach()
# robust_acc = (model(pgd(model, x, y)).argmax(1) == y).float().mean()(Here the model is assumed to include its own input normalisation, so attacks operate in $[0, 1]$ pixel space.) On an undefended CIFAR-10 classifier, PGD with $\epsilon = 8/255$ typically drives accuracy close to zero.
Why do adversarial examples exist?#
- Linearity hypothesis (Goodfellow et al.): in high dimensions, many tiny coordinated changes add up. For a linear score $\mathbf{w}^\top\mathbf{x}$, a perturbation $\epsilon\,\text{sign}(\mathbf{w})$ changes the score by $\epsilon\|\mathbf{w}\|_1$, which grows with dimension โ even though each pixel barely changes.
- Non-robust features (Ilyas et al., 2019): datasets contain genuinely predictive but imperceptible patterns; standard training exploits them, and adversaries flip them. Adversarial vulnerability is partly a property of the data, not only of the model.
- Geometry: decision boundaries lie close to most data points in some directions of high-dimensional space.
Beyond the digital lab#
- Transferability: adversarial examples crafted on one model often fool others, enabling black-box transfer attacks via a surrogate model.
- Query-based attacks estimate gradients from outputs alone.
- Physical attacks: printed adversarial patches, stickers on stop signs, or adversarial eyeglass frames have fooled classifiers and detectors in real-world conditions.
- Universal perturbations: a single perturbation that fools a model on most images.
Defences#
Many proposed defences were later broken. Athalye, Carlini and Wagner (2018) showed that defences relying on obfuscated gradients (non-differentiable preprocessing, randomness) gave a false sense of security and were defeated by adaptive attacks.
The most reliable empirical defence is adversarial training โ train on adversarial examples generated on the fly, solving the minโmax problem:
It works, but it multiplies training cost (several attack steps per batch), lowers clean accuracy, and robustness typically holds only for the threat model used in training. Variants such as TRADES balance clean and robust accuracy. Certified defences (e.g. randomised smoothing) provide provable guarantees within a radius, at a cost in accuracy.
Why it matters beyond security#
Adversarial robustness connects to general reliability: robust models often have more interpretable gradients and features aligned with human perception. And adversarial thinking extends to other modalities โ adversarial audio, text perturbations, and prompt injection and jailbreaks against language models, which we meet in the LLM and Ethics tracks.