๐Ÿ‘๏ธ Computer Vision ยท Lecture 27 of 27

Adversarial Examples: Fooling Neural Networks and Defending Them

Imperceptible perturbations can make a network confidently wrong. We derive FGSM and PGD attacks, explain why adversarial examples exist, cover physical and black-box attacks, and evaluate defences including adversarial training.

In 2013 Szegedy and colleagues discovered something unsettling: adding a tiny, carefully crafted perturbation โ€” invisible to humans โ€” to an image could make a state-of-the-art network misclassify it with high confidence. A panda becomes a "gibbon"; a stop sign becomes a "speed limit". These adversarial examples reveal that networks do not see the world the way we do, and they raise security concerns for any system where an attacker might manipulate inputs.

Threat model#

An adversarial example $\mathbf{x}' = \mathbf{x} + \boldsymbol{\delta}$ changes the model's prediction while the perturbation is small under some norm:

$$ \|\boldsymbol{\delta}\|_p \le \epsilon, \qquad f(\mathbf{x} + \boldsymbol{\delta}) \ne y $$

Commonly $\ell_\infty$ with $\epsilon = 8/255$ (each pixel changes by at most about 3%), or $\ell_2$. Attacks can be untargeted (any wrong class) or targeted (a chosen class).

Knowledge levels:

  • White-box โ€” the attacker knows the model and its gradients.
  • Black-box โ€” only queries (or nothing at all) are available.

FGSM: the fast gradient sign method#

Goodfellow, Shlens and Szegedy (2015) took one step in the direction that increases the loss the most under an $\ell_\infty$ constraint:

$$ \mathbf{x}' = \mathbf{x} + \epsilon\cdot\text{sign}\big(\nabla_{\mathbf{x}}\mathcal{L}(f(\mathbf{x}), y)\big) $$

PGD: projected gradient descent#

Madry et al. (2018) iterate smaller steps and project back into the allowed ball after each:

$$ \mathbf{x}^{t+1} = \Pi_{\mathcal{B}_\epsilon(\mathbf{x})}\Big(\mathbf{x}^t + \alpha\cdot\text{sign}\big(\nabla_{\mathbf{x}}\mathcal{L}(f(\mathbf{x}^t), y)\big)\Big) $$

starting from a random point in the ball. PGD is a strong, standard first-order attack.

python
import torch
import torch.nn.functional as F

def fgsm(model, x, y, eps=8 / 255):
    x = x.clone().requires_grad_(True)
    F.cross_entropy(model(x), y).backward()
    return (x + eps * x.grad.sign()).clamp(0, 1).detach()

def pgd(model, x, y, eps=8 / 255, alpha=2 / 255, steps=10):
    x_adv = (x + torch.empty_like(x).uniform_(-eps, eps)).clamp(0, 1)
    for _ in range(steps):
        x_adv.requires_grad_(True)
        grad, = torch.autograd.grad(F.cross_entropy(model(x_adv), y), x_adv)
        x_adv = x_adv.detach() + alpha * grad.sign()
        x_adv = torch.min(torch.max(x_adv, x - eps), x + eps).clamp(0, 1)   # project to the eps-ball
    return x_adv.detach()

# robust_acc = (model(pgd(model, x, y)).argmax(1) == y).float().mean()

(Here the model is assumed to include its own input normalisation, so attacks operate in $[0, 1]$ pixel space.) On an undefended CIFAR-10 classifier, PGD with $\epsilon = 8/255$ typically drives accuracy close to zero.

Why do adversarial examples exist?#

  • Linearity hypothesis (Goodfellow et al.): in high dimensions, many tiny coordinated changes add up. For a linear score $\mathbf{w}^\top\mathbf{x}$, a perturbation $\epsilon\,\text{sign}(\mathbf{w})$ changes the score by $\epsilon\|\mathbf{w}\|_1$, which grows with dimension โ€” even though each pixel barely changes.
  • Non-robust features (Ilyas et al., 2019): datasets contain genuinely predictive but imperceptible patterns; standard training exploits them, and adversaries flip them. Adversarial vulnerability is partly a property of the data, not only of the model.
  • Geometry: decision boundaries lie close to most data points in some directions of high-dimensional space.

Beyond the digital lab#

  • Transferability: adversarial examples crafted on one model often fool others, enabling black-box transfer attacks via a surrogate model.
  • Query-based attacks estimate gradients from outputs alone.
  • Physical attacks: printed adversarial patches, stickers on stop signs, or adversarial eyeglass frames have fooled classifiers and detectors in real-world conditions.
  • Universal perturbations: a single perturbation that fools a model on most images.

Defences#

Many proposed defences were later broken. Athalye, Carlini and Wagner (2018) showed that defences relying on obfuscated gradients (non-differentiable preprocessing, randomness) gave a false sense of security and were defeated by adaptive attacks.

The most reliable empirical defence is adversarial training โ€” train on adversarial examples generated on the fly, solving the minโ€“max problem:

$$ \min_{\boldsymbol{\theta}}\;\mathbb{E}_{(\mathbf{x}, y)}\left[\max_{\|\boldsymbol{\delta}\| \le \epsilon}\mathcal{L}\big(f_{\boldsymbol{\theta}}(\mathbf{x} + \boldsymbol{\delta}), y\big)\right] $$

It works, but it multiplies training cost (several attack steps per batch), lowers clean accuracy, and robustness typically holds only for the threat model used in training. Variants such as TRADES balance clean and robust accuracy. Certified defences (e.g. randomised smoothing) provide provable guarantees within a radius, at a cost in accuracy.

Why it matters beyond security#

Adversarial robustness connects to general reliability: robust models often have more interpretable gradients and features aligned with human perception. And adversarial thinking extends to other modalities โ€” adversarial audio, text perturbations, and prompt injection and jailbreaks against language models, which we meet in the LLM and Ethics tracks.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

Explaining Vision Models: Saliency Maps, Grad-CAM and Their Limits

Which pixels made the model decide? We study gradient saliency, Grad-CAM, integrated gradients and occlusion, show how they reveal shortcuts, and discuss sanity checks that expose unreliable explanations.

Intermediateโฑ 5 min#160
๐Ÿ‘๏ธ Computer Vision

Vision Foundation Models: Segment Anything and Promptable Vision

Vision is following language towards general-purpose foundation models. We study the Segment Anything Model's promptable design and data engine, open-vocabulary detection, and how foundation models change vision workflows.

Advancedโฑ 5 min#159
๐Ÿ‘๏ธ Computer Vision

AI in Medical Imaging: Opportunities, Pitfalls and Validation

Deep learning can detect disease in X-rays, retinal scans and pathology slides. We survey modalities and tasks, discuss data and labelling challenges, shortcut learning, rigorous clinical validation, and deployment responsibilities.

Intermediateโฑ 5 min#158