The basic GAN generates random samples. Researchers quickly extended the idea to control what is generated, to translate images between domains, and to reach photorealistic quality. This lecture surveys the architectures that defined the GAN era and whose ideas still shape image synthesis.
Conditional GANs#
Mirza and Osindero (2014) fed a condition $\mathbf{y}$ — e.g. a class label — to both generator and discriminator: $G(\mathbf{z}, \mathbf{y})$ and $D(\mathbf{x}, \mathbf{y})$. The discriminator now judges whether $\mathbf{x}$ is a real example of class $\mathbf{y}$. Refinements such as the projection discriminator and class-conditional batch normalisation led to BigGAN (Brock et al., 2019), which generated diverse, high-fidelity ImageNet images at scale, and introduced the truncation trick (sampling $\mathbf{z}$ from a truncated distribution to trade diversity for fidelity).
Pix2Pix: paired image-to-image translation#
Isola et al. (2017) framed many problems as translating one image into another: sketches → photos, segmentation maps → street scenes, day → night, aerial photo → map. With paired training data $(\mathbf{x}, \mathbf{y})$:
- Generator: a U-Net mapping $\mathbf{x}$ to $\hat{\mathbf{y}}$ (skip connections carry low-level detail).
- Discriminator: a PatchGAN that classifies each $N \times N$ patch as real or fake, focusing on local texture.
- Loss: adversarial loss + an L1 reconstruction term:
L1 captures low-frequency correctness; the adversarial term makes outputs sharp.
CycleGAN: unpaired translation#
Paired data is often unavailable (there are no photos of the same scene painted by Monet). Zhu et al. (2017) learned mappings $G: X \to Y$ and $F: Y \to X$ from unpaired collections, with adversarial losses in both domains plus a cycle-consistency loss:
Translating a horse to a zebra and back should return the original horse — preventing the generator from producing arbitrary zebras unrelated to the input. CycleGAN powered photo↔painting style transfer, summer↔winter, and domain adaptation (e.g. synthetic → realistic driving scenes).
Progressive GAN and StyleGAN#
Progressive growing (Karras et al., 2018) started training at $4 \times 4$ resolution and progressively added layers up to $1024 \times 1024$, stabilising high-resolution training.
StyleGAN (Karras, Laine & Aila, 2019) redesigned the generator:
- A mapping network (8-layer MLP) transforms $\mathbf{z}$ into an intermediate latent $\mathbf{w}$, which is less entangled than $\mathbf{z}$.
- The synthesis network starts from a learned constant, and $\mathbf{w}$ controls each layer through adaptive instance normalisation (AdaIN) — injecting "style" at every resolution.
- Per-pixel noise inputs add stochastic detail (hair strands, freckles).
Styles at coarse layers control pose and face shape; middle layers control features; fine layers control colour and micro-texture. Style mixing combines coarse styles from one image with fine styles from another. StyleGAN2 (2020) removed characteristic artefacts (weight demodulation replaced AdaIN, path-length regularisation), and StyleGAN3 (2021) tackled "texture sticking" with alias-free operations. The results — faces of people who do not exist — were strikingly photorealistic.
GAN inversion and editing#
Because StyleGAN's $\mathcal{W}$ space is well organised, you can invert a real image (find the latent that reproduces it) and then edit it by moving along semantic directions (age, smile, lighting) discovered with labels or unsupervised methods (e.g. PCA in latent space, GANSpace). This enabled powerful photo-editing tools.
Summary of variants#
| Model | Year | Key contribution |
|---|---|---|
| DCGAN | 2016 | Stable convolutional architecture guidelines |
| Conditional GAN | 2014 | Class/attribute control |
| Pix2Pix | 2017 | Paired translation: U-Net + PatchGAN + L1 |
| CycleGAN | 2017 | Unpaired translation via cycle consistency |
| WGAN-GP | 2017 | Wasserstein loss with gradient penalty |
| Progressive GAN | 2018 | Grow resolution during training |
| BigGAN | 2019 | Large-scale class-conditional generation; truncation trick |
| StyleGAN 1–3 | 2019–2021 | Style-based generator; photoreal faces; editable latent space |
Deepfakes and responsibility#
GAN-based face synthesis and face swapping enabled "deepfakes" — realistic fabricated images and videos of real people. Harms include non-consensual intimate imagery, fraud, harassment and political misinformation; the mere possibility of fakes also lets real evidence be dismissed ("the liar's dividend"). Countermeasures include detection models (an arms race), content provenance standards such as C2PA that cryptographically sign media at capture or creation, watermarking, platform policies and laws against specific abuses.