✨ Generative AI & LLMs · Lecture 4 of 30

GAN Variants: Conditional GANs, Pix2Pix, CycleGAN and StyleGAN

A tour of influential GAN architectures — conditional GANs, paired and unpaired image-to-image translation, progressive growing and StyleGAN's style-based generator — plus the deepfake concerns they raised.

The basic GAN generates random samples. Researchers quickly extended the idea to control what is generated, to translate images between domains, and to reach photorealistic quality. This lecture surveys the architectures that defined the GAN era and whose ideas still shape image synthesis.

Conditional GANs#

Mirza and Osindero (2014) fed a condition $\mathbf{y}$ — e.g. a class label — to both generator and discriminator: $G(\mathbf{z}, \mathbf{y})$ and $D(\mathbf{x}, \mathbf{y})$. The discriminator now judges whether $\mathbf{x}$ is a real example of class $\mathbf{y}$. Refinements such as the projection discriminator and class-conditional batch normalisation led to BigGAN (Brock et al., 2019), which generated diverse, high-fidelity ImageNet images at scale, and introduced the truncation trick (sampling $\mathbf{z}$ from a truncated distribution to trade diversity for fidelity).

Pix2Pix: paired image-to-image translation#

Isola et al. (2017) framed many problems as translating one image into another: sketches → photos, segmentation maps → street scenes, day → night, aerial photo → map. With paired training data $(\mathbf{x}, \mathbf{y})$:

  • Generator: a U-Net mapping $\mathbf{x}$ to $\hat{\mathbf{y}}$ (skip connections carry low-level detail).
  • Discriminator: a PatchGAN that classifies each $N \times N$ patch as real or fake, focusing on local texture.
  • Loss: adversarial loss + an L1 reconstruction term:
$$ \mathcal{L} = \mathcal{L}_{\text{cGAN}}(G, D) + \lambda\,\mathbb{E}\big[\|\mathbf{y} - G(\mathbf{x})\|_1\big] $$

L1 captures low-frequency correctness; the adversarial term makes outputs sharp.

CycleGAN: unpaired translation#

Paired data is often unavailable (there are no photos of the same scene painted by Monet). Zhu et al. (2017) learned mappings $G: X \to Y$ and $F: Y \to X$ from unpaired collections, with adversarial losses in both domains plus a cycle-consistency loss:

$$ \mathcal{L}_{\text{cyc}} = \mathbb{E}_{\mathbf{x}}\big[\|F(G(\mathbf{x})) - \mathbf{x}\|_1\big] + \mathbb{E}_{\mathbf{y}}\big[\|G(F(\mathbf{y})) - \mathbf{y}\|_1\big] $$

Translating a horse to a zebra and back should return the original horse — preventing the generator from producing arbitrary zebras unrelated to the input. CycleGAN powered photo↔painting style transfer, summer↔winter, and domain adaptation (e.g. synthetic → realistic driving scenes).

Progressive GAN and StyleGAN#

Progressive growing (Karras et al., 2018) started training at $4 \times 4$ resolution and progressively added layers up to $1024 \times 1024$, stabilising high-resolution training.

StyleGAN (Karras, Laine & Aila, 2019) redesigned the generator:

  1. A mapping network (8-layer MLP) transforms $\mathbf{z}$ into an intermediate latent $\mathbf{w}$, which is less entangled than $\mathbf{z}$.
  2. The synthesis network starts from a learned constant, and $\mathbf{w}$ controls each layer through adaptive instance normalisation (AdaIN) — injecting "style" at every resolution.
  3. Per-pixel noise inputs add stochastic detail (hair strands, freckles).

Styles at coarse layers control pose and face shape; middle layers control features; fine layers control colour and micro-texture. Style mixing combines coarse styles from one image with fine styles from another. StyleGAN2 (2020) removed characteristic artefacts (weight demodulation replaced AdaIN, path-length regularisation), and StyleGAN3 (2021) tackled "texture sticking" with alias-free operations. The results — faces of people who do not exist — were strikingly photorealistic.

GAN inversion and editing#

Because StyleGAN's $\mathcal{W}$ space is well organised, you can invert a real image (find the latent that reproduces it) and then edit it by moving along semantic directions (age, smile, lighting) discovered with labels or unsupervised methods (e.g. PCA in latent space, GANSpace). This enabled powerful photo-editing tools.

Summary of variants#

ModelYearKey contribution
DCGAN2016Stable convolutional architecture guidelines
Conditional GAN2014Class/attribute control
Pix2Pix2017Paired translation: U-Net + PatchGAN + L1
CycleGAN2017Unpaired translation via cycle consistency
WGAN-GP2017Wasserstein loss with gradient penalty
Progressive GAN2018Grow resolution during training
BigGAN2019Large-scale class-conditional generation; truncation trick
StyleGAN 1–32019–2021Style-based generator; photoreal faces; editable latent space

Deepfakes and responsibility#

GAN-based face synthesis and face swapping enabled "deepfakes" — realistic fabricated images and videos of real people. Harms include non-consensual intimate imagery, fraud, harassment and political misinformation; the mere possibility of fakes also lets real evidence be dismissed ("the liar's dividend"). Countermeasures include detection models (an arms race), content provenance standards such as C2PA that cryptographically sign media at capture or creation, watermarking, platform policies and laws against specific abuses.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Generative Adversarial Networks: The Generator–Discriminator Game

GANs train a generator to fool a discriminator in a two-player game. We derive the minimax objective and its optimum, study training dynamics, mode collapse and the non-saturating loss, and build a DCGAN.

Advanced⏱ 5 min#193
✨ Generative AI & LLMs

Normalising Flows: Exact Likelihood with Invertible Networks

Normalising flows transform a simple distribution into a complex one through invertible mappings, giving exact likelihoods and fast sampling. We derive the change-of-variables formula, coupling layers, RealNVP and Glow, and continuous flows.

Advanced⏱ 5 min#195
✨ Generative AI & LLMs

Variational Autoencoders: Probabilistic Latent Spaces

VAEs turn autoencoders into generative models by learning a smooth, probabilistic latent space. We derive the evidence lower bound, the reparameterisation trick and the KL term, and discuss blurriness, posterior collapse and β-VAE.

Advanced⏱ 5 min#192