Welcome to the Generative AI track. So far, most of our models have been discriminative: given an input, predict a label. Generative models do something more ambitious: they learn the distribution of the data itself, so they can create new examples — images, text, audio, molecules — that look like they came from the real world. This is the technology behind text-to-image systems and large language models.
Two ways to model data#
- A discriminative model learns $p(y \mid \mathbf{x})$ — or just a decision boundary. Logistic regression, SVMs and most classifiers are discriminative.
- A generative model learns $p(\mathbf{x})$ (unconditional) or $p(\mathbf{x}, y)$ / $p(\mathbf{x} \mid y)$ (conditional). Naive Bayes, Gaussian mixtures, HMMs, VAEs, GANs, diffusion models and language models are generative.
A generative model can, in principle, be turned into a classifier with Bayes' rule, $p(y \mid \mathbf{x}) \propto p(\mathbf{x} \mid y)p(y)$. But modelling all of $\mathbf{x}$ is much harder than modelling a boundary: a $256 \times 256$ colour image has almost 200,000 dimensions, and the model must capture everything — textures, lighting, object shapes, their relationships.
What can generative models do?#
- Sampling — generate new, realistic data.
- Density estimation — evaluate how likely a data point is (useful for anomaly detection).
- Conditional generation — generate given a condition: a class, a caption, a sketch, a prompt.
- Representation learning — latent variables often capture meaningful factors.
- Imputation, editing and restoration — fill in missing parts, super-resolve, denoise.
- Simulation and data augmentation — synthetic training data (with caution).
The families of deep generative models#
| Family | Core idea | Exact likelihood? | Sample quality | Sampling speed |
|---|---|---|---|---|
| Autoregressive (PixelCNN, GPT) | Factorise $p(\mathbf{x}) = \prod_i p(x_i \mid x_{<i})$ | Yes | High | Slow (sequential) |
| Variational autoencoders | Latent variable model trained with a lower bound (ELBO) | Lower bound | Moderate (blurry for images) | Fast |
| GANs | Generator vs discriminator game | No | High (sharp) | Fast |
| Normalising flows | Invertible transformations with tractable Jacobians | Yes | Moderate | Fast |
| Diffusion models | Learn to reverse a gradual noising process | Lower bound | Very high, diverse | Slower (many steps; improving) |
| Energy-based models | Unnormalised density $e^{-E(\mathbf{x})}$ | No (intractable $Z$) | Varies | Slow (MCMC) |
Each makes a different trade-off among quality, diversity (mode coverage), speed and likelihood evaluation. We study each in turn.
Evaluating generative models#
This is notoriously hard:
- Likelihood / bits per dimension: measures fit to data, but high likelihood does not guarantee good-looking samples (and vice versa).
- Fréchet Inception Distance (FID): compares the mean and covariance of Inception-network features of real and generated images:
Lower is better. It captures both quality and diversity roughly, but depends on the feature network and sample size.
- Precision and recall for distributions: fidelity (samples look real) vs coverage (all modes represented).
- CLIP score: text–image alignment for text-to-image models.
- Human evaluation: still the gold standard for perceived quality.
A tiny generative model#
Even a Gaussian mixture is a generative model — fit, then sample:
import numpy as np
from sklearn.mixture import GaussianMixture
from sklearn.datasets import make_moons
X, _ = make_moons(n_samples=2000, noise=0.06, random_state=0)
gmm = GaussianMixture(n_components=12, covariance_type="full", random_state=0).fit(X)
samples, _ = gmm.sample(1000) # generate new data
print("avg log-likelihood of real data:", round(gmm.score(X), 3))
print("avg log-likelihood of noise: ", round(gmm.score(np.random.uniform(-2, 3, (1000, 2))), 3))Deep generative models do the same things — learn, sample, score — but in extremely high dimensions.