VAEs optimise a lower bound; GANs provide no likelihood at all. Normalising flows offer something rare: exact log-likelihood evaluation and efficient sampling, both through the same invertible network. They are valuable when precise densities matter — for anomaly detection, scientific inference and as building blocks in other models — and their continuous-time relatives connect directly to modern flow-matching generators.
The change-of-variables formula#
Take a simple base distribution $p_Z(\mathbf{z})$, e.g. a standard Gaussian, and an invertible, differentiable function $f$ with $\mathbf{x} = f(\mathbf{z})$ and $\mathbf{z} = f^{-1}(\mathbf{x})$. Then
or in log form, writing the inverse as a composition of $K$ steps $\mathbf{x} = \mathbf{h}_0 \to \mathbf{h}_1 \to \dots \to \mathbf{h}_K = \mathbf{z}$:
The Jacobian determinant accounts for how the transformation stretches or compresses volume — probability mass is conserved. A sequence of simple invertible steps "flows" the Gaussian into a complex distribution, hence the name.
- Training: maximise $\log p_X(\mathbf{x})$ on data (exact maximum likelihood).
- Sampling: draw $\mathbf{z} \sim p_Z$ and compute $\mathbf{x} = f(\mathbf{z})$.
The design challenge#
A general $d \times d$ Jacobian determinant costs $O(d^3)$ — impossible for images with $d$ in the hundreds of thousands. Flow layers are therefore designed so that (1) they are easily invertible and (2) their Jacobians are triangular, whose determinant is the product of the diagonal.
Coupling layers: RealNVP#
Dinh et al.'s affine coupling layer (RealNVP, 2017) splits the input into two halves $(\mathbf{x}_a, \mathbf{x}_b)$:
where $s$ and $t$ are arbitrary neural networks. Properties:
- Inverse is trivial: $\mathbf{x}_b = (\mathbf{y}_b - t(\mathbf{y}_a))\odot\exp(-s(\mathbf{y}_a))$ — no need to invert $s$ or $t$.
- Jacobian is triangular, so $\log|\det| = \sum_j s(\mathbf{x}_a)_j$ — essentially free.
Alternating which half is transformed (and permuting dimensions) lets all dimensions interact across layers.
import torch
import torch.nn as nn
class AffineCoupling(nn.Module):
def __init__(self, dim, hidden=128, flip=False):
super().__init__()
self.flip, self.d = flip, dim // 2
self.net = nn.Sequential(nn.Linear(self.d, hidden), nn.ReLU(), nn.Linear(hidden, hidden), nn.ReLU(),
nn.Linear(hidden, 2 * (dim - self.d)))
def forward(self, x): # data -> latent direction
xa, xb = x[:, :self.d], x[:, self.d:]
if self.flip: xa, xb = xb, xa
s, t = self.net(xa).chunk(2, dim=1)
s = torch.tanh(s) # keep scales stable
yb = xb * torch.exp(s) + t
y = torch.cat([yb, xa], 1) if self.flip else torch.cat([xa, yb], 1)
return y, s.sum(1) # output and log|det J|
class Flow(nn.Module):
def __init__(self, dim=2, n_layers=8):
super().__init__()
self.layers = nn.ModuleList([AffineCoupling(dim, flip=i % 2 == 1) for i in range(n_layers)])
self.base = torch.distributions.MultivariateNormal(torch.zeros(dim), torch.eye(dim))
def log_prob(self, x):
logdet = 0.0
for layer in self.layers:
x, ld = layer(x); logdet = logdet + ld
return self.base.log_prob(x) + logdet
from sklearn.datasets import make_moons
X = torch.tensor(make_moons(4000, noise=0.05)[0], dtype=torch.float32)
flow = Flow(); opt = torch.optim.Adam(flow.parameters(), 1e-3)
for step in range(2000):
idx = torch.randint(0, len(X), (256,))
loss = -flow.log_prob(X[idx]).mean() # exact negative log-likelihood
opt.zero_grad(); loss.backward(); opt.step()
print("final NLL:", round(loss.item(), 3))(Sampling requires implementing each layer's inverse — a good exercise.)
Other flow architectures#
- Glow (Kingma & Dhariwal, 2018): adds invertible $1 \times 1$ convolutions (learned channel permutations, with LU-decomposed weights for cheap determinants) and ActNorm; generated high-resolution faces with exact likelihood.
- Autoregressive flows: MAF (Masked Autoregressive Flow) has fast density evaluation but slow sampling; IAF (Inverse Autoregressive Flow) has the opposite trade-off — IAF was used to make VAE posteriors more flexible.
- Neural spline flows: use monotonic rational-quadratic splines as element-wise transforms — more expressive than affine coupling.
- Continuous normalising flows (Neural ODEs, FFJORD): define the transformation as the solution of an ODE $d\mathbf{z}/dt = f_\theta(\mathbf{z}, t)$; the log-density evolves by the trace of the Jacobian, computable with stochastic estimators.
Flow matching#
Training continuous flows by maximum likelihood requires expensive ODE solves. Flow matching (Lipman et al., 2023) and rectified flow (Liu et al., 2023) instead regress the network's velocity field directly onto simple target velocities along paths between noise and data (e.g. straight lines $\mathbf{x}_t = (1 - t)\mathbf{z} + t\mathbf{x}$ with target velocity $\mathbf{x} - \mathbf{z}$). Training becomes simulation-free regression, closely related to diffusion models; several recent state-of-the-art image and video generators are trained with flow matching.
Strengths and weaknesses#
Strengths: exact likelihood; exact latent inference ($\mathbf{z} = f^{-1}(\mathbf{x})$); fast sampling (for coupling flows); stable maximum-likelihood training.
Weaknesses: invertibility forces the latent to have the same dimension as the data; architectures are constrained; sample quality of discrete flows lagged behind GANs and diffusion models for images.