✨ Generative AI & LLMs · Lecture 5 of 30

Normalising Flows: Exact Likelihood with Invertible Networks

Normalising flows transform a simple distribution into a complex one through invertible mappings, giving exact likelihoods and fast sampling. We derive the change-of-variables formula, coupling layers, RealNVP and Glow, and continuous flows.

VAEs optimise a lower bound; GANs provide no likelihood at all. Normalising flows offer something rare: exact log-likelihood evaluation and efficient sampling, both through the same invertible network. They are valuable when precise densities matter — for anomaly detection, scientific inference and as building blocks in other models — and their continuous-time relatives connect directly to modern flow-matching generators.

The change-of-variables formula#

Take a simple base distribution $p_Z(\mathbf{z})$, e.g. a standard Gaussian, and an invertible, differentiable function $f$ with $\mathbf{x} = f(\mathbf{z})$ and $\mathbf{z} = f^{-1}(\mathbf{x})$. Then

$$ p_X(\mathbf{x}) = p_Z\big(f^{-1}(\mathbf{x})\big)\left|\det\frac{\partial f^{-1}(\mathbf{x})}{\partial\mathbf{x}}\right| $$

or in log form, writing the inverse as a composition of $K$ steps $\mathbf{x} = \mathbf{h}_0 \to \mathbf{h}_1 \to \dots \to \mathbf{h}_K = \mathbf{z}$:

$$ \log p_X(\mathbf{x}) = \log p_Z(\mathbf{z}) + \sum_{k=1}^{K}\log\left|\det\frac{\partial\mathbf{h}_k}{\partial\mathbf{h}_{k-1}}\right| $$

The Jacobian determinant accounts for how the transformation stretches or compresses volume — probability mass is conserved. A sequence of simple invertible steps "flows" the Gaussian into a complex distribution, hence the name.

  • Training: maximise $\log p_X(\mathbf{x})$ on data (exact maximum likelihood).
  • Sampling: draw $\mathbf{z} \sim p_Z$ and compute $\mathbf{x} = f(\mathbf{z})$.

The design challenge#

A general $d \times d$ Jacobian determinant costs $O(d^3)$ — impossible for images with $d$ in the hundreds of thousands. Flow layers are therefore designed so that (1) they are easily invertible and (2) their Jacobians are triangular, whose determinant is the product of the diagonal.

Coupling layers: RealNVP#

Dinh et al.'s affine coupling layer (RealNVP, 2017) splits the input into two halves $(\mathbf{x}_a, \mathbf{x}_b)$:

$$ \mathbf{y}_a = \mathbf{x}_a, \qquad \mathbf{y}_b = \mathbf{x}_b\odot\exp\big(s(\mathbf{x}_a)\big) + t(\mathbf{x}_a) $$

where $s$ and $t$ are arbitrary neural networks. Properties:

  • Inverse is trivial: $\mathbf{x}_b = (\mathbf{y}_b - t(\mathbf{y}_a))\odot\exp(-s(\mathbf{y}_a))$ — no need to invert $s$ or $t$.
  • Jacobian is triangular, so $\log|\det| = \sum_j s(\mathbf{x}_a)_j$ — essentially free.

Alternating which half is transformed (and permuting dimensions) lets all dimensions interact across layers.

python
import torch
import torch.nn as nn

class AffineCoupling(nn.Module):
    def __init__(self, dim, hidden=128, flip=False):
        super().__init__()
        self.flip, self.d = flip, dim // 2
        self.net = nn.Sequential(nn.Linear(self.d, hidden), nn.ReLU(), nn.Linear(hidden, hidden), nn.ReLU(),
                                 nn.Linear(hidden, 2 * (dim - self.d)))
    def forward(self, x):                                  # data -> latent direction
        xa, xb = x[:, :self.d], x[:, self.d:]
        if self.flip: xa, xb = xb, xa
        s, t = self.net(xa).chunk(2, dim=1)
        s = torch.tanh(s)                                  # keep scales stable
        yb = xb * torch.exp(s) + t
        y = torch.cat([yb, xa], 1) if self.flip else torch.cat([xa, yb], 1)
        return y, s.sum(1)                                 # output and log|det J|

class Flow(nn.Module):
    def __init__(self, dim=2, n_layers=8):
        super().__init__()
        self.layers = nn.ModuleList([AffineCoupling(dim, flip=i % 2 == 1) for i in range(n_layers)])
        self.base = torch.distributions.MultivariateNormal(torch.zeros(dim), torch.eye(dim))
    def log_prob(self, x):
        logdet = 0.0
        for layer in self.layers:
            x, ld = layer(x); logdet = logdet + ld
        return self.base.log_prob(x) + logdet

from sklearn.datasets import make_moons
X = torch.tensor(make_moons(4000, noise=0.05)[0], dtype=torch.float32)
flow = Flow(); opt = torch.optim.Adam(flow.parameters(), 1e-3)
for step in range(2000):
    idx = torch.randint(0, len(X), (256,))
    loss = -flow.log_prob(X[idx]).mean()                   # exact negative log-likelihood
    opt.zero_grad(); loss.backward(); opt.step()
print("final NLL:", round(loss.item(), 3))

(Sampling requires implementing each layer's inverse — a good exercise.)

Other flow architectures#

  • Glow (Kingma & Dhariwal, 2018): adds invertible $1 \times 1$ convolutions (learned channel permutations, with LU-decomposed weights for cheap determinants) and ActNorm; generated high-resolution faces with exact likelihood.
  • Autoregressive flows: MAF (Masked Autoregressive Flow) has fast density evaluation but slow sampling; IAF (Inverse Autoregressive Flow) has the opposite trade-off — IAF was used to make VAE posteriors more flexible.
  • Neural spline flows: use monotonic rational-quadratic splines as element-wise transforms — more expressive than affine coupling.
  • Continuous normalising flows (Neural ODEs, FFJORD): define the transformation as the solution of an ODE $d\mathbf{z}/dt = f_\theta(\mathbf{z}, t)$; the log-density evolves by the trace of the Jacobian, computable with stochastic estimators.

Flow matching#

Training continuous flows by maximum likelihood requires expensive ODE solves. Flow matching (Lipman et al., 2023) and rectified flow (Liu et al., 2023) instead regress the network's velocity field directly onto simple target velocities along paths between noise and data (e.g. straight lines $\mathbf{x}_t = (1 - t)\mathbf{z} + t\mathbf{x}$ with target velocity $\mathbf{x} - \mathbf{z}$). Training becomes simulation-free regression, closely related to diffusion models; several recent state-of-the-art image and video generators are trained with flow matching.

Strengths and weaknesses#

Strengths: exact likelihood; exact latent inference ($\mathbf{z} = f^{-1}(\mathbf{x})$); fast sampling (for coupling flows); stable maximum-likelihood training.

Weaknesses: invertibility forces the latent to have the same dimension as the data; architectures are constrained; sample quality of discrete flows lagged behind GANs and diffusion models for images.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Generative vs Discriminative Models: Learning to Create

We open the Generative AI track by contrasting models that draw boundaries with models that learn the data distribution itself, and map the families of deep generative models — autoregressive, VAEs, GANs, flows and diffusion.

Beginner⏱ 5 min#191
✨ Generative AI & LLMs

GAN Variants: Conditional GANs, Pix2Pix, CycleGAN and StyleGAN

A tour of influential GAN architectures — conditional GANs, paired and unpaired image-to-image translation, progressive growing and StyleGAN's style-based generator — plus the deepfake concerns they raised.

Advanced⏱ 5 min#194
✨ Generative AI & LLMs

Diffusion Models: Generating by Learning to Denoise

Diffusion models gradually add noise to data and train a network to reverse the process. We derive the DDPM forward and reverse processes, the simple noise-prediction loss, sampling, the score-based view, and faster samplers like DDIM.

Advanced⏱ 6 min#196