∑ Mathematics for ML · Lecture 13 of 25

The Gaussian Distribution: Why It Is Everywhere

The bell curve appears in noise models, weight initialisation, VAEs and diffusion models. We study univariate and multivariate Gaussians, the central limit theorem, and the closure properties that make Gaussians so convenient.

If one distribution deserves a full lecture, it is the Gaussian (normal) distribution. It models measurement noise, underlies least squares, initialises neural network weights, defines the latent space of VAEs and drives the noise process of diffusion models. Understanding why it appears so often, and how to manipulate it, is essential.

The univariate Gaussian#

$$ \mathcal{N}(x \mid \mu, \sigma^2) = \frac{1}{\sqrt{2\pi\sigma^2}}\exp\left(-\frac{(x - \mu)^2}{2\sigma^2}\right) $$
  • Mean $\mu$, variance $\sigma^2$.
  • About 68% of the mass lies within $\mu \pm \sigma$, 95% within $\mu \pm 1.96\sigma$, and 99.7% within $\mu \pm 3\sigma$.
  • Standardisation: if $X \sim \mathcal{N}(\mu, \sigma^2)$ then $Z = (X - \mu)/\sigma \sim \mathcal{N}(0, 1)$.

The log-density is a quadratic in $x$:

$$ \ln \mathcal{N}(x \mid \mu, \sigma^2) = -\frac{(x - \mu)^2}{2\sigma^2} - \frac{1}{2}\ln(2\pi\sigma^2) $$

This is why maximising Gaussian likelihood is equivalent to minimising squared error.

Why Gaussians are everywhere#

1. The Central Limit Theorem#

If $X_1, \dots, X_n$ are independent and identically distributed with mean $\mu$ and finite variance $\sigma^2$, then

$$ \sqrt{n}\,\frac{\bar{X}_n - \mu}{\sigma} \xrightarrow{d} \mathcal{N}(0, 1) $$

Sums of many small independent effects are approximately Gaussian, whatever their individual distributions. Measurement errors, heights and many aggregate quantities behave this way.

2. Maximum entropy#

Among all distributions with a given mean and variance, the Gaussian has the highest entropy — it makes the fewest additional assumptions. Choosing a Gaussian is the most "honest" choice when you only know the first two moments.

3. Mathematical convenience#

Gaussians are closed under linear transformations, marginalisation, conditioning and products (up to normalisation). Everything stays Gaussian, and everything has a closed form.

python
import numpy as np

rng = np.random.default_rng(0)
# CLT demo: averages of skewed exponential variables become Gaussian
for n in [1, 2, 10, 50]:
    means = rng.exponential(1.0, size=(100_000, n)).mean(axis=1)
    z = (means - 1.0) / (1.0 / np.sqrt(n))
    skew = np.mean(z**3)
    print(f"n={n:>2}: skewness of standardised mean = {skew:.3f}")   # -> 0 as n grows

The multivariate Gaussian#

For $\mathbf{x} \in \mathbb{R}^d$ with mean $\boldsymbol{\mu}$ and covariance $\boldsymbol{\Sigma}$ (positive definite):

$$ \mathcal{N}(\mathbf{x} \mid \boldsymbol{\mu}, \boldsymbol{\Sigma}) = \frac{1}{(2\pi)^{d/2}|\boldsymbol{\Sigma}|^{1/2}}\exp\left(-\frac{1}{2}(\mathbf{x} - \boldsymbol{\mu})^\top\boldsymbol{\Sigma}^{-1}(\mathbf{x} - \boldsymbol{\mu})\right) $$

The quantity in the exponent is the squared Mahalanobis distance. Contours of constant density are ellipsoids whose axes are the eigenvectors of $\boldsymbol{\Sigma}$, with lengths proportional to $\sqrt{\lambda_i}$.

  • $\boldsymbol{\Sigma} = \sigma^2\mathbf{I}$: spherical contours (isotropic).
  • Diagonal $\boldsymbol{\Sigma}$: axis-aligned ellipses (independent coordinates).
  • Full $\boldsymbol{\Sigma}$: rotated ellipses (correlated coordinates).

Closure properties (memorise these)#

Linear transformation. If $\mathbf{x} \sim \mathcal{N}(\boldsymbol{\mu}, \boldsymbol{\Sigma})$ then $\mathbf{A}\mathbf{x} + \mathbf{b} \sim \mathcal{N}(\mathbf{A}\boldsymbol{\mu} + \mathbf{b}, \mathbf{A}\boldsymbol{\Sigma}\mathbf{A}^\top)$.

Sampling via the reparameterisation. To sample from $\mathcal{N}(\boldsymbol{\mu}, \boldsymbol{\Sigma})$, factor $\boldsymbol{\Sigma} = \mathbf{L}\mathbf{L}^\top$ (Cholesky), draw $\boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})$ and set $\mathbf{x} = \boldsymbol{\mu} + \mathbf{L}\boldsymbol{\epsilon}$. This reparameterisation trick is what lets VAEs backpropagate through random sampling.

Marginals and conditionals. Partition $\mathbf{x} = (\mathbf{x}_a, \mathbf{x}_b)$. The marginal $p(\mathbf{x}_a)$ is Gaussian with mean $\boldsymbol{\mu}_a$ and covariance $\boldsymbol{\Sigma}_{aa}$. The conditional is also Gaussian:

$$ \boldsymbol{\mu}_{a \mid b} = \boldsymbol{\mu}_a + \boldsymbol{\Sigma}_{ab}\boldsymbol{\Sigma}_{bb}^{-1}(\mathbf{x}_b - \boldsymbol{\mu}_b), \qquad \boldsymbol{\Sigma}_{a \mid b} = \boldsymbol{\Sigma}_{aa} - \boldsymbol{\Sigma}_{ab}\boldsymbol{\Sigma}_{bb}^{-1}\boldsymbol{\Sigma}_{ba} $$

These formulas are the engine of Gaussian process regression and the Kalman filter. Notice that the conditional covariance never exceeds the marginal: observing $\mathbf{x}_b$ can only reduce uncertainty about $\mathbf{x}_a$.

Sums of independent Gaussians. $\mathcal{N}(\mu_1, \sigma_1^2) + \mathcal{N}(\mu_2, \sigma_2^2) = \mathcal{N}(\mu_1 + \mu_2, \sigma_1^2 + \sigma_2^2)$. Diffusion models exploit this: adding Gaussian noise step after step still yields a Gaussian with a closed-form variance, so we can jump to any noise level in one step.

python
import numpy as np

mu = np.array([1.0, 2.0])
Sigma = np.array([[2.0, 1.2], [1.2, 1.0]])
L = np.linalg.cholesky(Sigma)
eps = np.random.default_rng(0).normal(size=(100_000, 2))
x = mu + eps @ L.T                                # reparameterised samples
print("sample mean:", x.mean(0).round(3))
print("sample cov:\n", np.cov(x, rowvar=False).round(3))

Where you will meet the Gaussian in this course#

  • Linear regression: Gaussian noise ⇒ MSE loss.
  • Weight initialisation: Xavier/He schemes draw weights from Gaussians with carefully chosen variances.
  • Gaussian Mixture Models: clusters as Gaussians, fitted by EM.
  • VAEs: Gaussian prior and posterior in latent space; KL divergence between Gaussians in closed form.
  • Diffusion models: forward process adds Gaussian noise; the model learns to remove it.
  • Bayesian optimisation: Gaussian processes model unknown functions.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Expectation, Variance, Covariance and Correlation

Summaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.

Beginner⏱ 5 min#036
∑ Mathematics for ML

Bayes' Theorem: Updating Beliefs with Evidence

Bayes' theorem is the mathematical rule for learning from evidence. We derive it, work through the famous medical-test example, and see how it underlies Naive Bayes, Bayesian inference and spam filters.

Beginner⏱ 5 min#035
∑ Mathematics for ML

Random Variables and Probability Distributions

A random variable turns outcomes into numbers. We study discrete and continuous distributions — Bernoulli, categorical, binomial, Poisson, uniform, exponential, Beta — and when ML uses each.

Beginner⏱ 5 min#034