In the last lecture we assigned probabilities to events. In practice we care about numbers: the number of clicks, the pixel intensity, the class label, the waiting time. A random variable attaches numbers to outcomes, and its distribution summarises how likely each value is. Choosing the right distribution is choosing the right assumptions for your model.
Random variables#
A random variable $X$ is a function from the sample space to the real numbers. We distinguish:
- Discrete random variables take countable values; described by a probability mass function (PMF) $p(x) = P(X = x)$, with $\sum_x p(x) = 1$.
- Continuous random variables take values in intervals; described by a probability density function (PDF) $f(x)$ with $P(a \le X \le b) = \int_a^b f(x)\,dx$ and $\int f = 1$.
Both have a cumulative distribution function (CDF) $F(x) = P(X \le x)$.
Discrete distributions#
| Distribution | PMF | Mean | Variance | ML use |
|---|---|---|---|---|
| Bernoulli($p$) | $p^x(1-p)^{1-x}$, $x \in \{0,1\}$ | $p$ | $p(1-p)$ | Binary labels; logistic regression output |
| Categorical($\boldsymbol{\pi}$) | $\pi_k$ for class $k$ | — | — | Multiclass labels; softmax output; next-token prediction |
| Binomial($n, p$) | $\binom{n}{x}p^x(1-p)^{n-x}$ | $np$ | $np(1-p)$ | Number of successes; accuracy on $n$ test items |
| Poisson($\lambda$) | $\frac{\lambda^x e^{-\lambda}}{x!}$ | $\lambda$ | $\lambda$ | Counts per interval: arrivals, clicks, events |
| Geometric($p$) | $(1-p)^{x-1}p$ | $1/p$ | $(1-p)/p^2$ | Trials until first success |
Continuous distributions#
| Distribution | Mean | Variance | ML use | |
|---|---|---|---|---|
| Uniform($a, b$) | $\frac{1}{b - a}$ on $[a, b]$ | $\frac{a+b}{2}$ | $\frac{(b-a)^2}{12}$ | Random initialisation; random search |
| Gaussian($\mu, \sigma^2$) | $\frac{1}{\sqrt{2\pi}\sigma}e^{-\frac{(x-\mu)^2}{2\sigma^2}}$ | $\mu$ | $\sigma^2$ | Noise models; VAEs; diffusion; weight init |
| Exponential($\lambda$) | $\lambda e^{-\lambda x}$, $x \ge 0$ | $1/\lambda$ | $1/\lambda^2$ | Waiting times; survival analysis |
| Beta($\alpha, \beta$) | $\propto x^{\alpha-1}(1-x)^{\beta-1}$ on $[0,1]$ | $\frac{\alpha}{\alpha+\beta}$ | — | Prior over probabilities; A/B testing; bandits |
| Laplace($\mu, b$) | $\frac{1}{2b}e^{-\lvert x-\mu \rvert/b}$ | $\mu$ | $2b^2$ | Robust regression; L1 regularisation prior; differential privacy noise |
The Dirichlet distribution generalises the Beta to probability vectors (e.g. topic proportions in LDA topic models).
Joint, marginal and conditional distributions#
For two variables, the joint distribution $p(x, y)$ describes them together. The marginal is obtained by summing or integrating out the other variable: $p(x) = \sum_y p(x, y)$. The conditional is $p(y \mid x) = p(x, y)/p(x)$. This vocabulary is how we describe every model:
- A discriminative classifier models $p(y \mid \mathbf{x})$.
- A generative model models $p(\mathbf{x}, y)$ or $p(\mathbf{x})$.
Transformations of random variables#
If $Y = g(X)$ with $g$ invertible and differentiable, the density transforms as
The derivative term accounts for how $g$ stretches or compresses space. In many dimensions it becomes the absolute value of a Jacobian determinant — the core of normalising flows.
Sampling in code#
import numpy as np
import matplotlib.pyplot as plt
rng = np.random.default_rng(0)
samples = {
"Bernoulli(0.3)": rng.binomial(1, 0.3, 10_000),
"Poisson(4)": rng.poisson(4, 10_000),
"Gaussian(0,1)": rng.normal(0, 1, 10_000),
"Exponential(1)": rng.exponential(1.0, 10_000),
"Beta(2,5)": rng.beta(2, 5, 10_000),
}
fig, axes = plt.subplots(1, 5, figsize=(18, 3))
for ax, (name, s) in zip(axes, samples.items()):
ax.hist(s, bins=40, density=True, color="#7c3aed", alpha=0.8)
ax.set_title(f"{name}\nmean={s.mean():.2f} var={s.var():.2f}")
plt.tight_layout(); plt.show()Choosing a distribution for your model#
Choosing the output distribution is choosing the loss function:
- Binary target → Bernoulli → binary cross-entropy.
- Multiclass target → Categorical → softmax cross-entropy.
- Real-valued target with Gaussian noise → mean squared error.
- Real-valued target with heavy-tailed noise → Laplace → mean absolute error.
- Count target → Poisson → Poisson loss.
We will prove this correspondence in the maximum likelihood lecture.