∑ Mathematics for ML · Lecture 10 of 25

Random Variables and Probability Distributions

A random variable turns outcomes into numbers. We study discrete and continuous distributions — Bernoulli, categorical, binomial, Poisson, uniform, exponential, Beta — and when ML uses each.

In the last lecture we assigned probabilities to events. In practice we care about numbers: the number of clicks, the pixel intensity, the class label, the waiting time. A random variable attaches numbers to outcomes, and its distribution summarises how likely each value is. Choosing the right distribution is choosing the right assumptions for your model.

Random variables#

A random variable $X$ is a function from the sample space to the real numbers. We distinguish:

  • Discrete random variables take countable values; described by a probability mass function (PMF) $p(x) = P(X = x)$, with $\sum_x p(x) = 1$.
  • Continuous random variables take values in intervals; described by a probability density function (PDF) $f(x)$ with $P(a \le X \le b) = \int_a^b f(x)\,dx$ and $\int f = 1$.

Both have a cumulative distribution function (CDF) $F(x) = P(X \le x)$.

Discrete distributions#

DistributionPMFMeanVarianceML use
Bernoulli($p$)$p^x(1-p)^{1-x}$, $x \in \{0,1\}$$p$$p(1-p)$Binary labels; logistic regression output
Categorical($\boldsymbol{\pi}$)$\pi_k$ for class $k$——Multiclass labels; softmax output; next-token prediction
Binomial($n, p$)$\binom{n}{x}p^x(1-p)^{n-x}$$np$$np(1-p)$Number of successes; accuracy on $n$ test items
Poisson($\lambda$)$\frac{\lambda^x e^{-\lambda}}{x!}$$\lambda$$\lambda$Counts per interval: arrivals, clicks, events
Geometric($p$)$(1-p)^{x-1}p$$1/p$$(1-p)/p^2$Trials until first success

Continuous distributions#

DistributionPDFMeanVarianceML use
Uniform($a, b$)$\frac{1}{b - a}$ on $[a, b]$$\frac{a+b}{2}$$\frac{(b-a)^2}{12}$Random initialisation; random search
Gaussian($\mu, \sigma^2$)$\frac{1}{\sqrt{2\pi}\sigma}e^{-\frac{(x-\mu)^2}{2\sigma^2}}$$\mu$$\sigma^2$Noise models; VAEs; diffusion; weight init
Exponential($\lambda$)$\lambda e^{-\lambda x}$, $x \ge 0$$1/\lambda$$1/\lambda^2$Waiting times; survival analysis
Beta($\alpha, \beta$)$\propto x^{\alpha-1}(1-x)^{\beta-1}$ on $[0,1]$$\frac{\alpha}{\alpha+\beta}$—Prior over probabilities; A/B testing; bandits
Laplace($\mu, b$)$\frac{1}{2b}e^{-\lvert x-\mu \rvert/b}$$\mu$$2b^2$Robust regression; L1 regularisation prior; differential privacy noise

The Dirichlet distribution generalises the Beta to probability vectors (e.g. topic proportions in LDA topic models).

Joint, marginal and conditional distributions#

For two variables, the joint distribution $p(x, y)$ describes them together. The marginal is obtained by summing or integrating out the other variable: $p(x) = \sum_y p(x, y)$. The conditional is $p(y \mid x) = p(x, y)/p(x)$. This vocabulary is how we describe every model:

  • A discriminative classifier models $p(y \mid \mathbf{x})$.
  • A generative model models $p(\mathbf{x}, y)$ or $p(\mathbf{x})$.

Transformations of random variables#

If $Y = g(X)$ with $g$ invertible and differentiable, the density transforms as

$$ f_Y(y) = f_X\big(g^{-1}(y)\big)\left|\frac{d}{dy}g^{-1}(y)\right| $$

The derivative term accounts for how $g$ stretches or compresses space. In many dimensions it becomes the absolute value of a Jacobian determinant — the core of normalising flows.

Sampling in code#

python
import numpy as np
import matplotlib.pyplot as plt

rng = np.random.default_rng(0)
samples = {
    "Bernoulli(0.3)": rng.binomial(1, 0.3, 10_000),
    "Poisson(4)":     rng.poisson(4, 10_000),
    "Gaussian(0,1)":  rng.normal(0, 1, 10_000),
    "Exponential(1)": rng.exponential(1.0, 10_000),
    "Beta(2,5)":      rng.beta(2, 5, 10_000),
}
fig, axes = plt.subplots(1, 5, figsize=(18, 3))
for ax, (name, s) in zip(axes, samples.items()):
    ax.hist(s, bins=40, density=True, color="#7c3aed", alpha=0.8)
    ax.set_title(f"{name}\nmean={s.mean():.2f} var={s.var():.2f}")
plt.tight_layout(); plt.show()

Choosing a distribution for your model#

Choosing the output distribution is choosing the loss function:

  • Binary target → Bernoulli → binary cross-entropy.
  • Multiclass target → Categorical → softmax cross-entropy.
  • Real-valued target with Gaussian noise → mean squared error.
  • Real-valued target with heavy-tailed noise → Laplace → mean absolute error.
  • Count target → Poisson → Poisson loss.

We will prove this correspondence in the maximum likelihood lecture.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Probability Fundamentals: Sample Spaces, Axioms and Conditional Probability

Machine learning is reasoning under uncertainty. We build probability from Kolmogorov's axioms, then master conditional probability, the product and sum rules, and independence.

Beginner⏱ 5 min#033
∑ Mathematics for ML

Bayes' Theorem: Updating Beliefs with Evidence

Bayes' theorem is the mathematical rule for learning from evidence. We derive it, work through the famous medical-test example, and see how it underlies Naive Bayes, Bayesian inference and spam filters.

Beginner⏱ 5 min#035
∑ Mathematics for ML

Expectation, Variance, Covariance and Correlation

Summaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.

Beginner⏱ 5 min#036