A full probability distribution can be complicated. Often we summarise it with a few numbers: its centre, its spread and how variables move together. These summaries — expectation, variance, covariance — are not just descriptive statistics. Loss functions are expectations, the bias–variance trade-off is about variance, and PCA is about covariance.
Expectation#
The expected value (mean) of a random variable is its probability-weighted average:
For a function $g$: $\mathbb{E}[g(X)] = \sum_x g(x)p(x)$ — the law of the unconscious statistician.
Linearity of expectation#
This holds always — even when $X$ and $Y$ are dependent. It is one of the most powerful tools in probability.
In ML, the risk of a model is an expectation: $R(f) = \mathbb{E}_{(\mathbf{x}, y)}[\ell(f(\mathbf{x}), y)]$. Training minimises the empirical average over the training set as an estimate of this expectation.
Variance and standard deviation#
The standard deviation $\sigma = \sqrt{\text{Var}(X)}$ has the same units as $X$. Properties:
- $\text{Var}(aX + b) = a^2\,\text{Var}(X)$ — shifting does not change spread.
- For independent $X, Y$: $\text{Var}(X + Y) = \text{Var}(X) + \text{Var}(Y)$.
- In general: $\text{Var}(X + Y) = \text{Var}(X) + \text{Var}(Y) + 2\,\text{Cov}(X, Y)$.
Why averaging reduces noise#
For $n$ independent samples each with variance $\sigma^2$, the sample mean $\bar{X}$ has
This $1/\sqrt{n}$ law explains why larger mini-batches give less noisy gradient estimates (with diminishing returns — quadrupling the batch only halves the noise), why ensembles of independent models reduce variance, and why test-set accuracy estimates improve slowly with test-set size.
Covariance and correlation#
Covariance measures how two variables vary together:
Positive covariance: they tend to be large together. Its magnitude depends on units, so we normalise to get the Pearson correlation:
The covariance matrix#
For a random vector $\mathbf{x} \in \mathbb{R}^d$ with mean $\boldsymbol{\mu}$:
Properties:
- Symmetric and positive semi-definite: $\mathbf{a}^\top\boldsymbol{\Sigma}\mathbf{a} = \text{Var}(\mathbf{a}^\top\mathbf{x}) \ge 0$.
- The variance of any linear projection $\mathbf{a}^\top\mathbf{x}$ is $\mathbf{a}^\top\boldsymbol{\Sigma}\mathbf{a}$ — so the direction of maximum variance is the top eigenvector of $\boldsymbol{\Sigma}$. That is PCA.
- Under a linear transform $\mathbf{y} = \mathbf{A}\mathbf{x}$: $\boldsymbol{\Sigma}_y = \mathbf{A}\boldsymbol{\Sigma}_x\mathbf{A}^\top$.
The sample covariance from a centred data matrix $\mathbf{X}_c \in \mathbb{R}^{n \times d}$ is $\hat{\boldsymbol{\Sigma}} = \frac{1}{n-1}\mathbf{X}_c^\top\mathbf{X}_c$. The $n - 1$ (Bessel's correction) makes it unbiased.
import numpy as np
rng = np.random.default_rng(0)
n = 5000
study_hours = rng.normal(10, 3, n)
exam_score = 40 + 4 * study_hours + rng.normal(0, 8, n)
sleep_hours = rng.normal(7, 1, n)
X = np.column_stack([study_hours, exam_score, sleep_hours])
print("means:", X.mean(axis=0).round(2))
print("covariance:\n", np.cov(X, rowvar=False).round(2))
print("correlation:\n", np.corrcoef(X, rowvar=False).round(2))
x = rng.normal(size=n); y = x**2
print("corr(x, x^2):", np.corrcoef(x, y)[0, 1].round(3)) # ~0 despite dependenceHigher moments#
- Skewness $\mathbb{E}[(X - \mu)^3]/\sigma^3$ measures asymmetry (incomes are right-skewed).
- Kurtosis $\mathbb{E}[(X - \mu)^4]/\sigma^4$ measures tail heaviness. Heavy-tailed data (financial returns, network traffic) produce outliers that break methods assuming Gaussian noise.