∑ Mathematics for ML · Lecture 12 of 25

Expectation, Variance, Covariance and Correlation

Summaries of distributions drive everything from loss functions to PCA. We define expectation, variance, covariance and correlation, prove linearity of expectation, and study the covariance matrix.

A full probability distribution can be complicated. Often we summarise it with a few numbers: its centre, its spread and how variables move together. These summaries — expectation, variance, covariance — are not just descriptive statistics. Loss functions are expectations, the bias–variance trade-off is about variance, and PCA is about covariance.

Expectation#

The expected value (mean) of a random variable is its probability-weighted average:

$$ \mathbb{E}[X] = \sum_x x\,p(x) \quad \text{(discrete)}, \qquad \mathbb{E}[X] = \int x\,f(x)\,dx \quad \text{(continuous)} $$

For a function $g$: $\mathbb{E}[g(X)] = \sum_x g(x)p(x)$ — the law of the unconscious statistician.

Linearity of expectation#

$$ \mathbb{E}[aX + bY + c] = a\,\mathbb{E}[X] + b\,\mathbb{E}[Y] + c $$

This holds always — even when $X$ and $Y$ are dependent. It is one of the most powerful tools in probability.

In ML, the risk of a model is an expectation: $R(f) = \mathbb{E}_{(\mathbf{x}, y)}[\ell(f(\mathbf{x}), y)]$. Training minimises the empirical average over the training set as an estimate of this expectation.

Variance and standard deviation#

$$ \text{Var}(X) = \mathbb{E}\big[(X - \mathbb{E}[X])^2\big] = \mathbb{E}[X^2] - \mathbb{E}[X]^2 $$

The standard deviation $\sigma = \sqrt{\text{Var}(X)}$ has the same units as $X$. Properties:

  • $\text{Var}(aX + b) = a^2\,\text{Var}(X)$ — shifting does not change spread.
  • For independent $X, Y$: $\text{Var}(X + Y) = \text{Var}(X) + \text{Var}(Y)$.
  • In general: $\text{Var}(X + Y) = \text{Var}(X) + \text{Var}(Y) + 2\,\text{Cov}(X, Y)$.

Why averaging reduces noise#

For $n$ independent samples each with variance $\sigma^2$, the sample mean $\bar{X}$ has

$$ \text{Var}(\bar{X}) = \frac{\sigma^2}{n}, \qquad \text{SD}(\bar{X}) = \frac{\sigma}{\sqrt{n}} $$

This $1/\sqrt{n}$ law explains why larger mini-batches give less noisy gradient estimates (with diminishing returns — quadrupling the batch only halves the noise), why ensembles of independent models reduce variance, and why test-set accuracy estimates improve slowly with test-set size.

Covariance and correlation#

Covariance measures how two variables vary together:

$$ \text{Cov}(X, Y) = \mathbb{E}\big[(X - \mu_X)(Y - \mu_Y)\big] = \mathbb{E}[XY] - \mu_X\mu_Y $$

Positive covariance: they tend to be large together. Its magnitude depends on units, so we normalise to get the Pearson correlation:

$$ \rho(X, Y) = \frac{\text{Cov}(X, Y)}{\sigma_X\,\sigma_Y} \in [-1, 1] $$

The covariance matrix#

For a random vector $\mathbf{x} \in \mathbb{R}^d$ with mean $\boldsymbol{\mu}$:

$$ \boldsymbol{\Sigma} = \mathbb{E}\big[(\mathbf{x} - \boldsymbol{\mu})(\mathbf{x} - \boldsymbol{\mu})^\top\big], \qquad \Sigma_{ij} = \text{Cov}(x_i, x_j) $$

Properties:

  • Symmetric and positive semi-definite: $\mathbf{a}^\top\boldsymbol{\Sigma}\mathbf{a} = \text{Var}(\mathbf{a}^\top\mathbf{x}) \ge 0$.
  • The variance of any linear projection $\mathbf{a}^\top\mathbf{x}$ is $\mathbf{a}^\top\boldsymbol{\Sigma}\mathbf{a}$ — so the direction of maximum variance is the top eigenvector of $\boldsymbol{\Sigma}$. That is PCA.
  • Under a linear transform $\mathbf{y} = \mathbf{A}\mathbf{x}$: $\boldsymbol{\Sigma}_y = \mathbf{A}\boldsymbol{\Sigma}_x\mathbf{A}^\top$.

The sample covariance from a centred data matrix $\mathbf{X}_c \in \mathbb{R}^{n \times d}$ is $\hat{\boldsymbol{\Sigma}} = \frac{1}{n-1}\mathbf{X}_c^\top\mathbf{X}_c$. The $n - 1$ (Bessel's correction) makes it unbiased.

python
import numpy as np

rng = np.random.default_rng(0)
n = 5000
study_hours = rng.normal(10, 3, n)
exam_score  = 40 + 4 * study_hours + rng.normal(0, 8, n)
sleep_hours = rng.normal(7, 1, n)
X = np.column_stack([study_hours, exam_score, sleep_hours])

print("means:", X.mean(axis=0).round(2))
print("covariance:\n", np.cov(X, rowvar=False).round(2))
print("correlation:\n", np.corrcoef(X, rowvar=False).round(2))

x = rng.normal(size=n); y = x**2
print("corr(x, x^2):", np.corrcoef(x, y)[0, 1].round(3))   # ~0 despite dependence

Higher moments#

  • Skewness $\mathbb{E}[(X - \mu)^3]/\sigma^3$ measures asymmetry (incomes are right-skewed).
  • Kurtosis $\mathbb{E}[(X - \mu)^4]/\sigma^4$ measures tail heaviness. Heavy-tailed data (financial returns, network traffic) produce outliers that break methods assuming Gaussian noise.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Bayes' Theorem: Updating Beliefs with Evidence

Bayes' theorem is the mathematical rule for learning from evidence. We derive it, work through the famous medical-test example, and see how it underlies Naive Bayes, Bayesian inference and spam filters.

Beginner⏱ 5 min#035
∑ Mathematics for ML

The Gaussian Distribution: Why It Is Everywhere

The bell curve appears in noise models, weight initialisation, VAEs and diffusion models. We study univariate and multivariate Gaussians, the central limit theorem, and the closure properties that make Gaussians so convenient.

Intermediate⏱ 5 min#037
∑ Mathematics for ML

Random Variables and Probability Distributions

A random variable turns outcomes into numbers. We study discrete and continuous distributions — Bernoulli, categorical, binomial, Poisson, uniform, exponential, Beta — and when ML uses each.

Beginner⏱ 5 min#034