Machine learning constantly asks "how big?" and "how similar?". How large are the weights (regularisation)? How far is this point from that cluster centre (k-means)? How similar are two sentences (semantic search)? The mathematical tools for these questions are norms, inner products and distances.
Norms#
A norm $\|\cdot\|$ assigns a length to each vector and satisfies:
- $\|\mathbf{x}\| \ge 0$, with equality only for $\mathbf{x} = \mathbf{0}$;
- $\|c\mathbf{x}\| = |c|\,\|\mathbf{x}\|$;
- Triangle inequality: $\|\mathbf{x} + \mathbf{y}\| \le \|\mathbf{x}\| + \|\mathbf{y}\|$.
The family of $L_p$ norms:
| Norm | Formula | Unit ball shape (2-D) | ML use |
|---|---|---|---|
| $L_1$ (Manhattan) | $\sum_i \lvert x_i \rvert$ | Diamond | Lasso — produces sparse weights |
| $L_2$ (Euclidean) | $\sqrt{\sum_i x_i^2}$ | Circle | Ridge / weight decay; distances |
| $L_\infty$ (max) | $\max_i \lvert x_i \rvert$ | Square | Adversarial perturbation budgets |
| $L_0$ ("norm") | number of non-zeros | — | Sparsity (not a true norm) |
For matrices, the Frobenius norm $\|\mathbf{A}\|_F = \sqrt{\sum_{ij} A_{ij}^2}$ treats the matrix as a long vector; the spectral norm $\|\mathbf{A}\|_2 = \sigma_{\max}$ is the maximum stretch factor.
Inner products#
The standard inner (dot) product on $\mathbb{R}^d$ is
It links algebra with geometry: the sign tells whether vectors point the same way; zero means orthogonal. The Cauchy–Schwarz inequality $|\mathbf{x}^\top\mathbf{y}| \le \|\mathbf{x}\|\,\|\mathbf{y}\|$ guarantees $|\cos\theta| \le 1$.
Generalised inner products $\langle \mathbf{x}, \mathbf{y}\rangle_{\mathbf{M}} = \mathbf{x}^\top\mathbf{M}\mathbf{y}$ with positive definite $\mathbf{M}$ define different geometries — the basis of the Mahalanobis distance and kernel methods.
Cosine similarity#
Cosine similarity ignores length and compares only direction. It is the standard similarity for text embeddings and TF-IDF vectors, where document length should not dominate. If vectors are L2-normalised, cosine similarity equals the dot product, and squared Euclidean distance equals $2 - 2\cos\theta$ — so nearest-neighbour search by either criterion gives the same ranking.
Orthogonal projection#
The projection of $\mathbf{y}$ onto the line spanned by $\mathbf{a}$ is
The projection onto the column space of a matrix $\mathbf{A}$ (full column rank) is $\mathbf{P}\mathbf{y}$ with
Linear regression is projection. The least-squares prediction $\hat{\mathbf{y}} = \mathbf{X}\hat{\mathbf{w}}$ is the orthogonal projection of the target vector $\mathbf{y}$ onto the column space of the feature matrix $\mathbf{X}$. The residual $\mathbf{y} - \hat{\mathbf{y}}$ is orthogonal to every feature column — that orthogonality condition is the normal equation $\mathbf{X}^\top(\mathbf{y} - \mathbf{X}\mathbf{w}) = \mathbf{0}$.
Distance metrics#
A metric $d(\mathbf{x}, \mathbf{y})$ is non-negative, zero only for identical points, symmetric, and obeys the triangle inequality. Common choices:
- Euclidean: $\|\mathbf{x} - \mathbf{y}\|_2$ — default for continuous features (scale them first!).
- Manhattan: $\|\mathbf{x} - \mathbf{y}\|_1$ — robust to outliers in single coordinates.
- Mahalanobis: $\sqrt{(\mathbf{x} - \mathbf{y})^\top\boldsymbol{\Sigma}^{-1}(\mathbf{x} - \mathbf{y})}$ — accounts for feature correlations and scales; used in anomaly detection.
- Hamming: number of differing positions — for binary codes and strings.
- Cosine distance $1 - \cos$ — not a true metric, but widely used for embeddings.
import numpy as np
x = np.array([3.0, 4.0]); y = np.array([4.0, 3.0])
print("L1:", np.linalg.norm(x, 1), "L2:", np.linalg.norm(x), "Linf:", np.linalg.norm(x, np.inf))
cos = x @ y / (np.linalg.norm(x) * np.linalg.norm(y))
print("cosine:", round(cos, 3))
a = np.array([1.0, 1.0])
proj = (a @ x) / (a @ a) * a
print("projection of x on a:", proj, " residual . a =", (x - proj) @ a) # ~0