∑ Mathematics for ML · Lecture 6 of 25

Norms, Inner Products, Projections and Distances

How long is a vector, and how similar are two vectors? We study L1, L2 and L∞ norms, dot products, cosine similarity, orthogonal projections and distance metrics used across ML.

Machine learning constantly asks "how big?" and "how similar?". How large are the weights (regularisation)? How far is this point from that cluster centre (k-means)? How similar are two sentences (semantic search)? The mathematical tools for these questions are norms, inner products and distances.

Norms#

A norm $\|\cdot\|$ assigns a length to each vector and satisfies:

  1. $\|\mathbf{x}\| \ge 0$, with equality only for $\mathbf{x} = \mathbf{0}$;
  2. $\|c\mathbf{x}\| = |c|\,\|\mathbf{x}\|$;
  3. Triangle inequality: $\|\mathbf{x} + \mathbf{y}\| \le \|\mathbf{x}\| + \|\mathbf{y}\|$.

The family of $L_p$ norms:

$$ \|\mathbf{x}\|_p = \left( \sum_{i=1}^d |x_i|^p \right)^{1/p} $$
NormFormulaUnit ball shape (2-D)ML use
$L_1$ (Manhattan)$\sum_i \lvert x_i \rvert$DiamondLasso — produces sparse weights
$L_2$ (Euclidean)$\sqrt{\sum_i x_i^2}$CircleRidge / weight decay; distances
$L_\infty$ (max)$\max_i \lvert x_i \rvert$SquareAdversarial perturbation budgets
$L_0$ ("norm")number of non-zeros—Sparsity (not a true norm)

For matrices, the Frobenius norm $\|\mathbf{A}\|_F = \sqrt{\sum_{ij} A_{ij}^2}$ treats the matrix as a long vector; the spectral norm $\|\mathbf{A}\|_2 = \sigma_{\max}$ is the maximum stretch factor.

Inner products#

The standard inner (dot) product on $\mathbb{R}^d$ is

$$ \langle \mathbf{x}, \mathbf{y} \rangle = \mathbf{x}^\top\mathbf{y} = \sum_i x_i y_i = \|\mathbf{x}\|\,\|\mathbf{y}\|\cos\theta $$

It links algebra with geometry: the sign tells whether vectors point the same way; zero means orthogonal. The Cauchy–Schwarz inequality $|\mathbf{x}^\top\mathbf{y}| \le \|\mathbf{x}\|\,\|\mathbf{y}\|$ guarantees $|\cos\theta| \le 1$.

Generalised inner products $\langle \mathbf{x}, \mathbf{y}\rangle_{\mathbf{M}} = \mathbf{x}^\top\mathbf{M}\mathbf{y}$ with positive definite $\mathbf{M}$ define different geometries — the basis of the Mahalanobis distance and kernel methods.

Cosine similarity#

$$ \cos(\mathbf{x}, \mathbf{y}) = \frac{\mathbf{x}^\top\mathbf{y}}{\|\mathbf{x}\|\,\|\mathbf{y}\|} $$

Cosine similarity ignores length and compares only direction. It is the standard similarity for text embeddings and TF-IDF vectors, where document length should not dominate. If vectors are L2-normalised, cosine similarity equals the dot product, and squared Euclidean distance equals $2 - 2\cos\theta$ — so nearest-neighbour search by either criterion gives the same ranking.

Orthogonal projection#

The projection of $\mathbf{y}$ onto the line spanned by $\mathbf{a}$ is

$$ \text{proj}_{\mathbf{a}}(\mathbf{y}) = \frac{\mathbf{a}^\top\mathbf{y}}{\mathbf{a}^\top\mathbf{a}}\,\mathbf{a} $$

The projection onto the column space of a matrix $\mathbf{A}$ (full column rank) is $\mathbf{P}\mathbf{y}$ with

$$ \mathbf{P} = \mathbf{A}(\mathbf{A}^\top\mathbf{A})^{-1}\mathbf{A}^\top $$

Linear regression is projection. The least-squares prediction $\hat{\mathbf{y}} = \mathbf{X}\hat{\mathbf{w}}$ is the orthogonal projection of the target vector $\mathbf{y}$ onto the column space of the feature matrix $\mathbf{X}$. The residual $\mathbf{y} - \hat{\mathbf{y}}$ is orthogonal to every feature column — that orthogonality condition is the normal equation $\mathbf{X}^\top(\mathbf{y} - \mathbf{X}\mathbf{w}) = \mathbf{0}$.

Distance metrics#

A metric $d(\mathbf{x}, \mathbf{y})$ is non-negative, zero only for identical points, symmetric, and obeys the triangle inequality. Common choices:

  • Euclidean: $\|\mathbf{x} - \mathbf{y}\|_2$ — default for continuous features (scale them first!).
  • Manhattan: $\|\mathbf{x} - \mathbf{y}\|_1$ — robust to outliers in single coordinates.
  • Mahalanobis: $\sqrt{(\mathbf{x} - \mathbf{y})^\top\boldsymbol{\Sigma}^{-1}(\mathbf{x} - \mathbf{y})}$ — accounts for feature correlations and scales; used in anomaly detection.
  • Hamming: number of differing positions — for binary codes and strings.
  • Cosine distance $1 - \cos$ — not a true metric, but widely used for embeddings.
python
import numpy as np

x = np.array([3.0, 4.0]); y = np.array([4.0, 3.0])
print("L1:", np.linalg.norm(x, 1), "L2:", np.linalg.norm(x), "Linf:", np.linalg.norm(x, np.inf))
cos = x @ y / (np.linalg.norm(x) * np.linalg.norm(y))
print("cosine:", round(cos, 3))

a = np.array([1.0, 1.0])
proj = (a @ x) / (a @ a) * a
print("projection of x on a:", proj, " residual . a =", (x - proj) @ a)   # ~0
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Singular Value Decomposition: The Swiss Army Knife of Linear Algebra

Every matrix — any shape, any rank — factors as rotation, scaling, rotation. We derive the SVD, prove the Eckart–Young low-rank theorem, and apply it to compression, PCA, recommenders and least squares.

Intermediate⏱ 4 min#029
∑ Mathematics for ML

Eigenvalues and Eigenvectors: The Natural Axes of a Transformation

Eigenvectors are directions a matrix merely stretches. We derive the characteristic equation, diagonalisation and the spectral theorem, and connect them to PCA, PageRank, Markov chains and training stability.

Intermediate⏱ 5 min#028
∑ Mathematics for ML

Matrices and Linear Transformations

A matrix is not just a table of numbers — it is a function that transforms space. We cover matrix multiplication four ways, rank, inverses, determinants and the fundamental subspaces.

Beginner⏱ 5 min#027