In machine learning, everything becomes a vector. A house becomes (area, rooms, age). A grayscale image of $28 \times 28$ pixels becomes a vector in $\mathbb{R}^{784}$. A word becomes a 768-dimensional embedding. Today we learn the algebra of these objects rigorously, because the geometry of vector spaces is the geometry of data.
Vectors#
A vector $\mathbf{x} \in \mathbb{R}^d$ is an ordered list of $d$ real numbers:
We can view it in two ways: as a point in $d$-dimensional space, or as an arrow (a displacement) from the origin. Both views are useful.
Two basic operations:
- Addition: $(\mathbf{x} + \mathbf{y})_i = x_i + y_i$ — place arrows head to tail.
- Scalar multiplication: $(c\mathbf{x})_i = c\,x_i$ — stretch or flip the arrow.
Vector spaces#
A vector space over $\mathbb{R}$ is a set $V$ with addition and scalar multiplication satisfying eight axioms (associativity, commutativity, zero vector, additive inverses, distributivity, etc.). The essential point: closed under linear combinations. $\mathbb{R}^d$ is the standard example, but polynomials, matrices and functions also form vector spaces — which is why techniques like kernel methods can treat functions as vectors.
A subspace is a subset that is itself a vector space: it contains the zero vector and is closed under addition and scaling. In $\mathbb{R}^3$, subspaces are the origin, lines through the origin, planes through the origin, and $\mathbb{R}^3$ itself.
Linear combinations and span#
A linear combination of vectors $\mathbf{v}_1, \dots, \mathbf{v}_k$ is
The span of the vectors is the set of all such combinations. Two non-parallel vectors in $\mathbb{R}^3$ span a plane.
Linear independence#
Vectors are linearly independent if the only solution to $c_1\mathbf{v}_1 + \dots + c_k\mathbf{v}_k = \mathbf{0}$ is $c_1 = \dots = c_k = 0$. Otherwise one of them can be written as a combination of the others — it is redundant.
In data terms: if one feature equals twice another plus a third (e.g. "total price" = "unit price × quantity" already stored), the feature vectors are dependent. This multicollinearity makes linear regression coefficients unstable, a problem we fix with regularisation.
Basis and dimension#
A basis of a space is a set of linearly independent vectors that span it. Every vector in the space can be written uniquely as a combination of basis vectors; the coefficients are its coordinates in that basis. The number of vectors in any basis is the dimension.
The standard basis of $\mathbb{R}^3$ is $\mathbf{e}_1 = (1,0,0)$, $\mathbf{e}_2 = (0,1,0)$, $\mathbf{e}_3 = (0,0,1)$. But other bases can be far more useful: PCA finds a basis aligned with the directions of greatest variance; the Fourier basis represents signals as frequencies.
import numpy as np
V = np.array([[1, 2, 3],
[2, 4, 6], # = 2 * first row -> dependent
[0, 1, 1]]).T # columns are our vectors
print("rank:", np.linalg.matrix_rank(V)) # 2 -> the three vectors span only a plane
# Coordinates of x in a new basis B (columns): solve B c = x
B = np.array([[1, 1], [1, -1]]).T
x = np.array([3, 1])
c = np.linalg.solve(B, x)
print("coordinates:", c) # [2, 1] because 2*(1,1) + 1*(1,-1) = (3,1)The geometry of high dimensions#
Your intuition, built in 2-D and 3-D, can mislead you in high dimensions:
- Most of the volume of a high-dimensional ball lies near its surface.
- Two random vectors in high dimensions are almost always nearly orthogonal.
- Distances between random points concentrate — the nearest and farthest neighbours become almost equally far.
These facts are the curse of dimensionality, which we will study in its own lecture. They also explain why embeddings can store so much information: high-dimensional spaces have room for many nearly independent directions.
Vectors as meaning: embeddings#
Word embeddings famously capture relationships as vector arithmetic, for example $\mathbf{v}_{\text{king}} - \mathbf{v}_{\text{man}} + \mathbf{v}_{\text{woman}} \approx \mathbf{v}_{\text{queen}}$. Directions in embedding space can correspond to concepts such as gender, tense or sentiment. This is linear algebra acting as a model of meaning — and a reason that biased training data produces biased directions, which we must detect and address.