Last lecture we treated vectors as data. Today we study the functions that act on them. A matrix is the concrete representation of a linear transformation, and every layer of a neural network begins with one. Understanding matrices as transformations — rotating, stretching, projecting space — is the single most useful intuition in linear algebra.
Linear transformations#
A function $T: \mathbb{R}^n \to \mathbb{R}^m$ is linear if
Every such $T$ can be written as $T(\mathbf{x}) = \mathbf{A}\mathbf{x}$ for a unique $m \times n$ matrix $\mathbf{A}$. The columns of $\mathbf{A}$ are simply where the standard basis vectors land: column $j$ is $T(\mathbf{e}_j)$.
Four ways to see matrix multiplication#
For $\mathbf{C} = \mathbf{A}\mathbf{B}$ with $\mathbf{A} \in \mathbb{R}^{m \times k}$, $\mathbf{B} \in \mathbb{R}^{k \times n}$:
- Entry view: $C_{ij} = \sum_{l} A_{il}B_{lj}$ — dot product of row $i$ of $\mathbf{A}$ and column $j$ of $\mathbf{B}$.
- Column view: each column of $\mathbf{C}$ is $\mathbf{A}$ times the corresponding column of $\mathbf{B}$ — a combination of $\mathbf{A}$'s columns.
- Row view: each row of $\mathbf{C}$ is a combination of $\mathbf{B}$'s rows.
- Outer-product view: $\mathbf{C} = \sum_{l} \mathbf{a}_{:l}\,\mathbf{b}_{l:}$ — a sum of rank-one matrices.
The outer-product view is the key to understanding low-rank approximations, SVD and LoRA fine-tuning.
Composition: applying $\mathbf{B}$ then $\mathbf{A}$ equals applying $\mathbf{AB}$. This is why a stack of linear layers without non-linearities collapses into a single linear layer — the reason neural networks need activation functions.
Key properties#
- Associative: $(\mathbf{AB})\mathbf{C} = \mathbf{A}(\mathbf{BC})$.
- Distributive: $\mathbf{A}(\mathbf{B} + \mathbf{C}) = \mathbf{AB} + \mathbf{AC}$.
- Not commutative: generally $\mathbf{AB} \ne \mathbf{BA}$.
- Transpose: $(\mathbf{AB})^\top = \mathbf{B}^\top\mathbf{A}^\top$.
Rank and the four fundamental subspaces#
The rank of $\mathbf{A}$ is the dimension of its column space — the set of all outputs $\mathbf{A}\mathbf{x}$. Gilbert Strang's four fundamental subspaces of an $m \times n$ matrix of rank $r$:
| Subspace | Lives in | Dimension |
|---|---|---|
| Column space $C(\mathbf{A})$ | $\mathbb{R}^m$ | $r$ |
| Null space $N(\mathbf{A})$: $\mathbf{Ax} = \mathbf{0}$ | $\mathbb{R}^n$ | $n - r$ |
| Row space $C(\mathbf{A}^\top)$ | $\mathbb{R}^n$ | $r$ |
| Left null space $N(\mathbf{A}^\top)$ | $\mathbb{R}^m$ | $m - r$ |
The rank–nullity theorem: $\text{rank} + \text{nullity} = n$. The row space and null space are orthogonal complements.
Inverses and determinants#
A square matrix $\mathbf{A}$ is invertible if there is $\mathbf{A}^{-1}$ with $\mathbf{A}\mathbf{A}^{-1} = \mathbf{I}$. Equivalent conditions: full rank, trivial null space, non-zero determinant, no zero eigenvalue.
The determinant measures how the transformation scales volume: $|\det \mathbf{A}|$ is the volume of the image of the unit cube; the sign says whether orientation flips. $\det(\mathbf{AB}) = \det\mathbf{A}\,\det\mathbf{B}$. Determinants appear in the Gaussian density and in normalising flows, where we must track how a transformation changes probability volume.
Special matrices you will meet constantly#
- Identity $\mathbf{I}$ and diagonal matrices — scale each axis independently.
- Symmetric $\mathbf{A} = \mathbf{A}^\top$ — covariance matrices, Hessians, kernel matrices.
- Orthogonal $\mathbf{Q}^\top\mathbf{Q} = \mathbf{I}$ — rotations and reflections; preserve lengths and angles.
- Positive (semi-)definite $\mathbf{x}^\top\mathbf{A}\mathbf{x} \ge 0$ — covariance matrices; convex quadratic forms.
- Sparse matrices — graphs, text term–document matrices.
import numpy as np
theta = np.pi / 4
R = np.array([[np.cos(theta), -np.sin(theta)],
[np.sin(theta), np.cos(theta)]])
x = np.array([1.0, 0.0])
print(R @ x) # rotated 45 degrees
print(np.allclose(R.T @ R, np.eye(2))) # orthogonal
print(np.linalg.det(R)) # 1.0 -> preserves area
A = np.array([[4., 1.], [2., 3.]])
b = np.array([1., 2.])
print(np.linalg.solve(A, b)) # preferred over inv(A) @ bMatrices in neural networks#
A fully connected layer computes $\mathbf{h} = \phi(\mathbf{W}\mathbf{x} + \mathbf{b})$. For a mini-batch stored as rows of $\mathbf{X} \in \mathbb{R}^{B \times d}$, it is $\mathbf{H} = \phi(\mathbf{X}\mathbf{W}^\top + \mathbf{b})$ — one matrix multiplication for the whole batch. GPUs are, above all, machines for fast matrix multiplication, and that is why deep learning became practical.