∑ Mathematics for ML · Lecture 3 of 25

Matrices and Linear Transformations

A matrix is not just a table of numbers — it is a function that transforms space. We cover matrix multiplication four ways, rank, inverses, determinants and the fundamental subspaces.

Last lecture we treated vectors as data. Today we study the functions that act on them. A matrix is the concrete representation of a linear transformation, and every layer of a neural network begins with one. Understanding matrices as transformations — rotating, stretching, projecting space — is the single most useful intuition in linear algebra.

Linear transformations#

A function $T: \mathbb{R}^n \to \mathbb{R}^m$ is linear if

$$ T(a\mathbf{x} + b\mathbf{y}) = aT(\mathbf{x}) + bT(\mathbf{y}) $$

Every such $T$ can be written as $T(\mathbf{x}) = \mathbf{A}\mathbf{x}$ for a unique $m \times n$ matrix $\mathbf{A}$. The columns of $\mathbf{A}$ are simply where the standard basis vectors land: column $j$ is $T(\mathbf{e}_j)$.

Four ways to see matrix multiplication#

For $\mathbf{C} = \mathbf{A}\mathbf{B}$ with $\mathbf{A} \in \mathbb{R}^{m \times k}$, $\mathbf{B} \in \mathbb{R}^{k \times n}$:

  1. Entry view: $C_{ij} = \sum_{l} A_{il}B_{lj}$ — dot product of row $i$ of $\mathbf{A}$ and column $j$ of $\mathbf{B}$.
  2. Column view: each column of $\mathbf{C}$ is $\mathbf{A}$ times the corresponding column of $\mathbf{B}$ — a combination of $\mathbf{A}$'s columns.
  3. Row view: each row of $\mathbf{C}$ is a combination of $\mathbf{B}$'s rows.
  4. Outer-product view: $\mathbf{C} = \sum_{l} \mathbf{a}_{:l}\,\mathbf{b}_{l:}$ — a sum of rank-one matrices.

The outer-product view is the key to understanding low-rank approximations, SVD and LoRA fine-tuning.

Composition: applying $\mathbf{B}$ then $\mathbf{A}$ equals applying $\mathbf{AB}$. This is why a stack of linear layers without non-linearities collapses into a single linear layer — the reason neural networks need activation functions.

Key properties#

  • Associative: $(\mathbf{AB})\mathbf{C} = \mathbf{A}(\mathbf{BC})$.
  • Distributive: $\mathbf{A}(\mathbf{B} + \mathbf{C}) = \mathbf{AB} + \mathbf{AC}$.
  • Not commutative: generally $\mathbf{AB} \ne \mathbf{BA}$.
  • Transpose: $(\mathbf{AB})^\top = \mathbf{B}^\top\mathbf{A}^\top$.

Rank and the four fundamental subspaces#

The rank of $\mathbf{A}$ is the dimension of its column space — the set of all outputs $\mathbf{A}\mathbf{x}$. Gilbert Strang's four fundamental subspaces of an $m \times n$ matrix of rank $r$:

SubspaceLives inDimension
Column space $C(\mathbf{A})$$\mathbb{R}^m$$r$
Null space $N(\mathbf{A})$: $\mathbf{Ax} = \mathbf{0}$$\mathbb{R}^n$$n - r$
Row space $C(\mathbf{A}^\top)$$\mathbb{R}^n$$r$
Left null space $N(\mathbf{A}^\top)$$\mathbb{R}^m$$m - r$

The rank–nullity theorem: $\text{rank} + \text{nullity} = n$. The row space and null space are orthogonal complements.

Inverses and determinants#

A square matrix $\mathbf{A}$ is invertible if there is $\mathbf{A}^{-1}$ with $\mathbf{A}\mathbf{A}^{-1} = \mathbf{I}$. Equivalent conditions: full rank, trivial null space, non-zero determinant, no zero eigenvalue.

The determinant measures how the transformation scales volume: $|\det \mathbf{A}|$ is the volume of the image of the unit cube; the sign says whether orientation flips. $\det(\mathbf{AB}) = \det\mathbf{A}\,\det\mathbf{B}$. Determinants appear in the Gaussian density and in normalising flows, where we must track how a transformation changes probability volume.

Special matrices you will meet constantly#

  • Identity $\mathbf{I}$ and diagonal matrices — scale each axis independently.
  • Symmetric $\mathbf{A} = \mathbf{A}^\top$ — covariance matrices, Hessians, kernel matrices.
  • Orthogonal $\mathbf{Q}^\top\mathbf{Q} = \mathbf{I}$ — rotations and reflections; preserve lengths and angles.
  • Positive (semi-)definite $\mathbf{x}^\top\mathbf{A}\mathbf{x} \ge 0$ — covariance matrices; convex quadratic forms.
  • Sparse matrices — graphs, text term–document matrices.
python
import numpy as np

theta = np.pi / 4
R = np.array([[np.cos(theta), -np.sin(theta)],
              [np.sin(theta),  np.cos(theta)]])
x = np.array([1.0, 0.0])
print(R @ x)                           # rotated 45 degrees
print(np.allclose(R.T @ R, np.eye(2))) # orthogonal
print(np.linalg.det(R))                # 1.0 -> preserves area

A = np.array([[4., 1.], [2., 3.]])
b = np.array([1., 2.])
print(np.linalg.solve(A, b))           # preferred over inv(A) @ b

Matrices in neural networks#

A fully connected layer computes $\mathbf{h} = \phi(\mathbf{W}\mathbf{x} + \mathbf{b})$. For a mini-batch stored as rows of $\mathbf{X} \in \mathbb{R}^{B \times d}$, it is $\mathbf{H} = \phi(\mathbf{X}\mathbf{W}^\top + \mathbf{b})$ — one matrix multiplication for the whole batch. GPUs are, above all, machines for fast matrix multiplication, and that is why deep learning became practical.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

∑ Mathematics for ML

Vectors and Vector Spaces: The Language of Data

Every data point, word and image becomes a vector. We define vector spaces, linear combinations, span, independence, basis and dimension — and see why embeddings live in them.

Beginner⏱ 5 min#026
∑ Mathematics for ML

Eigenvalues and Eigenvectors: The Natural Axes of a Transformation

Eigenvectors are directions a matrix merely stretches. We derive the characteristic equation, diagonalisation and the spectral theorem, and connect them to PCA, PageRank, Markov chains and training stability.

Intermediate⏱ 5 min#028
∑ Mathematics for ML

Why Mathematics Matters for Machine Learning

You can call library functions without mathematics — until something breaks. We explain which branches of mathematics ML uses, why, and how to learn them efficiently.

Beginner⏱ 4 min#025