Can a network learn useful features without any labels? One elegant answer is to ask it to reproduce its own input โ but through a narrow bottleneck that forces it to discover the data's essential structure. This is the autoencoder, an early and still-important form of unsupervised (self-supervised) representation learning, and the direct ancestor of variational autoencoders and latent diffusion models.
Architecture#
An autoencoder has two parts:
- an encoder $\mathbf{z} = f_\phi(\mathbf{x})$ mapping the input to a latent code (the bottleneck);
- a decoder $\hat{\mathbf{x}} = g_\theta(\mathbf{z})$ mapping the code back to input space.
Training minimises a reconstruction loss, e.g.
(or binary cross-entropy for inputs in $[0, 1]$). No labels are needed.
Why a bottleneck?#
If the latent dimension were at least as large as the input and the network flexible enough, it could learn the identity and nothing useful. An undercomplete autoencoder ($\dim\mathbf{z} \ll \dim\mathbf{x}$) must compress, so it keeps the factors that explain most of the data.
Regularised autoencoders#
Instead of (or in addition to) a narrow bottleneck, we can constrain the code in other ways:
- Sparse autoencoders add a penalty encouraging most latent units to be inactive (e.g. L1 on activations or a KL penalty towards a small target activation). Each input is explained by a few active features. Sparse autoencoders have become a key tool in mechanistic interpretability, used to decompose a language model's internal activations into more interpretable features.
- Denoising autoencoders (Vincent et al., 2008) corrupt the input โ add noise, mask pixels โ and train the network to reconstruct the clean input. To denoise, the model must learn the structure of the data manifold rather than copy pixels. This idea โ reconstruct what was corrupted โ anticipates masked language modelling (BERT), masked image modelling (MAE) and diffusion models.
- Contractive autoencoders penalise the Jacobian of the encoder, making the code insensitive to small input changes.
A convolutional denoising autoencoder#
import torch
import torch.nn as nn
import torch.nn.functional as F
from torchvision import datasets, transforms
class ConvAE(nn.Module):
def __init__(self, latent=32):
super().__init__()
self.enc = nn.Sequential(
nn.Conv2d(1, 32, 3, 2, 1), nn.ReLU(), # 28 -> 14
nn.Conv2d(32, 64, 3, 2, 1), nn.ReLU(), # 14 -> 7
nn.Flatten(), nn.Linear(64 * 7 * 7, latent))
self.dec = nn.Sequential(
nn.Linear(latent, 64 * 7 * 7), nn.ReLU(), nn.Unflatten(1, (64, 7, 7)),
nn.ConvTranspose2d(64, 32, 4, 2, 1), nn.ReLU(), # 7 -> 14
nn.ConvTranspose2d(32, 1, 4, 2, 1), nn.Sigmoid()) # 14 -> 28
def forward(self, x):
z = self.enc(x)
return self.dec(z), z
data = datasets.MNIST(".", train=True, download=True, transform=transforms.ToTensor())
loader = torch.utils.data.DataLoader(data, batch_size=128, shuffle=True)
model = ConvAE(); opt = torch.optim.Adam(model.parameters(), 1e-3)
for epoch in range(3):
for x, _ in loader: # labels unused!
noisy = (x + 0.4 * torch.randn_like(x)).clamp(0, 1)
recon, _ = model(noisy)
loss = F.binary_cross_entropy(recon, x) # reconstruct the CLEAN image
opt.zero_grad(); loss.backward(); opt.step()
print(f"epoch {epoch}: loss {loss.item():.4f}")After training, feeding a noisy digit returns a clean-looking one, and the 32-dimensional codes cluster by digit identity โ even though the model never saw a label.
Applications#
- Dimensionality reduction and visualisation โ non-linear alternative to PCA.
- Denoising โ images, audio, sensor signals.
- Anomaly detection โ train on normal data; high reconstruction error flags anomalies (defects on a production line, unusual network traffic).
- Pretraining / feature learning โ use the encoder as a feature extractor for downstream tasks with few labels.
- Compression โ learned image codecs use autoencoders with quantised latents.
- Latent spaces for generation โ latent diffusion models (e.g. Stable Diffusion) first train an autoencoder to compress images into a compact latent space, then run diffusion there.
Limitations#
A plain autoencoder is not a good generative model. Its latent space is not organised in any particular way: decoding a random point or the midpoint between two codes often produces garbage, because the model was never asked to make the whole latent space meaningful. Variational autoencoders fix this by imposing a probabilistic structure on the latent space โ the subject of a lecture in the Generative AI track.