๐Ÿ‘๏ธ Computer Vision ยท Lecture 16 of 27

Vision Transformers (ViT): Images as Sequences of Patches

Transformers conquered language, then vision. We dissect ViT's patch embeddings, class token and positional encodings, compare inductive biases with CNNs, and survey DeiT, Swin and hierarchical designs.

In 2020, Dosovitskiy and colleagues at Google published a paper with a memorable title: "An Image is Worth 16ร—16 Words". They applied a nearly unmodified Transformer โ€” the architecture dominating NLP โ€” directly to images, treating an image as a sequence of patches. Trained on enough data, the Vision Transformer (ViT) matched or beat the best CNNs. Today, transformer-based vision backbones are central to foundation models, multimodal systems and image generation.

(If you have not yet studied transformers, read the NLP track's lecture on the Transformer architecture alongside this one.)

From image to token sequence#

  1. Split the image $H \times W \times C$ into $N = HW/P^2$ non-overlapping patches of size $P \times P$ (e.g. $16 \times 16$; a $224 \times 224$ image gives $14 \times 14 = 196$ patches).
  2. Flatten each patch into a vector of length $P^2C$ (768 for RGB $16 \times 16$ patches) and apply a learned linear projection to dimension $D$. (Equivalently, a convolution with kernel size and stride $P$.)
  3. Prepend a learnable [CLS] token whose final representation summarises the image for classification.
  4. Add positional embeddings (learned, one per position), since self-attention alone is permutation-invariant and would otherwise ignore where each patch came from.
$$ \mathbf{z}_0 = [\mathbf{x}_{\text{cls}};\; \mathbf{x}_p^1\mathbf{E};\; \dots;\; \mathbf{x}_p^N\mathbf{E}] + \mathbf{E}_{\text{pos}} $$
  1. Pass through $L$ standard transformer encoder blocks (pre-norm multi-head self-attention + MLP, with residual connections).
  2. Classify from the final [CLS] representation (or from the mean of patch tokens).
python
import torch
import torch.nn as nn

class PatchEmbed(nn.Module):
    def __init__(self, img=224, patch=16, in_ch=3, dim=384):
        super().__init__()
        self.proj = nn.Conv2d(in_ch, dim, kernel_size=patch, stride=patch)
        self.n = (img // patch) ** 2
    def forward(self, x):
        return self.proj(x).flatten(2).transpose(1, 2)        # (B, N, dim)

class TinyViT(nn.Module):
    def __init__(self, n_classes=10, dim=384, depth=6, heads=6, img=224, patch=16):
        super().__init__()
        self.embed = PatchEmbed(img, patch, 3, dim)
        self.cls = nn.Parameter(torch.zeros(1, 1, dim))
        self.pos = nn.Parameter(torch.randn(1, self.embed.n + 1, dim) * 0.02)
        layer = nn.TransformerEncoderLayer(dim, heads, 4 * dim, activation="gelu",
                                           batch_first=True, norm_first=True)
        self.blocks = nn.TransformerEncoder(layer, depth)
        self.norm, self.head = nn.LayerNorm(dim), nn.Linear(dim, n_classes)
    def forward(self, x):
        t = self.embed(x)
        t = torch.cat([self.cls.expand(len(t), -1, -1), t], dim=1) + self.pos
        return self.head(self.norm(self.blocks(t))[:, 0])     # classify from [CLS]

print(TinyViT()(torch.randn(2, 3, 224, 224)).shape)

Inductive bias: why data scale matters#

CNNs build in locality and translation equivariance. ViT builds in almost nothing: apart from patch extraction, every patch can attend to every other from the first layer, and spatial relationships must be learned through positional embeddings and attention.

Consequence: on ImageNet-1k alone (1.3M images), the original ViT underperformed comparable ResNets. Pretrained on much larger datasets (e.g. ImageNet-21k with 14M images, or the proprietary JFT-300M), ViT surpassed them, and it scaled better with data and compute. With little data, strong priors help; with abundant data, flexible models win โ€” a recurring theme.

Making ViTs data-efficient: DeiT#

DeiT (Touvron et al., 2021) trained competitive ViTs on ImageNet-1k alone using a strong recipe โ€” heavy augmentation (RandAugment, Mixup, CutMix, random erasing), regularisation (stochastic depth), and distillation from a CNN teacher via an extra distillation token. Training recipes, once again, mattered as much as architecture.

What ViTs learn#

Visualisations show that some attention heads in early layers attend locally (like convolutions), while others attend globally from the start. Attention distance increases with depth. Learned positional embeddings reproduce the 2-D grid structure. ViTs tend to be more shape-biased and somewhat more robust to certain corruptions and occlusions than CNNs.

Hierarchical vision transformers: Swin#

Global self-attention costs $O(N^2)$ in the number of patches, which becomes prohibitive for high-resolution dense tasks like detection and segmentation. Swin Transformer (Liu et al., 2021):

  • computes attention within local windows (linear cost in image size);
  • shifts the windows between consecutive layers so information flows across window boundaries;
  • merges patches between stages to build a CNN-like pyramid of feature maps.

Swin became a strong general backbone for detection and segmentation. Other hybrids (ConViT, CoAtNet, MobileViT) blend convolution and attention.

Self-supervised ViTs#

ViTs pair naturally with self-supervised learning: MAE (masked autoencoders) masks ~75% of patches and reconstructs them; DINO / DINOv2 learn powerful general-purpose features by self-distillation, which transfer well to many tasks with little labelled data. These are covered in a later lecture.

Choosing between CNNs and ViTs#

SituationSuggestion
Small dataset, limited computePretrained CNN (ResNet, ConvNeXt, EfficientNet) or pretrained ViT fine-tuned
Large-scale pretrainingViT scales very well
Dense prediction at high resolutionHierarchical (Swin-like) or CNN backbones
Multimodal models (image + text)ViT encoders are the norm (CLIP, most vision-language models)
Edge devicesEfficient CNNs or mobile hybrids
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

Instance Segmentation: Mask R-CNN and Beyond

Instance segmentation separates each individual object with its own mask. We study Mask R-CNN's mask branch and RoIAlign, compare instance, semantic and panoptic segmentation, and survey query-based models like Mask2Former.

Advancedโฑ 4 min#149
๐Ÿ‘๏ธ Computer Vision

Self-Supervised Vision: SimCLR, MoCo, DINO and Masked Autoencoders

Labels are expensive, images are abundant. Self-supervised methods learn visual representations from unlabelled images through contrastive learning, self-distillation or masked reconstruction. We compare the main families and how to use them.

Advancedโฑ 5 min#151
๐Ÿ‘๏ธ Computer Vision

Semantic Segmentation: FCN, U-Net and DeepLab

Segmentation labels every pixel. We cover fully convolutional networks, the encoderโ€“decoder U-Net with skip connections, DeepLab's atrous convolutions, loss functions like Dice, and evaluation with IoU.

Intermediateโฑ 5 min#148