👁️ Computer Vision · Lecture 18 of 27

CLIP: Connecting Images and Language

CLIP learns a shared embedding space for images and text from hundreds of millions of image–caption pairs. We explain its contrastive training, zero-shot classification with prompts, retrieval, limitations and its role in generative models.

Traditional classifiers recognise a fixed list of classes chosen before training. To add "solar panel" or "water tank", you must collect labelled data and retrain. In 2021 OpenAI's CLIP (Contrastive Language–Image Pre-training) showed a different path: learn from natural-language captions of images on the web, and you can recognise almost any concept that can be described in words — zero-shot, with no task-specific training.

Training objective#

CLIP has two encoders:

  • an image encoder (ResNet or ViT) producing $\mathbf{u}_i$;
  • a text encoder (transformer) producing $\mathbf{v}_i$;

each followed by a linear projection into a shared space and L2 normalisation. It was trained on about 400 million image–text pairs collected from the internet.

For a batch of $N$ pairs, compute the $N \times N$ matrix of cosine similarities $s_{ij} = \mathbf{u}_i^\top\mathbf{v}_j / \tau$ (with learned temperature $\tau$). The $N$ matching pairs on the diagonal are positives; all $N^2 - N$ mismatched pairs are negatives. The loss is a symmetric cross-entropy: classify the correct caption for each image, and the correct image for each caption.

python
import torch
import torch.nn.functional as F

def clip_loss(img_emb, txt_emb, logit_scale):
    img, txt = F.normalize(img_emb, dim=-1), F.normalize(txt_emb, dim=-1)
    logits = logit_scale * img @ txt.T                      # (N, N)
    labels = torch.arange(len(img))
    return (F.cross_entropy(logits, labels) + F.cross_entropy(logits.T, labels)) / 2

print(clip_loss(torch.randn(16, 512), torch.randn(16, 512), logit_scale=torch.tensor(100.0)))

Captions describe objects, scenes, actions, styles and attributes, so the model learns a broad visual vocabulary aligned with language.

Zero-shot classification#

To classify an image among classes {cat, dog, tent}:

  1. Turn each class into a prompt: "a photo of a cat", "a photo of a dog", "a photo of a tent".
  2. Embed the prompts with the text encoder and the image with the image encoder.
  3. Predict the class whose text embedding has the highest cosine similarity with the image.

The text embeddings act as the weights of a classifier generated on the fly from language. CLIP's zero-shot ImageNet accuracy matched the original ResNet-50 trained with 1.28 million labelled examples — without using any of them.

python
# pip install open_clip_torch
import torch, open_clip
from PIL import Image

model, _, preprocess = open_clip.create_model_and_transforms("ViT-B-32", pretrained="laion2b_s34b_b79k")
tokenizer = open_clip.get_tokenizer("ViT-B-32")
classes = ["a flooded road", "a dry road", "a collapsed building", "an intact building"]
image = preprocess(Image.open("scene.jpg")).unsqueeze(0)
text = tokenizer([f"a satellite photo of {c}" for c in classes])
with torch.no_grad():
    img_f = model.encode_image(image); txt_f = model.encode_text(text)
    img_f /= img_f.norm(dim=-1, keepdim=True); txt_f /= txt_f.norm(dim=-1, keepdim=True)
    probs = (100 * img_f @ txt_f.T).softmax(dim=-1)
print(dict(zip(classes, probs[0].round(decimals=3).tolist())))

Prompt engineering and ensembling#

Wording matters. "A photo of a {label}" usually beats the bare label. For satellite images, "a satellite photo of {label}" helps. Prompt ensembling — averaging text embeddings over many templates ("a blurry photo of", "a close-up photo of", "an origami") — improves accuracy further. For some fine-grained domains, adding context ("a type of pet") resolves ambiguity.

Other capabilities#

  • Image–text retrieval: search images by text, or find captions for images.
  • Linear probes: CLIP image features with a simple classifier are strong for many tasks with few labels.
  • Robustness: CLIP's zero-shot models were notably more robust to certain natural distribution shifts (sketches, renditions) than ImageNet-trained models of similar accuracy.
  • Foundation for generation: CLIP-style text encoders guide text-to-image diffusion models; CLIP scores are used to evaluate image–text alignment.
  • Open-vocabulary detection and segmentation build on CLIP-like embeddings.

Limitations#

Successors#

Open reproductions (OpenCLIP trained on LAION datasets), ALIGN (noisier but larger data), SigLIP (a sigmoid pairwise loss that removes the need for batch-wide softmax normalisation and works well at small batch sizes), multilingual CLIP variants, and domain-specific versions for medicine and remote sensing. CLIP-style vision encoders are the visual front end of most modern multimodal large language models.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

👁️ Computer Vision

Self-Supervised Vision: SimCLR, MoCo, DINO and Masked Autoencoders

Labels are expensive, images are abundant. Self-supervised methods learn visual representations from unlabelled images through contrastive learning, self-distillation or masked reconstruction. We compare the main families and how to use them.

Advanced⏱ 5 min#151
👁️ Computer Vision

Human Pose Estimation

Pose estimation locates body joints in images and video. We cover keypoint heatmap regression, top-down versus bottom-up approaches, part affinity fields, evaluation with OKS, 3-D pose, and applications from health to sport.

Intermediate⏱ 5 min#153
👁️ Computer Vision

Vision Transformers (ViT): Images as Sequences of Patches

Transformers conquered language, then vision. We dissect ViT's patch embeddings, class token and positional encodings, compare inductive biases with CNNs, and survey DeiT, Swin and hierarchical designs.

Advanced⏱ 5 min#150