Traditional classifiers recognise a fixed list of classes chosen before training. To add "solar panel" or "water tank", you must collect labelled data and retrain. In 2021 OpenAI's CLIP (Contrastive Language–Image Pre-training) showed a different path: learn from natural-language captions of images on the web, and you can recognise almost any concept that can be described in words — zero-shot, with no task-specific training.
Training objective#
CLIP has two encoders:
- an image encoder (ResNet or ViT) producing $\mathbf{u}_i$;
- a text encoder (transformer) producing $\mathbf{v}_i$;
each followed by a linear projection into a shared space and L2 normalisation. It was trained on about 400 million image–text pairs collected from the internet.
For a batch of $N$ pairs, compute the $N \times N$ matrix of cosine similarities $s_{ij} = \mathbf{u}_i^\top\mathbf{v}_j / \tau$ (with learned temperature $\tau$). The $N$ matching pairs on the diagonal are positives; all $N^2 - N$ mismatched pairs are negatives. The loss is a symmetric cross-entropy: classify the correct caption for each image, and the correct image for each caption.
import torch
import torch.nn.functional as F
def clip_loss(img_emb, txt_emb, logit_scale):
img, txt = F.normalize(img_emb, dim=-1), F.normalize(txt_emb, dim=-1)
logits = logit_scale * img @ txt.T # (N, N)
labels = torch.arange(len(img))
return (F.cross_entropy(logits, labels) + F.cross_entropy(logits.T, labels)) / 2
print(clip_loss(torch.randn(16, 512), torch.randn(16, 512), logit_scale=torch.tensor(100.0)))Captions describe objects, scenes, actions, styles and attributes, so the model learns a broad visual vocabulary aligned with language.
Zero-shot classification#
To classify an image among classes {cat, dog, tent}:
- Turn each class into a prompt: "a photo of a cat", "a photo of a dog", "a photo of a tent".
- Embed the prompts with the text encoder and the image with the image encoder.
- Predict the class whose text embedding has the highest cosine similarity with the image.
The text embeddings act as the weights of a classifier generated on the fly from language. CLIP's zero-shot ImageNet accuracy matched the original ResNet-50 trained with 1.28 million labelled examples — without using any of them.
# pip install open_clip_torch
import torch, open_clip
from PIL import Image
model, _, preprocess = open_clip.create_model_and_transforms("ViT-B-32", pretrained="laion2b_s34b_b79k")
tokenizer = open_clip.get_tokenizer("ViT-B-32")
classes = ["a flooded road", "a dry road", "a collapsed building", "an intact building"]
image = preprocess(Image.open("scene.jpg")).unsqueeze(0)
text = tokenizer([f"a satellite photo of {c}" for c in classes])
with torch.no_grad():
img_f = model.encode_image(image); txt_f = model.encode_text(text)
img_f /= img_f.norm(dim=-1, keepdim=True); txt_f /= txt_f.norm(dim=-1, keepdim=True)
probs = (100 * img_f @ txt_f.T).softmax(dim=-1)
print(dict(zip(classes, probs[0].round(decimals=3).tolist())))Prompt engineering and ensembling#
Wording matters. "A photo of a {label}" usually beats the bare label. For satellite images, "a satellite photo of {label}" helps. Prompt ensembling — averaging text embeddings over many templates ("a blurry photo of", "a close-up photo of", "an origami") — improves accuracy further. For some fine-grained domains, adding context ("a type of pet") resolves ambiguity.
Other capabilities#
- Image–text retrieval: search images by text, or find captions for images.
- Linear probes: CLIP image features with a simple classifier are strong for many tasks with few labels.
- Robustness: CLIP's zero-shot models were notably more robust to certain natural distribution shifts (sketches, renditions) than ImageNet-trained models of similar accuracy.
- Foundation for generation: CLIP-style text encoders guide text-to-image diffusion models; CLIP scores are used to evaluate image–text alignment.
- Open-vocabulary detection and segmentation build on CLIP-like embeddings.
Limitations#
Successors#
Open reproductions (OpenCLIP trained on LAION datasets), ALIGN (noisier but larger data), SigLIP (a sigmoid pairwise loss that removes the need for batch-wide softmax normalisation and works well at small batch sizes), multilingual CLIP variants, and domain-specific versions for medicine and remote sensing. CLIP-style vision encoders are the visual front end of most modern multimodal large language models.