๐Ÿ‘๏ธ Computer Vision ยท Lecture 25 of 27

Vision Foundation Models: Segment Anything and Promptable Vision

Vision is following language towards general-purpose foundation models. We study the Segment Anything Model's promptable design and data engine, open-vocabulary detection, and how foundation models change vision workflows.

For most of deep learning's history, each vision task required its own labelled dataset and model. Foundation models change this: trained once on enormous data, they can be adapted or prompted for many tasks. The Segment Anything Model (SAM), released by Meta AI in 2023, brought this paradigm to segmentation, and open-vocabulary detectors brought it to detection. These tools are transforming how practitioners annotate data and build vision systems.

Promptable segmentation#

SAM defines a new task: given an image and a prompt โ€” a point, a box, a rough mask (and, in later versions, text or concept prompts) โ€” output a valid segmentation mask for the indicated object. Because a single point can be ambiguous (the shirt, or the whole person?), SAM outputs multiple masks with predicted quality scores.

Architecture#

  1. Image encoder โ€” a large ViT pretrained with MAE, run once per image to produce an embedding. This is the expensive part.
  2. Prompt encoder โ€” encodes points and boxes with positional encodings, and masks with convolutions.
  3. Mask decoder โ€” a lightweight transformer that combines image and prompt embeddings, using two-way attention, to predict masks and their IoU scores in milliseconds.

The design separates a heavy, cacheable image embedding from a fast, interactive decoder โ€” so a user can click repeatedly and see masks update in real time.

The data engine#

SAM's power comes from data: the SA-1B dataset with over 1 billion masks on 11 million licensed, privacy-protected images. It was built with a model-in-the-loop data engine:

  1. Assisted-manual stage: annotators click; SAM proposes masks; annotators correct them. The model is retrained as data grows.
  2. Semi-automatic stage: SAM pre-fills confident masks; annotators add missing objects.
  3. Fully automatic stage: SAM is prompted with a grid of points and generates masks, filtered by predicted quality and stability.

This virtuous cycle โ€” better model โ†’ faster annotation โ†’ more data โ†’ better model โ€” is a template for building datasets at scale.

Using SAM#

python
# pip install segment-anything  (and download a checkpoint, e.g. sam_vit_b)
import numpy as np
from PIL import Image
from segment_anything import sam_model_registry, SamPredictor

sam = sam_model_registry["vit_b"](checkpoint="sam_vit_b_01ec64.pth")
predictor = SamPredictor(sam)
image = np.array(Image.open("satellite_tile.jpg").convert("RGB"))
predictor.set_image(image)                               # heavy encoder runs once

masks, scores, _ = predictor.predict(point_coords=np.array([[420, 310]]),
                                     point_labels=np.array([1]),        # 1 = foreground click
                                     multimask_output=True)
best = masks[scores.argmax()]
print("mask area (pixels):", int(best.sum()), " predicted quality:", round(float(scores.max()), 3))

Later versions extended SAM to video (propagating masks across frames with memory) and to text/concept prompts, and efficient variants run on mobile devices.

Open-vocabulary detection and grounding#

Detectors traditionally recognise fixed class lists. Open-vocabulary models align visual regions with text embeddings, so you can detect categories specified at inference time:

  • OWL-ViT / OWLv2: CLIP-style models adapted for detection with text or image queries.
  • Grounding DINO: detects objects described by free-form phrases ("the blue water tank on the roof").
  • Grounded-SAM: combines a grounding detector (text โ†’ boxes) with SAM (boxes โ†’ masks) for text-prompted segmentation.

Foundation-model workflows#

Limitations#

  • Class-agnostic masks: SAM segments "things" but does not say what they are โ€” combine it with classifiers or detectors.
  • Fine structures and domain shift: thin structures, low-contrast medical images and unusual sensors may need fine-tuning (e.g. medical adaptations such as MedSAM).
  • Compute: large image encoders are heavy; use efficient variants for edge devices.
  • Biases and coverage of training data persist, as with all foundation models.

The bigger picture#

General-purpose vision backbones (DINOv2), visionโ€“language models (CLIP, SigLIP), promptable segmenters (SAM) and multimodal large language models that "see" are converging towards general visual intelligence. For practitioners, the skill shifts from building every model from scratch to selecting, prompting, adapting, evaluating and combining foundation models responsibly.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

AI in Medical Imaging: Opportunities, Pitfalls and Validation

Deep learning can detect disease in X-rays, retinal scans and pathology slides. We survey modalities and tasks, discuss data and labelling challenges, shortcut learning, rigorous clinical validation, and deployment responsibilities.

Intermediateโฑ 5 min#158
๐Ÿ‘๏ธ Computer Vision

Explaining Vision Models: Saliency Maps, Grad-CAM and Their Limits

Which pixels made the model decide? We study gradient saliency, Grad-CAM, integrated gradients and occlusion, show how they reveal shortcuts, and discuss sanity checks that expose unreliable explanations.

Intermediateโฑ 5 min#160
๐Ÿ‘๏ธ Computer Vision

OCR and Document AI: From Scanned Forms to Structured Data

Much of the world's information is locked in scanned forms, receipts and handwritten records. We cover text detection, recognition with CRNN and CTC, layout-aware models, multilingual challenges and end-to-end document understanding.

Intermediateโฑ 5 min#157