For most of deep learning's history, each vision task required its own labelled dataset and model. Foundation models change this: trained once on enormous data, they can be adapted or prompted for many tasks. The Segment Anything Model (SAM), released by Meta AI in 2023, brought this paradigm to segmentation, and open-vocabulary detectors brought it to detection. These tools are transforming how practitioners annotate data and build vision systems.
Promptable segmentation#
SAM defines a new task: given an image and a prompt โ a point, a box, a rough mask (and, in later versions, text or concept prompts) โ output a valid segmentation mask for the indicated object. Because a single point can be ambiguous (the shirt, or the whole person?), SAM outputs multiple masks with predicted quality scores.
Architecture#
- Image encoder โ a large ViT pretrained with MAE, run once per image to produce an embedding. This is the expensive part.
- Prompt encoder โ encodes points and boxes with positional encodings, and masks with convolutions.
- Mask decoder โ a lightweight transformer that combines image and prompt embeddings, using two-way attention, to predict masks and their IoU scores in milliseconds.
The design separates a heavy, cacheable image embedding from a fast, interactive decoder โ so a user can click repeatedly and see masks update in real time.
The data engine#
SAM's power comes from data: the SA-1B dataset with over 1 billion masks on 11 million licensed, privacy-protected images. It was built with a model-in-the-loop data engine:
- Assisted-manual stage: annotators click; SAM proposes masks; annotators correct them. The model is retrained as data grows.
- Semi-automatic stage: SAM pre-fills confident masks; annotators add missing objects.
- Fully automatic stage: SAM is prompted with a grid of points and generates masks, filtered by predicted quality and stability.
This virtuous cycle โ better model โ faster annotation โ more data โ better model โ is a template for building datasets at scale.
Using SAM#
# pip install segment-anything (and download a checkpoint, e.g. sam_vit_b)
import numpy as np
from PIL import Image
from segment_anything import sam_model_registry, SamPredictor
sam = sam_model_registry["vit_b"](checkpoint="sam_vit_b_01ec64.pth")
predictor = SamPredictor(sam)
image = np.array(Image.open("satellite_tile.jpg").convert("RGB"))
predictor.set_image(image) # heavy encoder runs once
masks, scores, _ = predictor.predict(point_coords=np.array([[420, 310]]),
point_labels=np.array([1]), # 1 = foreground click
multimask_output=True)
best = masks[scores.argmax()]
print("mask area (pixels):", int(best.sum()), " predicted quality:", round(float(scores.max()), 3))Later versions extended SAM to video (propagating masks across frames with memory) and to text/concept prompts, and efficient variants run on mobile devices.
Open-vocabulary detection and grounding#
Detectors traditionally recognise fixed class lists. Open-vocabulary models align visual regions with text embeddings, so you can detect categories specified at inference time:
- OWL-ViT / OWLv2: CLIP-style models adapted for detection with text or image queries.
- Grounding DINO: detects objects described by free-form phrases ("the blue water tank on the roof").
- Grounded-SAM: combines a grounding detector (text โ boxes) with SAM (boxes โ masks) for text-prompted segmentation.
Foundation-model workflows#
Limitations#
- Class-agnostic masks: SAM segments "things" but does not say what they are โ combine it with classifiers or detectors.
- Fine structures and domain shift: thin structures, low-contrast medical images and unusual sensors may need fine-tuning (e.g. medical adaptations such as MedSAM).
- Compute: large image encoders are heavy; use efficient variants for edge devices.
- Biases and coverage of training data persist, as with all foundation models.
The bigger picture#
General-purpose vision backbones (DINOv2), visionโlanguage models (CLIP, SigLIP), promptable segmenters (SAM) and multimodal large language models that "see" are converging towards general visual intelligence. For practitioners, the skill shifts from building every model from scratch to selecting, prompting, adapting, evaluating and combining foundation models responsibly.