Welcome to the Computer Vision track. Humans recognise a friend's face in a crowd in a fraction of a second, yet for decades this was one of the hardest problems in computing. In 1966, a famous MIT summer project aimed to "solve" vision in a few months; the field has been working on it ever since. Today we set out the problem and the map of the territory.
How computers see an image#
To a computer, an image is an array of numbers. A grayscale image of height $H$ and width $W$ is a matrix in $\mathbb{R}^{H \times W}$, with each pixel typically an integer from 0 (black) to 255 (white). A colour image adds a channel dimension: $H \times W \times 3$ for red, green and blue. A 1080p photo therefore contains about 6.2 million numbers.
Other representations matter too: HSV separates hue from brightness (useful for colour-based segmentation), depth images store distance, multispectral satellite images have many bands beyond visible light, and medical scans (CT, MRI) are 3-D volumes.
Why vision is hard#
The same object produces wildly different pixel arrays because of:
- Viewpoint — a chair seen from above vs from the side.
- Illumination — bright sunlight vs shadow change every pixel value.
- Scale — near vs far.
- Deformation — a sitting cat vs a stretching cat.
- Occlusion — objects partly hidden.
- Background clutter — the object blends into its surroundings.
- Intra-class variation — thousands of chair designs.
Meanwhile, tiny pixel changes can alter meaning (a "stop" sign vs a "slow" sign). A good vision system must be invariant to nuisance variation while remaining sensitive to meaningful differences.
The landscape of vision tasks#
| Task | Output | Example |
|---|---|---|
| Image classification | One label per image | "This X-ray shows pneumonia" |
| Object detection | Boxes + labels | Locate every vehicle in a street scene |
| Semantic segmentation | Label per pixel | Road, building, vegetation in satellite images |
| Instance segmentation | Mask per object instance | Separate each person in a crowd |
| Pose estimation | Keypoints | Body joints for physiotherapy apps |
| Depth / 3-D reconstruction | Distances, meshes | Robot navigation |
| Tracking | Identities over video | Following players in sports |
| OCR / document understanding | Text and layout | Digitising registration forms |
| Image generation | New images | Text-to-image models |
| Vision–language | Captions, answers | "What is the child holding?" |
Three eras of computer vision#
- Geometry and hand-crafted rules (1960s–1990s) — edge detection, shape from shading, stereo geometry.
- Hand-crafted features + machine learning (2000s) — SIFT, HOG and bag-of-visual-words features fed to SVMs; the Viola–Jones face detector.
- Deep learning (2012–present) — CNNs learn features end-to-end; since 2020, Vision Transformers and large pretrained vision–language models learned from web-scale image–text data.
The 2012 ImageNet moment was dramatic: AlexNet reduced the top-5 error from about 26% to about 15%. Within a few years, error on this benchmark fell below reported human estimates.
A first look in code#
import numpy as np
from PIL import Image
import matplotlib.pyplot as plt
img = np.asarray(Image.open("photo.jpg").convert("RGB")) # (H, W, 3), uint8
print(img.shape, img.dtype, img.min(), img.max())
gray = img.mean(axis=2) # naive grayscale
red = img.copy(); red[..., 1:] = 0 # keep only the red channel
fig, ax = plt.subplots(1, 3, figsize=(12, 4))
for a, im, t in zip(ax, [img, gray, red], ["RGB", "Gray", "Red channel"]):
a.imshow(im, cmap="gray" if im.ndim == 2 else None); a.set_title(t); a.axis("off")
plt.show()Vision for good — and with care#
Computer vision assists doctors reading scans, maps informal settlements and flood extents from satellites after disasters, monitors crops, reads handwritten records and helps visually impaired people navigate. It also powers surveillance and facial recognition with serious risks to privacy and civil liberties, and has shown unequal accuracy across skin tones and genders. As engineers, we must evaluate across diverse populations and consider how a system will be used — themes we revisit in the Ethics track.