👁️ Computer Vision · Lecture 1 of 27

Introduction to Computer Vision: From Pixels to Perception

We open the Computer Vision track by asking how a machine can see. We cover how images are represented, why vision is hard, the landscape of vision tasks, and how deep learning transformed the field.

Welcome to the Computer Vision track. Humans recognise a friend's face in a crowd in a fraction of a second, yet for decades this was one of the hardest problems in computing. In 1966, a famous MIT summer project aimed to "solve" vision in a few months; the field has been working on it ever since. Today we set out the problem and the map of the territory.

How computers see an image#

To a computer, an image is an array of numbers. A grayscale image of height $H$ and width $W$ is a matrix in $\mathbb{R}^{H \times W}$, with each pixel typically an integer from 0 (black) to 255 (white). A colour image adds a channel dimension: $H \times W \times 3$ for red, green and blue. A 1080p photo therefore contains about 6.2 million numbers.

Other representations matter too: HSV separates hue from brightness (useful for colour-based segmentation), depth images store distance, multispectral satellite images have many bands beyond visible light, and medical scans (CT, MRI) are 3-D volumes.

Why vision is hard#

The same object produces wildly different pixel arrays because of:

  • Viewpoint — a chair seen from above vs from the side.
  • Illumination — bright sunlight vs shadow change every pixel value.
  • Scale — near vs far.
  • Deformation — a sitting cat vs a stretching cat.
  • Occlusion — objects partly hidden.
  • Background clutter — the object blends into its surroundings.
  • Intra-class variation — thousands of chair designs.

Meanwhile, tiny pixel changes can alter meaning (a "stop" sign vs a "slow" sign). A good vision system must be invariant to nuisance variation while remaining sensitive to meaningful differences.

The landscape of vision tasks#

TaskOutputExample
Image classificationOne label per image"This X-ray shows pneumonia"
Object detectionBoxes + labelsLocate every vehicle in a street scene
Semantic segmentationLabel per pixelRoad, building, vegetation in satellite images
Instance segmentationMask per object instanceSeparate each person in a crowd
Pose estimationKeypointsBody joints for physiotherapy apps
Depth / 3-D reconstructionDistances, meshesRobot navigation
TrackingIdentities over videoFollowing players in sports
OCR / document understandingText and layoutDigitising registration forms
Image generationNew imagesText-to-image models
Vision–languageCaptions, answers"What is the child holding?"

Three eras of computer vision#

  1. Geometry and hand-crafted rules (1960s–1990s) — edge detection, shape from shading, stereo geometry.
  2. Hand-crafted features + machine learning (2000s) — SIFT, HOG and bag-of-visual-words features fed to SVMs; the Viola–Jones face detector.
  3. Deep learning (2012–present) — CNNs learn features end-to-end; since 2020, Vision Transformers and large pretrained vision–language models learned from web-scale image–text data.

The 2012 ImageNet moment was dramatic: AlexNet reduced the top-5 error from about 26% to about 15%. Within a few years, error on this benchmark fell below reported human estimates.

A first look in code#

python
import numpy as np
from PIL import Image
import matplotlib.pyplot as plt

img = np.asarray(Image.open("photo.jpg").convert("RGB"))     # (H, W, 3), uint8
print(img.shape, img.dtype, img.min(), img.max())

gray = img.mean(axis=2)                                        # naive grayscale
red = img.copy(); red[..., 1:] = 0                             # keep only the red channel
fig, ax = plt.subplots(1, 3, figsize=(12, 4))
for a, im, t in zip(ax, [img, gray, red], ["RGB", "Gray", "Red channel"]):
    a.imshow(im, cmap="gray" if im.ndim == 2 else None); a.set_title(t); a.axis("off")
plt.show()

Vision for good — and with care#

Computer vision assists doctors reading scans, maps informal settlements and flood extents from satellites after disasters, monitors crops, reads handwritten records and helps visually impaired people navigate. It also powers surveillance and facial recognition with serious risks to privacy and civil liberties, and has shown unequal accuracy across skin tones and genders. As engineers, we must evaluate across diverse populations and consider how a system will be used — themes we revisit in the Ethics track.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

👁️ Computer Vision

Image Processing Fundamentals: Filtering, Convolution and Edge Detection

Before deep learning, images were processed with hand-designed filters. We study histograms, blurring, sharpening, gradients, the Sobel and Canny edge detectors, and morphological operations — the vocabulary CNNs later learned.

Beginner⏱ 4 min#136
👁️ Computer Vision

Classical Features: Harris Corners, SIFT, HOG and Bag of Visual Words

Before CNNs, vision relied on carefully engineered features. We study corner detection, SIFT keypoints and descriptors, HOG for pedestrian detection and the bag-of-visual-words model — ideas still used in geometry and robotics.

Intermediate⏱ 5 min#137
👁️ Computer Vision

LeNet and AlexNet: The Birth of Deep Vision

Two architectures bookend the rise of CNNs. LeNet-5 read handwritten digits in the 1990s; AlexNet won ImageNet in 2012 and launched the deep learning era. We dissect both and the innovations that made AlexNet work.

Beginner⏱ 5 min#138