๐Ÿ‘๏ธ Computer Vision ยท Lecture 2 of 27

Image Processing Fundamentals: Filtering, Convolution and Edge Detection

Before deep learning, images were processed with hand-designed filters. We study histograms, blurring, sharpening, gradients, the Sobel and Canny edge detectors, and morphological operations โ€” the vocabulary CNNs later learned.

Classical image processing is not obsolete. It is used for preprocessing, data augmentation, fast embedded systems and interpretable pipelines โ€” and it provides the intuition for what convolutional networks learn in their first layers. Today we learn the essential operations with OpenCV.

Point operations and histograms#

A point operation changes each pixel independently: brightness ($I + b$), contrast ($aI$), gamma correction ($I^\gamma$), thresholding. An image's histogram counts pixel intensities. Histogram equalisation spreads intensities across the full range, improving contrast in dark or washed-out images. CLAHE (Contrast-Limited Adaptive Histogram Equalisation) does this locally and is widely used for medical and low-light images.

Filtering by convolution#

Linear filtering replaces each pixel with a weighted sum of its neighbourhood โ€” a convolution with a kernel:

$$ I'(x, y) = \sum_{u, v}K(u, v)\,I(x + u, y + v) $$
KernelEffect
Box / mean ($\frac{1}{9}$ everywhere in $3\times3$)Blur
Gaussian $\propto e^{-(u^2+v^2)/2\sigma^2}$Smooth blur, removes noise
Sharpen $\begin{bmatrix}0&-1&0\\-1&5&-1\\0&-1&0\end{bmatrix}$Enhances edges
Sobel-x $\begin{bmatrix}-1&0&1\\-2&0&2\\-1&0&1\end{bmatrix}$Horizontal gradient (vertical edges)
LaplacianSecond derivative; responds to edges and blobs

Non-linear filters: the median filter replaces each pixel by the median of its neighbourhood and removes salt-and-pepper noise while preserving edges; the bilateral filter smooths while respecting edges by weighting neighbours by both spatial and intensity similarity.

Image gradients and edges#

Edges are where intensity changes sharply. The gradient $\nabla I = (I_x, I_y)$ has magnitude and direction:

$$ |\nabla I| = \sqrt{I_x^2 + I_y^2}, \qquad \theta = \arctan\frac{I_y}{I_x} $$

Sobel filters approximate $I_x$ and $I_y$ with built-in smoothing.

The Canny edge detector#

Canny (1986) remains the classic edge detector:

  1. Smooth with a Gaussian to reduce noise.
  2. Compute gradient magnitude and direction.
  3. Non-maximum suppression โ€” keep only pixels that are local maxima along the gradient direction, producing thin edges.
  4. Double thresholding โ€” strong edges above a high threshold, weak edges between thresholds.
  5. Hysteresis โ€” keep weak edges only if connected to strong ones.
python
import cv2
import numpy as np

img = cv2.imread("photo.jpg", cv2.IMREAD_GRAYSCALE)
blur = cv2.GaussianBlur(img, (5, 5), sigmaX=1.4)
gx = cv2.Sobel(blur, cv2.CV_64F, 1, 0, ksize=3)
gy = cv2.Sobel(blur, cv2.CV_64F, 0, 1, ksize=3)
magnitude = np.hypot(gx, gy)
edges = cv2.Canny(img, threshold1=50, threshold2=150)
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8, 8)).apply(img)
denoised = cv2.medianBlur(img, 5)
cv2.imwrite("edges.png", edges)

Morphological operations#

For binary images (e.g. after thresholding), morphology cleans shapes using a structuring element:

  • Erosion shrinks foreground, removing small specks.
  • Dilation grows foreground, filling small gaps.
  • Opening (erode then dilate) removes noise; closing (dilate then erode) fills holes.

Combined with connected-component labelling and contour detection, these let you count cells in microscopy, segment printed characters, or measure objects.

Geometric transformations#

Resizing (with interpolation: nearest, bilinear, bicubic), rotation, affine and perspective (homography) transforms. A homography can "flatten" a photo of a document taken at an angle โ€” the first step of any document-scanning app:

python
src = np.float32([[120, 80], [880, 60], [920, 700], [90, 720]])   # document corners in photo
dst = np.float32([[0, 0], [800, 0], [800, 1000], [0, 1000]])
H = cv2.getPerspectiveTransform(src, dst)
flat = cv2.warpPerspective(cv2.imread("form_photo.jpg"), H, (800, 1000))

From hand-designed to learned filters#

The first layer of a trained CNN learns filters that look strikingly like Gaussian derivatives, oriented edge detectors and colour-opponent blobs โ€” similar to classical filters and to receptive fields found in the mammalian visual cortex. Deep learning did not discard these ideas; it learned them, and then learned far more complex ones on top.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

Introduction to Computer Vision: From Pixels to Perception

We open the Computer Vision track by asking how a machine can see. We cover how images are represented, why vision is hard, the landscape of vision tasks, and how deep learning transformed the field.

Beginnerโฑ 5 min#135
๐Ÿ‘๏ธ Computer Vision

Classical Features: Harris Corners, SIFT, HOG and Bag of Visual Words

Before CNNs, vision relied on carefully engineered features. We study corner detection, SIFT keypoints and descriptors, HOG for pedestrian detection and the bag-of-visual-words model โ€” ideas still used in geometry and robotics.

Intermediateโฑ 5 min#137
๐Ÿ‘๏ธ Computer Vision

LeNet and AlexNet: The Birth of Deep Vision

Two architectures bookend the rise of CNNs. LeNet-5 read handwritten digits in the 1990s; AlexNet won ImageNet in 2012 and launched the deep learning era. We dissect both and the innovations that made AlexNet work.

Beginnerโฑ 5 min#138