Classical image processing is not obsolete. It is used for preprocessing, data augmentation, fast embedded systems and interpretable pipelines โ and it provides the intuition for what convolutional networks learn in their first layers. Today we learn the essential operations with OpenCV.
Point operations and histograms#
A point operation changes each pixel independently: brightness ($I + b$), contrast ($aI$), gamma correction ($I^\gamma$), thresholding. An image's histogram counts pixel intensities. Histogram equalisation spreads intensities across the full range, improving contrast in dark or washed-out images. CLAHE (Contrast-Limited Adaptive Histogram Equalisation) does this locally and is widely used for medical and low-light images.
Filtering by convolution#
Linear filtering replaces each pixel with a weighted sum of its neighbourhood โ a convolution with a kernel:
| Kernel | Effect |
|---|---|
| Box / mean ($\frac{1}{9}$ everywhere in $3\times3$) | Blur |
| Gaussian $\propto e^{-(u^2+v^2)/2\sigma^2}$ | Smooth blur, removes noise |
| Sharpen $\begin{bmatrix}0&-1&0\\-1&5&-1\\0&-1&0\end{bmatrix}$ | Enhances edges |
| Sobel-x $\begin{bmatrix}-1&0&1\\-2&0&2\\-1&0&1\end{bmatrix}$ | Horizontal gradient (vertical edges) |
| Laplacian | Second derivative; responds to edges and blobs |
Non-linear filters: the median filter replaces each pixel by the median of its neighbourhood and removes salt-and-pepper noise while preserving edges; the bilateral filter smooths while respecting edges by weighting neighbours by both spatial and intensity similarity.
Image gradients and edges#
Edges are where intensity changes sharply. The gradient $\nabla I = (I_x, I_y)$ has magnitude and direction:
Sobel filters approximate $I_x$ and $I_y$ with built-in smoothing.
The Canny edge detector#
Canny (1986) remains the classic edge detector:
- Smooth with a Gaussian to reduce noise.
- Compute gradient magnitude and direction.
- Non-maximum suppression โ keep only pixels that are local maxima along the gradient direction, producing thin edges.
- Double thresholding โ strong edges above a high threshold, weak edges between thresholds.
- Hysteresis โ keep weak edges only if connected to strong ones.
import cv2
import numpy as np
img = cv2.imread("photo.jpg", cv2.IMREAD_GRAYSCALE)
blur = cv2.GaussianBlur(img, (5, 5), sigmaX=1.4)
gx = cv2.Sobel(blur, cv2.CV_64F, 1, 0, ksize=3)
gy = cv2.Sobel(blur, cv2.CV_64F, 0, 1, ksize=3)
magnitude = np.hypot(gx, gy)
edges = cv2.Canny(img, threshold1=50, threshold2=150)
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8, 8)).apply(img)
denoised = cv2.medianBlur(img, 5)
cv2.imwrite("edges.png", edges)Morphological operations#
For binary images (e.g. after thresholding), morphology cleans shapes using a structuring element:
- Erosion shrinks foreground, removing small specks.
- Dilation grows foreground, filling small gaps.
- Opening (erode then dilate) removes noise; closing (dilate then erode) fills holes.
Combined with connected-component labelling and contour detection, these let you count cells in microscopy, segment printed characters, or measure objects.
Geometric transformations#
Resizing (with interpolation: nearest, bilinear, bicubic), rotation, affine and perspective (homography) transforms. A homography can "flatten" a photo of a document taken at an angle โ the first step of any document-scanning app:
src = np.float32([[120, 80], [880, 60], [920, 700], [90, 720]]) # document corners in photo
dst = np.float32([[0, 0], [800, 0], [800, 1000], [0, 1000]])
H = cv2.getPerspectiveTransform(src, dst)
flat = cv2.warpPerspective(cv2.imread("form_photo.jpg"), H, (800, 1000))From hand-designed to learned filters#
The first layer of a trained CNN learns filters that look strikingly like Gaussian derivatives, oriented edge detectors and colour-opponent blobs โ similar to classical filters and to receptive fields found in the mammalian visual cortex. Deep learning did not discard these ideas; it learned them, and then learned far more complex ones on top.