👁️

Computer Vision

From pixels to perception: classification, detection, segmentation, ViTs and 3D vision.

  1. 01Introduction to Computer Vision: From Pixels to PerceptionWe open the Computer Vision track by asking how a machine can see. We cover how images are represented, why vision is hard, the landscape of vision tasks, and how deep learning transformed the field.Beginner5 min
  2. 02Image Processing Fundamentals: Filtering, Convolution and Edge DetectionBefore deep learning, images were processed with hand-designed filters. We study histograms, blurring, sharpening, gradients, the Sobel and Canny edge detectors, and morphological operations — the vocabulary CNNs later learned.Beginner4 min
  3. 03Classical Features: Harris Corners, SIFT, HOG and Bag of Visual WordsBefore CNNs, vision relied on carefully engineered features. We study corner detection, SIFT keypoints and descriptors, HOG for pedestrian detection and the bag-of-visual-words model — ideas still used in geometry and robotics.Intermediate5 min
  4. 04LeNet and AlexNet: The Birth of Deep VisionTwo architectures bookend the rise of CNNs. LeNet-5 read handwritten digits in the 1990s; AlexNet won ImageNet in 2012 and launched the deep learning era. We dissect both and the innovations that made AlexNet work.Beginner5 min
  5. 05VGG and GoogLeNet/Inception — Depth and Multi-Scale DesignIn 2014 two architectures pushed CNNs deeper in opposite styles: VGG with uniform stacks of 3×3 convolutions, GoogLeNet with parallel multi-scale Inception modules and 1×1 bottlenecks. We compare their designs and lessons.Intermediate5 min
  6. 06ResNet in Depth: Architecture, Bottlenecks and VariantsResNet's residual blocks enabled 152-layer networks and became the default vision backbone. We study basic and bottleneck blocks, the full ResNet-50 layout, training recipe, and descendants such as ResNeXt and ConvNeXt.Intermediate5 min
  7. 07EfficientNet and Principled Model ScalingHow should a network grow when you have more compute — deeper, wider or higher resolution? EfficientNet's compound scaling answers "all three, in balance". We cover MBConv blocks, the B0–B7 family and EfficientNetV2.Intermediate5 min
  8. 08MobileNet and Efficient Architectures for Edge DevicesPhones, drones and microcontrollers need vision models that are small and fast. We derive the cost savings of depthwise separable convolutions and study MobileNet V1–V3, ShuffleNet and design principles for efficient inference.Intermediate5 min
  9. 09Building an Image Classification Pipeline End to EndA practical walkthrough of a real image classifier — collecting and splitting data, preprocessing, choosing a pretrained backbone, training, evaluating per class, inspecting errors and exporting the model.Beginner5 min
  10. 10Data Augmentation for Computer VisionAugmentation multiplies your data by encoding known invariances. We survey geometric, photometric and occlusion augmentations, automated policies like RandAugment, mixing methods, and augmentation for detection and segmentation.Intermediate5 min
  11. 11Object Detection I: R-CNN, Fast R-CNN and Faster R-CNNDetection asks what objects are in an image and where. We define bounding boxes, IoU and mAP, then trace the two-stage R-CNN family from selective search to region proposal networks and feature pyramids.Intermediate5 min
  12. 12Object Detection II: YOLO and Real-Time Detection"You Only Look Once" reframed detection as a single regression problem, enabling real-time performance. We study the original YOLO grid formulation, its evolution, anchor-free heads, and practical training with modern tools.Intermediate5 min
  13. 13Object Detection III: SSD, RetinaNet and the Focal LossOne-stage detectors face an extreme imbalance between background and objects. We study SSD's multi-scale default boxes, then derive RetinaNet's focal loss, which let one-stage detectors match two-stage accuracy.Advanced5 min
  14. 14Semantic Segmentation: FCN, U-Net and DeepLabSegmentation labels every pixel. We cover fully convolutional networks, the encoder–decoder U-Net with skip connections, DeepLab's atrous convolutions, loss functions like Dice, and evaluation with IoU.Intermediate5 min
  15. 15Instance Segmentation: Mask R-CNN and BeyondInstance segmentation separates each individual object with its own mask. We study Mask R-CNN's mask branch and RoIAlign, compare instance, semantic and panoptic segmentation, and survey query-based models like Mask2Former.Advanced4 min
  16. 16Vision Transformers (ViT): Images as Sequences of PatchesTransformers conquered language, then vision. We dissect ViT's patch embeddings, class token and positional encodings, compare inductive biases with CNNs, and survey DeiT, Swin and hierarchical designs.Advanced5 min
  17. 17Self-Supervised Vision: SimCLR, MoCo, DINO and Masked AutoencodersLabels are expensive, images are abundant. Self-supervised methods learn visual representations from unlabelled images through contrastive learning, self-distillation or masked reconstruction. We compare the main families and how to use them.Advanced5 min
  18. 18CLIP: Connecting Images and LanguageCLIP learns a shared embedding space for images and text from hundreds of millions of image–caption pairs. We explain its contrastive training, zero-shot classification with prompts, retrieval, limitations and its role in generative models.Intermediate5 min
  19. 19Human Pose EstimationPose estimation locates body joints in images and video. We cover keypoint heatmap regression, top-down versus bottom-up approaches, part affinity fields, evaluation with OKS, 3-D pose, and applications from health to sport.Intermediate5 min
  20. 20Face Recognition: Metric Learning, ArcFace and Responsible UseFace recognition maps faces to embeddings where the same person is close. We study the pipeline, triplet and angular-margin losses, verification versus identification, evaluation — and the serious ethical questions this technology raises.Advanced6 min
  21. 21Video Understanding: Action Recognition and Temporal ModellingVideo adds time to vision. We study optical flow, two-stream networks, 3-D convolutions, (2+1)-D factorisation, video transformers and tracking, plus the computational tricks that make video models practical.Advanced5 min
  22. 223-D Vision: Stereo, Depth Estimation, Point Clouds and NeRFImages are 2-D projections of a 3-D world. We cover camera geometry, stereo and monocular depth, structure from motion, point-cloud networks like PointNet, and neural scene representations such as NeRF and Gaussian splatting.Advanced5 min
  23. 23OCR and Document AI: From Scanned Forms to Structured DataMuch of the world's information is locked in scanned forms, receipts and handwritten records. We cover text detection, recognition with CRNN and CTC, layout-aware models, multilingual challenges and end-to-end document understanding.Intermediate5 min
  24. 24AI in Medical Imaging: Opportunities, Pitfalls and ValidationDeep learning can detect disease in X-rays, retinal scans and pathology slides. We survey modalities and tasks, discuss data and labelling challenges, shortcut learning, rigorous clinical validation, and deployment responsibilities.Intermediate5 min
  25. 25Vision Foundation Models: Segment Anything and Promptable VisionVision is following language towards general-purpose foundation models. We study the Segment Anything Model's promptable design and data engine, open-vocabulary detection, and how foundation models change vision workflows.Advanced5 min
  26. 26Explaining Vision Models: Saliency Maps, Grad-CAM and Their LimitsWhich pixels made the model decide? We study gradient saliency, Grad-CAM, integrated gradients and occlusion, show how they reveal shortcuts, and discuss sanity checks that expose unreliable explanations.Intermediate5 min
  27. 27Adversarial Examples: Fooling Neural Networks and Defending ThemImperceptible perturbations can make a network confidently wrong. We derive FGSM and PGD attacks, explain why adversarial examples exist, cover physical and black-box attacks, and evaluate defences including adversarial training.Advanced5 min