๐Ÿ‘๏ธ Computer Vision ยท Lecture 22 of 27

3-D Vision: Stereo, Depth Estimation, Point Clouds and NeRF

Images are 2-D projections of a 3-D world. We cover camera geometry, stereo and monocular depth, structure from motion, point-cloud networks like PointNet, and neural scene representations such as NeRF and Gaussian splatting.

Robots must know how far away obstacles are; augmented-reality apps must place virtual objects on real tables; drones map terrain; archaeologists and disaster responders reconstruct buildings in 3-D. 3-D vision recovers geometric structure from images and other sensors. It combines classical projective geometry โ€” which remains essential โ€” with deep learning.

The pinhole camera model#

A 3-D point $\mathbf{X} = (X, Y, Z)$ in camera coordinates projects to pixel $(u, v)$:

$$ \begin{bmatrix} u \\ v \\ 1 \end{bmatrix} \sim \mathbf{K}\begin{bmatrix} X/Z \\ Y/Z \\ 1 \end{bmatrix}, \qquad \mathbf{K} = \begin{bmatrix} f_x & 0 & c_x \\ 0 & f_y & c_y \\ 0 & 0 & 1 \end{bmatrix} $$

$\mathbf{K}$ holds the intrinsic parameters (focal lengths, principal point); the camera's rotation $\mathbf{R}$ and translation $\mathbf{t}$ relative to the world are the extrinsics. Division by $Z$ is why depth is lost in a single image: every point along a ray projects to the same pixel. Camera calibration (e.g. with a checkerboard) estimates $\mathbf{K}$ and lens distortion.

Stereo vision#

Two cameras separated by a baseline $B$ see a point at slightly different horizontal positions. The difference, the disparity $d = u_L - u_R$, gives depth for rectified cameras:

$$ Z = \frac{f\,B}{d} $$

Nearby objects have large disparity; distant ones small. Depth precision therefore degrades with distance. The hard part is stereo matching โ€” finding corresponding pixels โ€” which is ambiguous in textureless regions, reflections and occlusions. Classical semi-global matching and learned stereo networks (e.g. RAFT-Stereo) address it.

python
import cv2
left = cv2.imread("left_rectified.png", cv2.IMREAD_GRAYSCALE)
right = cv2.imread("right_rectified.png", cv2.IMREAD_GRAYSCALE)
sgbm = cv2.StereoSGBM_create(minDisparity=0, numDisparities=128, blockSize=5)
disparity = sgbm.compute(left, right).astype("float32") / 16.0
f, B = 700.0, 0.12                                    # focal length (px), baseline (m)
depth = f * B / (disparity + 1e-6)                    # metres, where disparity > 0

Monocular depth estimation#

From a single image, depth is ambiguous in principle, but humans use cues โ€” perspective, relative size, texture gradients, occlusion, shading. Deep networks learn these cues from data:

  • Supervised with depth sensors (LiDAR, RGB-D cameras).
  • Self-supervised from video or stereo pairs: predict depth and camera motion so that one frame can be warped to reconstruct its neighbour (Monodepth).
  • Large models trained on diverse mixed datasets (e.g. MiDaS, Depth Anything) produce robust relative depth for arbitrary images; metric depth requires scale information.

Structure from Motion and SLAM#

Structure from Motion (SfM) reconstructs 3-D points and camera poses from many overlapping photos: detect and match features (SIFT), estimate relative poses (essential matrix + RANSAC), triangulate points, and refine everything jointly with bundle adjustment โ€” minimising total reprojection error. COLMAP is the standard open-source tool. SLAM (Simultaneous Localisation and Mapping) does this incrementally and in real time for robots and AR devices.

Point clouds#

LiDAR and depth cameras produce point clouds โ€” unordered sets of 3-D points. PointNet (Qi et al., 2017) processes them directly: apply a shared MLP to each point, then aggregate with a symmetric function (max pooling), guaranteeing permutation invariance:

$$ f(\{\mathbf{x}_1, \dots, \mathbf{x}_n\}) = \gamma\left(\max_{i}h(\mathbf{x}_i)\right) $$

PointNet++ adds hierarchical local grouping; other approaches voxelise space (sparse 3-D convolutions) or project to bird's-eye views, as in autonomous-driving detectors.

Neural scene representations#

NeRF (Neural Radiance Fields, Mildenhall et al., 2020) represents a scene as an MLP mapping a 3-D position and viewing direction to colour and density:

$$ F_\Theta: (x, y, z, \theta, \phi) \mapsto (\mathbf{c}, \sigma) $$

Images are rendered by volume rendering along camera rays:

$$ \hat{C}(\mathbf{r}) = \sum_{i=1}^{N}T_i\,\big(1 - e^{-\sigma_i\delta_i}\big)\,\mathbf{c}_i, \qquad T_i = \exp\Big(-\sum_{j<i}\sigma_j\delta_j\Big) $$

Training minimises the difference between rendered and real pixels from posed photos. Positional encoding of inputs with sinusoids lets the MLP represent fine detail. NeRF produced photorealistic novel views but was slow; hash-grid encodings (Instant-NGP) made it fast.

3-D Gaussian Splatting (2023) represents scenes as millions of anisotropic 3-D Gaussians with colours and opacities, rendered by fast rasterisation โ€” real-time photorealistic rendering, rapidly adopted in graphics, mapping and cultural-heritage digitisation.

Applications#

Autonomous driving and robotics (depth, 3-D detection), AR/VR, drone mapping and photogrammetry for damage assessment and settlement planning, 3-D medical imaging, digital preservation of heritage sites, and e-commerce product visualisation.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

Video Understanding: Action Recognition and Temporal Modelling

Video adds time to vision. We study optical flow, two-stream networks, 3-D convolutions, (2+1)-D factorisation, video transformers and tracking, plus the computational tricks that make video models practical.

Advancedโฑ 5 min#155
๐Ÿ‘๏ธ Computer Vision

OCR and Document AI: From Scanned Forms to Structured Data

Much of the world's information is locked in scanned forms, receipts and handwritten records. We cover text detection, recognition with CRNN and CTC, layout-aware models, multilingual challenges and end-to-end document understanding.

Intermediateโฑ 5 min#157
๐Ÿ‘๏ธ Computer Vision

Face Recognition: Metric Learning, ArcFace and Responsible Use

Face recognition maps faces to embeddings where the same person is close. We study the pipeline, triplet and angular-margin losses, verification versus identification, evaluation โ€” and the serious ethical questions this technology raises.

Advancedโฑ 6 min#154