Robots must know how far away obstacles are; augmented-reality apps must place virtual objects on real tables; drones map terrain; archaeologists and disaster responders reconstruct buildings in 3-D. 3-D vision recovers geometric structure from images and other sensors. It combines classical projective geometry โ which remains essential โ with deep learning.
The pinhole camera model#
A 3-D point $\mathbf{X} = (X, Y, Z)$ in camera coordinates projects to pixel $(u, v)$:
$\mathbf{K}$ holds the intrinsic parameters (focal lengths, principal point); the camera's rotation $\mathbf{R}$ and translation $\mathbf{t}$ relative to the world are the extrinsics. Division by $Z$ is why depth is lost in a single image: every point along a ray projects to the same pixel. Camera calibration (e.g. with a checkerboard) estimates $\mathbf{K}$ and lens distortion.
Stereo vision#
Two cameras separated by a baseline $B$ see a point at slightly different horizontal positions. The difference, the disparity $d = u_L - u_R$, gives depth for rectified cameras:
Nearby objects have large disparity; distant ones small. Depth precision therefore degrades with distance. The hard part is stereo matching โ finding corresponding pixels โ which is ambiguous in textureless regions, reflections and occlusions. Classical semi-global matching and learned stereo networks (e.g. RAFT-Stereo) address it.
import cv2
left = cv2.imread("left_rectified.png", cv2.IMREAD_GRAYSCALE)
right = cv2.imread("right_rectified.png", cv2.IMREAD_GRAYSCALE)
sgbm = cv2.StereoSGBM_create(minDisparity=0, numDisparities=128, blockSize=5)
disparity = sgbm.compute(left, right).astype("float32") / 16.0
f, B = 700.0, 0.12 # focal length (px), baseline (m)
depth = f * B / (disparity + 1e-6) # metres, where disparity > 0Monocular depth estimation#
From a single image, depth is ambiguous in principle, but humans use cues โ perspective, relative size, texture gradients, occlusion, shading. Deep networks learn these cues from data:
- Supervised with depth sensors (LiDAR, RGB-D cameras).
- Self-supervised from video or stereo pairs: predict depth and camera motion so that one frame can be warped to reconstruct its neighbour (Monodepth).
- Large models trained on diverse mixed datasets (e.g. MiDaS, Depth Anything) produce robust relative depth for arbitrary images; metric depth requires scale information.
Structure from Motion and SLAM#
Structure from Motion (SfM) reconstructs 3-D points and camera poses from many overlapping photos: detect and match features (SIFT), estimate relative poses (essential matrix + RANSAC), triangulate points, and refine everything jointly with bundle adjustment โ minimising total reprojection error. COLMAP is the standard open-source tool. SLAM (Simultaneous Localisation and Mapping) does this incrementally and in real time for robots and AR devices.
Point clouds#
LiDAR and depth cameras produce point clouds โ unordered sets of 3-D points. PointNet (Qi et al., 2017) processes them directly: apply a shared MLP to each point, then aggregate with a symmetric function (max pooling), guaranteeing permutation invariance:
PointNet++ adds hierarchical local grouping; other approaches voxelise space (sparse 3-D convolutions) or project to bird's-eye views, as in autonomous-driving detectors.
Neural scene representations#
NeRF (Neural Radiance Fields, Mildenhall et al., 2020) represents a scene as an MLP mapping a 3-D position and viewing direction to colour and density:
Images are rendered by volume rendering along camera rays:
Training minimises the difference between rendered and real pixels from posed photos. Positional encoding of inputs with sinusoids lets the MLP represent fine detail. NeRF produced photorealistic novel views but was slow; hash-grid encodings (Instant-NGP) made it fast.
3-D Gaussian Splatting (2023) represents scenes as millions of anisotropic 3-D Gaussians with colours and opacities, rendered by fast rasterisation โ real-time photorealistic rendering, rapidly adopted in graphics, mapping and cultural-heritage digitisation.
Applications#
Autonomous driving and robotics (depth, 3-D detection), AR/VR, drone mapping and photogrammetry for damage assessment and settlement planning, 3-D medical imaging, digital preservation of heritage sites, and e-commerce product visualisation.