Face recognition is among the most commercially deployed โ and most controversial โ applications of computer vision. It unlocks phones and verifies identity for services, but it also enables mass surveillance and has shown unequal accuracy across demographic groups. As engineers, we must understand both how it works and why its use demands exceptional care.
The pipeline#
- Detection โ find faces (e.g. RetinaFace, MTCNN).
- Alignment โ detect landmarks (eyes, nose, mouth corners) and warp the face to a canonical pose.
- Embedding โ a deep network maps the aligned face to a vector (e.g. 512 dimensions), normalised to unit length.
- Comparison โ cosine similarity between embeddings, with a threshold.
Verification vs identification#
- Verification (1:1): "Is this person who they claim to be?" Compare one probe with one enrolled template. Used for phone unlock or document checks.
- Identification (1:N): "Who is this person?" Compare a probe against a gallery of $N$ identities. Error rates grow with $N$, because more gallery faces mean more chances of a false match.
Why not ordinary classification?#
A face system must recognise people not seen during training (new users enrol later). We therefore learn an embedding space where distances reflect identity โ metric learning โ rather than a fixed classifier.
Triplet loss#
FaceNet (Schroff et al., 2015) trained with triplets: an anchor $a$, a positive $p$ (same identity) and a negative $n$ (different identity):
The margin $m$ forces negatives to be farther than positives by at least $m$. Performance depends heavily on mining informative (semi-hard) triplets.
Angular-margin softmax losses#
Later methods train a classifier over training identities but modify the softmax to enforce margins on the hypersphere. With normalised embedding $\mathbf{x}$ and normalised class weights $\mathbf{W}_j$, the logit is $s\cos\theta_j$. ArcFace (Deng et al., 2019) adds an additive angular margin $m$ to the true class:
This forces embeddings of the same identity to cluster tightly with clear angular gaps between identities. After training, the classifier is discarded and the embeddings are used. (Related losses: SphereFace, CosFace.)
import torch
import torch.nn as nn
import torch.nn.functional as F
class ArcFaceHead(nn.Module):
def __init__(self, dim, n_ids, s=64.0, m=0.5):
super().__init__()
self.W = nn.Parameter(torch.randn(n_ids, dim) * 0.01)
self.s, self.m = s, m
def forward(self, emb, labels):
cos = F.linear(F.normalize(emb), F.normalize(self.W)).clamp(-1 + 1e-7, 1 - 1e-7)
theta = torch.acos(cos)
target = torch.cos(theta + self.m) # add margin to the true class angle
onehot = F.one_hot(labels, cos.size(1)).bool()
logits = torch.where(onehot, target, cos) * self.s
return F.cross_entropy(logits, labels)
head = ArcFaceHead(512, 1000)
print(head(torch.randn(8, 512), torch.randint(0, 1000, (8,))))Evaluation#
Operating points are defined by thresholds on similarity:
- False Match Rate (FMR / FAR) โ different people accepted as the same.
- False Non-Match Rate (FNMR / FRR) โ same person rejected.
Systems report, for example, FNMR at FMR = $10^{-6}$. Benchmarks range from LFW (now saturated) to harder datasets with pose, age and quality variation. Independent evaluations such as NIST's Face Recognition Vendor Tests report accuracy โ including demographic differentials โ across many commercial algorithms.
Bias and fairness#
Studies including Buolamwini and Gebru's Gender Shades (2018) found commercial facial-analysis systems had much higher error rates for darker-skinned women than for lighter-skinned men. NIST's 2019 demographic study found that many algorithms showed higher false-positive rates for some demographic groups, with large variation between algorithms. Causes include unbalanced training data and image-capture conditions. A false match in a policing context can lead to wrongful suspicion or arrest โ documented cases exist.
Responsible use#
Many cities and organisations have restricted facial recognition use by public authorities; several technology companies paused or limited sales to law enforcement. Responsible engineers treat this technology as high-risk by default.