Anomaly detection asks: which observations do not fit the pattern of the rest? A sudden spike in network traffic, a transaction far from a customer's habits, a sensor reading drifting before a machine fails, a registration record with impossible values. Anomalies are rare, often unlabelled, and diverse โ tomorrow's fraud may look nothing like yesterday's. This makes anomaly detection one of the most practically important and conceptually subtle areas of ML.
Types of anomalies#
- Point anomalies โ a single observation is unusual (a withdrawal of 50ร the usual amount).
- Contextual anomalies โ unusual in context (30 ยฐC is normal in summer, anomalous in winter).
- Collective anomalies โ a sequence or group is unusual even if each point is not (a slow, steady data exfiltration).
Settings#
- Supervised: labelled anomalies exist โ treat as imbalanced classification.
- Semi-supervised (novelty detection): train only on normal data; flag anything different.
- Unsupervised (outlier detection): unlabelled data containing some anomalies; assume they are few and different.
Statistical methods#
For a roughly Gaussian feature, flag points with $|z| = |x - \mu|/\sigma > 3$. More robust: the modified z-score using the median and median absolute deviation (MAD), since outliers inflate the mean and standard deviation. For multivariate data, the Mahalanobis distance
accounts for correlations โ a point can be normal in each feature separately yet anomalous jointly (tall and very light). Use a robust covariance estimate (Minimum Covariance Determinant, EllipticEnvelope).
Distance- and density-based methods#
- k-NN distance โ the distance to the $k$-th nearest neighbour; large = isolated.
- Local Outlier Factor (LOF) โ compares a point's local density to its neighbours' densities. A point in a sparse region next to a dense cluster is flagged, even if globally the sparse region is normal. LOF handles clusters of varying density.
Isolation Forest#
Isolation Forest (Liu, Ting & Zhou, 2008) inverts the usual logic: instead of modelling normal data, it isolates points. Build random trees by choosing a random feature and a random split value. Anomalies โ few and different โ get isolated in few splits (short paths); normal points require many. The anomaly score is based on the average path length $E[h(\mathbf{x})]$ across trees:
where $c(n)$ normalises by the average path length of an unsuccessful search in a binary tree. Scores near 1 indicate anomalies. It is fast, scales to large data and works well in moderately high dimensions โ an excellent default.
One-Class SVM#
Learns a boundary around normal data in a kernel feature space; points outside are anomalies. Powerful but sensitive to kernel parameters and slow on large datasets.
Reconstruction-based methods: autoencoders#
For images, sequences and other complex data, train an autoencoder on normal data. It learns to compress and reconstruct normal patterns; anomalous inputs reconstruct poorly. The reconstruction error $\|\mathbf{x} - \hat{\mathbf{x}}\|^2$ is the anomaly score. Variants use VAEs, predictive models for time series (forecast error as the score), or embeddings from pretrained networks combined with k-NN.
Comparing detectors#
import numpy as np
from sklearn.ensemble import IsolationForest
from sklearn.neighbors import LocalOutlierFactor
from sklearn.svm import OneClassSVM
from sklearn.covariance import EllipticEnvelope
from sklearn.metrics import roc_auc_score, average_precision_score
from sklearn.preprocessing import StandardScaler
rng = np.random.default_rng(0)
normal = np.vstack([rng.normal([0, 0], 1, (900, 2)), rng.normal([6, 6], 0.5, (300, 2))])
anomalies = rng.uniform(-6, 12, (40, 2))
X = StandardScaler().fit_transform(np.vstack([normal, anomalies]))
y = np.r_[np.zeros(len(normal)), np.ones(len(anomalies))] # 1 = anomaly (used only to evaluate)
detectors = {
"Isolation Forest": IsolationForest(n_estimators=300, random_state=0).fit(X),
"LOF": LocalOutlierFactor(n_neighbors=30, novelty=True).fit(X),
"One-Class SVM": OneClassSVM(gamma=0.5, nu=0.05).fit(X),
"Robust covariance": EllipticEnvelope(contamination=0.05, random_state=0).fit(X),
}
for name, d in detectors.items():
score = -d.score_samples(X) # higher = more anomalous
print(f"{name:<18} ROC-AUC={roc_auc_score(y, score):.3f} AP={average_precision_score(y, score):.3f}")The robust-covariance detector assumes a single elliptical cluster and struggles with two normal clusters; Isolation Forest and LOF adapt to the structure.
Evaluating without many labels#
- If some labelled anomalies exist, use precision at k, PR-AUC and recall at a fixed alert budget.
- Otherwise, have domain experts review the top-ranked alerts and measure the fraction that are genuine (precision at k) โ this is how most real systems are evaluated.
- The
contaminationparameter sets the decision threshold, not the ranking; choose it from the alert capacity your team can review.
Practical tips#
- Scale features; engineer features that express "normal behaviour" (e.g. deviation from a user's own history).
- Handle time: use rolling baselines and seasonality for time series.
- Combine detectors (average normalised ranks) for robustness.
- Monitor and retrain โ normal behaviour drifts.