๐Ÿ“ˆ Machine Learning ยท Lecture 39 of 47

Anomaly Detection: Finding the Unusual

Fraud, faults and intrusions are rare and varied. We cover statistical, distance-, density- and tree-based detectors, autoencoders for complex data, and how to evaluate detectors when labels are scarce.

Anomaly detection asks: which observations do not fit the pattern of the rest? A sudden spike in network traffic, a transaction far from a customer's habits, a sensor reading drifting before a machine fails, a registration record with impossible values. Anomalies are rare, often unlabelled, and diverse โ€” tomorrow's fraud may look nothing like yesterday's. This makes anomaly detection one of the most practically important and conceptually subtle areas of ML.

Types of anomalies#

  • Point anomalies โ€” a single observation is unusual (a withdrawal of 50ร— the usual amount).
  • Contextual anomalies โ€” unusual in context (30 ยฐC is normal in summer, anomalous in winter).
  • Collective anomalies โ€” a sequence or group is unusual even if each point is not (a slow, steady data exfiltration).

Settings#

  • Supervised: labelled anomalies exist โ†’ treat as imbalanced classification.
  • Semi-supervised (novelty detection): train only on normal data; flag anything different.
  • Unsupervised (outlier detection): unlabelled data containing some anomalies; assume they are few and different.

Statistical methods#

For a roughly Gaussian feature, flag points with $|z| = |x - \mu|/\sigma > 3$. More robust: the modified z-score using the median and median absolute deviation (MAD), since outliers inflate the mean and standard deviation. For multivariate data, the Mahalanobis distance

$$ D_M(\mathbf{x}) = \sqrt{(\mathbf{x} - \boldsymbol{\mu})^\top\boldsymbol{\Sigma}^{-1}(\mathbf{x} - \boldsymbol{\mu})} $$

accounts for correlations โ€” a point can be normal in each feature separately yet anomalous jointly (tall and very light). Use a robust covariance estimate (Minimum Covariance Determinant, EllipticEnvelope).

Distance- and density-based methods#

  • k-NN distance โ€” the distance to the $k$-th nearest neighbour; large = isolated.
  • Local Outlier Factor (LOF) โ€” compares a point's local density to its neighbours' densities. A point in a sparse region next to a dense cluster is flagged, even if globally the sparse region is normal. LOF handles clusters of varying density.

Isolation Forest#

Isolation Forest (Liu, Ting & Zhou, 2008) inverts the usual logic: instead of modelling normal data, it isolates points. Build random trees by choosing a random feature and a random split value. Anomalies โ€” few and different โ€” get isolated in few splits (short paths); normal points require many. The anomaly score is based on the average path length $E[h(\mathbf{x})]$ across trees:

$$ s(\mathbf{x}) = 2^{-E[h(\mathbf{x})]/c(n)} $$

where $c(n)$ normalises by the average path length of an unsuccessful search in a binary tree. Scores near 1 indicate anomalies. It is fast, scales to large data and works well in moderately high dimensions โ€” an excellent default.

One-Class SVM#

Learns a boundary around normal data in a kernel feature space; points outside are anomalies. Powerful but sensitive to kernel parameters and slow on large datasets.

Reconstruction-based methods: autoencoders#

For images, sequences and other complex data, train an autoencoder on normal data. It learns to compress and reconstruct normal patterns; anomalous inputs reconstruct poorly. The reconstruction error $\|\mathbf{x} - \hat{\mathbf{x}}\|^2$ is the anomaly score. Variants use VAEs, predictive models for time series (forecast error as the score), or embeddings from pretrained networks combined with k-NN.

Comparing detectors#

python
import numpy as np
from sklearn.ensemble import IsolationForest
from sklearn.neighbors import LocalOutlierFactor
from sklearn.svm import OneClassSVM
from sklearn.covariance import EllipticEnvelope
from sklearn.metrics import roc_auc_score, average_precision_score
from sklearn.preprocessing import StandardScaler

rng = np.random.default_rng(0)
normal = np.vstack([rng.normal([0, 0], 1, (900, 2)), rng.normal([6, 6], 0.5, (300, 2))])
anomalies = rng.uniform(-6, 12, (40, 2))
X = StandardScaler().fit_transform(np.vstack([normal, anomalies]))
y = np.r_[np.zeros(len(normal)), np.ones(len(anomalies))]      # 1 = anomaly (used only to evaluate)

detectors = {
    "Isolation Forest": IsolationForest(n_estimators=300, random_state=0).fit(X),
    "LOF": LocalOutlierFactor(n_neighbors=30, novelty=True).fit(X),
    "One-Class SVM": OneClassSVM(gamma=0.5, nu=0.05).fit(X),
    "Robust covariance": EllipticEnvelope(contamination=0.05, random_state=0).fit(X),
}
for name, d in detectors.items():
    score = -d.score_samples(X)                       # higher = more anomalous
    print(f"{name:<18} ROC-AUC={roc_auc_score(y, score):.3f}  AP={average_precision_score(y, score):.3f}")

The robust-covariance detector assumes a single elliptical cluster and struggles with two normal clusters; Isolation Forest and LOF adapt to the structure.

Evaluating without many labels#

  • If some labelled anomalies exist, use precision at k, PR-AUC and recall at a fixed alert budget.
  • Otherwise, have domain experts review the top-ranked alerts and measure the fraction that are genuine (precision at k) โ€” this is how most real systems are evaluated.
  • The contamination parameter sets the decision threshold, not the ranking; choose it from the alert capacity your team can review.

Practical tips#

  • Scale features; engineer features that express "normal behaviour" (e.g. deviation from a user's own history).
  • Handle time: use rolling baselines and seasonality for time series.
  • Combine detectors (average normalised ranks) for robustness.
  • Monitor and retrain โ€” normal behaviour drifts.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ“ˆ Machine Learning

DBSCAN and Density-Based Clustering

DBSCAN defines clusters as dense regions separated by sparse ones. It finds arbitrarily shaped clusters, labels outliers as noise and needs no k. We study core points, parameter selection and HDBSCAN.

Intermediateโฑ 5 min#076
๐Ÿ“ˆ Machine Learning

Hyperparameter Tuning: Grid, Random, Bayesian and Early-Stopping Methods

Hyperparameters control how models learn. We compare grid search, random search, Bayesian optimisation and successive halving, explain why random search beats grid search, and tune efficiently with Optuna.

Intermediateโฑ 5 min#087
๐Ÿ“ˆ Machine Learning

Ensemble Learning III: Voting, Stacking and Blending

Different models make different mistakes. We combine them with hard and soft voting, weighted averaging, and stacking with a meta-learner trained on out-of-fold predictions โ€” and discuss when the extra complexity pays off.

Intermediateโฑ 5 min#089