๐Ÿ“ˆ Machine Learning ยท Lecture 34 of 47

Handling Missing Data: Mechanisms, Imputation and Indicators

Missing values are rarely random. We classify missingness as MCAR, MAR or MNAR, compare deletion and imputation strategies from simple to iterative, and show why a missingness indicator is often a feature in itself.

Real datasets have holes. A survey respondent skips the income question; a sensor fails during a storm; a clinic did not record a test. How you handle these gaps can bias your conclusions, leak information or discard valuable signal. Before choosing a technique, you must ask why the data is missing.

Missingness mechanisms (Rubin's taxonomy)#

  1. MCAR โ€” Missing Completely At Random. The probability of missingness is unrelated to any data, observed or not. Example: a random subset of forms was lost. Deleting incomplete rows loses efficiency but does not bias results.
  2. MAR โ€” Missing At Random. Missingness depends only on observed variables. Example: younger respondents skip the income question more often, and age is recorded. Imputation methods that condition on observed variables can correct for this.
  3. MNAR โ€” Missing Not At Random. Missingness depends on the missing value itself. Example: people with very high or very low incomes decline to report income. No method can fully correct this without assumptions or extra data; it requires domain reasoning and sensitivity analysis.

Strategy 1: deletion#

  • Listwise deletion โ€” drop rows with any missing value. Simple; acceptable if missingness is rare and MCAR. Dangerous otherwise: with 20 features each missing 5% independently, you lose about 64% of rows.
  • Column deletion โ€” drop features that are mostly missing (e.g. > 60โ€“80%), unless the missingness itself is informative.

Strategy 2: simple imputation#

Replace missing values with a statistic computed on the training set:

  • Mean โ€” for roughly symmetric numeric features.
  • Median โ€” robust for skewed features (income).
  • Most frequent / constant "Missing" category โ€” for categorical features.

Simple imputation shrinks variance and weakens correlations, but combined with an indicator (below) it is often surprisingly effective for predictive modelling.

Strategy 3: missingness indicators#

Add a binary feature x_is_missing. This lets the model learn that missingness itself is informative โ€” a missing lab test may mean the doctor did not consider it necessary, which says something about the patient. In predictive tasks, indicator + simple imputation frequently beats sophisticated imputation.

Strategy 4: model-based imputation#

  • k-NN imputation โ€” fill a value from similar rows.
  • Iterative imputation (MICE-style) โ€” model each feature with missing values as a function of the others, cycling until convergence.
  • Multiple imputation โ€” create several plausible completed datasets, analyse each, and combine the results (Rubin's rules). This is the gold standard for statistical inference, because it propagates uncertainty about the missing values into standard errors.

Strategy 5: models that handle missingness natively#

Gradient boosting libraries (XGBoost, LightGBM, scikit-learn's HistGradientBoosting) learn a default direction for missing values at each split โ€” often the simplest and best option for tabular prediction.

python
import numpy as np
import pandas as pd
from sklearn.experimental import enable_iterative_imputer  # noqa: F401
from sklearn.impute import SimpleImputer, KNNImputer, IterativeImputer
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import cross_val_score
from sklearn.datasets import load_breast_cancer
from sklearn.preprocessing import StandardScaler

X, y = load_breast_cancer(return_X_y=True)
rng = np.random.default_rng(0)
X_miss = X.copy()
# MAR-style missingness: feature 0 missing more often when feature 1 is high
p = 0.1 + 0.5 * (X[:, 1] > np.median(X[:, 1]))
X_miss[rng.random(len(X)) < p, 0] = np.nan
X_miss[rng.random(X.shape) < 0.05] = np.nan            # plus some random holes

strategies = {
    "mean":            SimpleImputer(strategy="mean"),
    "median+indicator": SimpleImputer(strategy="median", add_indicator=True),
    "kNN":             KNNImputer(n_neighbors=5),
    "iterative":       IterativeImputer(random_state=0, max_iter=10),
}
for name, imp in strategies.items():
    pipe = make_pipeline(imp, StandardScaler(), LogisticRegression(max_iter=3000))
    print(f"{name:<17} acc = {cross_val_score(pipe, X_miss, y, cv=5).mean():.3f}")
print(f"{'native (HGB)':<17} acc = {cross_val_score(HistGradientBoostingClassifier(), X_miss, y, cv=5).mean():.3f}")

Rules for doing it right#

Choosing an approach#

GoalRecommended
Quick, strong predictive baselineGradient boosting with native handling, or median imputation + indicators
Linear models / neural networksImputation + indicators inside a pipeline
Statistical inference (effects, confidence intervals)Multiple imputation
Suspected MNARDomain reasoning, sensitivity analysis, collect more data
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ“ˆ Machine Learning

Feature Scaling and Normalisation: Standardisation, Minโ€“Max and Robust Scaling

Many algorithms silently assume features share a scale. We explain which models need scaling and why, compare standardisation, minโ€“max, robust and quantile scaling, and show how to apply them without leakage.

Beginnerโฑ 4 min#082
๐Ÿ“ˆ Machine Learning

Encoding Categorical Variables: One-Hot, Ordinal, Target and Beyond

Models need numbers, but many features are categories. We compare one-hot, ordinal, frequency, target and hashing encoders, handle high cardinality and unseen categories, and avoid target-encoding leakage.

Beginnerโฑ 5 min#084
๐Ÿ“ˆ Machine Learning

Feature Engineering: Turning Raw Data into Signal

Better features beat better algorithms. We survey the craft โ€” transformations, interactions, aggregations, date and text features, domain-driven ratios โ€” and how to engineer features without leaking the target.

Intermediateโฑ 5 min#081