Real datasets have holes. A survey respondent skips the income question; a sensor fails during a storm; a clinic did not record a test. How you handle these gaps can bias your conclusions, leak information or discard valuable signal. Before choosing a technique, you must ask why the data is missing.
Missingness mechanisms (Rubin's taxonomy)#
- MCAR โ Missing Completely At Random. The probability of missingness is unrelated to any data, observed or not. Example: a random subset of forms was lost. Deleting incomplete rows loses efficiency but does not bias results.
- MAR โ Missing At Random. Missingness depends only on observed variables. Example: younger respondents skip the income question more often, and age is recorded. Imputation methods that condition on observed variables can correct for this.
- MNAR โ Missing Not At Random. Missingness depends on the missing value itself. Example: people with very high or very low incomes decline to report income. No method can fully correct this without assumptions or extra data; it requires domain reasoning and sensitivity analysis.
Strategy 1: deletion#
- Listwise deletion โ drop rows with any missing value. Simple; acceptable if missingness is rare and MCAR. Dangerous otherwise: with 20 features each missing 5% independently, you lose about 64% of rows.
- Column deletion โ drop features that are mostly missing (e.g. > 60โ80%), unless the missingness itself is informative.
Strategy 2: simple imputation#
Replace missing values with a statistic computed on the training set:
- Mean โ for roughly symmetric numeric features.
- Median โ robust for skewed features (income).
- Most frequent / constant "Missing" category โ for categorical features.
Simple imputation shrinks variance and weakens correlations, but combined with an indicator (below) it is often surprisingly effective for predictive modelling.
Strategy 3: missingness indicators#
Add a binary feature x_is_missing. This lets the model learn that missingness itself is informative โ a missing lab test may mean the doctor did not consider it necessary, which says something about the patient. In predictive tasks, indicator + simple imputation frequently beats sophisticated imputation.
Strategy 4: model-based imputation#
- k-NN imputation โ fill a value from similar rows.
- Iterative imputation (MICE-style) โ model each feature with missing values as a function of the others, cycling until convergence.
- Multiple imputation โ create several plausible completed datasets, analyse each, and combine the results (Rubin's rules). This is the gold standard for statistical inference, because it propagates uncertainty about the missing values into standard errors.
Strategy 5: models that handle missingness natively#
Gradient boosting libraries (XGBoost, LightGBM, scikit-learn's HistGradientBoosting) learn a default direction for missing values at each split โ often the simplest and best option for tabular prediction.
import numpy as np
import pandas as pd
from sklearn.experimental import enable_iterative_imputer # noqa: F401
from sklearn.impute import SimpleImputer, KNNImputer, IterativeImputer
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import cross_val_score
from sklearn.datasets import load_breast_cancer
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True)
rng = np.random.default_rng(0)
X_miss = X.copy()
# MAR-style missingness: feature 0 missing more often when feature 1 is high
p = 0.1 + 0.5 * (X[:, 1] > np.median(X[:, 1]))
X_miss[rng.random(len(X)) < p, 0] = np.nan
X_miss[rng.random(X.shape) < 0.05] = np.nan # plus some random holes
strategies = {
"mean": SimpleImputer(strategy="mean"),
"median+indicator": SimpleImputer(strategy="median", add_indicator=True),
"kNN": KNNImputer(n_neighbors=5),
"iterative": IterativeImputer(random_state=0, max_iter=10),
}
for name, imp in strategies.items():
pipe = make_pipeline(imp, StandardScaler(), LogisticRegression(max_iter=3000))
print(f"{name:<17} acc = {cross_val_score(pipe, X_miss, y, cv=5).mean():.3f}")
print(f"{'native (HGB)':<17} acc = {cross_val_score(HistGradientBoostingClassifier(), X_miss, y, cv=5).mean():.3f}")Rules for doing it right#
Choosing an approach#
| Goal | Recommended |
|---|---|
| Quick, strong predictive baseline | Gradient boosting with native handling, or median imputation + indicators |
| Linear models / neural networks | Imputation + indicators inside a pipeline |
| Statistical inference (effects, confidence intervals) | Multiple imputation |
| Suspected MNAR | Domain reasoning, sensitivity analysis, collect more data |