A model is only as trustworthy as its evaluation. I have reviewed student projects with 99% accuracy that turned out to be worthless because the test data had leaked into training. Today we learn how to split data so that the numbers you report actually predict real-world performance.
Three sets, three purposes#
| Set | Used for | Touched how often |
|---|---|---|
| Training | Fitting model parameters | Constantly |
| Validation (development) | Choosing hyperparameters, features, models | Many times |
| Test | Final unbiased estimate of performance | Once |
Every time you make a decision based on a dataset, you fit to it a little. The validation set absorbs that bias, protecting the test set. If you tune on the test set, your reported score becomes optimistically biased.
Typical splits: 60/20/20 or 80/10/10 for moderate data; for very large datasets (millions of examples), validation and test sets of a few thousand to tens of thousands may be enough.
k-fold cross-validation#
With limited data, a single validation split is noisy and wasteful. k-fold cross-validation:
- Split the training data into $k$ equal folds (commonly 5 or 10).
- For each fold $i$: train on the other $k - 1$ folds, validate on fold $i$.
- Report the mean (and standard deviation) of the $k$ validation scores.
Every example is used for validation exactly once. The cost is $k$ training runs. Leave-one-out CV ($k = n$) has low bias but high variance and high cost; 5 or 10 folds is the usual compromise.
Choosing the right splitter#
- Stratified k-fold — preserves class proportions in each fold. Essential for imbalanced classification.
- Group k-fold — keeps all records of a group (patient, user, household, camp) in the same fold. Without it, the model can "recognise" individuals and look better than it is.
- Time-series split — train on the past, validate on the future, with an expanding or sliding window. Random shuffling of time series leaks future information.
- Repeated k-fold — repeat with different shuffles to reduce the variance of the estimate.
import numpy as np
from sklearn.model_selection import (KFold, StratifiedKFold, GroupKFold,
TimeSeriesSplit, cross_val_score)
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification
X, y = make_classification(n_samples=600, weights=[0.9, 0.1], random_state=0)
groups = np.repeat(np.arange(100), 6) # e.g. 6 records per patient
model = LogisticRegression(max_iter=1000)
for name, cv in [("KFold", KFold(5, shuffle=True, random_state=0)),
("Stratified", StratifiedKFold(5, shuffle=True, random_state=0)),
("Group", GroupKFold(5)),
("TimeSeries", TimeSeriesSplit(5))]:
s = cross_val_score(model, X, y, cv=cv, groups=groups if name == "Group" else None, scoring="f1")
print(f"{name:<11} F1 = {s.mean():.3f} ± {s.std():.3f}")Data leakage: the silent killer#
Leakage occurs when information unavailable at prediction time influences training or evaluation. Common forms:
- Preprocessing leakage — fitting a scaler, imputer, PCA or feature selector on the full dataset before splitting. The validation folds influence the transformation.
- Target leakage — a feature derived from the outcome ("discharge date" when predicting hospital admission; "account closed flag" when predicting churn).
- Temporal leakage — using future data to predict the past.
- Group leakage — the same entity in train and test.
- Duplicate leakage — near-identical examples across splits (common in scraped image and text datasets).
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import SelectKBest, f_classif
# Leakage demo: pure noise features, random labels -> true accuracy is 50%
rng = np.random.default_rng(0)
Xn, yn = rng.normal(size=(100, 10_000)), rng.integers(0, 2, 100)
# WRONG: select features on all data, then cross-validate
X_sel = SelectKBest(f_classif, k=20).fit_transform(Xn, yn)
print("leaky CV accuracy:", cross_val_score(LogisticRegression(), X_sel, yn, cv=5).mean())
# RIGHT: selection inside the pipeline
pipe = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression())
print("honest CV accuracy:", cross_val_score(pipe, Xn, yn, cv=5).mean())The leaky version reports impressive accuracy on pure noise; the honest version reports about 50%. This exact mistake has appeared in published research.
Nested cross-validation#
If you tune hyperparameters with cross-validation and then report the best CV score, that score is optimistically biased — you selected the maximum of noisy estimates. Nested CV fixes this:
- The inner loop tunes hyperparameters (e.g.
GridSearchCV). - The outer loop evaluates the entire tuning procedure on held-out folds.
from sklearn.model_selection import GridSearchCV
inner = GridSearchCV(make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)),
{"logisticregression__C": [0.01, 0.1, 1, 10]}, cv=3)
outer_scores = cross_val_score(inner, X, y, cv=StratifiedKFold(5, shuffle=True, random_state=1))
print("nested CV estimate:", outer_scores.mean().round(3))Practical guidance#
- Decide the split strategy by asking: how will the model be used? If it will predict for new patients, split by patient. If it will predict next month, split by time.
- Keep a final test set locked away; use CV within the training portion.
- Report mean ± standard deviation across folds.
- After choosing everything, retrain on all training data before the final test.