📈 Machine Learning · Lecture 11 of 47

Train, Validation and Test Splits — and Cross-Validation Done Right

Trustworthy evaluation starts with correct data splits. We cover hold-out validation, k-fold, stratified, group and time-series cross-validation, nested CV for tuning, and the leakage traps that invalidate results.

A model is only as trustworthy as its evaluation. I have reviewed student projects with 99% accuracy that turned out to be worthless because the test data had leaked into training. Today we learn how to split data so that the numbers you report actually predict real-world performance.

Three sets, three purposes#

SetUsed forTouched how often
TrainingFitting model parametersConstantly
Validation (development)Choosing hyperparameters, features, modelsMany times
TestFinal unbiased estimate of performanceOnce

Every time you make a decision based on a dataset, you fit to it a little. The validation set absorbs that bias, protecting the test set. If you tune on the test set, your reported score becomes optimistically biased.

Typical splits: 60/20/20 or 80/10/10 for moderate data; for very large datasets (millions of examples), validation and test sets of a few thousand to tens of thousands may be enough.

k-fold cross-validation#

With limited data, a single validation split is noisy and wasteful. k-fold cross-validation:

  1. Split the training data into $k$ equal folds (commonly 5 or 10).
  2. For each fold $i$: train on the other $k - 1$ folds, validate on fold $i$.
  3. Report the mean (and standard deviation) of the $k$ validation scores.

Every example is used for validation exactly once. The cost is $k$ training runs. Leave-one-out CV ($k = n$) has low bias but high variance and high cost; 5 or 10 folds is the usual compromise.

Choosing the right splitter#

  • Stratified k-fold — preserves class proportions in each fold. Essential for imbalanced classification.
  • Group k-fold — keeps all records of a group (patient, user, household, camp) in the same fold. Without it, the model can "recognise" individuals and look better than it is.
  • Time-series split — train on the past, validate on the future, with an expanding or sliding window. Random shuffling of time series leaks future information.
  • Repeated k-fold — repeat with different shuffles to reduce the variance of the estimate.
python
import numpy as np
from sklearn.model_selection import (KFold, StratifiedKFold, GroupKFold,
                                     TimeSeriesSplit, cross_val_score)
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification

X, y = make_classification(n_samples=600, weights=[0.9, 0.1], random_state=0)
groups = np.repeat(np.arange(100), 6)            # e.g. 6 records per patient

model = LogisticRegression(max_iter=1000)
for name, cv in [("KFold", KFold(5, shuffle=True, random_state=0)),
                 ("Stratified", StratifiedKFold(5, shuffle=True, random_state=0)),
                 ("Group", GroupKFold(5)),
                 ("TimeSeries", TimeSeriesSplit(5))]:
    s = cross_val_score(model, X, y, cv=cv, groups=groups if name == "Group" else None, scoring="f1")
    print(f"{name:<11} F1 = {s.mean():.3f} ± {s.std():.3f}")

Data leakage: the silent killer#

Leakage occurs when information unavailable at prediction time influences training or evaluation. Common forms:

  1. Preprocessing leakage — fitting a scaler, imputer, PCA or feature selector on the full dataset before splitting. The validation folds influence the transformation.
  2. Target leakage — a feature derived from the outcome ("discharge date" when predicting hospital admission; "account closed flag" when predicting churn).
  3. Temporal leakage — using future data to predict the past.
  4. Group leakage — the same entity in train and test.
  5. Duplicate leakage — near-identical examples across splits (common in scraped image and text datasets).
python
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import SelectKBest, f_classif

# Leakage demo: pure noise features, random labels -> true accuracy is 50%
rng = np.random.default_rng(0)
Xn, yn = rng.normal(size=(100, 10_000)), rng.integers(0, 2, 100)

# WRONG: select features on all data, then cross-validate
X_sel = SelectKBest(f_classif, k=20).fit_transform(Xn, yn)
print("leaky CV accuracy:", cross_val_score(LogisticRegression(), X_sel, yn, cv=5).mean())

# RIGHT: selection inside the pipeline
pipe = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression())
print("honest CV accuracy:", cross_val_score(pipe, Xn, yn, cv=5).mean())

The leaky version reports impressive accuracy on pure noise; the honest version reports about 50%. This exact mistake has appeared in published research.

Nested cross-validation#

If you tune hyperparameters with cross-validation and then report the best CV score, that score is optimistically biased — you selected the maximum of noisy estimates. Nested CV fixes this:

  • The inner loop tunes hyperparameters (e.g. GridSearchCV).
  • The outer loop evaluates the entire tuning procedure on held-out folds.
python
from sklearn.model_selection import GridSearchCV
inner = GridSearchCV(make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)),
                     {"logisticregression__C": [0.01, 0.1, 1, 10]}, cv=3)
outer_scores = cross_val_score(inner, X, y, cv=StratifiedKFold(5, shuffle=True, random_state=1))
print("nested CV estimate:", outer_scores.mean().round(3))

Practical guidance#

  • Decide the split strategy by asking: how will the model be used? If it will predict for new patients, split by patient. If it will predict next month, split by time.
  • Keep a final test set locked away; use CV within the training portion.
  • Report mean ± standard deviation across folds.
  • After choosing everything, retrain on all training data before the final test.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

📈 Machine Learning

Overfitting and Underfitting: Diagnosis with Learning Curves

Before fixing a model you must diagnose it. We learn to read learning curves and validation curves, recognise high bias and high variance, and choose the right remedy instead of guessing.

Beginner⏱ 5 min#059
📈 Machine Learning

Evaluation Metrics for Classification: Accuracy, Precision, Recall and F1

Accuracy can be dangerously misleading. We build the confusion matrix, define precision, recall, specificity, F-scores and balanced accuracy, and learn to choose metrics from the costs of errors.

Beginner⏱ 5 min#061
📈 Machine Learning

The Bias–Variance Trade-off: Derivation and Intuition

Why do simple models underfit and complex models overfit? We derive the bias–variance decomposition of expected squared error, visualise it, and discuss how modern deep learning complicates the classical picture.

Intermediate⏱ 5 min#058