We have met two ensemble families: bagging averages many copies of one high-variance model, and boosting builds a sequence of weak learners. A third family combines different kinds of models: a gradient-boosted tree ensemble, a regularised logistic regression, a k-NN model and a neural network each capture different aspects of the data. Combining them can outperform any single model โ the approach behind many competition-winning solutions.
Why diverse models help#
Recall the ensemble variance formula: averaging reduces error most when models' errors are weakly correlated. Models from different families make different kinds of mistakes: trees capture interactions but produce step-shaped predictions; linear models extrapolate smoothly but miss interactions; k-NN captures local structure. Diversity is the key ingredient.
Voting and averaging#
- Hard voting โ majority vote of predicted classes.
- Soft voting โ average predicted probabilities, then take the argmax. Usually better, because confident models get more say. Requires reasonably calibrated probabilities.
- Weighted averaging โ weights chosen on validation data (e.g. by a simple search or by minimising validation log-loss).
For regression, average the predictions (possibly weighted).
Stacking#
Stacked generalisation (Wolpert, 1992) learns how to combine models. A meta-learner is trained on the base models' predictions.
The critical detail is how to generate training data for the meta-learner:
- Split the training data into $K$ folds.
- For each base model and each fold, train on the other $K - 1$ folds and predict the held-out fold. This yields out-of-fold (OOF) predictions for every training example โ predictions made by models that never saw that example.
- Train the meta-learner on the OOF predictions (optionally plus original features).
- Refit each base model on the full training data to produce test-time inputs for the meta-learner.
The meta-learner should usually be simple โ logistic or ridge regression with non-negative weights โ to avoid overfitting the meta-level.
Blending#
Blending is a simpler variant: hold out a validation set, train base models on the rest, and train the meta-learner on predictions for the holdout. It is easier to implement and avoids leakage, but uses less data for each stage.
Implementation#
import numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.ensemble import (StackingClassifier, VotingClassifier, RandomForestClassifier,
HistGradientBoostingClassifier)
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.svm import SVC
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True)
cv = StratifiedKFold(5, shuffle=True, random_state=0)
base = [
("lr", make_pipeline(StandardScaler(), LogisticRegression(max_iter=2000))),
("svm", make_pipeline(StandardScaler(), SVC(probability=True, random_state=0))),
("knn", make_pipeline(StandardScaler(), KNeighborsClassifier(15))),
("rf", RandomForestClassifier(n_estimators=300, random_state=0)),
("gb", HistGradientBoostingClassifier(random_state=0)),
]
for name, m in base:
print(f"{name:<6} {cross_val_score(m, X, y, cv=cv, scoring='roc_auc').mean():.4f}")
voting = VotingClassifier(base, voting="soft")
stack = StackingClassifier(base, final_estimator=LogisticRegression(max_iter=2000),
cv=5, stack_method="predict_proba", passthrough=False)
print(f"soft voting {cross_val_score(voting, X, y, cv=cv, scoring='roc_auc').mean():.4f}")
print(f"stacking {cross_val_score(stack, X, y, cv=cv, scoring='roc_auc').mean():.4f}")Scikit-learn's StackingClassifier generates out-of-fold predictions internally using its cv argument. Note that we evaluate the whole stack with an outer cross-validation โ the same nested principle as for hyperparameter tuning.
Checking diversity#
Before stacking, inspect the correlation of OOF predictions between base models. If two models' predictions correlate at 0.99, keeping both adds cost but little value. Aim for strong individual models that disagree on different examples.
Costs and trade-offs#
| Benefit | Cost |
|---|---|
| Often 1โ3% improvement in competition-style metrics | Many models to train, serve and monitor |
| Robustness: one model's failure mode is diluted | Higher latency and memory at inference |
| Uses complementary strengths | Harder to explain and debug |