Adding irrelevant features makes models slower, harder to interpret, more expensive to maintain and โ especially with limited data โ less accurate. Feature selection chooses a subset of informative features. Unlike PCA, which creates new combined features, selection keeps original features, preserving interpretability: "the model uses these twelve survey questions" is easier to explain than "the model uses principal component 7".
Why select features?#
- Generalisation โ fewer irrelevant inputs reduce variance and the curse of dimensionality.
- Interpretability โ simpler models are easier to explain and audit.
- Cost โ each feature may cost money to collect (a lab test, a survey question, a sensor).
- Speed โ training and inference get faster.
- Robustness โ fewer features means fewer things to break when data pipelines change.
Three families of methods#
1. Filter methods#
Score each feature independently of any model, then keep the top ones.
- Variance threshold โ remove near-constant features.
- Correlation with the target (linear relationships only).
- Statistical tests โ ANOVA F-test for numeric features vs class; chi-squared for count features vs class.
- Mutual information โ captures non-linear dependence.
Fast and model-agnostic, but they ignore interactions (a feature useless alone may be powerful in combination) and redundancy (ten copies of the same informative feature all score highly).
2. Wrapper methods#
Search over feature subsets using a model's validated performance as the score.
- Sequential forward selection โ start empty; repeatedly add the feature that improves CV score most.
- Sequential backward elimination โ start with all; repeatedly remove the least useful.
- Recursive Feature Elimination (RFE) โ fit a model, drop the least important features by coefficient or importance, repeat. RFECV chooses the number of features by cross-validation.
Wrappers account for interactions and the specific model, but are computationally expensive ($O(d^2)$ model fits for sequential search) and can overfit the validation data when many subsets are compared.
3. Embedded methods#
Selection happens during training.
- Lasso / L1-regularised models drive irrelevant coefficients to exactly zero.
- Tree-based importance from random forests or gradient boosting.
- Elastic net for correlated feature groups.
Efficient and usually a good compromise.
Comparing methods#
import numpy as np
from sklearn.datasets import make_classification
from sklearn.feature_selection import (SelectKBest, mutual_info_classif, f_classif, RFECV,
SelectFromModel)
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import cross_val_score, StratifiedKFold
# 8 informative, 4 redundant, 88 pure-noise features
X, y = make_classification(n_samples=600, n_features=100, n_informative=8, n_redundant=4,
n_repeated=0, shuffle=False, random_state=0)
cv = StratifiedKFold(5, shuffle=True, random_state=0)
base = LogisticRegression(max_iter=3000)
pipes = {
"all features": make_pipeline(StandardScaler(), base),
"filter: F-test": make_pipeline(StandardScaler(), SelectKBest(f_classif, k=12), base),
"filter: MI": make_pipeline(StandardScaler(), SelectKBest(mutual_info_classif, k=12), base),
"embedded: L1": make_pipeline(StandardScaler(),
SelectFromModel(LogisticRegression(penalty="l1", C=0.1, solver="liblinear")),
base),
"wrapper: RFECV": make_pipeline(StandardScaler(), RFECV(LogisticRegression(max_iter=3000), step=5, cv=3), base),
}
for name, p in pipes.items():
print(f"{name:<16} CV accuracy = {cross_val_score(p, X, y, cv=cv).mean():.3f}")The first 12 columns are the useful ones (because shuffle=False); inspecting which features each method keeps shows how well it recovers them.
The selection-bias trap#
Stability#
Selected subsets can change dramatically with small data perturbations, especially with correlated features. Stability selection (Meinshausen & Bรผhlmann) runs lasso on many bootstrap subsamples and keeps features selected in a large fraction of runs โ a more trustworthy basis for scientific claims about which variables matter.
Practical guidance#
| Situation | Suggested approach |
|---|---|
| Thousands of features, quick reduction | Variance threshold + univariate filter (MI) |
| Linear model, want sparsity | Lasso / elastic net |
| Tree models | Usually no selection needed; prune with permutation importance if desired |
| Few features, compute available | RFECV or sequential selection |
| Scientific claims about variables | Stability selection; report uncertainty |
| Features are costly to collect | Wrapper with a cost-aware objective |