๐Ÿ“ˆ Machine Learning ยท Lecture 37 of 47

Feature Selection: Filter, Wrapper and Embedded Methods

More features are not always better. We compare filter methods (correlation, mutual information), wrappers (RFE, sequential selection) and embedded methods (lasso, tree importance), and learn to select without leaking.

Adding irrelevant features makes models slower, harder to interpret, more expensive to maintain and โ€” especially with limited data โ€” less accurate. Feature selection chooses a subset of informative features. Unlike PCA, which creates new combined features, selection keeps original features, preserving interpretability: "the model uses these twelve survey questions" is easier to explain than "the model uses principal component 7".

Why select features?#

  • Generalisation โ€” fewer irrelevant inputs reduce variance and the curse of dimensionality.
  • Interpretability โ€” simpler models are easier to explain and audit.
  • Cost โ€” each feature may cost money to collect (a lab test, a survey question, a sensor).
  • Speed โ€” training and inference get faster.
  • Robustness โ€” fewer features means fewer things to break when data pipelines change.

Three families of methods#

1. Filter methods#

Score each feature independently of any model, then keep the top ones.

  • Variance threshold โ€” remove near-constant features.
  • Correlation with the target (linear relationships only).
  • Statistical tests โ€” ANOVA F-test for numeric features vs class; chi-squared for count features vs class.
  • Mutual information โ€” captures non-linear dependence.

Fast and model-agnostic, but they ignore interactions (a feature useless alone may be powerful in combination) and redundancy (ten copies of the same informative feature all score highly).

2. Wrapper methods#

Search over feature subsets using a model's validated performance as the score.

  • Sequential forward selection โ€” start empty; repeatedly add the feature that improves CV score most.
  • Sequential backward elimination โ€” start with all; repeatedly remove the least useful.
  • Recursive Feature Elimination (RFE) โ€” fit a model, drop the least important features by coefficient or importance, repeat. RFECV chooses the number of features by cross-validation.

Wrappers account for interactions and the specific model, but are computationally expensive ($O(d^2)$ model fits for sequential search) and can overfit the validation data when many subsets are compared.

3. Embedded methods#

Selection happens during training.

  • Lasso / L1-regularised models drive irrelevant coefficients to exactly zero.
  • Tree-based importance from random forests or gradient boosting.
  • Elastic net for correlated feature groups.

Efficient and usually a good compromise.

Comparing methods#

python
import numpy as np
from sklearn.datasets import make_classification
from sklearn.feature_selection import (SelectKBest, mutual_info_classif, f_classif, RFECV,
                                       SelectFromModel)
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import cross_val_score, StratifiedKFold

# 8 informative, 4 redundant, 88 pure-noise features
X, y = make_classification(n_samples=600, n_features=100, n_informative=8, n_redundant=4,
                           n_repeated=0, shuffle=False, random_state=0)
cv = StratifiedKFold(5, shuffle=True, random_state=0)
base = LogisticRegression(max_iter=3000)

pipes = {
    "all features":   make_pipeline(StandardScaler(), base),
    "filter: F-test": make_pipeline(StandardScaler(), SelectKBest(f_classif, k=12), base),
    "filter: MI":     make_pipeline(StandardScaler(), SelectKBest(mutual_info_classif, k=12), base),
    "embedded: L1":   make_pipeline(StandardScaler(),
                                    SelectFromModel(LogisticRegression(penalty="l1", C=0.1, solver="liblinear")),
                                    base),
    "wrapper: RFECV": make_pipeline(StandardScaler(), RFECV(LogisticRegression(max_iter=3000), step=5, cv=3), base),
}
for name, p in pipes.items():
    print(f"{name:<16} CV accuracy = {cross_val_score(p, X, y, cv=cv).mean():.3f}")

The first 12 columns are the useful ones (because shuffle=False); inspecting which features each method keeps shows how well it recovers them.

The selection-bias trap#

Stability#

Selected subsets can change dramatically with small data perturbations, especially with correlated features. Stability selection (Meinshausen & Bรผhlmann) runs lasso on many bootstrap subsamples and keeps features selected in a large fraction of runs โ€” a more trustworthy basis for scientific claims about which variables matter.

Practical guidance#

SituationSuggested approach
Thousands of features, quick reductionVariance threshold + univariate filter (MI)
Linear model, want sparsityLasso / elastic net
Tree modelsUsually no selection needed; prune with permutation importance if desired
Few features, compute availableRFECV or sequential selection
Scientific claims about variablesStability selection; report uncertainty
Features are costly to collectWrapper with a cost-aware objective
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ“ˆ Machine Learning

Regularisation: Ridge, Lasso and Elastic Net

Penalising large weights tames overfitting. We derive ridge regression's closed form, explain why lasso yields sparse models, combine them in elastic net, and tune the penalty by cross-validation.

Intermediateโฑ 5 min#055
๐Ÿ“ˆ Machine Learning

Learning from Imbalanced Data

When one class is rare, naive models ignore it. We cover the right metrics, class weighting, over- and under-sampling, SMOTE, threshold moving and calibration โ€” and when each is appropriate.

Intermediateโฑ 5 min#085
๐Ÿ“ˆ Machine Learning

Hyperparameter Tuning: Grid, Random, Bayesian and Early-Stopping Methods

Hyperparameters control how models learn. We compare grid search, random search, Bayesian optimisation and successive halving, explain why random search beats grid search, and tune efficiently with Optuna.

Intermediateโฑ 5 min#087