๐Ÿ“ˆ Machine Learning ยท Lecture 6 of 47

Regularisation: Ridge, Lasso and Elastic Net

Penalising large weights tames overfitting. We derive ridge regression's closed form, explain why lasso yields sparse models, combine them in elastic net, and tune the penalty by cross-validation.

In the last lecture a degree-14 polynomial fit the training points perfectly and produced absurd predictions โ€” with gigantic coefficients. Regularisation attacks the problem directly: we add a penalty that discourages large or numerous weights. It is the most widely used defence against overfitting in all of machine learning, from linear models to billion-parameter networks (where it appears as weight decay).

The regularised objective#

$$ J(\mathbf{w}) = \underbrace{\frac{1}{n}\|\mathbf{y} - \mathbf{X}\mathbf{w}\|^2}_{\text{fit the data}} + \underbrace{\lambda\,\Omega(\mathbf{w})}_{\text{keep it simple}} $$

The hyperparameter $\lambda \ge 0$ controls the trade-off. $\lambda = 0$ gives ordinary least squares; $\lambda \to \infty$ forces the weights to zero. By convention, the intercept is not penalised.

Ridge regression (L2)#

$\Omega(\mathbf{w}) = \|\mathbf{w}\|_2^2 = \sum_j w_j^2$. The closed-form solution is

$$ \hat{\mathbf{w}}_{\text{ridge}} = (\mathbf{X}^\top\mathbf{X} + \lambda'\mathbf{I})^{-1}\mathbf{X}^\top\mathbf{y} $$

(with $\lambda' = n\lambda$ for our scaling). Adding $\lambda'\mathbf{I}$ makes the matrix always invertible and well-conditioned โ€” ridge was originally invented (Hoerl & Kennard, 1970) precisely to stabilise regression with correlated features.

SVD view. With $\mathbf{X} = \mathbf{U}\boldsymbol{\Sigma}\mathbf{V}^\top$, ridge shrinks the component along each singular direction by the factor

$$ \frac{\sigma_j^2}{\sigma_j^2 + \lambda'} $$

Directions with large singular values (strongly supported by data) are barely shrunk; directions with small singular values (poorly determined, noise-prone) are shrunk heavily. That is exactly the right behaviour.

Ridge shrinks all coefficients towards zero but rarely makes any exactly zero.

Lasso (L1)#

$\Omega(\mathbf{w}) = \|\mathbf{w}\|_1 = \sum_j |w_j|$. Lasso (Tibshirani, 1996 โ€” "least absolute shrinkage and selection operator") has no closed form but is convex; it is solved by coordinate descent.

Its defining property is sparsity: many coefficients become exactly zero, performing automatic feature selection. For a single coefficient with orthonormal features, the lasso solution is the soft-thresholding operator:

$$ \hat{w}_j = \text{sign}(w_j^{\text{OLS}})\,\max\big(|w_j^{\text{OLS}}| - \tfrac{\lambda'}{2},\; 0\big) $$

Any OLS coefficient smaller than the threshold is set to zero; larger ones are shrunk by a constant. Compare ridge, which scales every coefficient by the same factor $1/(1 + \lambda')$ and never reaches zero.

Elastic net#

Lasso has two weaknesses: with groups of highly correlated features it tends to pick one arbitrarily, and when $d > n$ it selects at most $n$ features. Elastic net (Zou & Hastie, 2005) mixes both penalties:

$$ \Omega(\mathbf{w}) = \alpha\|\mathbf{w}\|_1 + \frac{1 - \alpha}{2}\|\mathbf{w}\|_2^2 $$

It keeps lasso's sparsity while selecting correlated features together, and is a robust default for high-dimensional data such as genomics or text.

Comparison#

RidgeLassoElastic net
Penalty$\|\mathbf{w}\|_2^2$$\|\mathbf{w}\|_1$mix
Closed formYesNoNo
Sparse solutionNoYesYes
Correlated featuresShares weight among themPicks oneGroups them
Bayesian priorGaussianLaplaceโ€”

Practical rules#

  1. Standardise features before regularising. Penalties treat all coefficients equally, so scale matters.
  2. Tune $\lambda$ by cross-validation over a logarithmic grid (e.g. $10^{-4}$ to $10^{2}$).
  3. Consider the one-standard-error rule: choose the simplest model whose CV error is within one standard error of the minimum.
python
import numpy as np
from sklearn.datasets import make_regression
from sklearn.linear_model import RidgeCV, LassoCV, ElasticNetCV, LinearRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import cross_val_score

# 100 samples, 200 features, only 10 truly informative
X, y, coef = make_regression(n_samples=100, n_features=200, n_informative=10,
                             noise=10, coef=True, random_state=0)
alphas = np.logspace(-3, 3, 50)
models = {
    "OLS": LinearRegression(),
    "Ridge": RidgeCV(alphas=alphas),
    "Lasso": LassoCV(alphas=alphas, max_iter=20000, cv=5),
    "ElasticNet": ElasticNetCV(l1_ratio=[.2, .5, .8], alphas=alphas, max_iter=20000, cv=5),
}
for name, m in models.items():
    pipe = make_pipeline(StandardScaler(), m)
    r2 = cross_val_score(pipe, X, y, cv=5, scoring="r2").mean()
    pipe.fit(X, y)
    nz = int((np.abs(pipe[-1].coef_) > 1e-6).sum())
    print(f"{name:<10} CV R^2 = {r2:6.3f}   non-zero coefficients = {nz}")

With 200 features and only 100 samples, OLS overfits badly, ridge helps, and lasso/elastic net do best while keeping a small number of features โ€” close to the 10 that truly matter.

Regularisation beyond linear models#

The same idea appears everywhere: weight decay in neural networks, the $C$ parameter in SVMs and logistic regression (inverse regularisation strength), tree depth limits, dropout, early stopping, and data augmentation. All express a preference for simpler explanations โ€” a mathematical Occam's razor.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ“ˆ Machine Learning

Polynomial Regression and Basis Functions: Non-Linearity with Linear Models

Linear models can fit curves if we transform the inputs. We study polynomial, spline and radial basis features, watch overfitting happen as degree grows, and connect it to model selection.

Beginnerโฑ 5 min#054
๐Ÿ“ˆ Machine Learning

Feature Selection: Filter, Wrapper and Embedded Methods

More features are not always better. We compare filter methods (correlation, mutual information), wrappers (RFE, sequential selection) and embedded methods (lasso, tree importance), and learn to select without leaking.

Intermediateโฑ 5 min#086
๐Ÿ“ˆ Machine Learning

Overfitting and Underfitting: Diagnosis with Learning Curves

Before fixing a model you must diagnose it. We learn to read learning curves and validation curves, recognise high bias and high variance, and choose the right remedy instead of guessing.

Beginnerโฑ 5 min#059