In the last lecture a degree-14 polynomial fit the training points perfectly and produced absurd predictions โ with gigantic coefficients. Regularisation attacks the problem directly: we add a penalty that discourages large or numerous weights. It is the most widely used defence against overfitting in all of machine learning, from linear models to billion-parameter networks (where it appears as weight decay).
The regularised objective#
The hyperparameter $\lambda \ge 0$ controls the trade-off. $\lambda = 0$ gives ordinary least squares; $\lambda \to \infty$ forces the weights to zero. By convention, the intercept is not penalised.
Ridge regression (L2)#
$\Omega(\mathbf{w}) = \|\mathbf{w}\|_2^2 = \sum_j w_j^2$. The closed-form solution is
(with $\lambda' = n\lambda$ for our scaling). Adding $\lambda'\mathbf{I}$ makes the matrix always invertible and well-conditioned โ ridge was originally invented (Hoerl & Kennard, 1970) precisely to stabilise regression with correlated features.
SVD view. With $\mathbf{X} = \mathbf{U}\boldsymbol{\Sigma}\mathbf{V}^\top$, ridge shrinks the component along each singular direction by the factor
Directions with large singular values (strongly supported by data) are barely shrunk; directions with small singular values (poorly determined, noise-prone) are shrunk heavily. That is exactly the right behaviour.
Ridge shrinks all coefficients towards zero but rarely makes any exactly zero.
Lasso (L1)#
$\Omega(\mathbf{w}) = \|\mathbf{w}\|_1 = \sum_j |w_j|$. Lasso (Tibshirani, 1996 โ "least absolute shrinkage and selection operator") has no closed form but is convex; it is solved by coordinate descent.
Its defining property is sparsity: many coefficients become exactly zero, performing automatic feature selection. For a single coefficient with orthonormal features, the lasso solution is the soft-thresholding operator:
Any OLS coefficient smaller than the threshold is set to zero; larger ones are shrunk by a constant. Compare ridge, which scales every coefficient by the same factor $1/(1 + \lambda')$ and never reaches zero.
Elastic net#
Lasso has two weaknesses: with groups of highly correlated features it tends to pick one arbitrarily, and when $d > n$ it selects at most $n$ features. Elastic net (Zou & Hastie, 2005) mixes both penalties:
It keeps lasso's sparsity while selecting correlated features together, and is a robust default for high-dimensional data such as genomics or text.
Comparison#
| Ridge | Lasso | Elastic net | |
|---|---|---|---|
| Penalty | $\|\mathbf{w}\|_2^2$ | $\|\mathbf{w}\|_1$ | mix |
| Closed form | Yes | No | No |
| Sparse solution | No | Yes | Yes |
| Correlated features | Shares weight among them | Picks one | Groups them |
| Bayesian prior | Gaussian | Laplace | โ |
Practical rules#
- Standardise features before regularising. Penalties treat all coefficients equally, so scale matters.
- Tune $\lambda$ by cross-validation over a logarithmic grid (e.g. $10^{-4}$ to $10^{2}$).
- Consider the one-standard-error rule: choose the simplest model whose CV error is within one standard error of the minimum.
import numpy as np
from sklearn.datasets import make_regression
from sklearn.linear_model import RidgeCV, LassoCV, ElasticNetCV, LinearRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import cross_val_score
# 100 samples, 200 features, only 10 truly informative
X, y, coef = make_regression(n_samples=100, n_features=200, n_informative=10,
noise=10, coef=True, random_state=0)
alphas = np.logspace(-3, 3, 50)
models = {
"OLS": LinearRegression(),
"Ridge": RidgeCV(alphas=alphas),
"Lasso": LassoCV(alphas=alphas, max_iter=20000, cv=5),
"ElasticNet": ElasticNetCV(l1_ratio=[.2, .5, .8], alphas=alphas, max_iter=20000, cv=5),
}
for name, m in models.items():
pipe = make_pipeline(StandardScaler(), m)
r2 = cross_val_score(pipe, X, y, cv=5, scoring="r2").mean()
pipe.fit(X, y)
nz = int((np.abs(pipe[-1].coef_) > 1e-6).sum())
print(f"{name:<10} CV R^2 = {r2:6.3f} non-zero coefficients = {nz}")With 200 features and only 100 samples, OLS overfits badly, ridge helps, and lasso/elastic net do best while keeping a small number of features โ close to the 10 that truly matter.
Regularisation beyond linear models#
The same idea appears everywhere: weight decay in neural networks, the $C$ parameter in SVMs and logistic regression (inverse regularisation strength), tree depth limits, dropout, early stopping, and data augmentation. All express a preference for simpler explanations โ a mathematical Occam's razor.