📈 Machine Learning · Lecture 10 of 47

Overfitting and Underfitting: Diagnosis with Learning Curves

Before fixing a model you must diagnose it. We learn to read learning curves and validation curves, recognise high bias and high variance, and choose the right remedy instead of guessing.

Andrew Ng often tells students that the most valuable ML skill is knowing what to try next. When a model disappoints, should you collect more data, add features, use a bigger model, or regularise more? Each option costs time and money, and choosing wrongly can waste months. Learning curves are the diagnostic instrument that tells you which option will help.

Symptoms#

DiagnosisTraining errorValidation errorGap
Underfitting (high bias)HighHighSmall
Overfitting (high variance)LowHighLarge
Good fitLowLow (close to target)Small

You need a target to judge "high": human-level performance, a published result, or the business requirement. The gap between training error and the target is avoidable bias; the gap between training and validation error is variance.

Learning curves#

A learning curve plots training and validation error against the number of training examples.

  • High bias: both curves converge quickly to a high error plateau. Adding data does not help — the curves are already flat and close together.
  • High variance: training error is low, validation error is much higher, and the gap narrows slowly as data increases. More data will help.
python
import numpy as np
import matplotlib.pyplot as plt
from sklearn.model_selection import learning_curve
from sklearn.datasets import load_digits
from sklearn.naive_bayes import GaussianNB
from sklearn.svm import SVC

X, y = load_digits(return_X_y=True)
fig, axes = plt.subplots(1, 2, figsize=(11, 4))
for ax, (name, model) in zip(axes, [("Gaussian Naive Bayes (high bias)", GaussianNB()),
                                    ("RBF SVM, gamma=0.01 (high variance)", SVC(gamma=0.01))]):
    sizes, tr, va = learning_curve(model, X, y, cv=5, train_sizes=np.linspace(0.1, 1.0, 8))
    ax.plot(sizes, 1 - tr.mean(1), "o-", label="training error")
    ax.plot(sizes, 1 - va.mean(1), "o-", label="validation error")
    ax.set_title(name); ax.set_xlabel("training examples"); ax.legend()
plt.tight_layout(); plt.show()

Validation curves#

A validation curve plots training and validation error against a hyperparameter that controls complexity (tree depth, regularisation strength, polynomial degree).

  • On the low-complexity side, both errors are high → underfitting.
  • On the high-complexity side, training error keeps falling while validation error rises → overfitting.
  • Choose the value at the minimum of validation error.
python
from sklearn.model_selection import validation_curve
from sklearn.tree import DecisionTreeClassifier

depths = range(1, 21)
tr, va = validation_curve(DecisionTreeClassifier(random_state=0), X, y,
                          param_name="max_depth", param_range=depths, cv=5)
best = list(depths)[va.mean(1).argmax()]
print("best depth by validation:", best)

Choosing the remedy#

If you have…TryDo NOT waste time on
High biasBigger/more flexible model; more or better features; less regularisation; train longer; better architectureCollecting more data
High varianceMore data; data augmentation; regularisation; dropout; simpler model; feature selection; ensembling; early stoppingBigger model (usually)
BothAddress bias first, then variance—
Train/validation mismatch (different distributions)Make training data more like deployment data; domain adaptationTuning on the wrong distribution

Loss curves during training#

For iterative learners (neural networks, boosting), plot training and validation loss against epochs:

  • Both decreasing → keep training.
  • Validation loss starts rising while training loss keeps falling → overfitting begins; early stopping saves the best checkpoint.
  • Training loss not decreasing → optimisation problem (learning rate, bugs, bad initialisation), not a generalisation problem.
  • Validation loss lower than training loss → possible if training uses dropout/augmentation, or a sign of leakage or an unrepresentative split.

A diagnostic checklist#

  1. Can the model overfit a tiny subset (e.g. 50 examples) to near-zero training error? If not, there is a bug or the model is too weak.
  2. Compare training error with the target → avoidable bias.
  3. Compare validation with training error → variance.
  4. Plot learning curves → will more data help?
  5. Perform error analysis on validation mistakes → which categories of errors dominate?
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

📈 Machine Learning

Regularisation: Ridge, Lasso and Elastic Net

Penalising large weights tames overfitting. We derive ridge regression's closed form, explain why lasso yields sparse models, combine them in elastic net, and tune the penalty by cross-validation.

Intermediate⏱ 5 min#055
📈 Machine Learning

Polynomial Regression and Basis Functions: Non-Linearity with Linear Models

Linear models can fit curves if we transform the inputs. We study polynomial, spline and radial basis features, watch overfitting happen as degree grows, and connect it to model selection.

Beginner⏱ 5 min#054
📈 Machine Learning

The Bias–Variance Trade-off: Derivation and Intuition

Why do simple models underfit and complex models overfit? We derive the bias–variance decomposition of expected squared error, visualise it, and discuss how modern deep learning complicates the classical picture.

Intermediate⏱ 5 min#058