📈 Machine Learning · Lecture 14 of 47

Regression Metrics: MSE, RMSE, MAE, R² and Beyond

How good is a numeric prediction? We compare MSE, RMSE, MAE, R², MAPE and quantile loss, explain how each responds to outliers and scale, and match metrics to real decisions.

When predicting numbers — prices, temperatures, travel times, food-aid quantities — the question "how accurate is the model?" has several answers. Different metrics punish different kinds of error. Choosing the wrong one can make a model look good while it fails exactly where it matters.

Mean Squared Error and RMSE#

$$ \text{MSE} = \frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2, \qquad \text{RMSE} = \sqrt{\text{MSE}} $$
  • Squaring punishes large errors heavily: one error of 10 costs as much as 100 errors of 1.
  • RMSE is in the same units as the target, making it interpretable ("typically off by about 3.2 °C").
  • The prediction minimising expected squared error is the conditional mean $\mathbb{E}[y \mid \mathbf{x}]$.
  • Sensitive to outliers — a few extreme values can dominate.

Mean Absolute Error#

$$ \text{MAE} = \frac{1}{n}\sum_{i=1}^{n}|y_i - \hat{y}_i| $$
  • Treats all errors linearly; more robust to outliers.
  • Minimised by the conditional median.
  • Also in target units; often easier to explain to non-technical stakeholders.

Coefficient of determination R²#

$$ R^2 = 1 - \frac{\sum_i(y_i - \hat{y}_i)^2}{\sum_i(y_i - \bar{y})^2} $$

$R^2$ compares the model with the trivial predictor that always outputs the mean.

  • $R^2 = 1$: perfect predictions.
  • $R^2 = 0$: no better than the mean.
  • $R^2 < 0$: worse than the mean — possible on test data.

Caveats: $R^2$ depends on the variance of the target in the evaluation set (a narrow test range lowers $R^2$ even for a good model), it never decreases when features are added on the training set (use adjusted $R^2$ or validation data), and a high $R^2$ does not mean predictions are accurate enough for the decision at hand.

Percentage errors#

MAPE (Mean Absolute Percentage Error):

$$ \text{MAPE} = \frac{100\%}{n}\sum_{i=1}^{n}\left|\frac{y_i - \hat{y}_i}{y_i}\right| $$

Scale-free and popular in business forecasting, but it explodes when $y_i$ is near zero and penalises over-prediction more than under-prediction. Alternatives: sMAPE (symmetric), WAPE $\frac{\sum|y_i - \hat{y}_i|}{\sum|y_i|}$ (weighted — robust for intermittent demand), and MASE (scaled by a naive forecast's error — recommended for comparing forecasts across series).

Log-scale errors#

For targets spanning orders of magnitude (population, income, counts), use RMSLE:

$$ \text{RMSLE} = \sqrt{\frac{1}{n}\sum_i\big(\ln(1 + \hat{y}_i) - \ln(1 + y_i)\big)^2} $$

It measures relative error: predicting 110 for 100 costs about the same as predicting 1,100 for 1,000.

Quantile (pinball) loss and prediction intervals#

Often a single number is not enough — a planner needs to know "how much stock covers 90% of scenarios?". The quantile loss for quantile $\tau$ is

$$ L_\tau(y, \hat{y}) = \max\big(\tau(y - \hat{y}),\; (\tau - 1)(y - \hat{y})\big) $$

Minimising it yields the $\tau$-th conditional quantile. Training models for $\tau = 0.05$ and $0.95$ gives a 90% prediction interval, which you evaluate by coverage (does 90% of the truth fall inside?) and width.

Comparing metrics on data with outliers#

python
import numpy as np
from sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score

rng = np.random.default_rng(0)
y = rng.normal(100, 15, 500)
pred_a = y + rng.normal(0, 5, 500)            # consistently a little off
pred_b = y.copy(); pred_b[:5] += 120          # perfect except five huge errors

for name, p in [("A: small errors everywhere", pred_a), ("B: few huge errors", pred_b)]:
    rmse = mean_squared_error(y, p) ** 0.5
    mae = mean_absolute_error(y, p)
    print(f"{name:<28} RMSE={rmse:6.2f}  MAE={mae:5.2f}  R2={r2_score(y, p):.3f}")

RMSE prefers model A; MAE prefers model B. Which is better depends on the application: for aid delivery, five catastrophic under-estimates might mean five communities without supplies — the RMSE view matters. For a recommendation estimate, occasional large misses may be tolerable.

Choosing a metric#

SituationRecommended metric
Large errors are disproportionately harmfulRMSE / MSE
Outliers in target; want typical errorMAE
Communicate improvement over the mean baseline$R^2$ (alongside an absolute metric)
Targets span orders of magnitudeRMSLE, or MAE on log-scale
Forecasts across many series of different scaleMASE, WAPE
Decisions need uncertaintyQuantile loss, interval coverage and width
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

📈 Machine Learning

k-Nearest Neighbours: Learning by Similarity

The simplest learning algorithm stores the data and asks the neighbours. We analyse k-NN's bias–variance behaviour, distance choices, scaling, efficient search structures and its surprising theoretical guarantees.

Beginner⏱ 5 min#064
📈 Machine Learning

Evaluation Metrics for Classification: Accuracy, Precision, Recall and F1

Accuracy can be dangerously misleading. We build the confusion matrix, define precision, recall, specificity, F-scores and balanced accuracy, and learn to choose metrics from the costs of errors.

Beginner⏱ 5 min#061
📈 Machine Learning

Polynomial Regression and Basis Functions: Non-Linearity with Linear Models

Linear models can fit curves if we transform the inputs. We study polynomial, spline and radial basis features, watch overfitting happen as degree grows, and connect it to model selection.

Beginner⏱ 5 min#054