When predicting numbers — prices, temperatures, travel times, food-aid quantities — the question "how accurate is the model?" has several answers. Different metrics punish different kinds of error. Choosing the wrong one can make a model look good while it fails exactly where it matters.
Mean Squared Error and RMSE#
- Squaring punishes large errors heavily: one error of 10 costs as much as 100 errors of 1.
- RMSE is in the same units as the target, making it interpretable ("typically off by about 3.2 °C").
- The prediction minimising expected squared error is the conditional mean $\mathbb{E}[y \mid \mathbf{x}]$.
- Sensitive to outliers — a few extreme values can dominate.
Mean Absolute Error#
- Treats all errors linearly; more robust to outliers.
- Minimised by the conditional median.
- Also in target units; often easier to explain to non-technical stakeholders.
Coefficient of determination R²#
$R^2$ compares the model with the trivial predictor that always outputs the mean.
- $R^2 = 1$: perfect predictions.
- $R^2 = 0$: no better than the mean.
- $R^2 < 0$: worse than the mean — possible on test data.
Caveats: $R^2$ depends on the variance of the target in the evaluation set (a narrow test range lowers $R^2$ even for a good model), it never decreases when features are added on the training set (use adjusted $R^2$ or validation data), and a high $R^2$ does not mean predictions are accurate enough for the decision at hand.
Percentage errors#
MAPE (Mean Absolute Percentage Error):
Scale-free and popular in business forecasting, but it explodes when $y_i$ is near zero and penalises over-prediction more than under-prediction. Alternatives: sMAPE (symmetric), WAPE $\frac{\sum|y_i - \hat{y}_i|}{\sum|y_i|}$ (weighted — robust for intermittent demand), and MASE (scaled by a naive forecast's error — recommended for comparing forecasts across series).
Log-scale errors#
For targets spanning orders of magnitude (population, income, counts), use RMSLE:
It measures relative error: predicting 110 for 100 costs about the same as predicting 1,100 for 1,000.
Quantile (pinball) loss and prediction intervals#
Often a single number is not enough — a planner needs to know "how much stock covers 90% of scenarios?". The quantile loss for quantile $\tau$ is
Minimising it yields the $\tau$-th conditional quantile. Training models for $\tau = 0.05$ and $0.95$ gives a 90% prediction interval, which you evaluate by coverage (does 90% of the truth fall inside?) and width.
Comparing metrics on data with outliers#
import numpy as np
from sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score
rng = np.random.default_rng(0)
y = rng.normal(100, 15, 500)
pred_a = y + rng.normal(0, 5, 500) # consistently a little off
pred_b = y.copy(); pred_b[:5] += 120 # perfect except five huge errors
for name, p in [("A: small errors everywhere", pred_a), ("B: few huge errors", pred_b)]:
rmse = mean_squared_error(y, p) ** 0.5
mae = mean_absolute_error(y, p)
print(f"{name:<28} RMSE={rmse:6.2f} MAE={mae:5.2f} R2={r2_score(y, p):.3f}")RMSE prefers model A; MAE prefers model B. Which is better depends on the application: for aid delivery, five catastrophic under-estimates might mean five communities without supplies — the RMSE view matters. For a recommendation estimate, occasional large misses may be tolerable.
Choosing a metric#
| Situation | Recommended metric |
|---|---|
| Large errors are disproportionately harmful | RMSE / MSE |
| Outliers in target; want typical error | MAE |
| Communicate improvement over the mean baseline | $R^2$ (alongside an absolute metric) |
| Targets span orders of magnitude | RMSLE, or MAE on log-scale |
| Forecasts across many series of different scale | MASE, WAPE |
| Decisions need uncertainty | Quantile loss, interval coverage and width |