📈 Machine Learning · Lecture 43 of 47

Time Series Forecasting: Stationarity, ARIMA and Evaluation

Forecasting demand, rainfall or arrivals needs models that respect time. We decompose series into trend and seasonality, test for stationarity, build ARIMA and SARIMA models, and evaluate with rolling-origin backtesting.

Forecasting is everywhere: electricity demand, crop prices, hospital admissions, the number of new arrivals a reception centre should prepare for next month. Time series differ from ordinary tabular data in one crucial way: observations are ordered and dependent. Shuffling them destroys information, and random train–test splits leak the future. Today we learn the classical statistical approach — still a strong baseline that every forecaster must know.

Components of a time series#

A series $y_t$ is often described as a combination of:

  • Trend — long-term increase or decrease;
  • Seasonality — regular patterns with a fixed period (weekly, yearly);
  • Cycles — longer, irregular rises and falls (economic cycles);
  • Noise — the unpredictable remainder.

Decomposition (e.g. STL: Seasonal–Trend decomposition using Loess) separates these components and is the first plot to make for any new series.

Stationarity#

A series is (weakly) stationary if its mean, variance and autocovariance do not change over time. Most classical models assume stationarity. Trends and seasonality violate it.

Tools:

  • Plots of the series and its rolling mean/variance.
  • Augmented Dickey–Fuller (ADF) test — null hypothesis: a unit root (non-stationary).
  • Differencing — $\nabla y_t = y_t - y_{t-1}$ removes trends; seasonal differencing $y_t - y_{t-s}$ removes seasonality of period $s$.
  • Log or Box–Cox transform — stabilises growing variance.

Autocorrelation#

The autocorrelation function (ACF) measures correlation between $y_t$ and $y_{t-k}$ for each lag $k$. The partial autocorrelation function (PACF) measures correlation at lag $k$ after removing the effect of shorter lags. Their patterns guide model choice.

The ARIMA family#

AR($p$) — autoregressive: today depends linearly on the last $p$ values:

$$ y_t = c + \phi_1y_{t-1} + \dots + \phi_py_{t-p} + \varepsilon_t $$

MA($q$) — moving average: today depends on the last $q$ forecast errors:

$$ y_t = c + \varepsilon_t + \theta_1\varepsilon_{t-1} + \dots + \theta_q\varepsilon_{t-q} $$

ARIMA($p, d, q$) applies an ARMA($p, q$) model to the series differenced $d$ times. SARIMA($p,d,q$)($P,D,Q$)$_s$ adds seasonal AR, differencing and MA terms at lag $s$. Adding external regressors (holidays, rainfall, prices) gives SARIMAX.

Identification heuristics:

PatternSuggests
PACF cuts off after lag $p$, ACF decaysAR($p$)
ACF cuts off after lag $q$, PACF decaysMA($q$)
Both decayARMA; select by AIC

In practice, fit a few candidates and choose by AIC/BIC, or use automatic search (e.g. pmdarima.auto_arima). Then check that residuals look like white noise (Ljung–Box test, ACF of residuals).

python
import numpy as np
import pandas as pd
from statsmodels.tsa.statespace.sarimax import SARIMAX
from statsmodels.tsa.stattools import adfuller

rng = np.random.default_rng(0)
t = np.arange(120)                                  # 10 years of monthly data
y = 200 + 1.5 * t + 25 * np.sin(2 * np.pi * t / 12) + rng.normal(0, 8, 120)
series = pd.Series(y, index=pd.date_range("2016-01-01", periods=120, freq="MS"))

print("ADF p-value (raw):", round(adfuller(series)[1], 3))
print("ADF p-value (diff):", round(adfuller(series.diff().dropna())[1], 3))

train, test = series[:-12], series[-12:]
model = SARIMAX(train, order=(1, 1, 1), seasonal_order=(1, 1, 1, 12)).fit(disp=False)
fc = model.get_forecast(12)
pred, ci = fc.predicted_mean, fc.conf_int(alpha=0.05)

naive = train.iloc[-12:].values                     # seasonal naive: same month last year
mae = lambda a, b: np.mean(np.abs(np.asarray(a) - np.asarray(b)))
print(f"SARIMA MAE: {mae(test, pred):.2f}   seasonal-naive MAE: {mae(test, naive):.2f}")
print("95% interval for first forecast month:", ci.iloc[0].round(1).tolist())

Evaluating forecasts properly#

Always compare against simple baselines: the naive forecast (last value), the seasonal naive forecast (same period last season) and a moving average. Surprisingly often, sophisticated models barely beat them. Metrics: MAE, RMSE, and scale-free MASE (error relative to the naive forecast) for comparing across series.

Beyond ARIMA#

  • Exponential smoothing (ETS / Holt–Winters) — weighted averages with trend and seasonality; excellent for many business series.
  • Prophet — additive model with trend changepoints and holidays; easy to use.
  • Machine learning on lag features — gradient boosting with lags, rolling statistics and calendar features; strong when many related series are available (global models).
  • Deep learning — RNNs, temporal convolution, transformer forecasters and pretrained time-series foundation models.

Large forecasting competitions (the M-competitions) have repeatedly shown that combinations of methods and well-tuned global ML models perform very well, while simple statistical methods remain hard to beat for individual short series.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

📈 Machine Learning

Recommender Systems II: Matrix Factorisation and Latent Factors

Matrix factorisation represents users and items as vectors in a shared latent space. We derive the regularised objective, train it with SGD and ALS, add biases and implicit feedback, and connect it to modern embedding models.

Advanced⏱ 5 min#091
📈 Machine Learning

Semi-Supervised Learning: Learning from Few Labels and Many Unlabelled Examples

Labels are expensive; unlabelled data is cheap. We study the assumptions that make unlabelled data useful and the main techniques — self-training, label propagation, consistency regularisation and FixMatch.

Intermediate⏱ 5 min#093
📈 Machine Learning

Recommender Systems I: Content-Based and Collaborative Filtering

Recommendation engines drive what billions of people watch, read and buy. We compare content-based and collaborative approaches, build user- and item-based neighbourhood models, and confront cold start and popularity bias.

Intermediate⏱ 5 min#090