Time Series & Forecasting

60 min intermediate Lesson 5

Learning Outcomes

  • Explain what makes time series different from i.i.d. data, and why that breaks naive train/test splits.
  • Build a baseline forecast and beat it with classical models (exponential smoothing, ARIMA) in statsmodels.
  • Fit a Prophet model on seasonal business data and read its trend and seasonal components.
  • Detect anomalies in a metric stream with a residual rule and an Isolation Forest.
  • Decide when a deep model fits a forecasting problem — and why an LLM is the wrong tool for numbers.

Lesson Plan

Segment Duration Topic
Intro 3 min Why temporal data needs its own toolbox
The shape of the problem 8 min Trend, seasonality, stationarity, evaluation
Baselines & backtesting 9 min Naive baselines, walk-forward validation
Classical models 12 min Exponential smoothing, ARIMA
Prophet 10 min Decomposable trend + seasonality
Anomaly detection 9 min Residual rules, Isolation Forest
Deep models & the LLM trap 6 min When DL helps; why LLMs fail
Wrap-up 3 min Picking the right model

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ in a virtual environment.
  • pip install pandas numpy statsmodels scikit-learn prophet

1 Why time series is its own discipline

Most ML assumes rows are i.i.d. — independent and identically distributed: each row is its own little world, and shuffling rows changes nothing. Time series breaks this on purpose. Time series data is a sequence of measurements indexed by time (sales per day, CPU temperature per second) where order matters and today depends on yesterday.

That dependence has named structure: trend (the long drift up or down), seasonality (patterns repeating on a fixed period, like December spikes), stationarity (mean and variance don't change over time — many classical models assume this), and autocorrelation (correlation with its own past; a spike at lag 7 in daily data means weekly seasonality).

WARNING
The cardinal sin
Never shuffle a time series before splitting train and test. If your model trains on Friday to predict Tuesday, you've leaked the future into the past. Always split by time: train on the earliest data, test on the latest.

Decomposing a series lets you see the structure, using a classic monthly airline-passengers dataset (any monthly series works):

import pandas as pd
from statsmodels.tsa.seasonal import seasonal_decompose

url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
df = pd.read_csv(url, parse_dates=["Month"], index_col="Month")
df = df.asfreq("MS")  # set an explicit frequency or forecasts go silently wrong

seasonal_decompose(df["Passengers"], model="multiplicative", period=12).plot()

The four panels — observed, trend, seasonal, residual — are the mental model for everything below.

TIP
Set the frequency
The asfreq call is not optional — statsmodels and Prophet silently mishandle ambiguous indices and gaps. Always pin an explicit frequency or DatetimeIndex.

2 Baselines and honest backtesting

Build a baseline you must beat: naive (tomorrow equals today) and seasonal-naive (this December equals last December). If a deep model can't beat seasonal-naive, it isn't learning anything.

A single split gives one noisy number. Walk-forward validation (a time-series cross-validation — retraining on a growing window and testing the next chunk) is more honest. Cross-validation just means repeatedly scoring on data the model never trained on.

import numpy as np

def mae(y_true, y_pred):
    return np.mean(np.abs(np.asarray(y_true) - np.asarray(y_pred)))

series = df["Passengers"]
horizon = 12              # forecast one year ahead
train, test = series[:-horizon], series[-horizon:]

seasonal_naive = series.shift(12)[-horizon:]   # value 12 months earlier
print("Seasonal-naive MAE:", round(mae(test, seasonal_naive), 2))

We score with MAE (mean absolute error, the average miss in original units); RMSE punishes large misses harder; MAPE (percentage error) is intuitive but treacherous.

WARNING
MAPE lies on small numbers
MAPE divides by the actual value, so a near-zero actual makes the percentage explode. For demand that includes zeros, prefer MAE or a scaled error like MASE.

3 Exponential smoothing — the workhorse

Exponential smoothing forecasts by weighting recent observations more than old ones, weights shrinking exponentially into the past. Holt-Winters adds explicit trend and seasonality — a fast, hard-to-beat baseline. statsmodels exposes it as ExponentialSmoothing.

from statsmodels.tsa.holtwinters import ExponentialSmoothing

model = ExponentialSmoothing(
    train,
    trend="add",        # additive trend
    seasonal="mul",     # multiplicative seasonality (swing grows with level)
    seasonal_periods=12,
)
fit = model.fit()
forecast = fit.forecast(horizon)
print("Holt-Winters MAE:", round(mae(test, forecast), 2))

Use multiplicative seasonality when swings grow with the level (airline passengers), additive when the swing is roughly constant.

TIP
Start here, not at deep learning
On clean business data, Holt-Winters often matches models a hundred times more complex — milliseconds to train, no GPU, explainable to a finance team. And it beats an LLM cold: it encodes the data's generating structure and emits calibrated numbers, where an LLM only pattern-matches text.

4 ARIMA — modelling the dependence directly

ARIMA is AutoRegressive Integrated Moving Average. AR (p) predicts the next value from p past values. I (d) differences the series d times to make it stationary (removing trend). MA (q) uses the last q forecast errors. For seasonal data, SARIMA adds a (P, D, Q, m) block at lag m; statsmodels implements it as SARIMAX.

from statsmodels.tsa.statespace.sarimax import SARIMAX

model = SARIMAX(
    train,
    order=(1, 1, 1),                 # (p, d, q)
    seasonal_order=(1, 1, 1, 12),    # (P, D, Q, m)
    enforce_stationarity=False,
    enforce_invertibility=False,
)
fit = model.fit(disp=False)

pred = fit.get_forecast(steps=horizon)
forecast = pred.predicted_mean
conf_int = pred.conf_int()           # uncertainty bands, not just a point
print("SARIMA MAE:", round(mae(test, forecast), 2))

Pick (p, d, q) from ACF/PACF plots, or let pmdarima's auto_arima search, then check residuals. Rules of thumb: visible trend means increase d; a spike at the season length means add the seasonal term; autocorrelated residuals means increase p or q; poor on test means simplify.

WARNING
Confidence intervals are not optional
A point forecast of 410 means little without a band — get_forecast gives a prediction interval, so show it. An LLM cannot produce a calibrated interval.

5 Prophet for messy business data

Prophet is an open-source forecasting tool for business time series with strong seasonality, holiday effects, and missing data. It models a series as a sum of three readable pieces — trend, seasonal terms, holiday bumps — so non-experts can read why the forecast moves. It wants two columns, ds (datestamp) and y (value).

from prophet import Prophet

pdf = df.reset_index().rename(columns={"Month": "ds", "Passengers": "y"})
train_pdf = pdf.iloc[:-horizon]

m = Prophet(seasonality_mode="multiplicative", yearly_seasonality=True)
m.fit(train_pdf)

future = m.make_future_dataframe(periods=horizon, freq="MS")
fcst = m.predict(future)
# yhat is the forecast; yhat_lower / yhat_upper are the uncertainty band
print(fcst[["ds", "yhat", "yhat_lower", "yhat_upper"]].tail(horizon))
m.plot_components(fcst)   # trend, yearly seasonality, holidays if present

Add domain knowledge cheaply: call m.add_country_holidays("US") before fitting to capture holiday effects.

TIP
Prophet vs ARIMA
Reach for Prophet with multiple seasonalities, irregular gaps, known holidays, and stakeholders who need interpretable components. Reach for SARIMA when the series is cleaner and you want rigorous control. Let backtesting decide.

6 Anomaly detection on a metric stream

Forecasting and anomaly detection are siblings. Anomaly detection flags points that don't fit the pattern — a fraud spike, a sensor reading before a machine fails (predictive maintenance), a drop in conversions. The lightest approach is a residual rule: fit a forecast, subtract it from reality, and flag points whose error is far outside normal (resid = (fit.fittedvalues - train).dropna(), then where abs(resid) > 3 * resid.std()).

A stronger, feature-based approach is the Isolation Forest — an unsupervised model (no labelled anomalies needed) that isolates outliers by randomly partitioning the feature space, since anomalies separate in fewer cuts. scikit-learn ships it as IsolationForest:

from sklearn.ensemble import IsolationForest

feat = pd.DataFrame({"value": series})
feat["roll_mean"] = series.rolling(12, min_periods=1).mean()
feat["roll_std"] = series.rolling(12, min_periods=1).std().fillna(0)

iso = IsolationForest(contamination=0.02, random_state=42)
feat["flag"] = iso.fit_predict(feat)   # -1 = anomaly, 1 = normal
print(feat[feat["flag"] == -1])

contamination is your prior on how rare anomalies are — set it from domain knowledge. Evaluate with precision (of flagged points, the fraction truly anomalous — matters when false alarms are expensive) and recall (of true anomalies, the fraction caught — matters when misses are costly).

WARNING
Precision/recall is a tradeoff
Lower the threshold and you catch more real anomalies (higher recall) but page people for noise (lower precision). Tune it to the cost of a false alarm versus a missed event in your domain.

7 Deep models — and why LLMs are the wrong tool

Reach for deep learning only when classical models hit a ceiling: thousands of related series forecast jointly (every SKU in every store), rich covariates (weather, price, promotions), or long nonlinear dependencies. The workhorse is the LSTM (a recurrent network carrying a hidden "memory" across timesteps), alongside temporal convolutional networks and transformers. You rarely build these by hand — libraries like PyTorch Forecasting and Darts wrap them:

from pytorch_forecasting import TemporalFusionTransformer, TimeSeriesDataSet
# Build a TimeSeriesDataSet from a long-format DataFrame, then call
# TemporalFusionTransformer.from_dataset(...) and train via Lightning.

Now the headline: why are LLMs poor at numerical forecasting? They model tokens, not numbers — "412" is characters, not a quantity to difference or extrapolate. They have no parameters fitted to your series and no calibrated uncertainty, drift on long horizons, invent seasonality that isn't there, and cost far more than a millisecond SARIMA fit.

Situation Right tool
One clean monthly/weekly series Holt-Winters or SARIMA
Seasonality + holidays + interpretability Prophet
Thousands of related series, rich covariates Deep model (LSTM/TCN/transformer)
Outliers in a stream Residual rule or Isolation Forest
"Summarise why sales dipped" An LLM (reasoning around the numbers)
NOTE
The honest division of labour
Use a numerical model to produce the forecast and interval; use an LLM to explain the result or query the data in natural language. The LLM narrates; it does not forecast — the hybrid is what Lesson 9 builds on. Don't skip baselines: deep forecasters often lose to seasonal-naive on a single short series.

Questions & Answers

Q: My boss said to "just use the LLM" for our demand forecast. How do I push back?
Run the experiment instead of arguing. Fit Holt-Winters or Prophet, backtest with walk-forward MAE, and compare it to the LLM on the same held-out months. You'll almost always show lower error, calibrated intervals, and near-zero cost — a chart, not an opinion. Then offer the hybrid: numbers from the forecaster, narrative from the LLM.
Q: I split my data 80/20 randomly and got great scores, but production is terrible. Why?
You leaked the future. Random splitting lets the model train on data that comes after your test points, so it "saw the answers." Split by time and validate walk-forward — the gap between your score and reality is that leakage.
Q: Prophet and SARIMA disagree a lot on the same series. Which do I trust?
Neither, until you backtest. Run both through identical walk-forward validation and let held-out error decide. Disagreement means one model captures structure the other misses — inspect Prophet's components and SARIMA's residuals. An ensemble (average of the two) often beats either alone.
Q: My anomaly detector floods the on-call channel with false positives. Fix?
You're at the wrong point on the precision/recall curve. Raise the threshold (or lower contamination), add features that capture normal seasonality so weekday spikes aren't flagged, and require an anomaly to persist for several points before alerting. Tune to the real cost: a missed fraud event vs. a needless page.
Q: Do I really need deep learning if I have hundreds of product series?
Not necessarily. Hundreds of independent classical fits are cheap and parallelisable, and often win when each series is short. Deep "global" models pay off when the series are related (sharing learning across SKUs) with rich covariates. Start classical; graduate only when baselines are exhausted.

Key Takeaways

  1. Order is the whole point. Time series breaks the i.i.d. assumption — never shuffle, always split and validate by time, or you'll leak the future and ship a model that fails in production.

  2. Earn complexity with baselines. Seasonal-naive and Holt-Winters are fast, explainable, and frequently unbeaten — benchmark against them before reaching for ARIMA or deep learning.

  3. Classical models encode real structure. ARIMA/SARIMA model dependence directly; Prophet decomposes trend, seasonality, and holidays into readable components. Both give calibrated intervals, which LLMs cannot.

  4. Anomaly detection is forecasting's sibling. Residual rules and Isolation Forest catch fraud, outages, and failing machines — and live or die on the precision/recall tradeoff you tune to your costs.

  5. Deep learning is for scale, not default. LSTMs, TCNs, and transformer forecasters shine on many related series with rich covariates; they're overkill on a single short series.

  6. LLMs narrate, they don't forecast. Token models can't add numbers reliably, have no calibrated intervals, and drift on long horizons. Use numerical models for the forecast; let the LLM explain it.

Next Steps: Lesson 6: Recommendation Systems