Time Series & Forecasting
Learning Outcomes
- Explain what makes time series different from i.i.d. data, and why that breaks naive train/test splits.
- Build a baseline forecast and beat it with classical models (exponential smoothing, ARIMA) in statsmodels.
- Fit a Prophet model on seasonal business data and read its trend and seasonal components.
- Detect anomalies in a metric stream with a residual rule and an Isolation Forest.
- Decide when a deep model fits a forecasting problem — and why an LLM is the wrong tool for numbers.
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why temporal data needs its own toolbox |
| The shape of the problem | 8 min | Trend, seasonality, stationarity, evaluation |
| Baselines & backtesting | 9 min | Naive baselines, walk-forward validation |
| Classical models | 12 min | Exponential smoothing, ARIMA |
| Prophet | 10 min | Decomposable trend + seasonality |
| Anomaly detection | 9 min | Residual rules, Isolation Forest |
| Deep models & the LLM trap | 6 min | When DL helps; why LLMs fail |
| Wrap-up | 3 min | Picking the right model |
Before You Begin
Pre-work:
- Complete Lesson 2: Machine Learning Fundamentals — you'll reuse overfitting and train/test discipline.
- Skim Lesson 1: The AI Landscape — Not Just LLMs for context.
- Be comfortable with pandas indexing.
Shopping List:
- Python 3.10+ in a virtual environment.
pip install pandas numpy statsmodels scikit-learn prophet
Most ML assumes rows are i.i.d. — independent and identically distributed: each row is its own little world, and shuffling rows changes nothing. Time series breaks this on purpose. Time series data is a sequence of measurements indexed by time (sales per day, CPU temperature per second) where order matters and today depends on yesterday.
That dependence has named structure: trend (the long drift up or down), seasonality (patterns repeating on a fixed period, like December spikes), stationarity (mean and variance don't change over time — many classical models assume this), and autocorrelation (correlation with its own past; a spike at lag 7 in daily data means weekly seasonality).
Decomposing a series lets you see the structure, using a classic monthly airline-passengers dataset (any monthly series works):
import pandas as pd
from statsmodels.tsa.seasonal import seasonal_decompose
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
df = pd.read_csv(url, parse_dates=["Month"], index_col="Month")
df = df.asfreq("MS") # set an explicit frequency or forecasts go silently wrong
seasonal_decompose(df["Passengers"], model="multiplicative", period=12).plot()
The four panels — observed, trend, seasonal, residual — are the mental model for everything below.
asfreq call is not optional — statsmodels and Prophet silently mishandle ambiguous indices and gaps. Always pin an explicit frequency or DatetimeIndex.Build a baseline you must beat: naive (tomorrow equals today) and seasonal-naive (this December equals last December). If a deep model can't beat seasonal-naive, it isn't learning anything.
A single split gives one noisy number. Walk-forward validation (a time-series cross-validation — retraining on a growing window and testing the next chunk) is more honest. Cross-validation just means repeatedly scoring on data the model never trained on.
import numpy as np
def mae(y_true, y_pred):
return np.mean(np.abs(np.asarray(y_true) - np.asarray(y_pred)))
series = df["Passengers"]
horizon = 12 # forecast one year ahead
train, test = series[:-horizon], series[-horizon:]
seasonal_naive = series.shift(12)[-horizon:] # value 12 months earlier
print("Seasonal-naive MAE:", round(mae(test, seasonal_naive), 2))
We score with MAE (mean absolute error, the average miss in original units); RMSE punishes large misses harder; MAPE (percentage error) is intuitive but treacherous.
Exponential smoothing forecasts by weighting recent observations more than old ones, weights shrinking exponentially into the past. Holt-Winters adds explicit trend and seasonality — a fast, hard-to-beat baseline. statsmodels exposes it as ExponentialSmoothing.
from statsmodels.tsa.holtwinters import ExponentialSmoothing
model = ExponentialSmoothing(
train,
trend="add", # additive trend
seasonal="mul", # multiplicative seasonality (swing grows with level)
seasonal_periods=12,
)
fit = model.fit()
forecast = fit.forecast(horizon)
print("Holt-Winters MAE:", round(mae(test, forecast), 2))
Use multiplicative seasonality when swings grow with the level (airline passengers), additive when the swing is roughly constant.
ARIMA is AutoRegressive Integrated Moving Average. AR (p) predicts the next value from p past values. I (d) differences the series d times to make it stationary (removing trend). MA (q) uses the last q forecast errors. For seasonal data, SARIMA adds a (P, D, Q, m) block at lag m; statsmodels implements it as SARIMAX.
from statsmodels.tsa.statespace.sarimax import SARIMAX
model = SARIMAX(
train,
order=(1, 1, 1), # (p, d, q)
seasonal_order=(1, 1, 1, 12), # (P, D, Q, m)
enforce_stationarity=False,
enforce_invertibility=False,
)
fit = model.fit(disp=False)
pred = fit.get_forecast(steps=horizon)
forecast = pred.predicted_mean
conf_int = pred.conf_int() # uncertainty bands, not just a point
print("SARIMA MAE:", round(mae(test, forecast), 2))
Pick (p, d, q) from ACF/PACF plots, or let pmdarima's auto_arima search, then check residuals. Rules of thumb: visible trend means increase d; a spike at the season length means add the seasonal term; autocorrelated residuals means increase p or q; poor on test means simplify.
get_forecast gives a prediction interval, so show it. An LLM cannot produce a calibrated interval.Prophet is an open-source forecasting tool for business time series with strong seasonality, holiday effects, and missing data. It models a series as a sum of three readable pieces — trend, seasonal terms, holiday bumps — so non-experts can read why the forecast moves. It wants two columns, ds (datestamp) and y (value).
from prophet import Prophet
pdf = df.reset_index().rename(columns={"Month": "ds", "Passengers": "y"})
train_pdf = pdf.iloc[:-horizon]
m = Prophet(seasonality_mode="multiplicative", yearly_seasonality=True)
m.fit(train_pdf)
future = m.make_future_dataframe(periods=horizon, freq="MS")
fcst = m.predict(future)
# yhat is the forecast; yhat_lower / yhat_upper are the uncertainty band
print(fcst[["ds", "yhat", "yhat_lower", "yhat_upper"]].tail(horizon))
m.plot_components(fcst) # trend, yearly seasonality, holidays if present
Add domain knowledge cheaply: call m.add_country_holidays("US") before fitting to capture holiday effects.
Forecasting and anomaly detection are siblings. Anomaly detection flags points that don't fit the pattern — a fraud spike, a sensor reading before a machine fails (predictive maintenance), a drop in conversions. The lightest approach is a residual rule: fit a forecast, subtract it from reality, and flag points whose error is far outside normal (resid = (fit.fittedvalues - train).dropna(), then where abs(resid) > 3 * resid.std()).
A stronger, feature-based approach is the Isolation Forest — an unsupervised model (no labelled anomalies needed) that isolates outliers by randomly partitioning the feature space, since anomalies separate in fewer cuts. scikit-learn ships it as IsolationForest:
from sklearn.ensemble import IsolationForest
feat = pd.DataFrame({"value": series})
feat["roll_mean"] = series.rolling(12, min_periods=1).mean()
feat["roll_std"] = series.rolling(12, min_periods=1).std().fillna(0)
iso = IsolationForest(contamination=0.02, random_state=42)
feat["flag"] = iso.fit_predict(feat) # -1 = anomaly, 1 = normal
print(feat[feat["flag"] == -1])
contamination is your prior on how rare anomalies are — set it from domain knowledge. Evaluate with precision (of flagged points, the fraction truly anomalous — matters when false alarms are expensive) and recall (of true anomalies, the fraction caught — matters when misses are costly).
Reach for deep learning only when classical models hit a ceiling: thousands of related series forecast jointly (every SKU in every store), rich covariates (weather, price, promotions), or long nonlinear dependencies. The workhorse is the LSTM (a recurrent network carrying a hidden "memory" across timesteps), alongside temporal convolutional networks and transformers. You rarely build these by hand — libraries like PyTorch Forecasting and Darts wrap them:
from pytorch_forecasting import TemporalFusionTransformer, TimeSeriesDataSet
# Build a TimeSeriesDataSet from a long-format DataFrame, then call
# TemporalFusionTransformer.from_dataset(...) and train via Lightning.
Now the headline: why are LLMs poor at numerical forecasting? They model tokens, not numbers — "412" is characters, not a quantity to difference or extrapolate. They have no parameters fitted to your series and no calibrated uncertainty, drift on long horizons, invent seasonality that isn't there, and cost far more than a millisecond SARIMA fit.
| Situation | Right tool |
|---|---|
| One clean monthly/weekly series | Holt-Winters or SARIMA |
| Seasonality + holidays + interpretability | Prophet |
| Thousands of related series, rich covariates | Deep model (LSTM/TCN/transformer) |
| Outliers in a stream | Residual rule or Isolation Forest |
| "Summarise why sales dipped" | An LLM (reasoning around the numbers) |
Questions & Answers
contamination), add features that capture normal seasonality so weekday spikes aren't flagged, and require an anomaly to persist for several points before alerting. Tune to the real cost: a missed fraud event vs. a needless page.Key Takeaways
-
Order is the whole point. Time series breaks the i.i.d. assumption — never shuffle, always split and validate by time, or you'll leak the future and ship a model that fails in production.
-
Earn complexity with baselines. Seasonal-naive and Holt-Winters are fast, explainable, and frequently unbeaten — benchmark against them before reaching for ARIMA or deep learning.
-
Classical models encode real structure. ARIMA/SARIMA model dependence directly; Prophet decomposes trend, seasonality, and holidays into readable components. Both give calibrated intervals, which LLMs cannot.
-
Anomaly detection is forecasting's sibling. Residual rules and Isolation Forest catch fraud, outages, and failing machines — and live or die on the precision/recall tradeoff you tune to your costs.
-
Deep learning is for scale, not default. LSTMs, TCNs, and transformer forecasters shine on many related series with rich covariates; they're overkill on a single short series.
-
LLMs narrate, they don't forecast. Token models can't add numbers reliably, have no calibrated intervals, and drift on long horizons. Use numerical models for the forecast; let the LLM explain it.
Next Steps: Lesson 6: Recommendation Systems