Troubleshooting ML Pipelines
A classical ML pipeline fails differently from an LLM call: no prompt to tweak, failures hide in the numbers. Organised problem to cause to fix. See decision patterns and tools reference. First law: trust the metrics; split off a test set first.
Model Will Not Converge
"Converge" means the training loss stops dropping and settles low. Flat, NaN, or oscillating loss means the optimiser is not learning.
| Symptom | Likely cause | Fix |
|---|---|---|
| Loss flat from step 0 | LR too low; bad init | Raise LR 10x; check init |
| Loss explodes to NaN | LR too high; unscaled features | Lower LR; scale; clip gradients |
| Loss oscillates | Batch too small; LR too high | Larger batch; add LR decay |
| sklearn "did not converge" | Too few iters; unscaled | Scale; raise max_iter |
Scaling is the most common fix: linear models, SVMs and neural nets assume features share similar ranges, so wrap a StandardScaler before the estimator (see Data Leakage). In PyTorch, clip gradients with clip_grad_norm_ before optimizer.step().
Strong Train, Poor Test (Overfitting)
Overfitting = the model memorises the training set instead of the pattern, scoring high on seen data, low on unseen. The tell: a large train-test gap (e.g. 0.99 vs 0.71).
| Cause | Diagnostic | Fix |
|---|---|---|
| Model too complex | Train >> test score | Regularise; reduce capacity |
| Too few training rows | High variance across folds | More data; simpler model |
| No regularisation | Large weights | Add L1/L2; dropout |
Cross-validation trains and tests on several splits and averages the scores, far more honest than one; a wide std signals an unstable estimate (too little data). A frozen LLM cannot overfit your data because it never trains on it, which is why classical ML wins when you have labels and need calibrated scores.
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score
model = RandomForestClassifier(max_depth=8, n_estimators=200)
scores = cross_val_score(model, X, y, cv=5, scoring="f1")
print(f"F1: {scores.mean():.3f} +/- {scores.std():.3f}")
Data Leakage
Leakage = test-set (or future) info sneaks into training, so offline metrics look brilliant and production collapses. The most dangerous ML bug: it hides as success.
| Leak source | Symptom | Fix |
|---|---|---|
| Scaler/encoder fit on full data | Score too good to be true | Fit transforms in CV, on train only |
| Target-derived feature | One feature dominates importance | Drop features computed from the label |
| Random split of time series | Great offline, bad live | Split chronologically |
Wrap preprocessing in a Pipeline so transforms learn from training folds only. For time series, never shuffle; use TimeSeriesSplit.
from sklearn.model_selection import cross_val_score, TimeSeriesSplit
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
pipe = make_pipeline(StandardScaler(), SVC()) # scaler re-fit per fold
cross_val_score(pipe, X, y, cv=TimeSeriesSplit(n_splits=5))
Severe Class Imbalance
When fraud is ~0.1% of transactions, a model that always predicts "not fraud" hits 99.9% accuracy catching none, so accuracy is the wrong metric. Precision = of cases flagged positive, how many were truly positive (false alarms). Recall = of all true positives, how many were caught (misses).
| Cause | Diagnostic | Fix |
|---|---|---|
| Model ignores minority class | High accuracy, recall ~0 | Class weights; resample |
| Wrong metric optimised | Accuracy looks great | Track precision/recall, PR-AUC |
| Default 0.5 threshold | Few/no positives flagged | Tune the threshold |
from sklearn.metrics import classification_report
from sklearn.linear_model import LogisticRegression
# class_weight rebalances the loss toward the rare class
clf = LogisticRegression(class_weight="balanced", max_iter=1000)
clf.fit(X_train, y_train)
print(classification_report(y_test, clf.predict(X_test)))
SMOTE (train only) synthesises minority rows.
Deployment & Serving Errors
The model trained fine but the API returns 500s or wrong shapes: usually environment or contract mismatches.
| Symptom | Cause | Fix |
|---|---|---|
ValueError: shape mismatch |
Columns reordered/missing | Validate schema; persist column order |
| Unpickling error on load | Library version differs from train | Pin versions; record environment |
| Predictions all identical | Scaler/encoder not loaded | Save the full pipeline |
| High p99 latency | Cold model load per request | Load at startup; batch requests |
Persist the entire pipeline so preprocessing ships with the model; validate the contract first:
import joblib, pandas as pd
loaded = joblib.load("model.joblib") # saved via joblib.dump(pipe, ...)
EXPECTED = ["amount", "merchant_id", "hour", "country"]
def predict(row: dict):
missing = [c for c in EXPECTED if c not in row]
if missing:
raise ValueError(f"missing features: {missing}")
return loaded.predict(pd.DataFrame([row])[EXPECTED])[0]
Drift-Detection Alerts
Drift = live data moves away from training data, so a once-accurate model quietly degrades (e.g. a forecaster after a product launch).
| Alert type | Meaning | Action |
|---|---|---|
| Data drift | Input distribution shifted | Inspect features; consider retrain |
| Concept drift | Input-to-label link changed | Retrain with fresh labels |
| Prediction drift | Output mix shifted | Check upstream data, then labels |
| Label delay | No ground truth yet | Monitor proxies; backfill |
Compare recent inputs to a training baseline; small p-value = they differ:
from scipy.stats import ks_2samp
# Kolmogorov-Smirnov: same distribution as training?
stat, p = ks_2samp(train_feature, live_feature)
if p < 0.01:
print("Drift: investigate, then consider retraining")
Do not auto-retrain on every alert: confirm the drift is real and you have fresh, trustworthy labels, or you retrain on noise.
Training-Serving Skew & Feature-Store Inconsistency
Skew = a feature computed one way in training (batch SQL over history) and another at serving (live code). Fine offline, underperforms live, no error.
| Cause | Diagnostic | Fix |
|---|---|---|
| Two code paths for one feature | Offline good, online bad | Share one definition |
| Different time windows | Stats differ live vs train | Align aggregation windows |
| Stale features at serve | Predictions lag reality | Set TTLs; monitor freshness |
| Null handling differs | Spikes in null rate | Identical imputation both sides |
The durable fix is a feature store: one feature definition served to both training and inference. Spot-check by logging served vectors and recomputing them:
import numpy as np
served = np.load("served_features.npy") # logged at inference
offline = np.load("recomputed_features.npy") # recomputed offline
assert np.abs(served - offline).max() < 1e-6, "training-serving skew"
Debugging order: (1) confirm no leakage, (2) check scaling, (3) pick a metric matching the problem (precision/recall over accuracy for imbalance), (4) verify serving matches training. See the rules and cheat sheet, or the course home.