Troubleshooting ML Pipelines

Reference intermediate

A classical ML pipeline fails differently from an LLM call: no prompt to tweak, failures hide in the numbers. Organised problem to cause to fix. See decision patterns and tools reference. First law: trust the metrics; split off a test set first.

Model Will Not Converge

"Converge" means the training loss stops dropping and settles low. Flat, NaN, or oscillating loss means the optimiser is not learning.

Symptom Likely cause Fix
Loss flat from step 0 LR too low; bad init Raise LR 10x; check init
Loss explodes to NaN LR too high; unscaled features Lower LR; scale; clip gradients
Loss oscillates Batch too small; LR too high Larger batch; add LR decay
sklearn "did not converge" Too few iters; unscaled Scale; raise max_iter

Scaling is the most common fix: linear models, SVMs and neural nets assume features share similar ranges, so wrap a StandardScaler before the estimator (see Data Leakage). In PyTorch, clip gradients with clip_grad_norm_ before optimizer.step().

Strong Train, Poor Test (Overfitting)

Overfitting = the model memorises the training set instead of the pattern, scoring high on seen data, low on unseen. The tell: a large train-test gap (e.g. 0.99 vs 0.71).

Cause Diagnostic Fix
Model too complex Train >> test score Regularise; reduce capacity
Too few training rows High variance across folds More data; simpler model
No regularisation Large weights Add L1/L2; dropout

Cross-validation trains and tests on several splits and averages the scores, far more honest than one; a wide std signals an unstable estimate (too little data). A frozen LLM cannot overfit your data because it never trains on it, which is why classical ML wins when you have labels and need calibrated scores.

from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score

model = RandomForestClassifier(max_depth=8, n_estimators=200)
scores = cross_val_score(model, X, y, cv=5, scoring="f1")
print(f"F1: {scores.mean():.3f} +/- {scores.std():.3f}")

Data Leakage

Leakage = test-set (or future) info sneaks into training, so offline metrics look brilliant and production collapses. The most dangerous ML bug: it hides as success.

Leak source Symptom Fix
Scaler/encoder fit on full data Score too good to be true Fit transforms in CV, on train only
Target-derived feature One feature dominates importance Drop features computed from the label
Random split of time series Great offline, bad live Split chronologically

Wrap preprocessing in a Pipeline so transforms learn from training folds only. For time series, never shuffle; use TimeSeriesSplit.

from sklearn.model_selection import cross_val_score, TimeSeriesSplit
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC

pipe = make_pipeline(StandardScaler(), SVC())  # scaler re-fit per fold
cross_val_score(pipe, X, y, cv=TimeSeriesSplit(n_splits=5))

Severe Class Imbalance

When fraud is ~0.1% of transactions, a model that always predicts "not fraud" hits 99.9% accuracy catching none, so accuracy is the wrong metric. Precision = of cases flagged positive, how many were truly positive (false alarms). Recall = of all true positives, how many were caught (misses).

Cause Diagnostic Fix
Model ignores minority class High accuracy, recall ~0 Class weights; resample
Wrong metric optimised Accuracy looks great Track precision/recall, PR-AUC
Default 0.5 threshold Few/no positives flagged Tune the threshold
from sklearn.metrics import classification_report
from sklearn.linear_model import LogisticRegression

# class_weight rebalances the loss toward the rare class
clf = LogisticRegression(class_weight="balanced", max_iter=1000)
clf.fit(X_train, y_train)
print(classification_report(y_test, clf.predict(X_test)))

SMOTE (train only) synthesises minority rows.

Deployment & Serving Errors

The model trained fine but the API returns 500s or wrong shapes: usually environment or contract mismatches.

Symptom Cause Fix
ValueError: shape mismatch Columns reordered/missing Validate schema; persist column order
Unpickling error on load Library version differs from train Pin versions; record environment
Predictions all identical Scaler/encoder not loaded Save the full pipeline
High p99 latency Cold model load per request Load at startup; batch requests

Persist the entire pipeline so preprocessing ships with the model; validate the contract first:

import joblib, pandas as pd

loaded = joblib.load("model.joblib")  # saved via joblib.dump(pipe, ...)
EXPECTED = ["amount", "merchant_id", "hour", "country"]

def predict(row: dict):
    missing = [c for c in EXPECTED if c not in row]
    if missing:
        raise ValueError(f"missing features: {missing}")
    return loaded.predict(pd.DataFrame([row])[EXPECTED])[0]

Drift-Detection Alerts

Drift = live data moves away from training data, so a once-accurate model quietly degrades (e.g. a forecaster after a product launch).

Alert type Meaning Action
Data drift Input distribution shifted Inspect features; consider retrain
Concept drift Input-to-label link changed Retrain with fresh labels
Prediction drift Output mix shifted Check upstream data, then labels
Label delay No ground truth yet Monitor proxies; backfill

Compare recent inputs to a training baseline; small p-value = they differ:

from scipy.stats import ks_2samp

# Kolmogorov-Smirnov: same distribution as training?
stat, p = ks_2samp(train_feature, live_feature)
if p < 0.01:
    print("Drift: investigate, then consider retraining")

Do not auto-retrain on every alert: confirm the drift is real and you have fresh, trustworthy labels, or you retrain on noise.

Training-Serving Skew & Feature-Store Inconsistency

Skew = a feature computed one way in training (batch SQL over history) and another at serving (live code). Fine offline, underperforms live, no error.

Cause Diagnostic Fix
Two code paths for one feature Offline good, online bad Share one definition
Different time windows Stats differ live vs train Align aggregation windows
Stale features at serve Predictions lag reality Set TTLs; monitor freshness
Null handling differs Spikes in null rate Identical imputation both sides

The durable fix is a feature store: one feature definition served to both training and inference. Spot-check by logging served vectors and recomputing them:

import numpy as np

served  = np.load("served_features.npy")      # logged at inference
offline = np.load("recomputed_features.npy")  # recomputed offline
assert np.abs(served - offline).max() < 1e-6, "training-serving skew"

Debugging order: (1) confirm no leakage, (2) check scaling, (3) pick a metric matching the problem (precision/recall over accuracy for imbalance), (4) verify serving matches training. See the rules and cheat sheet, or the course home.