Classical ML vs GenAI Decision Patterns
A lookup table for one question: given a problem, what should I reach for? The LLM-only instinct is to send everything to a chat model, often the slowest and costliest option. Classical machine learning (ML), models that learn from labelled examples instead of being prompted, often wins on cost, latency, and accuracy for narrow tasks. Find the closest pattern, read the why, apply the rule on boundary cases.
Quick decision table
| Problem type | Recommended approach | Why |
|---|---|---|
| High-throughput classification | Classical ML (gradient boosting) | Millisecond inference, near-zero cost |
| Structured / tabular prediction | Classical ML (XGBoost, LightGBM) | Trees beat deep nets on spreadsheets |
| Anomaly / fraud detection | Classical ML (isolation forest) | Learns "normal" from unlabelled data |
| Time-series forecasting | Statistical (Prophet, ARIMA) | Handles seasonality without GPUs |
| Ranking & recommendation | Classical + embeddings (ranker) | Sub-second ranking at scale |
| Image classification / QC | Deep learning (CNN + transfer learning) | Pretrained backbones make pixels cheap |
| Object detection / localisation | Deep learning (YOLO / Ultralytics) | Finds and boxes objects in real time |
| OCR / document layout | Hybrid (OCR + LLM cleanup) | OCR extracts; LLM cleans up |
| Zero-shot text extraction | GenAI (LLM, structured output) | No labels, schema known, low volume |
| Open-ended reasoning | GenAI (LLM) | Language understanding, no training set |
| Semantic search / clustering | Embeddings + classical retrieval | Vector similarity is fast and cheap |
| Customer-support routing | Start GenAI, distil to classical | Bootstrap with an LLM, then distil |
Terms, in plain language: Classification = predict which bucket (spam/not-spam). Regression = predict a number (house price). Embedding = numbers representing meaning, so "similar" things sit close together. Transfer learning = reuse a model trained on a huge dataset, fine-tuned on your small one. Zero-shot = handle a task from the prompt, with no training.
High-throughput classification & tabular prediction
Approach: classical ML. To label millions of items per day (transactions, logs), an LLM at hundreds of milliseconds plus a per-call cost is a non-starter; a gradient-boosted tree predicts in microseconds on CPU. Precision = of items flagged positive, how many really were? Recall = of all true positives, how many caught? On tabular data, trees (XGBoost, LightGBM) beat deep nets and LLMs.
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report
from sklearn.datasets import load_breast_cancer
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2)
clf = GradientBoostingClassifier().fit(X_tr, y_tr)
print(classification_report(y_te, clf.predict(X_te))) # precision/recall/f1
Decision rule: if volume times per-call LLM cost is meaningful, or latency must be under 50 ms, train a classifier. Use the LLM only with no labels.
Anomaly & fraud detection
Approach: classical ML. Fraud is rare and rarely cleanly labelled. Unsupervised models (no labels) learn what "normal" looks like and score how far each point strays.
from sklearn.ensemble import IsolationForest
model = IsolationForest(contamination=0.01).fit(X_train) # X_train = "normal" traffic
flags = model.predict(X_new) # -1 = anomaly, 1 = normal
Decision rule: rare events, few or no labels: isolation forest or one-class SVM. With a labelled history, supervised classification wins.
Time-series forecasting
Approach: statistical / classical. Demand, traffic, and sensor forecasts have seasonality (repeating daily/weekly/yearly cycles); Prophet and ARIMA model these directly without GPUs. LLMs forecast numbers badly.
from prophet import Prophet
m = Prophet(yearly_seasonality=True).fit(df) # df: ds=timestamp, y=metric
forecast = m.predict(m.make_future_dataframe(periods=30))
Decision rule: clear seasonality, limited history: Prophet/ARIMA. Many related series with rich features: gradient boosting.
Vision: classification, QC, and detection
Approach: deep learning. Detecting defects on a production line needs a convolutional neural network (CNN), built to read pixels. Rarely train from scratch: start from a pretrained backbone and fine-tune (transfer learning). When you need where an object is, a detector like YOLO returns boxes live.
import torchvision.models as models
import torch.nn as nn
net = models.resnet18(weights="IMAGENET1K_V1") # pretrained backbone
for p in net.parameters():
p.requires_grad = False # freeze, reuse features
net.fc = nn.Linear(net.fc.in_features, 2) # fine-tune: pass / fail
# Detection (boxes) instead of classification:
from ultralytics import YOLO
for box in YOLO("yolo11n.pt")("line.jpg")[0].boxes:
print(box.cls, box.conf, box.xyxy) # class, conf, coords
Decision rule: fixed label set per image: CNN. Need coordinates, counts, or tracking: YOLO. Open-ended "describe this": vision LLM.
OCR & document understanding
Approach: hybrid. A dedicated OCR engine pulls text from scans far cheaper than an LLM. Pipeline: (1) OCR produces raw text plus word boxes; (2) rules and regex grab obvious fields (dates, totals); (3) an LLM, on the leftovers only, structures them into JSON.
Decision rule: clean digital PDFs: parse directly, no LLM. Noisy scans: OCR, then LLM for the rest.
Ranking, recommendation, and semantic search
Approach: classical + embeddings. Recommendation engines score thousands of candidates per request in milliseconds; matrix factorisation or a learning-to-rank model does this, where an LLM cannot. For "find similar" or "group these", embed text once, store the vectors, then run nearest-neighbour search or k-means.
from sklearn.cluster import KMeans
labels = KMeans(n_clusters=8, n_init="auto").fit_predict(vectors) # vectors = embeddings
Decision rule: real-time ordering: classical ranker over precomputed features. "Answer a question about these docs": embeddings for retrieval, LLM for the answer (RAG). Use an LLM to explain picks, not to rank.
Where GenAI wins, plus the bootstrap pattern
Approach: GenAI. With no labelled data, a known schema, and modest volume, an LLM with structured output is the fastest path, no training loop. Summarising, drafting, free-form Q&A, and reasoning over language lean on the model's world knowledge; do not force a classical model here. A strong hybrid lifecycle combines both: Day 1, ship a zero-shot LLM and log its I/O; Month 2, those logs are a free labelled dataset; Month 3+, train a classifier to replace most calls, keeping the LLM for hard cases.
Decision rule: fluent language or novel reasoning: LLM. Fixed category or number: classical. No data and need to ship: LLM, then distil to a classifier once the task is stable, with the LLM as fallback.
Reading these patterns together
The rule of thumb: the narrower and more repetitive the task, the more classical ML wins; the more open-ended and language-driven, the more GenAI wins. Most production systems are hybrids.
See rules, cheat sheet, Lesson 10, tools-reference.