When to Use Classical ML vs GenAI
Learning Outcomes
- Apply a six-dimension framework — latency, cost, interpretability, data, accuracy, maintenance — to any AI problem
- Identify the signatures where an LLM is the wrong tool, and explain why with numbers
- Measure the per-prediction cost and latency gap between a scikit-learn model and an LLM call
- Choose between classical ML, deep learning, and generative AI for concrete case studies
- Defend a tool choice using interpretability and regulatory constraints
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | "Just use an LLM" is a default, not a decision |
| Framework | 7 min | The six dimensions that decide the tool |
| Measure | 8 min | Cost and latency: a classifier vs an LLM |
| Case study | 7 min | Fraud scoring — where LLMs lose |
| Case study | 7 min | Open-ended triage — where LLMs win |
| Interpretability | 6 min | Regulated decisions you can defend |
| Decision tree | 5 min | A repeatable flowchart |
| Wrap-up | 2 min | Key takeaways and next steps |
Before You Begin
Pre-work:
- Be comfortable with supervised learning basics from Lesson 2: Machine Learning Fundamentals
- Skim the production concerns in Lesson 7: MLOps — Models to Production
- Have a mental model of an LLM API call (tokens in, tokens out, billed per token)
Shopping List:
- Python 3.10+ and a virtual environment
pip install scikit-learn pandas numpyfor the worked classifier- Optional, to reproduce the LLM-side numbers: any LLM SDK and an API key
Since chat models got good, the reflex is "throw an LLM at it" — a default, not a decision. A generative AI / LLM produces free-form output and generalises to tasks it was never trained on. Classical ML means smaller, specialised models — logistic regression, gradient-boosted trees, SVMs — trained on your labels to do one thing. Deep learning sits between: neural nets (CNNs, encoder transformers) for a specific task.
Score your problem on six dimensions first:
| Dimension | Favours classical ML | Favours LLM/GenAI |
|---|---|---|
| Latency | Sub-millisecond, high QPS | Hundreds of ms ok |
| Cost | Fractions of a cent, millions/day | A few cents, low volume |
| Interpretability | Regulated, audited | Explanation optional |
| Data | Thousands of labels | Few or zero labels |
| Accuracy | Fixed schema, stable | Open-ended, fuzzy |
| Maintenance | Team retrains | Prompt tweaks |
These often conflict — a problem can scream "LLM" on data (zero labels) and "classical" on cost. The skill is weighing them, not finding a clean winner.
Numbers settle arguments. Classification assigns an input to one of a fixed set of categories; training fits a model to labelled examples so it generalises. A runnable scikit-learn pipeline:
import time
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.metrics import classification_report
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y)
# StandardScaler centres features; LogisticRegression is a linear classifier.
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(X_tr, y_tr)
print(classification_report(y_te, model.predict(X_te)))
start = time.perf_counter()
for _ in range(10_000):
model.predict(X_te[:1])
print(f"Latency: {(time.perf_counter() - start) / 10_000 * 1e6:.1f} us/pred")
On a laptop CPU this prints accuracy in the high 0.90s and latency in the low tens of microseconds — no network, no key, no token bill. Classifying the same row with an LLM means prompting 30 features into a model that returns in 300-1500 ms, billed per token — four orders of magnitude slower, and per-call cost versus free:
| Property | LogisticRegression | LLM API call |
|---|---|---|
| Latency | tens of microseconds | hundreds of ms |
| Marginal cost | ~0 (your CPU) | per-token, recurring |
| Throughput/box | hundreds of thousands/sec | vendor rate-limited |
| Determinism | identical every run | can vary |
Fraud detection scores each transaction live: a processor handles tens of thousands per second, and the score must return before checkout — often under 50 ms. Every dimension points to classical ML.
A gradient-boosted tree — an ensemble of small decision trees, each correcting the previous one's errors — is the standard: fast, accurate on tabular data, explainable.
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import precision_recall_curve, roc_auc_score
# y: 1 = fraud, 0 = legitimate. Fraud is rare, so accuracy alone misleads.
clf = HistGradientBoostingClassifier(max_iter=300, learning_rate=0.1)
clf.fit(X_tr, y_tr)
scores = clf.predict_proba(X_te)[:, 1]
print("ROC-AUC:", round(roc_auc_score(y_te, scores), 4))
precision, recall, thresholds = precision_recall_curve(y_te, scores)
Accuracy is the wrong metric when fraud is 0.1% of traffic. Precision: of what you flagged, how much was really fraud? Recall: of all real fraud, how much did you catch? A model that flags nothing scores 99.9% accuracy and catches zero fraud. You tune the threshold to balance the two against business cost.
The LLM loses every round: it misses the latency budget, the per-call cost across billions of transactions is ruinous, and "the model felt it was fraudulent" does not survive an audit.
Now flip it. A support team receives free-form emails — bugs, refunds, feature ideas, security disclosures. You route each to a queue and draft a one-line summary. New categories appear monthly, and you have almost no labelled data.
| Dimension | This problem | Verdict |
|---|---|---|
| Latency | A few seconds is fine | LLM ok |
| Cost | Thousands/day, not billions | LLM ok |
| Data | Few labels, categories shift | LLM strong |
| Accuracy | Fuzzy, open-ended | LLM strong |
| Maintenance | Change a prompt, not data | LLM strong |
This is where the LLM's superpower — zero-shot capability, doing a task from a description alone with no task-specific training examples — earns its keep. A classifier would mean labelling thousands of emails and re-labelling on every change; the LLM handles a new category by editing one prompt line:
# REQUIRES an API key — illustrates the zero-shot pattern.
CATEGORIES = ["bug", "refund", "feature_request", "security", "other"]
def triage(client, email_text):
prompt = (
"Classify the email into exactly one of: " + ", ".join(CATEGORIES)
+ ". Then give a one-line summary. "
"Return JSON with keys category and summary.\n\nEmail:\n" + email_text
)
return client.responses.create(model="your-model", input=prompt).output_text
Adding a billing_dispute category is one word in the list — no retraining, no new labels, no redeploy. The LLM also produces the summary in the same call, a generative task a classifier cannot do.
In regulated domains — lending, insurance, hiring, healthcare — you must often produce a reason code: an auditable explanation of why a decision was made. "A huge transformer output 'deny'" is not defensible before a regulator.
Classical models give this almost for free. Feature importance quantifies how much each input drives predictions overall; SHAP (SHapley Additive exPlanations) attributes a single prediction to its features:
import pandas as pd
import shap
from sklearn.ensemble import RandomForestClassifier
clf = RandomForestClassifier(n_estimators=200, random_state=0).fit(X_tr, y_tr)
# Global: which features matter overall.
imp = pd.Series(clf.feature_importances_, index=feature_names)
print(imp.sort_values(ascending=False).head(5))
# Per-decision: why THIS applicant was denied.
shap_values = shap.TreeExplainer(clf).shap_values(X_te[:1])
That turns "denied" into "denied because credit_utilisation was 0.92 and there were 3 missed payments." LLMs can write an explanation, but it is a plausible narrative generated after the fact — not a faithful account of what the model did.
| Need | Classical ML | LLM |
|---|---|---|
| Reason code | SHAP, coefficients — faithful | Narrative — not faithful |
| Reproducible | Deterministic | Can vary |
| Regulatory acceptance | Established | Often insufficient |
Before any other dimension, ask: what labelled data do I have? It decides more architectures than latency or cost.
- Thousands of labels, stable target → classical ML or task-specific deep learning likely beats an LLM on accuracy, cost, and speed.
- A few hundred labels → use transfer learning: take a model pre-trained on a huge generic dataset and fine-tune its final layers on your small one, so it arrives already knowing general features.
- Zero or a handful → an LLM's zero-shot ability is hard to beat as a start.
Transfer learning with a Hugging Face BERT-family encoder (a transformer trained to understand text, not generate it):
from transformers import AutoTokenizer, AutoModelForSequenceClassification
# Add a classification head to a pre-trained encoder, then fine-tune only
# on our own labels — far fewer than training from scratch.
name = "distilbert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name, num_labels=3)
# Fine-tune with the Trainer API on a few thousand rows -> a tiny, fast,
# cheap model you own and can run on a CPU.
A fine-tuned DistilBERT classifier runs in milliseconds, costs nothing per call, and on a narrow task often matches or beats a general LLM — specialised for exactly your labels.
Walk this top to bottom; the first hard constraint that fires usually wins.
1. Hard latency/throughput SLA (sub-10ms, high QPS)? -> Classical ML. STOP.
2. Per-call cost dominates at volume (millions+/day)? -> Classical ML. STOP.
3. Regulated, needs faithful reason codes? -> Classical ML. STOP.
4. Unstructured text/images AND labels scarce? -> LLM / foundation model.
5. Thousands of labels, stable narrow target? -> Classical / deep learning.
6. Open-ended, generative, fast-changing? -> LLM.
Otherwise -> prototype with an LLM, measure, decide.
Worked example — "predict next month's demand from 5 years of daily sales": no hard SLA, modest volume, not regulated (1-3 pass); structured numeric series (4 → no); thousands of points, stable target (5 → YES). Verdict: classical time-series methods (see Lesson 5: Time Series & Forecasting), not an LLM. A pocket lookup:
| Problem signature | Right tool |
|---|---|
| Tabular, high-volume, regulated | Gradient-boosted trees |
| Numeric forecasting | ARIMA / Prophet / temporal nets |
| Object detection on a camera feed | YOLO / task-specific CV model |
| Open-ended text triage, few labels | LLM (zero-shot) |
| Few hundred labels, narrow task | Transfer learning (fine-tune BERT) |
Questions & Answers
Key Takeaways
- An LLM is a default, not a decision. Score every problem on latency, cost, interpretability, data, accuracy, and maintenance — "just use an LLM" carries a real, recurring tax.
- The latency and cost gaps are architectural. A classical classifier predicts in microseconds for free; an LLM call takes hundreds of milliseconds and bills per token. You cannot prompt across four orders of magnitude.
- LLMs lose on high-throughput, tabular, regulated work. Fraud scoring wants fast, cheap, explainable gradient-boosted trees — classical ML is strictly better, not a compromise.
- LLMs win on open-ended, unstructured, low-label work. Zero-shot triage of free-form text with shifting categories is where editing one prompt line beats labelling thousands of examples.
- Data availability is the quiet deciding factor. Many labels favour classical ML; a few hundred favour transfer learning; near-zero favours an LLM start that later graduates to a specialised model.
- Faithful explanations only come from classical models. SHAP values and coefficients give auditable reason codes; an LLM's self-explanation is a plausible narrative, not proof.
Next Steps: Lesson 9: Combining Classical ML + LLMs