When to Use Classical ML vs GenAI

45 min advanced Lesson 8

Learning Outcomes

  • Apply a six-dimension framework — latency, cost, interpretability, data, accuracy, maintenance — to any AI problem
  • Identify the signatures where an LLM is the wrong tool, and explain why with numbers
  • Measure the per-prediction cost and latency gap between a scikit-learn model and an LLM call
  • Choose between classical ML, deep learning, and generative AI for concrete case studies
  • Defend a tool choice using interpretability and regulatory constraints

Lesson Plan

Segment Duration Topic
Intro 3 min "Just use an LLM" is a default, not a decision
Framework 7 min The six dimensions that decide the tool
Measure 8 min Cost and latency: a classifier vs an LLM
Case study 7 min Fraud scoring — where LLMs lose
Case study 7 min Open-ended triage — where LLMs win
Interpretability 6 min Regulated decisions you can defend
Decision tree 5 min A repeatable flowchart
Wrap-up 2 min Key takeaways and next steps

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ and a virtual environment
  • pip install scikit-learn pandas numpy for the worked classifier
  • Optional, to reproduce the LLM-side numbers: any LLM SDK and an API key

1 Stop Defaulting — The Six Dimensions That Decide

Since chat models got good, the reflex is "throw an LLM at it" — a default, not a decision. A generative AI / LLM produces free-form output and generalises to tasks it was never trained on. Classical ML means smaller, specialised models — logistic regression, gradient-boosted trees, SVMs — trained on your labels to do one thing. Deep learning sits between: neural nets (CNNs, encoder transformers) for a specific task.

Score your problem on six dimensions first:

Dimension Favours classical ML Favours LLM/GenAI
Latency Sub-millisecond, high QPS Hundreds of ms ok
Cost Fractions of a cent, millions/day A few cents, low volume
Interpretability Regulated, audited Explanation optional
Data Thousands of labels Few or zero labels
Accuracy Fixed schema, stable Open-ended, fuzzy
Maintenance Team retrains Prompt tweaks

These often conflict — a problem can scream "LLM" on data (zero labels) and "classical" on cost. The skill is weighing them, not finding a clean winner.

NOTE
The default tax
An LLM call is a network round-trip to a billion-parameter model, billed per token, with latency you do not control — for a task a 50 KB logistic regression could do in microseconds on your CPU. Know when that tax is worth paying.

2 Measure the Gap — A Classifier vs an LLM

Numbers settle arguments. Classification assigns an input to one of a fixed set of categories; training fits a model to labelled examples so it generalises. A runnable scikit-learn pipeline:

import time
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.metrics import classification_report

X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y)

# StandardScaler centres features; LogisticRegression is a linear classifier.
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(X_tr, y_tr)
print(classification_report(y_te, model.predict(X_te)))

start = time.perf_counter()
for _ in range(10_000):
    model.predict(X_te[:1])
print(f"Latency: {(time.perf_counter() - start) / 10_000 * 1e6:.1f} us/pred")

On a laptop CPU this prints accuracy in the high 0.90s and latency in the low tens of microseconds — no network, no key, no token bill. Classifying the same row with an LLM means prompting 30 features into a model that returns in 300-1500 ms, billed per token — four orders of magnitude slower, and per-call cost versus free:

Property LogisticRegression LLM API call
Latency tens of microseconds hundreds of ms
Marginal cost ~0 (your CPU) per-token, recurring
Throughput/box hundreds of thousands/sec vendor rate-limited
Determinism identical every run can vary
WARNING
Latency is not a tuning problem
You cannot prompt your way to microseconds — the gap is architectural, a 50 KB linear model versus a billion-parameter model behind a network hop. If your SLA needs sub-10ms at high QPS, an LLM is disqualified before accuracy enters the conversation.

3 Case Study — Fraud Scoring (LLM Loses)

Fraud detection scores each transaction live: a processor handles tens of thousands per second, and the score must return before checkout — often under 50 ms. Every dimension points to classical ML.

A gradient-boosted tree — an ensemble of small decision trees, each correcting the previous one's errors — is the standard: fast, accurate on tabular data, explainable.

from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import precision_recall_curve, roc_auc_score

# y: 1 = fraud, 0 = legitimate. Fraud is rare, so accuracy alone misleads.
clf = HistGradientBoostingClassifier(max_iter=300, learning_rate=0.1)
clf.fit(X_tr, y_tr)

scores = clf.predict_proba(X_te)[:, 1]
print("ROC-AUC:", round(roc_auc_score(y_te, scores), 4))
precision, recall, thresholds = precision_recall_curve(y_te, scores)

Accuracy is the wrong metric when fraud is 0.1% of traffic. Precision: of what you flagged, how much was really fraud? Recall: of all real fraud, how much did you catch? A model that flags nothing scores 99.9% accuracy and catches zero fraud. You tune the threshold to balance the two against business cost.

The LLM loses every round: it misses the latency budget, the per-call cost across billions of transactions is ruinous, and "the model felt it was fraudulent" does not survive an audit.

NOTE
Tabular data is classical ML's home turf
For structured rows of numbers and categories — transactions, sensor readings, click logs — gradient-boosted trees (XGBoost, LightGBM) routinely beat deep nets and LLMs on accuracy while being faster and cheaper. LLMs shine on unstructured text and images, not spreadsheets.

4 Case Study — Open-Ended Triage (LLM Wins)

Now flip it. A support team receives free-form emails — bugs, refunds, feature ideas, security disclosures. You route each to a queue and draft a one-line summary. New categories appear monthly, and you have almost no labelled data.

Dimension This problem Verdict
Latency A few seconds is fine LLM ok
Cost Thousands/day, not billions LLM ok
Data Few labels, categories shift LLM strong
Accuracy Fuzzy, open-ended LLM strong
Maintenance Change a prompt, not data LLM strong

This is where the LLM's superpower — zero-shot capability, doing a task from a description alone with no task-specific training examples — earns its keep. A classifier would mean labelling thousands of emails and re-labelling on every change; the LLM handles a new category by editing one prompt line:

# REQUIRES an API key — illustrates the zero-shot pattern.
CATEGORIES = ["bug", "refund", "feature_request", "security", "other"]

def triage(client, email_text):
    prompt = (
        "Classify the email into exactly one of: " + ", ".join(CATEGORIES)
        + ". Then give a one-line summary. "
        "Return JSON with keys category and summary.\n\nEmail:\n" + email_text
    )
    return client.responses.create(model="your-model", input=prompt).output_text

Adding a billing_dispute category is one word in the list — no retraining, no new labels, no redeploy. The LLM also produces the summary in the same call, a generative task a classifier cannot do.

TIP
The signals for an LLM
Reach for an LLM when at least two hold: input is unstructured language or images, labelled data is scarce, the task is open-ended or changes often, or you need generated output (a summary, draft, explanation) not just a label. Triage hits all four.

5 Interpretability — What You Can Defend

In regulated domains — lending, insurance, hiring, healthcare — you must often produce a reason code: an auditable explanation of why a decision was made. "A huge transformer output 'deny'" is not defensible before a regulator.

Classical models give this almost for free. Feature importance quantifies how much each input drives predictions overall; SHAP (SHapley Additive exPlanations) attributes a single prediction to its features:

import pandas as pd
import shap
from sklearn.ensemble import RandomForestClassifier

clf = RandomForestClassifier(n_estimators=200, random_state=0).fit(X_tr, y_tr)

# Global: which features matter overall.
imp = pd.Series(clf.feature_importances_, index=feature_names)
print(imp.sort_values(ascending=False).head(5))

# Per-decision: why THIS applicant was denied.
shap_values = shap.TreeExplainer(clf).shap_values(X_te[:1])

That turns "denied" into "denied because credit_utilisation was 0.92 and there were 3 missed payments." LLMs can write an explanation, but it is a plausible narrative generated after the fact — not a faithful account of what the model did.

Need Classical ML LLM
Reason code SHAP, coefficients — faithful Narrative — not faithful
Reproducible Deterministic Can vary
Regulatory acceptance Established Often insufficient
WARNING
Generated explanations are not faithful
When an LLM tells you why it answered, treat it as a hypothesis — the text is generated to be plausible, not to report the model's reasoning. Where the explanation must be auditable, that gap disqualifies LLM-only decisions.

6 Data Availability — The Quiet Deciding Factor

Before any other dimension, ask: what labelled data do I have? It decides more architectures than latency or cost.

  • Thousands of labels, stable target → classical ML or task-specific deep learning likely beats an LLM on accuracy, cost, and speed.
  • A few hundred labels → use transfer learning: take a model pre-trained on a huge generic dataset and fine-tune its final layers on your small one, so it arrives already knowing general features.
  • Zero or a handful → an LLM's zero-shot ability is hard to beat as a start.

Transfer learning with a Hugging Face BERT-family encoder (a transformer trained to understand text, not generate it):

from transformers import AutoTokenizer, AutoModelForSequenceClassification

# Add a classification head to a pre-trained encoder, then fine-tune only
# on our own labels — far fewer than training from scratch.
name = "distilbert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name, num_labels=3)
# Fine-tune with the Trainer API on a few thousand rows -> a tiny, fast,
# cheap model you own and can run on a CPU.

A fine-tuned DistilBERT classifier runs in milliseconds, costs nothing per call, and on a narrow task often matches or beats a general LLM — specialised for exactly your labels.

NOTE
A useful sequence
Many teams bootstrap with the LLM, then graduate: prototype zero-shot, use the LLM to help label a few thousand examples, then train a small specialised model. The LLM's fast start, classical ML's cheap steady state — exactly what the next lesson builds.

7 A Decision Tree You Can Apply on Monday

Walk this top to bottom; the first hard constraint that fires usually wins.

1. Hard latency/throughput SLA (sub-10ms, high QPS)?  -> Classical ML. STOP.
2. Per-call cost dominates at volume (millions+/day)? -> Classical ML. STOP.
3. Regulated, needs faithful reason codes?            -> Classical ML. STOP.
4. Unstructured text/images AND labels scarce?        -> LLM / foundation model.
5. Thousands of labels, stable narrow target?         -> Classical / deep learning.
6. Open-ended, generative, fast-changing?             -> LLM.
   Otherwise -> prototype with an LLM, measure, decide.

Worked example — "predict next month's demand from 5 years of daily sales": no hard SLA, modest volume, not regulated (1-3 pass); structured numeric series (4 → no); thousands of points, stable target (5 → YES). Verdict: classical time-series methods (see Lesson 5: Time Series & Forecasting), not an LLM. A pocket lookup:

Problem signature Right tool
Tabular, high-volume, regulated Gradient-boosted trees
Numeric forecasting ARIMA / Prophet / temporal nets
Object detection on a camera feed YOLO / task-specific CV model
Open-ended text triage, few labels LLM (zero-shot)
Few hundred labels, narrow task Transfer learning (fine-tune BERT)
TIP
Re-run the tree when the world changes
A decision is only valid for the constraints you scored it against. When volume jumps 100x, an SLA tightens, or a regulator gets involved, re-walk the tree — last quarter's answer can be this quarter's mistake.

Questions & Answers

Q: My LLM prototype already works. Why rebuild it as a classical model?
You might not. If volume stays low and there is no SLA or audit requirement, the prototype can be the product. Rebuild only when a constraint bites: per-call cost becomes a real line item, latency misses an SLA, or a regulator needs reason codes — a measured number crossing a threshold, not aesthetics.
Q: Isn't an LLM more accurate than a little logistic regression?
Not on its home turf. For structured tabular data and narrow, well-labelled tasks, a tuned gradient-boosted tree routinely matches or beats a general LLM — it was specialised for your features. The LLM's advantage is breadth (open-ended tasks, unstructured input, zero labels), not accuracy on a fixed schema. Measure both with precision/recall, not accuracy, for imbalanced data.
Q: How do I estimate per-prediction LLM cost before committing?
Count input and output tokens with the SDK's tokenizer, multiply by your provider's per-token rates, then by projected daily volume. Do the same for a classical model: roughly zero marginal cost plus the amortised cost of the box. The two usually differ by orders of magnitude at scale — build this estimate before the architecture is locked.
Q: We have no labelled data at all. Does that force an LLM forever?
No — it is a starting point, not a destination. Use the LLM zero-shot now, and use it to help label a few thousand examples cheaply. Once you have labels, transfer-learn a small specialised model for the steady state. You get the LLM's instant start and classical ML's cheap, ownable production path — Lesson 9 builds this hybrid.
Q: Can't I ask the LLM to explain its decision and use that as a reason code?
Not for anything audited. An LLM-generated explanation is text optimised to sound plausible, not a faithful report of the computation — the two can diverge and you cannot prove they agree. For regulated decisions you need an explanation that provably reflects the model, which is what SHAP values and coefficients give you. Treat generated explanations as hypotheses, not compliance artefacts.

Key Takeaways

  1. An LLM is a default, not a decision. Score every problem on latency, cost, interpretability, data, accuracy, and maintenance — "just use an LLM" carries a real, recurring tax.
  2. The latency and cost gaps are architectural. A classical classifier predicts in microseconds for free; an LLM call takes hundreds of milliseconds and bills per token. You cannot prompt across four orders of magnitude.
  3. LLMs lose on high-throughput, tabular, regulated work. Fraud scoring wants fast, cheap, explainable gradient-boosted trees — classical ML is strictly better, not a compromise.
  4. LLMs win on open-ended, unstructured, low-label work. Zero-shot triage of free-form text with shifting categories is where editing one prompt line beats labelling thousands of examples.
  5. Data availability is the quiet deciding factor. Many labels favour classical ML; a few hundred favour transfer learning; near-zero favours an LLM start that later graduates to a specialised model.
  6. Faithful explanations only come from classical models. SHAP values and coefficients give auditable reason codes; an LLM's self-explanation is a plausible narrative, not proof.

Next Steps: Lesson 9: Combining Classical ML + LLMs