Combining Classical ML + LLMs
Learning Outcomes
- Identify the four core hybrid patterns and match each to a real production problem
- Build a fast classifier that routes easy cases to cheap ML and hard cases to an LLM
- Use an LLM as a feature extractor whose embeddings feed a classical scikit-learn model
- Pre-filter a high-volume stream with ML so the LLM only sees what matters
- Bootstrap a labelled dataset with LLM-generated training data and validate it honestly
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why hybrid beats "LLM for everything" |
| Explain | 8 min | The four hybrid patterns and when each fits |
| Build | 11 min | Pattern 1: routing with a fast classifier |
| Build | 11 min | Pattern 2: LLM feature extraction into a classical model |
| Build | 10 min | Pattern 3: ML pre-filtering before the LLM |
| Build | 10 min | Pattern 4: LLM-generated training data |
| Explain | 4 min | Evaluating the whole hybrid system |
| Wrap-up | 3 min | Key takeaways and what comes next |
Before You Begin
Pre-work:
- Be comfortable with the ML workflow and metrics from Lesson 2: Machine Learning Fundamentals
- Understand encoder models from Lesson 4: NLP Beyond Chat
- Know the decision dimensions from Lesson 8: When to Use Classical ML vs GenAI
Shopping List:
- Python 3.10+ and a virtual environment
pip install scikit-learn sentence-transformers pandas numpy— and one LLM client (openai,anthropic, orollama)- An LLM you can call: a hosted API key, or a local model from the Local LLMs course if you prefer to stay offline
- A CPU is fine for everything here; the embedding model downloads ~90 MB on first run
The instinct after a year of chatbots is to throw an LLM at every text problem. It works in a demo and bankrupts you in production. An LLM call is orders of magnitude slower and more expensive than a logistic regression — the gap between milliseconds and seconds, between fractions of a cent and a real per-request cost at scale.
A hybrid architecture uses each tool where it is strongest: classical ML for the high-volume, well-defined, latency-sensitive 95%, and the LLM for the messy, reasoning-heavy 5% that classical models cannot handle. Two terms, used throughout:
| Term | Plain-English meaning |
|---|---|
| Classical ML | A model trained on labelled examples to map inputs to outputs — e.g. logistic regression, gradient boosting. Fast, cheap, narrow. |
| Embedding | A list of floats that captures the meaning of an input, so similar inputs land near each other in vector space. |
Four patterns are worth knowing. They are not mutually exclusive — real systems stack them.
| Pattern | What it does | Wins you get |
|---|---|---|
| Routing | Cheap classifier decides which requests need the LLM | Lower cost, lower latency |
| Feature extraction | LLM/encoder turns text into embeddings; classical model decides | Small labelled data goes far |
| Pre-filtering | ML discards or triages the firehose before the LLM | LLM only sees what matters |
| Synthetic data | LLM generates labelled examples to train a classical model | Bootstrap with no labels |
Imagine a customer-support inbox. Most tickets are routine ("reset my password", "where's my order") and a cheap classifier handles them with a canned action. A minority are genuinely complex and deserve an LLM. A router is a fast model that decides per request: handle it locally, or escalate.
We will train a router on TF-IDF features. TF-IDF (term frequency–inverse document frequency) turns text into a sparse numeric vector that weights words by how distinctive they are — common words count for little, rare-but-meaningful words count for a lot. It runs in microseconds and is shockingly strong for routing.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
# Toy labelled tickets: 0 = routine (handle locally), 1 = complex (send to LLM)
tickets = [
"reset my password please", "where is my order", "change my email address",
"cancel subscription", "update billing card", "track my package",
"your product damaged my equipment and I want legal options",
"I was charged three times and the refund logic seems broken across currencies",
"explain why my enterprise SSO integration fails intermittently under load",
"I need a custom data export combining usage, billing, and audit logs",
]
labels = [0, 0, 0, 0, 0, 0, 1, 1, 1, 1]
router = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2)),
LogisticRegression(max_iter=1000),
)
router.fit(tickets, labels)
The router exposes a probability, not just a label. That probability is the dial you tune: route confidently-easy tickets to ML, confidently-hard ones to the LLM, and send the uncertain middle to the LLM too (when in doubt, pay for quality).
def handle(text, escalate_above=0.5):
p_complex = router.predict_proba([text])[0][1]
if p_complex >= escalate_above:
return call_llm(text) # the expensive path
return canned_response(text) # the near-free path
def canned_response(text):
return "Resolved by automated workflow."
def call_llm(text):
# Stand-in for your real LLM client; see Step 3 for a concrete call
return f"[LLM reasoning over]: {text}"
print(handle("where is my order")) # near-free
print(handle("refund logic seems broken across currencies")) # escalated
escalate_above and more traffic hits the LLM — higher quality, higher bill. Raise it and you save money but risk fumbling a hard ticket. Plot escalation rate against cost on real traffic and pick the knee of the curve, not a round number.The router is only half the pattern; the escalation path must actually invoke an LLM and return something structured your system can act on. Ask the LLM for JSON so the response is machine-parseable rather than prose you have to scrape.
import json
from openai import OpenAI # swap for anthropic or ollama; the shape is identical
client = OpenAI() # reads OPENAI_API_KEY from the environment
SYSTEM = (
"You are a support triage assistant. Read the ticket and respond with "
"JSON only: an 'intent' string, an 'urgency' of low|medium|high, and a "
"one-sentence 'suggested_action'. No prose outside the JSON."
)
def call_llm(text):
resp = client.chat.completions.create(
model="gpt-4o-mini",
response_format={"type": "json_object"},
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": text},
],
)
return json.loads(resp.choices[0].message.content)
Here is the JSON contract the LLM must satisfy — define it explicitly so downstream code can rely on it:
{
"intent": "billing_dispute",
"urgency": "high",
"suggested_action": "Escalate to billing engineering and issue provisional refund."
}
Running locally? The same logic with Ollama is a one-line swap (from ollama import chat), keeping every byte on your machine. Either way, the LLM only sees the ~5% of traffic the router escalated, so a slow, pricey call becomes affordable because it is rare.
json.loads in a try/except and fall back to a safe default so one malformed response never crashes the pipeline.Sometimes you have a clean classification task but very few labels — say a few hundred examples. Training a deep model from scratch would overfit (memorise the training set and fail on new data) instantly. The fix is to let a pretrained encoder do the heavy lifting: it turns each text into a rich embedding, and a tiny classical model learns the decision boundary on top. This is transfer learning — reusing knowledge a big model already learned, instead of starting from zero.
We use sentence-transformers, a library of encoder models (BERT-family) that map a sentence to a fixed-length embedding. The encoder is frozen; only the cheap classifier trains.
from sentence_transformers import SentenceTransformer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
import numpy as np
encoder = SentenceTransformer("all-MiniLM-L6-v2") # ~90 MB, 384-dim vectors
texts = [
"The product arrived broken and support ignored me",
"Absolutely love this, best purchase of the year",
"It's fine, does the job, nothing special",
"Terrible experience, requesting a full refund",
"Fast shipping and exactly as described, very happy",
"Mediocre quality for the price, would not rebuy",
]
y = ["neg", "pos", "neutral", "neg", "pos", "neutral"]
X = encoder.encode(texts) # shape: (6, 384) — meaning vectors
clf = LogisticRegression(max_iter=1000)
The embeddings carry the semantic load, so logistic regression — a model with a few hundred parameters — is enough. To check it generalises rather than memorises, use cross-validation: split the data into folds, train on most, test on the held-out fold, rotate, and average. It is the honest estimate of real-world performance.
scores = cross_val_score(clf, X, y, cv=3)
print(f"CV accuracy: {scores.mean():.2f} (+/- {scores.std():.2f})")
clf.fit(X, y)
print(clf.predict(encoder.encode(["this thing is a complete disaster"])))
# -> ['neg']
| Approach for small labelled data | Inference cost per item | Typical fit |
|---|---|---|
| Pure LLM zero-shot prompt | High (one API call) | Prototypes, no labels at all |
| LLM embeddings + classical head | Very low (one matmul) | Hundreds of labels, high volume |
| Fine-tune a full transformer | Low at inference, high to train | Thousands of labels, in-house |
When the input is a firehose — millions of log lines, social posts, transactions — you cannot afford to LLM every record. Pre-filtering puts a cheap model in front that discards the obvious noise, so the LLM only reasons over the rare candidates that survive.
Two complementary filters do most of the work. First, an anomaly detector flags records that look unusual. Isolation Forest is a classic: it isolates outliers by randomly partitioning the data — anomalies need fewer splits to isolate, so they score as suspicious. It is unsupervised, so it needs no labels.
from sklearn.ensemble import IsolationForest
import numpy as np
# Each row: [amount, hour_of_day, num_prior_txns] — normal-ish traffic
normal = np.random.normal(loc=[50, 13, 8], scale=[15, 4, 3], size=(500, 3))
detector = IsolationForest(contamination=0.02, random_state=0)
detector.fit(normal)
incoming = np.array([
[52, 14, 9], # ordinary
[9999, 3, 1], # huge amount, 3am, first txn -> likely anomaly
])
flags = detector.predict(incoming) # -1 = anomaly, 1 = normal
print(flags) # -> [ 1 -1]
Only the rows the detector flags get the expensive treatment. That is the whole budget trick: contamination of 2% means roughly 98% of records never touch the LLM.
def review_pipeline(rows, raw_texts):
flags = detector.predict(rows)
results = []
for row, text, flag in zip(rows, raw_texts, flags):
if flag == -1: # anomalous -> reason about it
results.append(call_llm(text)) # from Step 3
else:
results.append({"intent": "normal", "urgency": "low",
"suggested_action": "auto-approve"})
return results
The same shape applies to text: a TF-IDF keyword filter or the Step-2 router can pre-screen so the LLM only sees plausible candidates. Pre-filtering and routing are cousins — routing chooses a path, pre-filtering discards before any path.
contamination generously and let the LLM be the precise second stage that rejects the false alarms.The classic chicken-and-egg of ML: you want a fast classifier, but have no labelled data to train it on. An LLM breaks the deadlock by generating synthetic training data — plausible labelled examples in your target categories — which you then use to train a cheap, fast classical model for production. You pay for the LLM once, at training time, not on every request.
import json
from openai import OpenAI
client = OpenAI()
def generate_examples(category, n=15):
prompt = (
f"Generate {n} short, realistic customer support messages with the "
f"intent '{category}'. Vary tone, length, and phrasing. "
"Return a JSON array of strings only."
)
resp = client.chat.completions.create(
model="gpt-4o-mini",
response_format={"type": "json_object"},
messages=[{"role": "user",
"content": prompt + " Wrap the array under key 'examples'."}],
)
return json.loads(resp.choices[0].message.content)["examples"]
categories = ["password_reset", "billing_dispute", "shipping_delay"]
texts, labels = [], []
for cat in categories:
for ex in generate_examples(cat):
texts.append(ex)
labels.append(cat)
Now train the same TF-IDF + logistic regression from Step 2 on this synthetic corpus, and you have a real-time classifier born from zero hand-labelled data.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
intent_clf = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2)),
LogisticRegression(max_iter=1000),
)
intent_clf.fit(texts, labels)
print(intent_clf.predict(["I never got the link to set a new password"]))
# -> ['password_reset']
The catch is distribution shift: LLM-generated text is cleaner and more templated than the messy reality of real users, so a model trained only on synthetic data can disappoint on production traffic. Treat synthetic data as a bootstrap, then mix in real examples as soon as you have them.
| Synthetic-data risk | Symptom | Mitigation |
|---|---|---|
| Too uniform | Model brittle to real phrasing | Crank up the LLM diversity prompt; raise temperature |
| Label noise | LLM mislabels its own examples | Spot-check a sample by hand before training |
| Distribution shift | Great offline, weak in prod | Validate on a small real labelled holdout set |
A hybrid system has more than one accuracy number — it has the classical stage's metrics, the LLM stage's quality, and the system metrics of cost and latency. Optimising one in isolation is how you ship something that scores well and loses money.
Start with the right per-class metrics. For an imbalanced router, accuracy lies; use precision and recall. Precision answers "of the items I flagged complex, how many really were?" Recall answers "of the truly complex items, how many did I catch?" The router's recall on the complex class is the one to guard — a missed escalation is a customer the system failed.
from sklearn.metrics import classification_report
import numpy as np
y_true = ["complex", "routine", "complex", "routine", "complex"]
y_pred = ["complex", "routine", "routine", "routine", "complex"] # one miss
print(classification_report(y_true, y_pred, digits=2))
Then layer the system economics on top — the numbers that decide whether the hybrid is worth it:
def system_report(n_requests, escalation_rate,
llm_cost=0.01, ml_cost=0.00001,
llm_latency_ms=1200, ml_latency_ms=2):
n_llm = n_requests * escalation_rate
n_ml = n_requests - n_llm
cost = n_llm * llm_cost + n_ml * ml_cost
all_llm_cost = n_requests * llm_cost
avg_latency = (n_llm * llm_latency_ms + n_ml * ml_latency_ms) / n_requests
return {
"hybrid_cost": round(cost, 2),
"all_llm_cost": round(all_llm_cost, 2),
"savings_pct": round(100 * (1 - cost / all_llm_cost), 1),
"avg_latency_ms": round(avg_latency, 1),
}
print(system_report(n_requests=100_000, escalation_rate=0.05))
With 5% escalation, the hybrid pays for the LLM on 5,000 requests instead of 100,000 — a roughly 95% cost cut and a fraction of the average latency, while the LLM still handles every genuinely hard case.
Questions & Answers
contamination or add discriminating features. If the filter is well tuned but the residual volume is still expensive, add a cheaper middle stage (a small classical classifier) between the filter and the LLM so only the hardest survivors reach the generalist.Key Takeaways
- Hybrid beats "LLM for everything" — pay the LLM tax only on the small slice of inputs that genuinely need a generalist; let trained specialists handle the high-volume rest.
- Routing trades a tunable threshold for cost — a fast TF-IDF classifier decides what escalates, and the threshold is your direct lever on the quality-vs-bill curve.
- Embeddings + a classical head stretch small data — let a frozen encoder carry the semantics so a tiny logistic regression generalises from hundreds of labels, with near-zero inference cost.
- Pre-filtering protects the budget — an unsupervised anomaly detector discards 98% of the firehose so the LLM only reasons over rare candidates; tune it for recall, not precision.
- Synthetic data bootstraps from zero labels — generate examples with the LLM once at training time, but always validate on a real holdout to catch distribution shift.
- Evaluate the system, not the stage — track hard-class precision/recall alongside cost per 1,000 requests and p95 latency; the hybrid only wins if quality holds while both drop.
Next Steps: Lesson 10: Careers in AI