Combining Classical ML + LLMs

60 min advanced Lesson 9

Learning Outcomes

  • Identify the four core hybrid patterns and match each to a real production problem
  • Build a fast classifier that routes easy cases to cheap ML and hard cases to an LLM
  • Use an LLM as a feature extractor whose embeddings feed a classical scikit-learn model
  • Pre-filter a high-volume stream with ML so the LLM only sees what matters
  • Bootstrap a labelled dataset with LLM-generated training data and validate it honestly

Lesson Plan

Segment Duration Topic
Intro 3 min Why hybrid beats "LLM for everything"
Explain 8 min The four hybrid patterns and when each fits
Build 11 min Pattern 1: routing with a fast classifier
Build 11 min Pattern 2: LLM feature extraction into a classical model
Build 10 min Pattern 3: ML pre-filtering before the LLM
Build 10 min Pattern 4: LLM-generated training data
Explain 4 min Evaluating the whole hybrid system
Wrap-up 3 min Key takeaways and what comes next

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ and a virtual environment
  • pip install scikit-learn sentence-transformers pandas numpy — and one LLM client (openai, anthropic, or ollama)
  • An LLM you can call: a hosted API key, or a local model from the Local LLMs course if you prefer to stay offline
  • A CPU is fine for everything here; the embedding model downloads ~90 MB on first run

1 Why Hybrid — The Cost-Performance Sweet Spot

The instinct after a year of chatbots is to throw an LLM at every text problem. It works in a demo and bankrupts you in production. An LLM call is orders of magnitude slower and more expensive than a logistic regression — the gap between milliseconds and seconds, between fractions of a cent and a real per-request cost at scale.

A hybrid architecture uses each tool where it is strongest: classical ML for the high-volume, well-defined, latency-sensitive 95%, and the LLM for the messy, reasoning-heavy 5% that classical models cannot handle. Two terms, used throughout:

Term Plain-English meaning
Classical ML A model trained on labelled examples to map inputs to outputs — e.g. logistic regression, gradient boosting. Fast, cheap, narrow.
Embedding A list of floats that captures the meaning of an input, so similar inputs land near each other in vector space.

Four patterns are worth knowing. They are not mutually exclusive — real systems stack them.

Pattern What it does Wins you get
Routing Cheap classifier decides which requests need the LLM Lower cost, lower latency
Feature extraction LLM/encoder turns text into embeddings; classical model decides Small labelled data goes far
Pre-filtering ML discards or triages the firehose before the LLM LLM only sees what matters
Synthetic data LLM generates labelled examples to train a classical model Bootstrap with no labels
NOTE
The core trade
An LLM is a generalist you rent by the token; a classical model is a specialist you train once and run nearly free. Hybrid systems pay the LLM tax only on the inputs that actually need a generalist, and let the trained specialist handle the rest.

2 Pattern 1 — Routing With a Fast Classifier

Imagine a customer-support inbox. Most tickets are routine ("reset my password", "where's my order") and a cheap classifier handles them with a canned action. A minority are genuinely complex and deserve an LLM. A router is a fast model that decides per request: handle it locally, or escalate.

We will train a router on TF-IDF features. TF-IDF (term frequency–inverse document frequency) turns text into a sparse numeric vector that weights words by how distinctive they are — common words count for little, rare-but-meaningful words count for a lot. It runs in microseconds and is shockingly strong for routing.

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline

# Toy labelled tickets: 0 = routine (handle locally), 1 = complex (send to LLM)
tickets = [
    "reset my password please", "where is my order", "change my email address",
    "cancel subscription", "update billing card", "track my package",
    "your product damaged my equipment and I want legal options",
    "I was charged three times and the refund logic seems broken across currencies",
    "explain why my enterprise SSO integration fails intermittently under load",
    "I need a custom data export combining usage, billing, and audit logs",
]
labels = [0, 0, 0, 0, 0, 0, 1, 1, 1, 1]

router = make_pipeline(
    TfidfVectorizer(ngram_range=(1, 2)),
    LogisticRegression(max_iter=1000),
)
router.fit(tickets, labels)

The router exposes a probability, not just a label. That probability is the dial you tune: route confidently-easy tickets to ML, confidently-hard ones to the LLM, and send the uncertain middle to the LLM too (when in doubt, pay for quality).

def handle(text, escalate_above=0.5):
    p_complex = router.predict_proba([text])[0][1]
    if p_complex >= escalate_above:
        return call_llm(text)          # the expensive path
    return canned_response(text)       # the near-free path

def canned_response(text):
    return "Resolved by automated workflow."

def call_llm(text):
    # Stand-in for your real LLM client; see Step 3 for a concrete call
    return f"[LLM reasoning over]: {text}"

print(handle("where is my order"))                       # near-free
print(handle("refund logic seems broken across currencies"))  # escalated
TIP
Tune the threshold to your budget
Lower escalate_above and more traffic hits the LLM — higher quality, higher bill. Raise it and you save money but risk fumbling a hard ticket. Plot escalation rate against cost on real traffic and pick the knee of the curve, not a round number.
WARNING
The router itself can be wrong
A misrouted complex ticket gets a canned reply and an angry customer. Measure router recall on the complex class specifically (Step 7), and bias the threshold toward escalation when the cost of a miss is high.

3 Make the Escalation Path a Real LLM Call

The router is only half the pattern; the escalation path must actually invoke an LLM and return something structured your system can act on. Ask the LLM for JSON so the response is machine-parseable rather than prose you have to scrape.

import json
from openai import OpenAI   # swap for anthropic or ollama; the shape is identical

client = OpenAI()   # reads OPENAI_API_KEY from the environment

SYSTEM = (
    "You are a support triage assistant. Read the ticket and respond with "
    "JSON only: an 'intent' string, an 'urgency' of low|medium|high, and a "
    "one-sentence 'suggested_action'. No prose outside the JSON."
)

def call_llm(text):
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        response_format={"type": "json_object"},
        messages=[
            {"role": "system", "content": SYSTEM},
            {"role": "user", "content": text},
        ],
    )
    return json.loads(resp.choices[0].message.content)

Here is the JSON contract the LLM must satisfy — define it explicitly so downstream code can rely on it:

{
  "intent": "billing_dispute",
  "urgency": "high",
  "suggested_action": "Escalate to billing engineering and issue provisional refund."
}

Running locally? The same logic with Ollama is a one-line swap (from ollama import chat), keeping every byte on your machine. Either way, the LLM only sees the ~5% of traffic the router escalated, so a slow, pricey call becomes affordable because it is rare.

NOTE
Always pin a schema
Forcing JSON output (and validating it on receipt) is what turns an LLM from a chatbot into a reliable system component. Wrap json.loads in a try/except and fall back to a safe default so one malformed response never crashes the pipeline.

4 Pattern 2 — LLM Feature Extraction Into a Classical Model

Sometimes you have a clean classification task but very few labels — say a few hundred examples. Training a deep model from scratch would overfit (memorise the training set and fail on new data) instantly. The fix is to let a pretrained encoder do the heavy lifting: it turns each text into a rich embedding, and a tiny classical model learns the decision boundary on top. This is transfer learning — reusing knowledge a big model already learned, instead of starting from zero.

We use sentence-transformers, a library of encoder models (BERT-family) that map a sentence to a fixed-length embedding. The encoder is frozen; only the cheap classifier trains.

from sentence_transformers import SentenceTransformer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
import numpy as np

encoder = SentenceTransformer("all-MiniLM-L6-v2")  # ~90 MB, 384-dim vectors

texts = [
    "The product arrived broken and support ignored me",
    "Absolutely love this, best purchase of the year",
    "It's fine, does the job, nothing special",
    "Terrible experience, requesting a full refund",
    "Fast shipping and exactly as described, very happy",
    "Mediocre quality for the price, would not rebuy",
]
y = ["neg", "pos", "neutral", "neg", "pos", "neutral"]

X = encoder.encode(texts)          # shape: (6, 384) — meaning vectors
clf = LogisticRegression(max_iter=1000)

The embeddings carry the semantic load, so logistic regression — a model with a few hundred parameters — is enough. To check it generalises rather than memorises, use cross-validation: split the data into folds, train on most, test on the held-out fold, rotate, and average. It is the honest estimate of real-world performance.

scores = cross_val_score(clf, X, y, cv=3)
print(f"CV accuracy: {scores.mean():.2f} (+/- {scores.std():.2f})")
clf.fit(X, y)
print(clf.predict(encoder.encode(["this thing is a complete disaster"])))
# -> ['neg']
Approach for small labelled data Inference cost per item Typical fit
Pure LLM zero-shot prompt High (one API call) Prototypes, no labels at all
LLM embeddings + classical head Very low (one matmul) Hundreds of labels, high volume
Fine-tune a full transformer Low at inference, high to train Thousands of labels, in-house
TIP
Cache the embeddings
Encoding is the slow part. Compute embeddings once, store the vectors (a numpy array or a parquet file), and retrain the classical head in seconds whenever labels change. You only re-encode when the raw text changes.
WARNING
Match encoder to domain
A general-purpose encoder may miss jargon in legal, medical, or code text. If accuracy stalls, try a domain-tuned encoder before reaching for a bigger model — the embedding quality, not the classifier, is usually the ceiling.

5 Pattern 3 — ML Pre-Filtering Before the LLM

When the input is a firehose — millions of log lines, social posts, transactions — you cannot afford to LLM every record. Pre-filtering puts a cheap model in front that discards the obvious noise, so the LLM only reasons over the rare candidates that survive.

Two complementary filters do most of the work. First, an anomaly detector flags records that look unusual. Isolation Forest is a classic: it isolates outliers by randomly partitioning the data — anomalies need fewer splits to isolate, so they score as suspicious. It is unsupervised, so it needs no labels.

from sklearn.ensemble import IsolationForest
import numpy as np

# Each row: [amount, hour_of_day, num_prior_txns]  — normal-ish traffic
normal = np.random.normal(loc=[50, 13, 8], scale=[15, 4, 3], size=(500, 3))
detector = IsolationForest(contamination=0.02, random_state=0)
detector.fit(normal)

incoming = np.array([
    [52, 14, 9],      # ordinary
    [9999, 3, 1],     # huge amount, 3am, first txn -> likely anomaly
])
flags = detector.predict(incoming)   # -1 = anomaly, 1 = normal
print(flags)   # -> [ 1 -1]

Only the rows the detector flags get the expensive treatment. That is the whole budget trick: contamination of 2% means roughly 98% of records never touch the LLM.

def review_pipeline(rows, raw_texts):
    flags = detector.predict(rows)
    results = []
    for row, text, flag in zip(rows, raw_texts, flags):
        if flag == -1:                       # anomalous -> reason about it
            results.append(call_llm(text))   # from Step 3
        else:
            results.append({"intent": "normal", "urgency": "low",
                            "suggested_action": "auto-approve"})
    return results

The same shape applies to text: a TF-IDF keyword filter or the Step-2 router can pre-screen so the LLM only sees plausible candidates. Pre-filtering and routing are cousins — routing chooses a path, pre-filtering discards before any path.

NOTE
Tune for recall on the rare class
A fraud or incident filter must rarely miss a true positive, even at the cost of extra false alarms (those just become a few more LLM calls). Set contamination generously and let the LLM be the precise second stage that rejects the false alarms.
WARNING
The filter sees a distribution, not a label
Isolation Forest flags unusual, which is not the same as bad. A legitimate Black Friday spike looks anomalous too. Always have the LLM (or a human) make the final call on flagged items — the filter only earns the right to ask the question.

6 Pattern 4 — LLM-Generated Training Data

The classic chicken-and-egg of ML: you want a fast classifier, but have no labelled data to train it on. An LLM breaks the deadlock by generating synthetic training data — plausible labelled examples in your target categories — which you then use to train a cheap, fast classical model for production. You pay for the LLM once, at training time, not on every request.

import json
from openai import OpenAI

client = OpenAI()

def generate_examples(category, n=15):
    prompt = (
        f"Generate {n} short, realistic customer support messages with the "
        f"intent '{category}'. Vary tone, length, and phrasing. "
        "Return a JSON array of strings only."
    )
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        response_format={"type": "json_object"},
        messages=[{"role": "user",
                   "content": prompt + " Wrap the array under key 'examples'."}],
    )
    return json.loads(resp.choices[0].message.content)["examples"]

categories = ["password_reset", "billing_dispute", "shipping_delay"]
texts, labels = [], []
for cat in categories:
    for ex in generate_examples(cat):
        texts.append(ex)
        labels.append(cat)

Now train the same TF-IDF + logistic regression from Step 2 on this synthetic corpus, and you have a real-time classifier born from zero hand-labelled data.

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline

intent_clf = make_pipeline(
    TfidfVectorizer(ngram_range=(1, 2)),
    LogisticRegression(max_iter=1000),
)
intent_clf.fit(texts, labels)
print(intent_clf.predict(["I never got the link to set a new password"]))
# -> ['password_reset']

The catch is distribution shift: LLM-generated text is cleaner and more templated than the messy reality of real users, so a model trained only on synthetic data can disappoint on production traffic. Treat synthetic data as a bootstrap, then mix in real examples as soon as you have them.

Synthetic-data risk Symptom Mitigation
Too uniform Model brittle to real phrasing Crank up the LLM diversity prompt; raise temperature
Label noise LLM mislabels its own examples Spot-check a sample by hand before training
Distribution shift Great offline, weak in prod Validate on a small real labelled holdout set
WARNING
Never validate on synthetic data alone
Synthetic train, synthetic test will flatter you with a fake high score. Hold out a small set of real labelled examples for evaluation, even if you trained entirely on synthetic. The real holdout is the only number you should trust.
TIP
Use the LLM to label, not just to generate
A variant: feed real, unlabelled production text to the LLM and have it assign labels (weak supervision). You get real-distribution inputs with LLM-quality labels — often a stronger bootstrap than fully synthetic text.

7 Evaluate the Hybrid System as a Whole

A hybrid system has more than one accuracy number — it has the classical stage's metrics, the LLM stage's quality, and the system metrics of cost and latency. Optimising one in isolation is how you ship something that scores well and loses money.

Start with the right per-class metrics. For an imbalanced router, accuracy lies; use precision and recall. Precision answers "of the items I flagged complex, how many really were?" Recall answers "of the truly complex items, how many did I catch?" The router's recall on the complex class is the one to guard — a missed escalation is a customer the system failed.

from sklearn.metrics import classification_report
import numpy as np

y_true = ["complex", "routine", "complex", "routine", "complex"]
y_pred = ["complex", "routine", "routine", "routine", "complex"]  # one miss
print(classification_report(y_true, y_pred, digits=2))

Then layer the system economics on top — the numbers that decide whether the hybrid is worth it:

def system_report(n_requests, escalation_rate,
                  llm_cost=0.01, ml_cost=0.00001,
                  llm_latency_ms=1200, ml_latency_ms=2):
    n_llm = n_requests * escalation_rate
    n_ml = n_requests - n_llm
    cost = n_llm * llm_cost + n_ml * ml_cost
    all_llm_cost = n_requests * llm_cost
    avg_latency = (n_llm * llm_latency_ms + n_ml * ml_latency_ms) / n_requests
    return {
        "hybrid_cost": round(cost, 2),
        "all_llm_cost": round(all_llm_cost, 2),
        "savings_pct": round(100 * (1 - cost / all_llm_cost), 1),
        "avg_latency_ms": round(avg_latency, 1),
    }

print(system_report(n_requests=100_000, escalation_rate=0.05))

With 5% escalation, the hybrid pays for the LLM on 5,000 requests instead of 100,000 — a roughly 95% cost cut and a fraction of the average latency, while the LLM still handles every genuinely hard case.

NOTE
Three numbers, one decision
Track quality (precision/recall on the hard class), cost per 1,000 requests, and p95 latency together. The hybrid wins only if quality holds while cost and latency drop. If quality sags, lower the escalation threshold; if cost balloons, raise it. You are steering, not setting-and-forgetting.

Questions & Answers

Q: Isn't a hybrid system just more moving parts to maintain than one LLM call?
Yes, and that is a real cost — you now own a classifier, its training data, and a routing threshold. The trade is worth it once volume is high enough that the LLM bill or latency hurts. Below a few thousand requests a day, a single LLM call may genuinely be the right, simpler choice. Hybrid is an optimisation you reach for when scale forces it, not a default.
Q: How do I keep the router from rotting as user behaviour changes?
Log a sample of routed requests and their outcomes, then periodically check whether the router's confidence still tracks reality. When the complex-class recall drifts down, retrain on fresh labelled examples. This is the drift monitoring covered in Lesson 7: MLOps — a hybrid system needs the same retraining discipline as any production model.
Q: If I'm using LLM embeddings as features, why not just prompt the LLM to classify directly?
Because the embedding + classical head is far cheaper and faster at inference, and it gives you a calibrated probability and an interpretable, retrainable decision boundary. Direct prompting is better when you have almost no labels or the categories change constantly. With a few hundred stable labels, the classical head on top of embeddings usually wins on cost, latency, and consistency.
Q: Won't synthetic training data just teach my model the LLM's biases and blind spots?
It can. Synthetic data inherits the generator's quirks and is cleaner than real text. That is exactly why you validate on a real labelled holdout and mix in real examples as soon as they exist. Treat the synthetic corpus as a cold-start bridge to get a v1 model live, not as a permanent substitute for real data.
Q: My pre-filter passes too much to the LLM and the bill is still high. What now?
First confirm the filter is the bottleneck by logging pass-through rate. If it's too high, your anomaly threshold is too loose or your features don't separate signal from noise — tighten contamination or add discriminating features. If the filter is well tuned but the residual volume is still expensive, add a cheaper middle stage (a small classical classifier) between the filter and the LLM so only the hardest survivors reach the generalist.

Key Takeaways

  1. Hybrid beats "LLM for everything" — pay the LLM tax only on the small slice of inputs that genuinely need a generalist; let trained specialists handle the high-volume rest.
  2. Routing trades a tunable threshold for cost — a fast TF-IDF classifier decides what escalates, and the threshold is your direct lever on the quality-vs-bill curve.
  3. Embeddings + a classical head stretch small data — let a frozen encoder carry the semantics so a tiny logistic regression generalises from hundreds of labels, with near-zero inference cost.
  4. Pre-filtering protects the budget — an unsupervised anomaly detector discards 98% of the firehose so the LLM only reasons over rare candidates; tune it for recall, not precision.
  5. Synthetic data bootstraps from zero labels — generate examples with the LLM once at training time, but always validate on a real holdout to catch distribution shift.
  6. Evaluate the system, not the stage — track hard-class precision/recall alongside cost per 1,000 requests and p95 latency; the hybrid only wins if quality holds while both drop.

Next Steps: Lesson 10: Careers in AI