NLP Beyond Chat

60 min intermediate Lesson 4

Learning Outcomes

  • Map the NLP task zoo and identify which problems do not need a chatbot
  • Explain the difference between encoder models (BERT family) and decoder LLMs in plain language
  • Extract named entities and run sentiment analysis with spaCy and Hugging Face pipelines
  • Build a TF-IDF text classifier and compare it against an LLM on cost, speed, and consistency
  • Choose between a focused model and an LLM for a real production NLP task

Lesson Plan

Segment Duration Topic
Intro 3 min NLP is bigger than chat
The Task Zoo 8 min Classification, NER, sentiment, topics
Encoders vs Decoders 9 min BERT vs GPT, and embeddings
NER with spaCy 10 min Runnable entity extraction
Sentiment with Transformers 9 min A Hugging Face pipeline end to end
TF-IDF vs BERT vs LLM 12 min Three classifiers, real metrics
Topic Modelling 6 min Finding themes with BERTopic
Wrap-up 3 min Decision guidance and next steps

Before You Begin

Pre-work:

Shopping List:

  • pip install scikit-learn spacy transformers torch bertopic sentence-transformers
  • The spaCy English model: python -m spacy download en_core_web_sm
  • A CPU is fine for everything here; a GPU only speeds up the transformer steps
  • About 2 GB of disk for downloaded models

1 The NLP Task Zoo — Most of It Is Not Chat

Natural Language Processing (NLP) is the field of getting computers to work with human language. The spotlight is on conversational LLMs, but most NLP in production is not a chatbot — it is narrow tasks that take text in and produce structured output.

Task Input → Output Real-world use
Text classification Document → one of N labels Routing tickets, spam filters
Sentiment analysis Sentence → positive/negative + score Brand monitoring, review triage
Named entity recognition (NER) Text → spans tagged person/org/date Contract analysis, redacting PII
Topic modelling Corpus → clusters of related docs Themes in surveys at scale
Question answering Question + passage → answer span Search, document lookup

Classification, sentiment, and NER produce a fixed, structured output — the same input always yields the same label, in milliseconds, for a fraction of a cent. That is what makes a focused model attractive.

NOTE
Why this matters
An LLM can do every task in this table. But for a high-volume narrow task — classify 10 million tickets a day — a 100-million-parameter model on a CPU is cheaper, faster, and more consistent than 10 million frontier-model calls. The skill is knowing which is which.

The decision boundary: a narrow, repetitive task with (or able to be given) labelled data favours a focused model; open-ended, low-volume, or cross-format reasoning favours an LLM.


2 Encoders vs Decoders — BERT vs GPT in Plain Language

BERT-family models and GPT-family LLMs are both built on the Transformer, the architecture that learns relationships between words using "attention." But they are trained for opposite jobs.

  • An encoder (BERT) reads the whole sentence at once, in both directions. It is optimised to understand text, not generate it.
  • A decoder (GPT, and LLMs generally) reads left-to-right and predicts the next token. It is optimised to generate text.

For classification, sentiment, and NER — understand a fixed input, emit a label — an encoder fits. For chat and reasoning, a decoder fits.

The other key idea is embeddings: a vector representing the meaning of text, where similar meanings sit close together. Encoders are excellent embedding machines.

# Sentence embeddings: turning text into vectors you can compare
from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("all-MiniLM-L6-v2")
sents = ["The invoice is overdue.", "Payment is late.", "I love this product!"]
vecs = model.encode(sents)

def cosine(a, b):
    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))

print("late vs overdue:", round(cosine(vecs[0], vecs[1]), 3))   # high
print("late vs love:   ", round(cosine(vecs[0], vecs[2]), 3))   # low

The two "late payment" sentences score close together; the cheerful one is far away — with no keyword overlap. That is meaning as geometry, the foundation under search, clustering, and classification by meaning.

TIP
Mental model
Encoders are readers; decoders are writers. A BERT model around 110M parameters and its distilled cousin around 66M are tiny next to a frontier LLM, yet read text superbly. Small-but-focused is a feature.

3 Named Entity Recognition with spaCy

Named entity recognition finds and labels the "things" in text — people, organisations, dates, money, locations. It is the workhorse behind contract analysis, PII redaction, and information extraction. spaCy is a fast, production-grade NLP library shipping pretrained pipelines with NER built in.

import spacy

nlp = spacy.load("en_core_web_sm")   # small English pipeline, downloaded earlier

text = ("Acme Corp signed a $4.2 million contract with the city of Leeds "
        "on March 3, 2026, naming Dr. Priya Shah as project lead.")

for ent in nlp(text).ents:
    print(f"{ent.text:<20} {ent.label_}")

# Acme Corp -> ORG | $4.2 million -> MONEY | Leeds -> GPE
# March 3, 2026 -> DATE | Priya Shah -> PERSON

GPE means "geopolitical entity" (country, city, or state). spaCy labels each entity from a fixed scheme and runs at thousands of documents per second on a CPU. Compare prompting an LLM to "extract all entities as JSON" — slower, costs per call, and the format can drift between runs.

WARNING
Off-the-shelf is a starting point
Pretrained NER is trained on general text (news, web). On your domain — medical notes, legal clauses, internal product names — accuracy drops. Fine-tune spaCy on a few hundred labelled examples, or add rule-based patterns, and always measure on your own data first. Reach for an LLM instead when entities are fuzzy, context-dependent, or you must extract and reason about them in one pass.

4 Sentiment Analysis with a Hugging Face Pipeline

Sentiment analysis classifies text by emotional tone. Hugging Face Transformers' one-line pipeline() helper downloads a fine-tuned model and handles tokenisation, inference, and decoding. The default sentiment model is a DistilBERT fine-tuned on SST-2, returning POSITIVE/NEGATIVE with a confidence score.

from transformers import pipeline

# First call downloads the default DistilBERT SST-2 model, then caches it
clf = pipeline("sentiment-analysis")

reviews = [
    "Shipping was fast and the quality exceeded my expectations.",
    "Broke after two days. Total waste of money.",
    "It's fine, does the job.",
]

for r, out in zip(reviews, clf(reviews)):
    print(f"{out['label']:<8} {out['score']:.3f}  | {r}")

The default model is binary — no native "neutral," so "It's fine" is forced into positive or negative with a middling score. For more classes, pass a model via pipeline("text-classification", model="..."). And it runs locally: no API key, no per-token billing.

NOTE
Calibrate the score, do not trust it blindly
The score is the model's confidence, not a probability of being correct. A 0.62 prediction is a coin-flip in disguise. In production, set a confidence threshold (say 0.85) and route low-confidence items to a human or an LLM fallback — cheap model first, expensive model only on the hard cases. That pattern is the heart of Lesson 9.

5 The Showdown — TF-IDF vs BERT vs LLM

Classify text three ways and compare on accuracy, speed, and cost. The task: routing support tickets into categories.

Approach A — TF-IDF + a classical model. TF-IDF (Term Frequency–Inverse Document Frequency) turns each document into a sparse vector of word weights: common words low, distinctive words high. Feed those to a Logistic Regression — a linear model that learns a weight per feature and outputs class probabilities.

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report

texts = [
    "my card was charged twice", "refund not received", "double billing issue",
    "app crashes on login", "page won't load", "getting a 500 error",
    "how do I reset my password", "where are the account settings",
    "cannot find the export button",
]
labels = ["billing", "billing", "billing", "bug", "bug", "bug",
          "how-to", "how-to", "how-to"]

X_tr, X_te, y_tr, y_te = train_test_split(texts, labels, test_size=0.34, random_state=0)
model = make_pipeline(TfidfVectorizer(), LogisticRegression(max_iter=1000))
model.fit(X_tr, y_tr)
print(classification_report(y_te, model.predict(X_te), zero_division=0))

Trains in milliseconds, predicts in microseconds. The trade-off: TF-IDF sees words, not meaning, so on small data it can miss links between phrasings that share no words.

Approach B — fine-tuned BERT. BERT understands meaning, so it generalises better on the same data. The mechanism is transfer learning: take a model already trained on huge text, then briefly continue training (fine-tune) it on your tickets — reusing general language knowledge and only teaching it your categories.

from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
                          TrainingArguments, Trainer)

tok = AutoTokenizer.from_pretrained("distilbert-base-uncased")
bert = AutoModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased", num_labels=3)
args = TrainingArguments(output_dir="out", num_train_epochs=3,
                         per_device_train_batch_size=8, eval_strategy="epoch")
trainer = Trainer(model=bert, args=args)  # add tokenised train/eval datasets
# trainer.train()

Approach C — an LLM, zero-shot. No training. Describe the categories in a prompt and let the model pick — Hugging Face's zero-shot-classification pipeline locally, or a hosted LLM:

You are a ticket router. Classify the ticket into exactly one of:
billing, bug, how-to. Reply with only the label.
Ticket: "my card was charged twice"

The honest comparison every engineer should internalise:

Dimension TF-IDF + LogReg Fine-tuned BERT LLM (zero-shot)
Labelled data needed Hundreds–thousands Hundreds–thousands None
Training cost Seconds, CPU Minutes–hours, GPU helps None
Inference latency Microseconds Milliseconds Hundreds of ms–seconds
Per-prediction cost ~Zero ~Zero Per-token API cost
Consistency Deterministic Deterministic Can vary run to run
Accuracy (narrow task) Good with data Best with data Good, no data needed
TIP
The decision rule
No labelled data and low volume? Start with the LLM today. Have data and high volume? A fine-tuned encoder pays for itself fast. Surprisingly often, plain TF-IDF + Logistic Regression is good enough — and the cheapest to run, debug, and explain. Always benchmark the simple baseline first.

6 Topic Modelling — Finding Themes Without Labels

The previous tasks were supervised — you had labels. Topic modelling is unsupervised: a pile of documents, no labels, and the question "what themes are in here?" Classic use: 50,000 free-text survey responses and "what are people complaining about?"

The modern approach is BERTopic. It embeds every document (Step 2's vectors), clusters the embeddings so similar documents group together, then names each cluster by its most distinctive words — meaning-aware, unlike older bag-of-words methods such as LDA (Latent Dirichlet Allocation).

from bertopic import BERTopic

# docs = a list of thousands of short texts (reviews, tickets, survey answers)
topic_model = BERTopic(min_topic_size=10)
topics, probs = topic_model.fit_transform(docs)

print(topic_model.get_topic_info().head())   # counts per topic
print(topic_model.get_topic(0))              # top words for topic 0

get_topic_info() returns topics with sizes; topic -1 is the "outlier" bucket for documents that fit no cluster. You get a data-driven map of your corpus in a few lines — no labelling, no prompt, fully reproducible. Prefer LDA on very large corpora where embedding compute is a concern; prefer BERTopic when semantic quality matters more than raw speed.

WARNING
Topics need human interpretation
BERTopic gives clusters and keywords, not meaning. A cluster of [refund, charge, card] is yours to read as 'billing complaints.' This is where classical and generative methods team up: cluster with BERTopic for speed and reproducibility, then ask an LLM to write a one-line label per cluster. Cheap structure first, generative polish second.

Questions & Answers

Q: LLMs can do all of this in one prompt. Why maintain a separate spaCy or BERT model?
Volume and unit economics. One LLM call is cheap; ten million a day is not, and it is slow and non-deterministic. A focused encoder runs on a CPU at thousands of documents per second for effectively zero marginal cost and gives the same answer every time. For narrow, high-throughput tasks it is the right choice — reserve the LLM for open-ended or low-volume work, or as a fallback on low-confidence cases.
Q: I don't have any labelled data. Is a classical classifier off the table?
For now, yes — TF-IDF and fine-tuned BERT both need labels. Start with an LLM zero-shot to ship today, then have it (or humans) label the traffic it sees. Once you have a few thousand labelled examples, train a focused model for the high-volume routing and keep the LLM as fallback. This "LLM bootstraps the data, small model serves it" loop is a standard production pattern.
Q: My fine-tuned model scores 99% on the test set. Should I trust it?
Be suspicious. Near-perfect scores often signal overfitting (the model memorised the training data) or data leakage (duplicate texts span the train and test splits). De-duplicate first, use a proper held-out test set, and ideally evaluate on data from a different time period. Check per-class precision and recall, not just overall accuracy — a tiny class can hide total failure behind a high average.
Q: The default Hugging Face sentiment model only gives positive/negative. How do I get neutral or custom classes?
The default is binary because it was fine-tuned on SST-2. For more classes, pass a model trained for your label set via pipeline("text-classification", model="..."), use a zero-shot-classification pipeline where you supply candidate labels at call time, or fine-tune your own head (Step 5, Approach B). Do not fake "neutral" with a score threshold on a binary model — that is fragile.
Q: BERTopic gave me 80 topics and a giant outlier bucket. Did I do something wrong?
Probably just defaults. Raise min_topic_size to force larger, fewer clusters, or reduce topics afterward with the model's topic-reduction methods. A big topic -1 bucket means many documents didn't cluster — common with short or diverse text. Topic modelling is exploratory: it surfaces structure for a human to interpret, not a final answer.

Key Takeaways

  1. Most production NLP is not chat. Classification, NER, sentiment, and topic modelling produce structured output, and that structure makes focused models fast, cheap, and consistent.

  2. Encoders read, decoders write. BERT-family encoders understand a fixed input and emit a label or embedding; GPT-family decoders generate. Match the architecture to the job.

  3. Embeddings turn meaning into geometry. Sentence vectors let you compare, cluster, and classify by meaning rather than keywords — the foundation under search, topic modelling, and hybrid pipelines.

  4. Benchmark the simple baseline first. TF-IDF + Logistic Regression is often good enough and the easiest model to run, debug, and explain. Escalate only when the data shows you need to.

  5. Match the tool to volume and data. No data and low volume favour an LLM; abundant data and high volume favour a focused encoder, with confidence-thresholded fallback for the best of both.

  6. Off-the-shelf is a starting point. Pretrained NER and sentiment models drop in accuracy on your domain. Always measure on your own data before trusting anything in production.

Next Steps: Lesson 5: Time Series & Forecasting