NLP Beyond Chat
Learning Outcomes
- Map the NLP task zoo and identify which problems do not need a chatbot
- Explain the difference between encoder models (BERT family) and decoder LLMs in plain language
- Extract named entities and run sentiment analysis with spaCy and Hugging Face pipelines
- Build a TF-IDF text classifier and compare it against an LLM on cost, speed, and consistency
- Choose between a focused model and an LLM for a real production NLP task
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | NLP is bigger than chat |
| The Task Zoo | 8 min | Classification, NER, sentiment, topics |
| Encoders vs Decoders | 9 min | BERT vs GPT, and embeddings |
| NER with spaCy | 10 min | Runnable entity extraction |
| Sentiment with Transformers | 9 min | A Hugging Face pipeline end to end |
| TF-IDF vs BERT vs LLM | 12 min | Three classifiers, real metrics |
| Topic Modelling | 6 min | Finding themes with BERTopic |
| Wrap-up | 3 min | Decision guidance and next steps |
Before You Begin
Pre-work:
- Complete Lesson 2: Machine Learning Fundamentals so train/test split, precision, and recall are familiar
- Skim Lesson 1: The AI Landscape for where NLP sits in the wider map of AI
- Have Python 3.10+ and a virtual environment ready
Shopping List:
pip install scikit-learn spacy transformers torch bertopic sentence-transformers- The spaCy English model:
python -m spacy download en_core_web_sm - A CPU is fine for everything here; a GPU only speeds up the transformer steps
- About 2 GB of disk for downloaded models
Natural Language Processing (NLP) is the field of getting computers to work with human language. The spotlight is on conversational LLMs, but most NLP in production is not a chatbot — it is narrow tasks that take text in and produce structured output.
| Task | Input → Output | Real-world use |
|---|---|---|
| Text classification | Document → one of N labels | Routing tickets, spam filters |
| Sentiment analysis | Sentence → positive/negative + score | Brand monitoring, review triage |
| Named entity recognition (NER) | Text → spans tagged person/org/date | Contract analysis, redacting PII |
| Topic modelling | Corpus → clusters of related docs | Themes in surveys at scale |
| Question answering | Question + passage → answer span | Search, document lookup |
Classification, sentiment, and NER produce a fixed, structured output — the same input always yields the same label, in milliseconds, for a fraction of a cent. That is what makes a focused model attractive.
The decision boundary: a narrow, repetitive task with (or able to be given) labelled data favours a focused model; open-ended, low-volume, or cross-format reasoning favours an LLM.
BERT-family models and GPT-family LLMs are both built on the Transformer, the architecture that learns relationships between words using "attention." But they are trained for opposite jobs.
- An encoder (BERT) reads the whole sentence at once, in both directions. It is optimised to understand text, not generate it.
- A decoder (GPT, and LLMs generally) reads left-to-right and predicts the next token. It is optimised to generate text.
For classification, sentiment, and NER — understand a fixed input, emit a label — an encoder fits. For chat and reasoning, a decoder fits.
The other key idea is embeddings: a vector representing the meaning of text, where similar meanings sit close together. Encoders are excellent embedding machines.
# Sentence embeddings: turning text into vectors you can compare
from sentence_transformers import SentenceTransformer
import numpy as np
model = SentenceTransformer("all-MiniLM-L6-v2")
sents = ["The invoice is overdue.", "Payment is late.", "I love this product!"]
vecs = model.encode(sents)
def cosine(a, b):
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
print("late vs overdue:", round(cosine(vecs[0], vecs[1]), 3)) # high
print("late vs love: ", round(cosine(vecs[0], vecs[2]), 3)) # low
The two "late payment" sentences score close together; the cheerful one is far away — with no keyword overlap. That is meaning as geometry, the foundation under search, clustering, and classification by meaning.
Named entity recognition finds and labels the "things" in text — people, organisations, dates, money, locations. It is the workhorse behind contract analysis, PII redaction, and information extraction. spaCy is a fast, production-grade NLP library shipping pretrained pipelines with NER built in.
import spacy
nlp = spacy.load("en_core_web_sm") # small English pipeline, downloaded earlier
text = ("Acme Corp signed a $4.2 million contract with the city of Leeds "
"on March 3, 2026, naming Dr. Priya Shah as project lead.")
for ent in nlp(text).ents:
print(f"{ent.text:<20} {ent.label_}")
# Acme Corp -> ORG | $4.2 million -> MONEY | Leeds -> GPE
# March 3, 2026 -> DATE | Priya Shah -> PERSON
GPE means "geopolitical entity" (country, city, or state). spaCy labels each entity from a fixed scheme and runs at thousands of documents per second on a CPU. Compare prompting an LLM to "extract all entities as JSON" — slower, costs per call, and the format can drift between runs.
Sentiment analysis classifies text by emotional tone. Hugging Face Transformers' one-line pipeline() helper downloads a fine-tuned model and handles tokenisation, inference, and decoding. The default sentiment model is a DistilBERT fine-tuned on SST-2, returning POSITIVE/NEGATIVE with a confidence score.
from transformers import pipeline
# First call downloads the default DistilBERT SST-2 model, then caches it
clf = pipeline("sentiment-analysis")
reviews = [
"Shipping was fast and the quality exceeded my expectations.",
"Broke after two days. Total waste of money.",
"It's fine, does the job.",
]
for r, out in zip(reviews, clf(reviews)):
print(f"{out['label']:<8} {out['score']:.3f} | {r}")
The default model is binary — no native "neutral," so "It's fine" is forced into positive or negative with a middling score. For more classes, pass a model via pipeline("text-classification", model="..."). And it runs locally: no API key, no per-token billing.
Classify text three ways and compare on accuracy, speed, and cost. The task: routing support tickets into categories.
Approach A — TF-IDF + a classical model. TF-IDF (Term Frequency–Inverse Document Frequency) turns each document into a sparse vector of word weights: common words low, distinctive words high. Feed those to a Logistic Regression — a linear model that learns a weight per feature and outputs class probabilities.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report
texts = [
"my card was charged twice", "refund not received", "double billing issue",
"app crashes on login", "page won't load", "getting a 500 error",
"how do I reset my password", "where are the account settings",
"cannot find the export button",
]
labels = ["billing", "billing", "billing", "bug", "bug", "bug",
"how-to", "how-to", "how-to"]
X_tr, X_te, y_tr, y_te = train_test_split(texts, labels, test_size=0.34, random_state=0)
model = make_pipeline(TfidfVectorizer(), LogisticRegression(max_iter=1000))
model.fit(X_tr, y_tr)
print(classification_report(y_te, model.predict(X_te), zero_division=0))
Trains in milliseconds, predicts in microseconds. The trade-off: TF-IDF sees words, not meaning, so on small data it can miss links between phrasings that share no words.
Approach B — fine-tuned BERT. BERT understands meaning, so it generalises better on the same data. The mechanism is transfer learning: take a model already trained on huge text, then briefly continue training (fine-tune) it on your tickets — reusing general language knowledge and only teaching it your categories.
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
TrainingArguments, Trainer)
tok = AutoTokenizer.from_pretrained("distilbert-base-uncased")
bert = AutoModelForSequenceClassification.from_pretrained(
"distilbert-base-uncased", num_labels=3)
args = TrainingArguments(output_dir="out", num_train_epochs=3,
per_device_train_batch_size=8, eval_strategy="epoch")
trainer = Trainer(model=bert, args=args) # add tokenised train/eval datasets
# trainer.train()
Approach C — an LLM, zero-shot. No training. Describe the categories in a prompt and let the model pick — Hugging Face's zero-shot-classification pipeline locally, or a hosted LLM:
You are a ticket router. Classify the ticket into exactly one of:
billing, bug, how-to. Reply with only the label.
Ticket: "my card was charged twice"
The honest comparison every engineer should internalise:
| Dimension | TF-IDF + LogReg | Fine-tuned BERT | LLM (zero-shot) |
|---|---|---|---|
| Labelled data needed | Hundreds–thousands | Hundreds–thousands | None |
| Training cost | Seconds, CPU | Minutes–hours, GPU helps | None |
| Inference latency | Microseconds | Milliseconds | Hundreds of ms–seconds |
| Per-prediction cost | ~Zero | ~Zero | Per-token API cost |
| Consistency | Deterministic | Deterministic | Can vary run to run |
| Accuracy (narrow task) | Good with data | Best with data | Good, no data needed |
The previous tasks were supervised — you had labels. Topic modelling is unsupervised: a pile of documents, no labels, and the question "what themes are in here?" Classic use: 50,000 free-text survey responses and "what are people complaining about?"
The modern approach is BERTopic. It embeds every document (Step 2's vectors), clusters the embeddings so similar documents group together, then names each cluster by its most distinctive words — meaning-aware, unlike older bag-of-words methods such as LDA (Latent Dirichlet Allocation).
from bertopic import BERTopic
# docs = a list of thousands of short texts (reviews, tickets, survey answers)
topic_model = BERTopic(min_topic_size=10)
topics, probs = topic_model.fit_transform(docs)
print(topic_model.get_topic_info().head()) # counts per topic
print(topic_model.get_topic(0)) # top words for topic 0
get_topic_info() returns topics with sizes; topic -1 is the "outlier" bucket for documents that fit no cluster. You get a data-driven map of your corpus in a few lines — no labelling, no prompt, fully reproducible. Prefer LDA on very large corpora where embedding compute is a concern; prefer BERTopic when semantic quality matters more than raw speed.
Questions & Answers
pipeline("text-classification", model="..."), use a zero-shot-classification pipeline where you supply candidate labels at call time, or fine-tune your own head (Step 5, Approach B). Do not fake "neutral" with a score threshold on a binary model — that is fragile.min_topic_size to force larger, fewer clusters, or reduce topics afterward with the model's topic-reduction methods. A big topic -1 bucket means many documents didn't cluster — common with short or diverse text. Topic modelling is exploratory: it surfaces structure for a human to interpret, not a final answer.Key Takeaways
-
Most production NLP is not chat. Classification, NER, sentiment, and topic modelling produce structured output, and that structure makes focused models fast, cheap, and consistent.
-
Encoders read, decoders write. BERT-family encoders understand a fixed input and emit a label or embedding; GPT-family decoders generate. Match the architecture to the job.
-
Embeddings turn meaning into geometry. Sentence vectors let you compare, cluster, and classify by meaning rather than keywords — the foundation under search, topic modelling, and hybrid pipelines.
-
Benchmark the simple baseline first. TF-IDF + Logistic Regression is often good enough and the easiest model to run, debug, and explain. Escalate only when the data shows you need to.
-
Match the tool to volume and data. No data and low volume favour an LLM; abundant data and high volume favour a focused encoder, with confidence-thresholded fallback for the best of both.
-
Off-the-shelf is a starting point. Pretrained NER and sentiment models drop in accuracy on your domain. Always measure on your own data before trusting anything in production.
Next Steps: Lesson 5: Time Series & Forecasting