The AI Landscape — Not Just LLMs
Learning Outcomes
- Map the full AI/ML landscape and place large language models (LLMs) as one branch among many.
- Distinguish the major families — classical ML, deep learning, computer vision, NLP, reinforcement learning — by what each optimises.
- Run a real scikit-learn classifier end to end and read its accuracy, precision, and recall.
- Contrast a specialised model with an LLM on one task to feel the cost, latency, and accuracy trade-offs.
- Choose the right tool for a problem using a simple, repeatable checklist.
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why "AI" is not a synonym for "LLM" |
| The landscape | 7 min | The nested map: AI ⊃ ML ⊃ deep learning ⊃ generative |
| Classical ML | 9 min | Run a real fraud-style classifier in scikit-learn |
| Deep-learning branches | 8 min | Computer vision, NLP, reinforcement learning |
| Specialised vs LLM | 8 min | Same task, two tools, trade-offs |
| Decision checklist | 6 min | Picking the right model |
| Wrap-up | 4 min | Recap, what's next |
Before You Begin
Pre-work:
- Confirm you can run Python 3.10+ and create a virtual environment.
- Skim the course landing page so you know where this lesson sits.
- If you have used an LLM, keep that mental model handy — we contrast it with the other tools.
Shopping List:
- A terminal and a Python virtual environment.
- The packages installed in Step 3:
scikit-learn,pandas,numpy. - ~50 MB of disk and an internet connection (one small dataset downloads on first run).
If your only exposure to AI is a chat box, you have seen one room of a large building. Artificial intelligence (AI) is the goal of getting computers to do things that normally need human judgement. Machine learning (ML) is the subset where, instead of hand-writing rules, you let an algorithm learn them from examples. Deep learning is the subset of ML using large neural networks; generative AI — the LLMs and image models you know — is the slice that produces content.
The relationship is nested, like Russian dolls:
AI
└── Machine Learning (learns patterns from data)
├── Classical ML (trees, linear models) ← runs most of industry
└── Deep Learning (neural networks)
├── Computer Vision ├── NLP (incl. LLMs)
├── Reinforcement Learning
└── Generative AI ← the part that got famous
The famous part is the smallest box. The money is in the larger ones — the fraud filter on your card, the model routing your order, the system flagging a defective part — and almost none are LLMs. This course tours those larger boxes.
Every ML family is defined by the shape of its problem — what goes in, what comes out, what "good" means:
| Branch | Input → Output | Example | Typical model |
|---|---|---|---|
| Classical ML | Tabular → label/number | Fraud, churn, credit | Gradient-boosted trees |
| Computer vision | Pixels → labels, boxes | Manufacturing QC | CNNs, YOLO, ViTs |
| NLP (non-generative) | Text → label/span | Spam, sentiment, NER | BERT-style transformers |
| Time series | Past → future values | Demand forecasting | Prophet, ARIMA |
| Recommendation | User × item → ranking | "You may also like" | Matrix factorisation |
| Reinforcement learning | State → action | Robotics, ad bidding | Policy gradients |
| Generative AI | Prompt → new content | Drafting, code, images | LLMs, diffusion |
A few terms to lock in before we write code:
- Supervised learning: train on labelled examples so the model predicts the answer for new rows. Fraud detection is supervised — past transactions are tagged fraud / legit.
- Unsupervised learning: no labels; the model finds structure on its own (clustering customers).
- Features: the input columns the model reads. Label (target): the column to predict.
We will train a classifier — a model that assigns each input to a category. Set up an environment (on Windows, activate with .venv\Scripts\activate):
python3 -m venv .venv && source .venv/bin/activate
pip install scikit-learn pandas numpy
The breast-cancer dataset in scikit-learn is a clean stand-in for any "classify this row as good/bad" task.
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, precision_score, recall_score
# 569 rows, 30 numeric features, binary label (0 = malignant, 1 = benign)
X, y = load_breast_cancer(return_X_y=True)
# Hold back 20% to test on data the model never saw during training
X_tr, X_te, y_tr, y_te = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y)
# A random forest = an ensemble of decision trees that vote
model = RandomForestClassifier(n_estimators=200, random_state=42).fit(X_tr, y_tr)
p = model.predict(X_te)
print(accuracy_score(y_te, p), precision_score(y_te, p), recall_score(y_te, p))
You should see accuracy around 0.96. Define those metrics now, because accuracy alone lies on imbalanced problems like fraud (99.9% of rows are legit):
- Precision = of everything flagged positive, how much really was.
- Recall = of the truly positive, how much you caught.
- Overfitting = the model memorised the training data and fails on new data; the test split catches it. Cross-validation (
cross_val_score(model, X, y, cv=5)) confirms the score isn't a fluke.
When the input is raw pixels, audio, or long free text, deep learning takes over — networks that learn their own features. Three branches matter most.
Computer vision turns pixels into structured output via transfer learning — reuse a network trained on millions of images. Detectors are a few lines:
from ultralytics import YOLO # pip install ultralytics
model = YOLO("yolov8n.pt") # pretrained detector, small/fast
for box in model("https://ultralytics.com/images/bus.jpg")[0].boxes:
print(box.cls, round(float(box.conf), 2)) # class id + confidence
That is how manufacturing QC flags a cracked weld — hundreds of frames a second.
NLP beyond chat is mostly classification and extraction, not generation. Embeddings — text turned into a vector so similar meanings sit close together — power search and routing. A lightweight classifier handles sentiment far cheaper than an LLM: pipeline("sentiment-analysis") from Hugging Face Transformers returns a label and confidence.
Reinforcement learning (RL) is the odd one out: no labelled dataset. An agent takes actions in an environment, receives a reward, and learns a policy maximising long-run reward — fitting robotics, ad bidding, and control.
Put both tools on one task — classifying a message as "complaint" or "not."
Option A — a specialised classifier. Train once on labelled examples; every prediction is then a fast, free CPU call:
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
texts = ["broken, want a refund", "thanks, arrived early",
"terrible service", "love it"]
labels = ["complaint", "ok", "complaint", "ok"]
clf = make_pipeline(TfidfVectorizer(), LogisticRegression()).fit(texts, labels)
print(clf.predict(["the item never showed up"])) # -> ['complaint']
Option B — an LLM. Zero training. Send a prompt — "Reply with one word, complaint or ok: ..." — and parse the reply. Both work; the difference is everything around accuracy:
| Dimension | Specialised classifier | LLM |
|---|---|---|
| Setup | Labelled data + training | Works from a prompt instantly |
| Per-call cost | ~Free, runs on CPU | Per-token API cost |
| Latency | Milliseconds | Hundreds of ms to seconds |
| Throughput | Millions/day on a laptop | Rate-limited, pricey at scale |
| Explainability | Probabilities, coefficients | Opaque free text |
| Cold-start | Needs examples first | Strong with zero examples |
The LLM is the Swiss Army knife: unbeatable when the task is fuzzy, novel, or low-volume. The specialised model is purpose-built: unbeatable when the task is narrow, high-volume, and cost- and latency-sensitive.
Run any new problem through these in order; the first strong signal decides it.
1. INPUT/OUTPUT? tabular -> classical ML | images -> vision
text+fixed labels -> NLP classifier | time series -> forecasting
user x item -> recommender | sequential+reward -> RL
open-ended NEW content -> generative AI / LLM
2. LABELLED DATA? lots -> specialised model | little+fuzzy -> LLM
3. VOLUME/LATENCY? high+tight -> specialised | occasional -> LLM ok
4. MUST EXPLAIN? yes (credit/medical/legal) -> interpretable classical ML
5. WORLD CHANGES? retrainable -> classical ML | shifting -> LLM/hybrid
Worked example — flag fraud, millions/day, must justify declines to regulators. Tabular (Q1 → classical ML), high volume + tight latency (Q3 → specialised), explainable (Q4). The checklist points at gradient-boosted trees, not an LLM.
Questions & Answers
Key Takeaways
- AI is nested, LLMs are a small box. AI ⊃ ML ⊃ deep learning ⊃ generative AI; most production AI lives in the larger boxes.
- Generative AI creates; everything else mostly decides. Classifiers, detectors, and forecasters make the decisions businesses pay for.
- Read the I/O first. Input type, output type, and what "correct" means eliminate most of the landscape.
- Accuracy alone lies — use precision and recall. On imbalanced problems, headline accuracy can hide a model that catches nothing.
- Don't default to the LLM. For high-volume, narrow, explainability-sensitive tasks a specialised model is usually cheaper, faster, and easier to audit.
- Specialised + generative wins. LLMs bootstrap and handle fuzzy cases; specialised models run the high-volume core.
Next Steps: Lesson 2: Machine Learning Fundamentals