The AI Landscape — Not Just LLMs

45 min beginner Lesson 1

Learning Outcomes

  • Map the full AI/ML landscape and place large language models (LLMs) as one branch among many.
  • Distinguish the major families — classical ML, deep learning, computer vision, NLP, reinforcement learning — by what each optimises.
  • Run a real scikit-learn classifier end to end and read its accuracy, precision, and recall.
  • Contrast a specialised model with an LLM on one task to feel the cost, latency, and accuracy trade-offs.
  • Choose the right tool for a problem using a simple, repeatable checklist.

Lesson Plan

Segment Duration Topic
Intro 3 min Why "AI" is not a synonym for "LLM"
The landscape 7 min The nested map: AI ⊃ ML ⊃ deep learning ⊃ generative
Classical ML 9 min Run a real fraud-style classifier in scikit-learn
Deep-learning branches 8 min Computer vision, NLP, reinforcement learning
Specialised vs LLM 8 min Same task, two tools, trade-offs
Decision checklist 6 min Picking the right model
Wrap-up 4 min Recap, what's next

Before You Begin

Pre-work:

  • Confirm you can run Python 3.10+ and create a virtual environment.
  • Skim the course landing page so you know where this lesson sits.
  • If you have used an LLM, keep that mental model handy — we contrast it with the other tools.

Shopping List:

  • A terminal and a Python virtual environment.
  • The packages installed in Step 3: scikit-learn, pandas, numpy.
  • ~50 MB of disk and an internet connection (one small dataset downloads on first run).

1 Reset your mental model: AI is a field, not a product

If your only exposure to AI is a chat box, you have seen one room of a large building. Artificial intelligence (AI) is the goal of getting computers to do things that normally need human judgement. Machine learning (ML) is the subset where, instead of hand-writing rules, you let an algorithm learn them from examples. Deep learning is the subset of ML using large neural networks; generative AI — the LLMs and image models you know — is the slice that produces content.

The relationship is nested, like Russian dolls:

AI
└── Machine Learning (learns patterns from data)
    ├── Classical ML (trees, linear models) ← runs most of industry
    └── Deep Learning (neural networks)
        ├── Computer Vision   ├── NLP (incl. LLMs)
        ├── Reinforcement Learning
        └── Generative AI ← the part that got famous

The famous part is the smallest box. The money is in the larger ones — the fraud filter on your card, the model routing your order, the system flagging a defective part — and almost none are LLMs. This course tours those larger boxes.

NOTE
The one-sentence summary
Generative AI creates content; the rest of ML mostly makes decisions and predictions — what businesses pay for.

2 What each branch actually optimises

Every ML family is defined by the shape of its problem — what goes in, what comes out, what "good" means:

Branch Input → Output Example Typical model
Classical ML Tabular → label/number Fraud, churn, credit Gradient-boosted trees
Computer vision Pixels → labels, boxes Manufacturing QC CNNs, YOLO, ViTs
NLP (non-generative) Text → label/span Spam, sentiment, NER BERT-style transformers
Time series Past → future values Demand forecasting Prophet, ARIMA
Recommendation User × item → ranking "You may also like" Matrix factorisation
Reinforcement learning State → action Robotics, ad bidding Policy gradients
Generative AI Prompt → new content Drafting, code, images LLMs, diffusion

A few terms to lock in before we write code:

  • Supervised learning: train on labelled examples so the model predicts the answer for new rows. Fraud detection is supervised — past transactions are tagged fraud / legit.
  • Unsupervised learning: no labels; the model finds structure on its own (clustering customers).
  • Features: the input columns the model reads. Label (target): the column to predict.
TIP
Read the I/O first
Write down the input type, output type, and what 'correct' means. That one sentence eliminates most of the landscape.

3 Run a real classical-ML model (the kind that runs the economy)

We will train a classifier — a model that assigns each input to a category. Set up an environment (on Windows, activate with .venv\Scripts\activate):

python3 -m venv .venv && source .venv/bin/activate
pip install scikit-learn pandas numpy

The breast-cancer dataset in scikit-learn is a clean stand-in for any "classify this row as good/bad" task.

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, precision_score, recall_score

# 569 rows, 30 numeric features, binary label (0 = malignant, 1 = benign)
X, y = load_breast_cancer(return_X_y=True)

# Hold back 20% to test on data the model never saw during training
X_tr, X_te, y_tr, y_te = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y)

# A random forest = an ensemble of decision trees that vote
model = RandomForestClassifier(n_estimators=200, random_state=42).fit(X_tr, y_tr)
p = model.predict(X_te)
print(accuracy_score(y_te, p), precision_score(y_te, p), recall_score(y_te, p))

You should see accuracy around 0.96. Define those metrics now, because accuracy alone lies on imbalanced problems like fraud (99.9% of rows are legit):

  • Precision = of everything flagged positive, how much really was.
  • Recall = of the truly positive, how much you caught.
  • Overfitting = the model memorised the training data and fails on new data; the test split catches it. Cross-validation (cross_val_score(model, X, y, cv=5)) confirms the score isn't a fluke.
WARNING
Accuracy is a trap on imbalanced data
If 1 in 1000 transactions is fraud, a model that says 'never fraud' is 99.9% accurate and useless. Read precision and recall — and ask which mistake costs more: a missed fraud or a blocked good customer.

4 The deep-learning branches: vision, NLP, and RL

When the input is raw pixels, audio, or long free text, deep learning takes over — networks that learn their own features. Three branches matter most.

Computer vision turns pixels into structured output via transfer learning — reuse a network trained on millions of images. Detectors are a few lines:

from ultralytics import YOLO          # pip install ultralytics

model = YOLO("yolov8n.pt")            # pretrained detector, small/fast
for box in model("https://ultralytics.com/images/bus.jpg")[0].boxes:
    print(box.cls, round(float(box.conf), 2))   # class id + confidence

That is how manufacturing QC flags a cracked weld — hundreds of frames a second.

NLP beyond chat is mostly classification and extraction, not generation. Embeddings — text turned into a vector so similar meanings sit close together — power search and routing. A lightweight classifier handles sentiment far cheaper than an LLM: pipeline("sentiment-analysis") from Hugging Face Transformers returns a label and confidence.

Reinforcement learning (RL) is the odd one out: no labelled dataset. An agent takes actions in an environment, receives a reward, and learns a policy maximising long-run reward — fitting robotics, ad bidding, and control.

NOTE
Why these are not LLMs
An LLM can describe an image or guess sentiment, but a YOLO detector returns exact pixel coordinates far more cheaply, and an RL policy optimises a reward an LLM has no concept of.

5 Same task, two tools: feel the trade-off

Put both tools on one task — classifying a message as "complaint" or "not."

Option A — a specialised classifier. Train once on labelled examples; every prediction is then a fast, free CPU call:

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline

texts  = ["broken, want a refund", "thanks, arrived early",
          "terrible service", "love it"]
labels = ["complaint", "ok", "complaint", "ok"]
clf = make_pipeline(TfidfVectorizer(), LogisticRegression()).fit(texts, labels)
print(clf.predict(["the item never showed up"]))   # -> ['complaint']

Option B — an LLM. Zero training. Send a prompt — "Reply with one word, complaint or ok: ..." — and parse the reply. Both work; the difference is everything around accuracy:

Dimension Specialised classifier LLM
Setup Labelled data + training Works from a prompt instantly
Per-call cost ~Free, runs on CPU Per-token API cost
Latency Milliseconds Hundreds of ms to seconds
Throughput Millions/day on a laptop Rate-limited, pricey at scale
Explainability Probabilities, coefficients Opaque free text
Cold-start Needs examples first Strong with zero examples

The LLM is the Swiss Army knife: unbeatable when the task is fuzzy, novel, or low-volume. The specialised model is purpose-built: unbeatable when the task is narrow, high-volume, and cost- and latency-sensitive.

TIP
The honest rule of thumb
Use the LLM to bootstrap — label your first examples and prototype. Then, if the task is high-volume and well-defined, train a small specialised model to run it cheaply. Lesson 9 covers this.

6 A repeatable checklist for picking the right tool

Run any new problem through these in order; the first strong signal decides it.

1. INPUT/OUTPUT?  tabular -> classical ML | images -> vision
   text+fixed labels -> NLP classifier | time series -> forecasting
   user x item -> recommender | sequential+reward -> RL
   open-ended NEW content -> generative AI / LLM
2. LABELLED DATA?  lots -> specialised model | little+fuzzy -> LLM
3. VOLUME/LATENCY?  high+tight -> specialised | occasional -> LLM ok
4. MUST EXPLAIN?  yes (credit/medical/legal) -> interpretable classical ML
5. WORLD CHANGES?  retrainable -> classical ML | shifting -> LLM/hybrid

Worked example — flag fraud, millions/day, must justify declines to regulators. Tabular (Q1 → classical ML), high volume + tight latency (Q3 → specialised), explainable (Q4). The checklist points at gradient-boosted trees, not an LLM.

WARNING
Resist 'LLM for everything'
An LLM can attempt almost any task, which makes it tempting as a default. But 'can' is not 'should': for high-volume, narrow, explainability-sensitive work a specialised model is usually cheaper, faster, and easier to audit.

Questions & Answers

Q: LLMs keep getting smarter — won't they absorb all these specialised models eventually?
For some fuzzy, low-volume tasks they already have. But the constraints favouring specialised models — millisecond latency, near-zero per-call cost at billions of predictions, exact coordinates, auditable decisions — are economic and physical, not "intelligence" gaps. A model costing cents per call won't run a fraud filter doing a million transactions a minute.
Q: I'm a strong Python developer but new to ML. Do I have to learn the math first?
No. Be productive with scikit-learn and Hugging Face the way you are with any library — understand the inputs, outputs, and metrics. Ship working models first; reach for the math when debugging why one underperforms. Lesson 2 builds that intuition.
Q: My model scored 99% accuracy. Am I done?
Almost certainly not. High accuracy on imbalanced data usually means the model learned to always predict the majority class. Check precision and recall, use cross-validation, and confirm you tested on unseen data.
Q: Is "classical ML" outdated? It sounds old next to deep learning and LLMs.
"Classical" describes the technique, not its relevance. On structured/tabular data — most enterprise data — gradient-boosted trees routinely beat deep networks and LLMs while training in seconds and explaining themselves. Matching the tool to the data is the skill that pays. Using an LLM and a classifier together is common too (Lesson 9).

Key Takeaways

  1. AI is nested, LLMs are a small box. AI ⊃ ML ⊃ deep learning ⊃ generative AI; most production AI lives in the larger boxes.
  2. Generative AI creates; everything else mostly decides. Classifiers, detectors, and forecasters make the decisions businesses pay for.
  3. Read the I/O first. Input type, output type, and what "correct" means eliminate most of the landscape.
  4. Accuracy alone lies — use precision and recall. On imbalanced problems, headline accuracy can hide a model that catches nothing.
  5. Don't default to the LLM. For high-volume, narrow, explainability-sensitive tasks a specialised model is usually cheaper, faster, and easier to audit.
  6. Specialised + generative wins. LLMs bootstrap and handle fuzzy cases; specialised models run the high-volume core.

Next Steps: Lesson 2: Machine Learning Fundamentals