Machine Learning Fundamentals

60 min beginner Lesson 2

Learning Outcomes

  • Distinguish supervised, unsupervised, and reinforcement learning by the data and feedback each needs.
  • Run the full train / evaluate / iterate workflow on a real dataset with scikit-learn.
  • Diagnose overfitting and underfitting from train-versus-test scores.
  • Evaluate a classifier with the right metric — accuracy, precision, recall, F1 — not the convenient one.
  • Decide when a small classical model beats reaching for an LLM, and justify it on cost and latency.

Lesson 1 mapped the AI landscape; this is your hands-on foundation for everything classical. You will train, break, and fix a real model — and understand why.

Lesson Plan

Segment Duration Topic
Intro 3 min Why ML literacy matters in the LLM era
Three paradigms 10 min Supervised, unsupervised, reinforcement
Features & splits 10 min Data vocabulary, the train/test split
First model + Pipeline 12 min A scikit-learn classifier, leak-free
Overfitting & evaluation 13 min Bias-variance, cross-validation, metrics
Unsupervised + tool choice 9 min Clustering, classical ML vs LLMs
Wrap-up 3 min What's next

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ and a virtual environment (venv or conda) — not system Python.
  • The packages below. No GPU, cloud, or API key — it all runs on a laptop in seconds.
python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install scikit-learn numpy

1 The three paradigms — what feedback does the model get?

Machine learning means writing a program that learns a rule from examples instead of hand-coding it. The paradigms differ in what feedback the algorithm gets.

  • Supervised learning — inputs and correct answers. Each example pairs features with a label. A categorical label is classification, a numeric one regression.
  • Unsupervised learning — no labels. The model finds structure itself: grouping similar items (clustering) or compressing columns (dimensionality reduction).
  • Reinforcement learning (RL) — no dataset of answers. An agent acts in an environment, gets a reward, and learns a policy maximising reward over time (robotics, trading).
Paradigm You provide Typical task
Supervised Features + labels Classification, regression (fraud)
Unsupervised Features only Clustering, anomaly detection
Reinforcement Environment + reward Sequential decisions

Most production ML is supervised.

NOTE
Where do LLMs fit?
An LLM is trained with self-supervised learning (predict the next token — labels come from the text itself), then refined with reinforcement learning from human feedback. Not a fourth paradigm — a huge application of the three.

2 Features, labels, and the train/test split

Every supervised problem is a table. Rows are samples; most columns are features (X), one is the label (y).

The most important habit in ML: never evaluate on the data the model learned from. A model can memorise its training data and look perfect yet fail on anything new. So split into a training set and a held-back test set, scored once. The breast-cancer dataset ships with scikit-learn — 569 samples, 30 features, binary label.

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

data = load_breast_cancer()
X, y = data.data, data.target          # X: (569, 30) features, y: (569,) labels

X_train, X_test, y_train, y_test = train_test_split(
    X, y,
    test_size=0.2,      # hold back 20% for honest evaluation
    stratify=y,         # keep the class ratio identical in both splits
    random_state=42,    # reproducible
)
print(X_train.shape, X_test.shape)      # (455, 30) (114, 30)

stratify=y keeps the class balance identical across splits, so a random draw can't hand you a lopsided test set; random_state fixes the shuffle.

WARNING
The cardinal sin: data leakage
If test-set information sneaks into training — e.g. scaling features using stats from the whole dataset before the split — your test score becomes a lie. Fit every transformation on training data only; Step 3 shows the safe way.

3 Your first model — train, predict, score (leak-free)

We'll use logistic regression — despite the name, a classifier that learns a weighted combination of features, squashes it to a probability, and predicts the likelier class. Models also misbehave when features span different scales, so we add feature scaling (StandardScaler: mean 0, variance 1), learned from training data only. A Pipeline does this safely: fit fits scaler then model; predict reuses the learned scaling — no leakage.

from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

pipe = make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000))
pipe.fit(X_train, y_train)            # scaler + model, fit on training data only
print(f"Test accuracy: {pipe.score(X_test, y_test):.3f}")   # ~0.97

That is the entire scikit-learn contract, identical across hundreds of models: fit to learn, predict/score to use. Swap LogisticRegression for RandomForestClassifier or SVC(kernel="rbf") and nothing else changes — that consistent API is why scikit-learn is the default toolkit. Distance-based models (SVMs, k-NN) need the scaling.

TIP
Read the model, not just the score
Linear models are interpretable: the fitted model exposes a coef_ array with one weight per feature — sort by absolute value to see what drove the decision. For lending or fraud, that is often a regulatory requirement.

4 Overfitting and the bias-variance trade-off

This explains most model failures. Overfitting is when a model memorises the training data — including its noise — scoring well on it but poorly on new data. Underfitting is the opposite: too simple, scoring poorly on both. The sign of overfitting is a large train-test gap.

This is the bias-variance trade-off: bias is error from oversimplifying, variance from over-sensitivity to the training sample. Sweep a decision tree's depth to watch it:

from sklearn.tree import DecisionTreeClassifier

for depth in [1, 3, 5, 10, 20, None]:     # None = grow until pure
    tree = DecisionTreeClassifier(max_depth=depth, random_state=42).fit(X_train, y_train)
    tr, te = tree.score(X_train, y_train), tree.score(X_test, y_test)
    print(f"depth={str(depth):<4} train={tr:.3f} test={te:.3f}")

At depth=1 both scores are mediocre — underfitting (high bias). As depth grows, train accuracy climbs toward 1.000 while test plateaus or dips — that widening gap is overfitting (high variance). The best test score is in the middle. Cures: regularisation (cap depth, raise a penalty term), more data, or fewer noisy features. Depth is a hyperparameter — a knob you tune, not a learned weight.

WARNING
100% training accuracy is a red flag, not a trophy
A model that perfectly fits its training data has almost certainly memorised noise. Always report the test score and check the train-test gap; if they diverge, you are overfitting.

5 Cross-validation and the right metric

A single train/test split is noisy — your score depends on which 20% landed in the test set. k-fold cross-validation fixes this: split the training data into k folds, train on k−1, validate on the held-out fold, rotate, average. A stabler estimate and its variability, without touching the test set — and the right place to tune hyperparameters.

from sklearn.model_selection import cross_val_score

scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring="f1")
print(f"5-fold F1: {scores.mean():.3f} +/- {scores.std():.3f}")

Why not just accuracy? Where 1% of transactions are fraud, a model that always predicts "not fraud" scores 99% and catches zero criminals. Accuracy (fraction correct) lies on imbalanced data. Use:

  • Precision — of cases flagged positive, how many really were? (Few false alarms.)
  • Recall — of all real positives, how many did we catch? (Few misses.)
  • F1 — harmonic mean of the two.
from sklearn.metrics import classification_report, confusion_matrix

preds = pipe.predict(X_test)
print(confusion_matrix(y_test, preds))
print(classification_report(y_test, preds, target_names=data.target_names))
Metric Optimise when...
Accuracy Classes balanced
Precision False alarms costly
Recall Misses costly (cancer, fraud)
F1 Imbalanced, need one number
TIP
Pick the metric from the business cost
A cancer screen maximises recall — a missed tumour beats a false alarm. A spam filter leans on precision — burying a real email is worse than one spam through. Choose it before you train.

6 Unsupervised learning, and when classical ML beats an LLM

With no labels you switch to clustering — grouping samples so members resemble each other more than outsiders. The classic is k-means: pick k, place k centres, assign each point to its nearest, move centres to their members' mean, repeat. One label per sample — the basis of segmentation.

from sklearn.datasets import make_blobs
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

X_cust, _ = make_blobs(n_samples=600, centers=4, random_state=42)
for k in range(2, 7):
    labels = KMeans(n_clusters=k, n_init=10, random_state=42).fit_predict(X_cust)
    print(f"k={k}  silhouette={silhouette_score(X_cust, labels):.3f}")

You can't measure accuracy without labels, so use an intrinsic score like the silhouette (cluster tightness and separation, −1 to 1). The highest-silhouette k is your best guess at the number of groups — here, 4.

Now the judgement call. LLMs excel at open-ended language. But for prediction over structured/tabular data — fraud, churn, demand, maintenance — a small classical model usually wins:

Dimension Boosted trees LLM API
Latency Sub-ms, on CPU 100s of ms to seconds
Cost at scale Negligible Per-token, adds up
Accuracy (tabular) State of the art Often weaker
Interpretability High (SHAP) Low (black box)

The rule: table in, number or class out → reach for classical ML first. Your strongest default is a gradient-boosted tree — HistGradientBoostingClassifier (or XGBoost / LightGBM), called with the same fit/score. These win consistently on tables.

NOTE
Embeddings: where classical ML meets modern AI
An embedding turns complex data (text, an image) into a fixed-length vector, placing similar things near each other. Run k-means on LLM-generated text embeddings to cluster support tickets — classical ML and GenAI together, as in Lesson 9.

Questions & Answers

Q: My test accuracy is 96% — am I done?
Not necessarily. Check the class balance: if 95% of samples are one class, 96% barely beats guessing the majority. Read the confusion matrix and per-class precision/recall, then confirm no leakage. A high number is the start of the investigation.
Q: How much data do I actually need?
A useful rule is hundreds-to-thousands of labelled samples per class for simple tabular classification — far less than the billions of tokens an LLM trains on. Plot a learning curve (score vs set size): if still rising, more data helps; if flat, improve features.
Q: Why bother with logistic regression when boosted trees usually win?
It is a fast baseline that proves the problem is learnable, it is highly interpretable (regulators often require it), and if a linear model already hits target, shipping it is cheaper and safer. Beating it is how you prove a complex model earns its keep.
Q: I changed random_state and my score moved. Which one is real?
Neither — that movement is the noise cross-validation measures. Don't trust a single split: report the mean and standard deviation across k folds. If two models overlap within a standard deviation, treat them as tied and ship the simpler one.
Q: Can't I just ask an LLM to predict churn from my customer table?
You can, but not for production: it is slower, costs more, gives no feature importances, and is usually less accurate than a gradient-boosted tree on the same table. LLMs shine on unstructured language; structured prediction is classical ML's home turf — more in Lesson 8.

Key Takeaways

  1. Three paradigms, defined by feedback. Supervised has labelled answers, unsupervised finds structure with none, reinforcement learns from a reward — most production ML is supervised.
  2. Never score on training data. Split into train and test, score the test set once, and prevent leakage by fitting every transformation inside a Pipeline on training data.
  3. The fit / predict / score loop is universal. scikit-learn's API means swapping logistic regression for a boosted tree changes one line.
  4. Overfitting is the central failure mode. Watch the train-test gap; close it with regularisation, more data, or fewer noisy features.
  5. Pick the metric from the cost of being wrong. Accuracy lies on imbalanced data — use precision, recall, or F1, confirmed with cross-validation. For tabular inputs, a small classical model beats an LLM on cost, latency, and interpretability.

Next Steps: Lesson 3: Computer Vision in 2026