Machine Learning Fundamentals
Learning Outcomes
- Distinguish supervised, unsupervised, and reinforcement learning by the data and feedback each needs.
- Run the full train / evaluate / iterate workflow on a real dataset with scikit-learn.
- Diagnose overfitting and underfitting from train-versus-test scores.
- Evaluate a classifier with the right metric — accuracy, precision, recall, F1 — not the convenient one.
- Decide when a small classical model beats reaching for an LLM, and justify it on cost and latency.
Lesson 1 mapped the AI landscape; this is your hands-on foundation for everything classical. You will train, break, and fix a real model — and understand why.
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why ML literacy matters in the LLM era |
| Three paradigms | 10 min | Supervised, unsupervised, reinforcement |
| Features & splits | 10 min | Data vocabulary, the train/test split |
| First model + Pipeline | 12 min | A scikit-learn classifier, leak-free |
| Overfitting & evaluation | 13 min | Bias-variance, cross-validation, metrics |
| Unsupervised + tool choice | 9 min | Clustering, classical ML vs LLMs |
| Wrap-up | 3 min | What's next |
Before You Begin
Pre-work:
- Complete Lesson 1: The AI Landscape — Not Just LLMs so the categories of AI are fresh.
- Be comfortable reading Python (functions, lists, dicts).
- Keep the course cheat-sheet open as a glossary.
Shopping List:
- Python 3.10+ and a virtual environment (
venvorconda) — not system Python. - The packages below. No GPU, cloud, or API key — it all runs on a laptop in seconds.
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install scikit-learn numpy
Machine learning means writing a program that learns a rule from examples instead of hand-coding it. The paradigms differ in what feedback the algorithm gets.
- Supervised learning — inputs and correct answers. Each example pairs features with a label. A categorical label is classification, a numeric one regression.
- Unsupervised learning — no labels. The model finds structure itself: grouping similar items (clustering) or compressing columns (dimensionality reduction).
- Reinforcement learning (RL) — no dataset of answers. An agent acts in an environment, gets a reward, and learns a policy maximising reward over time (robotics, trading).
| Paradigm | You provide | Typical task |
|---|---|---|
| Supervised | Features + labels | Classification, regression (fraud) |
| Unsupervised | Features only | Clustering, anomaly detection |
| Reinforcement | Environment + reward | Sequential decisions |
Most production ML is supervised.
Every supervised problem is a table. Rows are samples; most columns are features (X), one is the label (y).
The most important habit in ML: never evaluate on the data the model learned from. A model can memorise its training data and look perfect yet fail on anything new. So split into a training set and a held-back test set, scored once. The breast-cancer dataset ships with scikit-learn — 569 samples, 30 features, binary label.
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
data = load_breast_cancer()
X, y = data.data, data.target # X: (569, 30) features, y: (569,) labels
X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2, # hold back 20% for honest evaluation
stratify=y, # keep the class ratio identical in both splits
random_state=42, # reproducible
)
print(X_train.shape, X_test.shape) # (455, 30) (114, 30)
stratify=y keeps the class balance identical across splits, so a random draw can't hand you a lopsided test set; random_state fixes the shuffle.
We'll use logistic regression — despite the name, a classifier that learns a weighted combination of features, squashes it to a probability, and predicts the likelier class. Models also misbehave when features span different scales, so we add feature scaling (StandardScaler: mean 0, variance 1), learned from training data only. A Pipeline does this safely: fit fits scaler then model; predict reuses the learned scaling — no leakage.
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipe = make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000))
pipe.fit(X_train, y_train) # scaler + model, fit on training data only
print(f"Test accuracy: {pipe.score(X_test, y_test):.3f}") # ~0.97
That is the entire scikit-learn contract, identical across hundreds of models: fit to learn, predict/score to use. Swap LogisticRegression for RandomForestClassifier or SVC(kernel="rbf") and nothing else changes — that consistent API is why scikit-learn is the default toolkit. Distance-based models (SVMs, k-NN) need the scaling.
This explains most model failures. Overfitting is when a model memorises the training data — including its noise — scoring well on it but poorly on new data. Underfitting is the opposite: too simple, scoring poorly on both. The sign of overfitting is a large train-test gap.
This is the bias-variance trade-off: bias is error from oversimplifying, variance from over-sensitivity to the training sample. Sweep a decision tree's depth to watch it:
from sklearn.tree import DecisionTreeClassifier
for depth in [1, 3, 5, 10, 20, None]: # None = grow until pure
tree = DecisionTreeClassifier(max_depth=depth, random_state=42).fit(X_train, y_train)
tr, te = tree.score(X_train, y_train), tree.score(X_test, y_test)
print(f"depth={str(depth):<4} train={tr:.3f} test={te:.3f}")
At depth=1 both scores are mediocre — underfitting (high bias). As depth grows, train accuracy climbs toward 1.000 while test plateaus or dips — that widening gap is overfitting (high variance). The best test score is in the middle. Cures: regularisation (cap depth, raise a penalty term), more data, or fewer noisy features. Depth is a hyperparameter — a knob you tune, not a learned weight.
A single train/test split is noisy — your score depends on which 20% landed in the test set. k-fold cross-validation fixes this: split the training data into k folds, train on k−1, validate on the held-out fold, rotate, average. A stabler estimate and its variability, without touching the test set — and the right place to tune hyperparameters.
from sklearn.model_selection import cross_val_score
scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring="f1")
print(f"5-fold F1: {scores.mean():.3f} +/- {scores.std():.3f}")
Why not just accuracy? Where 1% of transactions are fraud, a model that always predicts "not fraud" scores 99% and catches zero criminals. Accuracy (fraction correct) lies on imbalanced data. Use:
- Precision — of cases flagged positive, how many really were? (Few false alarms.)
- Recall — of all real positives, how many did we catch? (Few misses.)
- F1 — harmonic mean of the two.
from sklearn.metrics import classification_report, confusion_matrix
preds = pipe.predict(X_test)
print(confusion_matrix(y_test, preds))
print(classification_report(y_test, preds, target_names=data.target_names))
| Metric | Optimise when... |
|---|---|
| Accuracy | Classes balanced |
| Precision | False alarms costly |
| Recall | Misses costly (cancer, fraud) |
| F1 | Imbalanced, need one number |
With no labels you switch to clustering — grouping samples so members resemble each other more than outsiders. The classic is k-means: pick k, place k centres, assign each point to its nearest, move centres to their members' mean, repeat. One label per sample — the basis of segmentation.
from sklearn.datasets import make_blobs
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
X_cust, _ = make_blobs(n_samples=600, centers=4, random_state=42)
for k in range(2, 7):
labels = KMeans(n_clusters=k, n_init=10, random_state=42).fit_predict(X_cust)
print(f"k={k} silhouette={silhouette_score(X_cust, labels):.3f}")
You can't measure accuracy without labels, so use an intrinsic score like the silhouette (cluster tightness and separation, −1 to 1). The highest-silhouette k is your best guess at the number of groups — here, 4.
Now the judgement call. LLMs excel at open-ended language. But for prediction over structured/tabular data — fraud, churn, demand, maintenance — a small classical model usually wins:
| Dimension | Boosted trees | LLM API |
|---|---|---|
| Latency | Sub-ms, on CPU | 100s of ms to seconds |
| Cost at scale | Negligible | Per-token, adds up |
| Accuracy (tabular) | State of the art | Often weaker |
| Interpretability | High (SHAP) | Low (black box) |
The rule: table in, number or class out → reach for classical ML first. Your strongest default is a gradient-boosted tree — HistGradientBoostingClassifier (or XGBoost / LightGBM), called with the same fit/score. These win consistently on tables.
Questions & Answers
Key Takeaways
- Three paradigms, defined by feedback. Supervised has labelled answers, unsupervised finds structure with none, reinforcement learns from a reward — most production ML is supervised.
- Never score on training data. Split into train and test, score the test set once, and prevent leakage by fitting every transformation inside a Pipeline on training data.
- The fit / predict / score loop is universal. scikit-learn's API means swapping logistic regression for a boosted tree changes one line.
- Overfitting is the central failure mode. Watch the train-test gap; close it with regularisation, more data, or fewer noisy features.
- Pick the metric from the cost of being wrong. Accuracy lies on imbalanced data — use precision, recall, or F1, confirmed with cross-validation. For tabular inputs, a small classical model beats an LLM on cost, latency, and interpretability.
Next Steps: Lesson 3: Computer Vision in 2026