Recommendation Systems
Learning Outcomes
- Distinguish collaborative from content-based filtering and explain when each fits
- Build an item-based recommender and a matrix-factorisation model from interaction data
- Evaluate recommendations with ranking metrics — precision@k, recall@k, NDCG, coverage
- Diagnose and mitigate the cold-start problem for new users and new items
- Decide where an LLM belongs in a recommender stack versus where classical ML wins
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why recommenders are ML's biggest commercial win |
| Explain | 8 min | The two families: collaborative vs content-based |
| Build | 12 min | Item-based collaborative filtering from scratch |
| Build | 12 min | Matrix factorisation with a real library |
| Explain | 10 min | Evaluation metrics unique to ranking |
| Build | 8 min | Content-based + hybrid recommenders |
| Explain | 4 min | Cold-start and where LLMs fit |
| Wrap-up | 3 min | Key takeaways, preview MLOps |
Before You Begin
Pre-work:
- Comfort with the supervised/unsupervised split from Lesson 2: ML Fundamentals
- How embeddings encode meaning from Lesson 4: NLP Beyond Chat — recommenders use the same idea on users and items.
Shopping List:
- Python 3.10+ in a virtual environment
pip install numpy pandas scikit-learn scipy implicit- A terminal and a code editor
A recommender system predicts what a user will want from a catalogue too large to browse — "you might also like" rows, the autoplay queue, a self-reordering homepage. It's a quiet workhorse behind much of ML's revenue, and two ideas power most of it.
Collaborative filtering looks only at behaviour: people who interacted with what you did also interacted with X. Content-based filtering looks at item attributes: you liked sci-fi, here's another.
| Approach | Signal | Strength | Weakness |
|---|---|---|---|
| Collaborative | User-item interactions | Finds cross-genre hits | Useless for brand-new items/users |
| Content-based | Item attributes/text | Works on day-one items | Stays in a filter bubble |
| Hybrid | Both | Covers each other's gaps | More moving parts |
The data is an interaction matrix: rows users, columns items, cells ratings or implicit signals (a play, click, purchase). It is almost entirely empty — a typical user has touched well under 1% of the catalogue. This sparsity is the central reality of recommenders.
The most intuitive recommender is item-based collaborative filtering: measure how similarly users behave toward each pair of items, then recommend ones similar to what a user liked. "Similar" means cosine similarity — items rated alike by the same users point the same direction (near 1).
On a tiny ratings table, item_cf.py:
import pandas as pd
from sklearn.metrics.pairwise import cosine_similarity
# user, item, rating (1-5)
ratings = pd.DataFrame([
(1, "Inception", 5), (1, "Interstellar", 5), (1, "Toy Story", 2),
(2, "Inception", 4), (2, "The Matrix", 5), (2, "Toy Story", 1),
(3, "Toy Story", 5), (3, "Up", 5), (3, "Inception", 2),
(4, "The Matrix", 5),(4, "Interstellar", 4), (4, "Up", 1),
(5, "Toy Story", 4), (5, "Up", 5), (5, "The Matrix", 2),
], columns=["user", "item", "rating"])
# user x item matrix; missing = 0 (no signal)
m = ratings.pivot_table(index="user", columns="item",
values="rating", fill_value=0)
# transpose so each ROW is an item's vector across users, then compare
sim = pd.DataFrame(cosine_similarity(m.T), index=m.columns, columns=m.columns)
def recommend(user, k=3):
rated = m.loc[user][m.loc[user] > 0].index
scores = sim[rated].dot(m.loc[user][rated]) # similarity-weighted
return scores.drop(rated).sort_values(ascending=False).head(k)
print(recommend(user=1)) # likes Inception/Interstellar -> expect The Matrix
User 1 loves cerebral sci-fi, so The Matrix surfaces above Toy Story — a few lines of linear algebra.
Neighbourhood similarity is brittle when two items share almost no common raters. Matrix factorisation fixes that by approximating the user-item matrix as the product of two skinny matrices, each with a few hidden columns called latent factors — learned dimensions of taste (one might behave like "arthouse vs blockbuster"). Each user and item becomes a short vector (an embedding); a predicted rating is their dot product, learned by minimising error on the ratings you have.
For implicit data the standard algorithm is ALS (Alternating Least Squares), in the implicit library, mf_als.py:
import numpy as np
from scipy.sparse import csr_matrix
from implicit.als import AlternatingLeastSquares
# implicit signals (play/click counts), NOT star ratings, sparse
data = np.array([5, 3, 1, 4, 2, 5, 2, 4, 1, 3, 5, 2])
rows = np.array([0, 0, 0, 1, 1, 1, 2, 2, 2, 3, 3, 3]) # users
cols = np.array([0, 1, 2, 0, 2, 3, 1, 3, 0, 1, 2, 3]) # items
ui = csr_matrix((data, (rows, cols)), shape=(4, 4))
model = AlternatingLeastSquares(factors=16, regularization=0.05,
iterations=20, random_state=42)
model.fit(ui)
ids, scores = model.recommend(0, ui[0], N=2)
print("Recommended:", ids, "scores:", scores.round(3))
print("User 0 embedding:", model.user_factors[0].round(2)) # reusable vector
Two dials matter. factors is the embedding size — too few can't capture taste nuance; too many overfits (memorises training data, generalises poorly). regularization curbs that by penalising large values, as in ridge regression.
The trap for engineers from classification: accuracy and RMSE are the wrong metrics here. A user sees a short ranked list and never scrolls to item 4,000 — what matters is whether good items sit near the top. Evaluate ranking, not rating:
- Precision@k — of the k items recommended, what fraction were relevant? (Is the top clean?)
- Recall@k — of all relevant items, what fraction made the top k? (Did we surface the good ones?)
- NDCG@k (Normalised Discounted Cumulative Gain) — like precision@k but rewards ranking the best items higher: a hit at rank 1 beats a hit at rank 5, discounted by position, normalised to 0–1.
- Coverage — what fraction of the catalogue ever gets recommended? Low coverage buries the long tail.
Compute the first three for a user's ranked list — metrics.py:
import numpy as np
def precision_at_k(rec, rel, k):
return sum(i in rel for i in rec[:k]) / k
def recall_at_k(rec, rel, k):
return sum(i in rel for i in rec[:k]) / len(rel) if rel else 0.0
def ndcg_at_k(rec, rel, k):
# each hit contributes 1/log2(rank+1) -> higher rank, more credit
dcg = sum(1.0 / np.log2(i + 2) for i, it in enumerate(rec[:k]) if it in rel)
idcg = sum(1.0 / np.log2(i + 2) for i in range(min(len(rel), k)))
return dcg / idcg if idcg > 0 else 0.0
rec = [5, 7, 12, 3, 42] # model's ranked item ids
rel = {12, 5, 99} # what the user actually liked
for fn in (precision_at_k, recall_at_k, ndcg_at_k):
print(f"{fn.__name__}: {fn(rec, rel, k=5):.3f}")
Precision and recall can disagree sharply; NDCG punishes item 12 at rank 3, not rank 1.
Collaborative filtering is blind to a brand-new item. Content-based filtering rescues it by recommending on item attributes: turn each item's text into a vector, then recommend ones close to what the user liked. We vectorise with TF-IDF (Term Frequency–Inverse Document Frequency) — a cheap classic that weights a word high if it's frequent in this item but rare across the catalogue — content.py:
import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
items = pd.DataFrame({
"item": ["Inception", "Interstellar", "The Matrix", "Toy Story", "Up"],
"tags": [
"sci-fi thriller dreams heist mind-bending",
"sci-fi space time travel emotional epic",
"sci-fi action virtual reality dystopia",
"animation family toys friendship comedy",
"animation family adventure balloons heartfelt",
]})
vecs = TfidfVectorizer().fit_transform(items["tags"]) # sparse TF-IDF
sim = cosine_similarity(vecs)
def similar_to(title, k=2):
i = items.index[items["item"] == title][0]
ranked = sorted(enumerate(sim[i]), key=lambda x: x[1], reverse=True)
return [(items.iloc[j]["item"], round(s, 3)) for j, s in ranked[1:k + 1]]
print("Because you watched Inception:", similar_to("Inception"))
This works on a day-one item with zero interactions — its tags alone place it near similar films. In production, swap TF-IDF for a sentence-transformer embedding.
A hybrid recommender blends both with a weighted score — alpha * cf_score + (1 - alpha) * content_score, alpha tuned per user (content for new users, collaborative as history grows). Netflix, Spotify, and Amazon are all hybrids run in two stages: a fast candidate generator retrieves a few hundred items, then a heavier ranker orders them.
Cold start is the defining failure mode — you can't use behaviour you don't have:
| Cold-start | Problem | Practical fix |
|---|---|---|
| New user | No interaction history | Content-based on a short onboarding survey; show trending |
| New item | No interactions yet | Content-based on attributes; inject into lists to gather signal |
| New system | Empty matrix | Bootstrap with popularity, editorial picks, imported metadata |
Use a fallback cascade — the strongest signal you have:
def recommend_with_fallback(user, n_interactions, k=10):
if n_interactions >= 20:
return collaborative(user, k) # rich history -> CF
elif n_interactions > 0:
cf, cb = collaborative(user, k), content_based(user, k)
return blend(cf, cb, alpha=n_interactions / 20) # ramp toward CF
return popular_items(k) # brand-new user
A pure-exploitation policy never shows new items enough to learn they're good; reserving a slice of each list for under-explored items (bandit-style exploration) keeps the tail alive.
Questions & Answers
Key Takeaways
- Two families, complementary gaps — collaborative filtering finds surprises but fails on new items; content-based works on day-one items but stays in a bubble. Combine them.
- Matrix factorisation turns sparsity into embeddings — short vectors over latent factors make predictions a cheap dot product and give reusable embeddings.
- Evaluate ranking, not rating — precision@k, recall@k, NDCG, and coverage measure what users experience; accuracy and RMSE miss the point, since users only see the top.
- Split by time, validate online — a chronological split prevents future leakage; only an A/B test on a real business metric confirms an offline win.
- Cold start is the hard part — fall back from collaborative to content-based to popularity as signal thins, and keep list space for exploration.
- LLMs sit at the edges, not the core — classical ML ranks cheaply; the LLM handles language, descriptions, and cold-start parsing.
Next Steps: Lesson 7: MLOps — Models to Production