Recommendation Systems

60 min intermediate Lesson 6

Learning Outcomes

  • Distinguish collaborative from content-based filtering and explain when each fits
  • Build an item-based recommender and a matrix-factorisation model from interaction data
  • Evaluate recommendations with ranking metrics — precision@k, recall@k, NDCG, coverage
  • Diagnose and mitigate the cold-start problem for new users and new items
  • Decide where an LLM belongs in a recommender stack versus where classical ML wins

Lesson Plan

Segment Duration Topic
Intro 3 min Why recommenders are ML's biggest commercial win
Explain 8 min The two families: collaborative vs content-based
Build 12 min Item-based collaborative filtering from scratch
Build 12 min Matrix factorisation with a real library
Explain 10 min Evaluation metrics unique to ranking
Build 8 min Content-based + hybrid recommenders
Explain 4 min Cold-start and where LLMs fit
Wrap-up 3 min Key takeaways, preview MLOps

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ in a virtual environment
  • pip install numpy pandas scikit-learn scipy implicit
  • A terminal and a code editor

1 The Two Families: Collaborative vs Content-Based

A recommender system predicts what a user will want from a catalogue too large to browse — "you might also like" rows, the autoplay queue, a self-reordering homepage. It's a quiet workhorse behind much of ML's revenue, and two ideas power most of it.

Collaborative filtering looks only at behaviour: people who interacted with what you did also interacted with X. Content-based filtering looks at item attributes: you liked sci-fi, here's another.

Approach Signal Strength Weakness
Collaborative User-item interactions Finds cross-genre hits Useless for brand-new items/users
Content-based Item attributes/text Works on day-one items Stays in a filter bubble
Hybrid Both Covers each other's gaps More moving parts

The data is an interaction matrix: rows users, columns items, cells ratings or implicit signals (a play, click, purchase). It is almost entirely empty — a typical user has touched well under 1% of the catalogue. This sparsity is the central reality of recommenders.

NOTE
Explicit vs implicit feedback
Explicit feedback is a deliberate rating (1–5 stars). Implicit feedback is behaviour you infer preference from — a play, click, watch-time. It's far more abundant but noisier (a click is not a thumbs-up), and most production systems run on it.

2 Build Item-Based Collaborative Filtering From Scratch

The most intuitive recommender is item-based collaborative filtering: measure how similarly users behave toward each pair of items, then recommend ones similar to what a user liked. "Similar" means cosine similarity — items rated alike by the same users point the same direction (near 1).

On a tiny ratings table, item_cf.py:

import pandas as pd
from sklearn.metrics.pairwise import cosine_similarity

# user, item, rating (1-5)
ratings = pd.DataFrame([
    (1, "Inception", 5), (1, "Interstellar", 5), (1, "Toy Story", 2),
    (2, "Inception", 4), (2, "The Matrix", 5),   (2, "Toy Story", 1),
    (3, "Toy Story", 5), (3, "Up", 5),           (3, "Inception", 2),
    (4, "The Matrix", 5),(4, "Interstellar", 4), (4, "Up", 1),
    (5, "Toy Story", 4), (5, "Up", 5),           (5, "The Matrix", 2),
], columns=["user", "item", "rating"])

# user x item matrix; missing = 0 (no signal)
m = ratings.pivot_table(index="user", columns="item",
                        values="rating", fill_value=0)

# transpose so each ROW is an item's vector across users, then compare
sim = pd.DataFrame(cosine_similarity(m.T), index=m.columns, columns=m.columns)

def recommend(user, k=3):
    rated = m.loc[user][m.loc[user] > 0].index
    scores = sim[rated].dot(m.loc[user][rated])  # similarity-weighted
    return scores.drop(rated).sort_values(ascending=False).head(k)

print(recommend(user=1))   # likes Inception/Interstellar -> expect The Matrix

User 1 loves cerebral sci-fi, so The Matrix surfaces above Toy Story — a few lines of linear algebra.

TIP
Why item-based beats user-based at scale
Users change daily and often outnumber items, but item-item similarities stay stable for years — precompute them nightly and serve from a fast lookup. Caveat: filling blanks with 0 conflates 'disliked' with 'never seen' — mean-center each user, or use an implicit model.

3 Matrix Factorisation — Learning Latent Tastes

Neighbourhood similarity is brittle when two items share almost no common raters. Matrix factorisation fixes that by approximating the user-item matrix as the product of two skinny matrices, each with a few hidden columns called latent factors — learned dimensions of taste (one might behave like "arthouse vs blockbuster"). Each user and item becomes a short vector (an embedding); a predicted rating is their dot product, learned by minimising error on the ratings you have.

For implicit data the standard algorithm is ALS (Alternating Least Squares), in the implicit library, mf_als.py:

import numpy as np
from scipy.sparse import csr_matrix
from implicit.als import AlternatingLeastSquares

# implicit signals (play/click counts), NOT star ratings, sparse
data = np.array([5, 3, 1, 4, 2, 5, 2, 4, 1, 3, 5, 2])
rows = np.array([0, 0, 0, 1, 1, 1, 2, 2, 2, 3, 3, 3])   # users
cols = np.array([0, 1, 2, 0, 2, 3, 1, 3, 0, 1, 2, 3])   # items
ui = csr_matrix((data, (rows, cols)), shape=(4, 4))

model = AlternatingLeastSquares(factors=16, regularization=0.05,
                                iterations=20, random_state=42)
model.fit(ui)

ids, scores = model.recommend(0, ui[0], N=2)
print("Recommended:", ids, "scores:", scores.round(3))
print("User 0 embedding:", model.user_factors[0].round(2))  # reusable vector

Two dials matter. factors is the embedding size — too few can't capture taste nuance; too many overfits (memorises training data, generalises poorly). regularization curbs that by penalising large values, as in ridge regression.

NOTE
Embeddings bridge to everything else
The vectors ALS learns are reusable embeddings — feed one into a classifier, run nearest-neighbour search, or combine them with LLM text embeddings. This is why Lesson 9 can fuse classical recommenders with generative models. (Implicit ALS expects counts; for star ratings use SVD.)

4 Evaluate With Ranking Metrics, Not Accuracy

The trap for engineers from classification: accuracy and RMSE are the wrong metrics here. A user sees a short ranked list and never scrolls to item 4,000 — what matters is whether good items sit near the top. Evaluate ranking, not rating:

  • Precision@k — of the k items recommended, what fraction were relevant? (Is the top clean?)
  • Recall@k — of all relevant items, what fraction made the top k? (Did we surface the good ones?)
  • NDCG@k (Normalised Discounted Cumulative Gain) — like precision@k but rewards ranking the best items higher: a hit at rank 1 beats a hit at rank 5, discounted by position, normalised to 0–1.
  • Coverage — what fraction of the catalogue ever gets recommended? Low coverage buries the long tail.

Compute the first three for a user's ranked list — metrics.py:

import numpy as np

def precision_at_k(rec, rel, k):
    return sum(i in rel for i in rec[:k]) / k

def recall_at_k(rec, rel, k):
    return sum(i in rel for i in rec[:k]) / len(rel) if rel else 0.0

def ndcg_at_k(rec, rel, k):
    # each hit contributes 1/log2(rank+1) -> higher rank, more credit
    dcg = sum(1.0 / np.log2(i + 2) for i, it in enumerate(rec[:k]) if it in rel)
    idcg = sum(1.0 / np.log2(i + 2) for i in range(min(len(rel), k)))
    return dcg / idcg if idcg > 0 else 0.0

rec = [5, 7, 12, 3, 42]   # model's ranked item ids
rel = {12, 5, 99}         # what the user actually liked
for fn in (precision_at_k, recall_at_k, ndcg_at_k):
    print(f"{fn.__name__}: {fn(rec, rel, k=5):.3f}")

Precision and recall can disagree sharply; NDCG punishes item 12 at rank 3, not rank 1.

WARNING
Never evaluate on a random split
Recommenders predict the future from the past, so split chronologically or hold out each user's last item — a random split leaks future behaviour and inflates every metric. Offline scores are a proxy; confirm the winner with an online A/B test.

5 Content-Based and Hybrid Recommenders

Collaborative filtering is blind to a brand-new item. Content-based filtering rescues it by recommending on item attributes: turn each item's text into a vector, then recommend ones close to what the user liked. We vectorise with TF-IDF (Term Frequency–Inverse Document Frequency) — a cheap classic that weights a word high if it's frequent in this item but rare across the catalogue — content.py:

import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

items = pd.DataFrame({
    "item": ["Inception", "Interstellar", "The Matrix", "Toy Story", "Up"],
    "tags": [
        "sci-fi thriller dreams heist mind-bending",
        "sci-fi space time travel emotional epic",
        "sci-fi action virtual reality dystopia",
        "animation family toys friendship comedy",
        "animation family adventure balloons heartfelt",
    ]})

vecs = TfidfVectorizer().fit_transform(items["tags"])   # sparse TF-IDF
sim = cosine_similarity(vecs)

def similar_to(title, k=2):
    i = items.index[items["item"] == title][0]
    ranked = sorted(enumerate(sim[i]), key=lambda x: x[1], reverse=True)
    return [(items.iloc[j]["item"], round(s, 3)) for j, s in ranked[1:k + 1]]

print("Because you watched Inception:", similar_to("Inception"))

This works on a day-one item with zero interactions — its tags alone place it near similar films. In production, swap TF-IDF for a sentence-transformer embedding.

A hybrid recommender blends both with a weighted score — alpha * cf_score + (1 - alpha) * content_score, alpha tuned per user (content for new users, collaborative as history grows). Netflix, Spotify, and Amazon are all hybrids run in two stages: a fast candidate generator retrieves a few hundred items, then a heavier ranker orders them.


6 The Cold-Start Problem and Where LLMs Fit

Cold start is the defining failure mode — you can't use behaviour you don't have:

Cold-start Problem Practical fix
New user No interaction history Content-based on a short onboarding survey; show trending
New item No interactions yet Content-based on attributes; inject into lists to gather signal
New system Empty matrix Bootstrap with popularity, editorial picks, imported metadata

Use a fallback cascade — the strongest signal you have:

def recommend_with_fallback(user, n_interactions, k=10):
    if n_interactions >= 20:
        return collaborative(user, k)             # rich history -> CF
    elif n_interactions > 0:
        cf, cb = collaborative(user, k), content_based(user, k)
        return blend(cf, cb, alpha=n_interactions / 20)  # ramp toward CF
    return popular_items(k)                        # brand-new user

A pure-exploitation policy never shows new items enough to learn they're good; reserving a slice of each list for under-explored items (bandit-style exploration) keeps the tail alive.

TIP
Where LLMs actually fit
LLMs are NOT the ranking engine — too slow and costly to score millions of candidates per request, and they don't learn from your interaction matrix. They shine at the cold edges: generating item descriptions for content-based similarity, turning free-text onboarding into a taste profile, and writing 'why we recommended this' copy. Classical ML ranks; the LLM does the language.

Questions & Answers

Q: Couldn't I just put my whole catalogue in an LLM prompt and ask it to recommend?
For a toy catalogue, yes. At scale, no: a million items won't fit a context window, the LLM has never seen your users' behaviour, and per-request cost and latency dwarf a dot product. Use it for language and cold start, not ranking.
Q: My offline NDCG is great but engagement didn't move in the A/B test. What happened?
Likely causes: your split leaked future data (use a chronological split), your "relevant" labels don't match what users value, or you optimised for clicks while the business cares about retention. Offline metrics filter bad models; the A/B test decides.
Q: How do I pick the number of latent factors for matrix factorisation?
Sweep it (16, 32, 64, 128) and evaluate ranking metrics on a held-out chronological split, watching for where validation NDCG stops improving — beyond that you're overfitting. More factors capture finer taste but cost memory and latency; tune jointly with regularization.
Q: Implicit feedback is so noisy — a click isn't a real preference. Doesn't that wreck the model?
It's noisy but plentiful. Implicit MF weights interactions by confidence (more plays = more signal) and treats unobserved cells as weak negatives. You lose precision per interaction but gain volume. Just don't pass raw counts into a model built for explicit star ratings.

Key Takeaways

  1. Two families, complementary gaps — collaborative filtering finds surprises but fails on new items; content-based works on day-one items but stays in a bubble. Combine them.
  2. Matrix factorisation turns sparsity into embeddings — short vectors over latent factors make predictions a cheap dot product and give reusable embeddings.
  3. Evaluate ranking, not rating — precision@k, recall@k, NDCG, and coverage measure what users experience; accuracy and RMSE miss the point, since users only see the top.
  4. Split by time, validate online — a chronological split prevents future leakage; only an A/B test on a real business metric confirms an offline win.
  5. Cold start is the hard part — fall back from collaborative to content-based to popularity as signal thins, and keep list space for exploration.
  6. LLMs sit at the edges, not the core — classical ML ranks cheaply; the LLM handles language, descriptions, and cold-start parsing.

Next Steps: Lesson 7: MLOps — Models to Production