Algorithm & Architecture Cheat Sheet

Reference intermediate

A one-page map of the algorithms that run most production AI. The shift from LLMs: an LLM is one huge general model you prompt, while these are small specialised models you train on your own data — cheaper, faster, offline-capable, and auditable. See Tools Reference.

Glossary (read this first)

Term Plain-language meaning
Feature A measurable input column (amount, day_of_week).
Label / target The answer you predict (fraud=1, price=42.0).
Overfitting Memorises training data; fails on new data.
Cross-validation Score across several splits, not one lucky split.
Transfer learning Start from a pre-trained model, fine-tune on your data.
Embedding Numbers representing meaning, comparable mathematically.
Precision / Recall Of flagged items how many right / of true positives how many caught.

Classical ML (scikit-learn)

Your default for tabular data (rows and columns). Trains in seconds, explains its decisions.

Algorithm Strength Weakness Typical use case
Linear regression Simple, fast, interpretable Only straight-line relationships Price / demand
Logistic regression Strong calibrated yes/no baseline Misses complex interactions Churn, click-through
Decision tree Human-readable if/else rules Overfits alone Interpretable triage
Random forest Robust, low-tuning, mixed data Larger, less transparent Fraud, medical risk
Gradient boosting Best accuracy on tabular Sensitive to tuning Credit scoring, ranking
k-means Fast unsupervised grouping Must pick k; round clusters Segmentation
PCA Compresses many features into few Components hard to name Visualising, denoising

For large tables (10k+ rows) prefer HistGradientBoostingClassifier: much faster than the classic booster and handles missing values natively.

from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import cross_val_score
from sklearn.datasets import load_breast_cancer

X, y = load_breast_cancer(return_X_y=True)
clf = HistGradientBoostingClassifier(max_iter=200, random_state=0)
# cross-validate = honest score across 5 splits, not one lucky one
print(cross_val_score(clf, X, y, cv=5, scoring="f1").mean())  # ~0.97

See Lesson 2: ML Fundamentals.

Computer Vision

Vision models read pixels. You rarely train from scratch — fine-tune a pre-trained model (transfer learning), needing hundreds of labels, not millions.

Architecture Strength Weakness Typical use case
CNN (e.g. ResNet) Solid backbone for whole-image labels Slower than tiny task models Manufacturing QC pass/fail
YOLO (Ultralytics) Real-time boxes around many objects Weaker on tiny/overlapping Safety cameras, counting
Mask R-CNN Pixel-exact masks per object Heavy, slower than YOLO Medical imaging, measurement
SAM 2 (Meta) Segments almost anything, zero training Large; needs a click/box prompt Annotation, video cut-out

YOLO11 handles detection, segmentation, classification, pose, and oriented boxes from one toolkit; the newer YOLO26 adds NMS-free end-to-end inference for edge deployment. SAM 2 segments objects across images and video from a click, box, or mask.

from ultralytics import YOLO
model = YOLO("yolo11n.pt")           # nano = fast, edge-friendly
results = model("factory_line.jpg")  # boxes + class + confidence

See Lesson 3: Computer Vision in 2026.

NLP (encoders — not chatbots)

For routing tickets or tagging documents at scale, a small encoder beats an LLM on speed, cost, and consistency. An encoder turns text into embeddings for classifying or comparing; it does not generate.

Tool / model Strength Weakness Typical use case
spaCy pipeline Fast NER + tagging, production-ready Training needed for niche entities Extract names, dates, orgs
BERT (fine-tuned) High-accuracy classification, cheap to run Needs labels + GPU to train Sentiment, intent, moderation
Sentence-Transformers Sentence embeddings for similarity Not for generation Semantic search, dedup, RAG

BERT is a bidirectional encoder; pair it with AutoModelForSequenceClassification to fine-tune a classifier. spaCy's trained pipelines extract entities via doc.ents.

See Lesson 4: NLP Beyond Chat.

Time Series & Forecasting

Numbers over time: sales, sensors, server metrics. LLMs forecast numbers poorly; these are purpose-built and explainable.

Method Strength Weakness Typical use case
ARIMA / SARIMAX Strong baseline, takes exogenous vars Needs stationarity + tuning Inventory, econometrics
Exponential smoothing Trend + seasonality, very fast No external drivers Short-term demand
Prophet (Meta) Easy, robust to gaps/holidays Weaker on high-frequency data Business metrics, capacity
Temporal CNN / transformer Learns complex multi-series patterns Data-hungry, needs GPU Predictive maintenance

statsmodels provides ARIMA, SARIMAX, and Holt-Winters ExponentialSmoothing. Prophet (pip install prophet) fits an additive seasonality + holidays model, takes columns ds/y, and returns yhat with uncertainty bounds.

See Lesson 5: Time Series & Forecasting.

Recommendation Systems

"What next?" engines. Choose by the data you have: behaviour, item attributes, or both.

Approach Strength Weakness Typical use case
Collaborative filtering Learns taste from behaviour alone Cold-start on new users/items "Others also liked"
Content-based Works day one via item features Filter bubble News, job matching
Hybrid Blends both signals More moving parts Most production systems
Deep / two-tower Scales to millions, learns embeddings Data-hungry, harder to debug Large e-commerce, streaming

Evaluate with ranking metrics, not accuracy: precision@k (relevant items in the top k), NDCG (rewards ranking the best items highest), coverage (catalogue breadth). Item-based filtering is cosine_similarity over a users-by-items matrix.

See Lesson 6: Recommendation Systems.

Pick the right tool in 30 seconds

Rule of thumb: if the answer is a number, label, box, or ranking, a specialised model is usually faster, cheaper, and more auditable. Reach for the LLM only when the task is open-ended language. Full criteria: Decision Patterns.