Algorithm & Architecture Cheat Sheet
A one-page map of the algorithms that run most production AI. The shift from LLMs: an LLM is one huge general model you prompt, while these are small specialised models you train on your own data — cheaper, faster, offline-capable, and auditable. See Tools Reference.
Glossary (read this first)
| Term | Plain-language meaning |
|---|---|
| Feature | A measurable input column (amount, day_of_week). |
| Label / target | The answer you predict (fraud=1, price=42.0). |
| Overfitting | Memorises training data; fails on new data. |
| Cross-validation | Score across several splits, not one lucky split. |
| Transfer learning | Start from a pre-trained model, fine-tune on your data. |
| Embedding | Numbers representing meaning, comparable mathematically. |
| Precision / Recall | Of flagged items how many right / of true positives how many caught. |
Classical ML (scikit-learn)
Your default for tabular data (rows and columns). Trains in seconds, explains its decisions.
| Algorithm | Strength | Weakness | Typical use case |
|---|---|---|---|
| Linear regression | Simple, fast, interpretable | Only straight-line relationships | Price / demand |
| Logistic regression | Strong calibrated yes/no baseline | Misses complex interactions | Churn, click-through |
| Decision tree | Human-readable if/else rules | Overfits alone | Interpretable triage |
| Random forest | Robust, low-tuning, mixed data | Larger, less transparent | Fraud, medical risk |
| Gradient boosting | Best accuracy on tabular | Sensitive to tuning | Credit scoring, ranking |
| k-means | Fast unsupervised grouping | Must pick k; round clusters |
Segmentation |
| PCA | Compresses many features into few | Components hard to name | Visualising, denoising |
For large tables (10k+ rows) prefer HistGradientBoostingClassifier: much faster than the classic booster and handles missing values natively.
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import cross_val_score
from sklearn.datasets import load_breast_cancer
X, y = load_breast_cancer(return_X_y=True)
clf = HistGradientBoostingClassifier(max_iter=200, random_state=0)
# cross-validate = honest score across 5 splits, not one lucky one
print(cross_val_score(clf, X, y, cv=5, scoring="f1").mean()) # ~0.97
See Lesson 2: ML Fundamentals.
Computer Vision
Vision models read pixels. You rarely train from scratch — fine-tune a pre-trained model (transfer learning), needing hundreds of labels, not millions.
| Architecture | Strength | Weakness | Typical use case |
|---|---|---|---|
| CNN (e.g. ResNet) | Solid backbone for whole-image labels | Slower than tiny task models | Manufacturing QC pass/fail |
| YOLO (Ultralytics) | Real-time boxes around many objects | Weaker on tiny/overlapping | Safety cameras, counting |
| Mask R-CNN | Pixel-exact masks per object | Heavy, slower than YOLO | Medical imaging, measurement |
| SAM 2 (Meta) | Segments almost anything, zero training | Large; needs a click/box prompt | Annotation, video cut-out |
YOLO11 handles detection, segmentation, classification, pose, and oriented boxes from one toolkit; the newer YOLO26 adds NMS-free end-to-end inference for edge deployment. SAM 2 segments objects across images and video from a click, box, or mask.
from ultralytics import YOLO
model = YOLO("yolo11n.pt") # nano = fast, edge-friendly
results = model("factory_line.jpg") # boxes + class + confidence
See Lesson 3: Computer Vision in 2026.
NLP (encoders — not chatbots)
For routing tickets or tagging documents at scale, a small encoder beats an LLM on speed, cost, and consistency. An encoder turns text into embeddings for classifying or comparing; it does not generate.
| Tool / model | Strength | Weakness | Typical use case |
|---|---|---|---|
| spaCy pipeline | Fast NER + tagging, production-ready | Training needed for niche entities | Extract names, dates, orgs |
| BERT (fine-tuned) | High-accuracy classification, cheap to run | Needs labels + GPU to train | Sentiment, intent, moderation |
| Sentence-Transformers | Sentence embeddings for similarity | Not for generation | Semantic search, dedup, RAG |
BERT is a bidirectional encoder; pair it with AutoModelForSequenceClassification to fine-tune a classifier. spaCy's trained pipelines extract entities via doc.ents.
See Lesson 4: NLP Beyond Chat.
Time Series & Forecasting
Numbers over time: sales, sensors, server metrics. LLMs forecast numbers poorly; these are purpose-built and explainable.
| Method | Strength | Weakness | Typical use case |
|---|---|---|---|
| ARIMA / SARIMAX | Strong baseline, takes exogenous vars | Needs stationarity + tuning | Inventory, econometrics |
| Exponential smoothing | Trend + seasonality, very fast | No external drivers | Short-term demand |
| Prophet (Meta) | Easy, robust to gaps/holidays | Weaker on high-frequency data | Business metrics, capacity |
| Temporal CNN / transformer | Learns complex multi-series patterns | Data-hungry, needs GPU | Predictive maintenance |
statsmodels provides ARIMA, SARIMAX, and Holt-Winters ExponentialSmoothing. Prophet (pip install prophet) fits an additive seasonality + holidays model, takes columns ds/y, and returns yhat with uncertainty bounds.
See Lesson 5: Time Series & Forecasting.
Recommendation Systems
"What next?" engines. Choose by the data you have: behaviour, item attributes, or both.
| Approach | Strength | Weakness | Typical use case |
|---|---|---|---|
| Collaborative filtering | Learns taste from behaviour alone | Cold-start on new users/items | "Others also liked" |
| Content-based | Works day one via item features | Filter bubble | News, job matching |
| Hybrid | Blends both signals | More moving parts | Most production systems |
| Deep / two-tower | Scales to millions, learns embeddings | Data-hungry, harder to debug | Large e-commerce, streaming |
Evaluate with ranking metrics, not accuracy: precision@k (relevant items in the top k), NDCG (rewards ranking the best items highest), coverage (catalogue breadth). Item-based filtering is cosine_similarity over a users-by-items matrix.
See Lesson 6: Recommendation Systems.
Pick the right tool in 30 seconds
Rule of thumb: if the answer is a number, label, box, or ranking, a specialised model is usually faster, cheaper, and more auditable. Reach for the LLM only when the task is open-ended language. Full criteria: Decision Patterns.