AI/ML Tools & Frameworks Reference
A categorised map of the tools that do the work beyond LLMs — what the rest of ML runs on. Key terms are defined inline.
The landscape at a glance
| Domain | Go-to tool(s) | Reach for it when... |
|---|---|---|
| Tabular ML | scikit-learn | Rows-and-columns data; fast, interpretable |
| Deep learning | PyTorch (research), TensorFlow/Keras (serving) | Neural nets on images, audio, text |
| NLP | Transformers, spaCy | Encoders, NER, classification, embeddings |
| Computer vision | Ultralytics/YOLO, OpenCV | Detection/segmentation; image transforms |
| Time series | statsmodels, Prophet, NeuralForecast | ARIMA, seasonal forecasting, many series |
| MLOps | MLflow, Weights & Biases | Log runs, compare metrics, version models |
General ML — scikit-learn
The standard Python library for classical ML, built on NumPy and SciPy: classification, regression, clustering, dimensionality reduction, model selection, preprocessing; no deep learning or GPU.
Cross-validation (rotate which fold is held out) estimates unseen performance honestly and catches overfitting. Precision = of items flagged positive, how many really were; recall = of truly positive ones, how many caught.
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score, train_test_split
from sklearn.metrics import classification_report
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
clf = RandomForestClassifier(n_estimators=200, random_state=0)
print("CV accuracy:", cross_val_score(clf, X_tr, y_tr, cv=5).mean())
clf.fit(X_tr, y_tr)
print(classification_report(y_te, clf.predict(X_te))) # precision / recall / F1
Reach for it first on any tabular problem: fast, interpretable, and gradient-boosted trees usually beat a neural net here. An LLM would be slower and hard to audit.
Deep learning — PyTorch and TensorFlow
PyTorch is a tensor library with GPU acceleration and automatic differentiation (autograd) for neural networks — the dominant framework for research and fine-tuning. TensorFlow is an end-to-end ML platform whose Keras API and TensorFlow Lite give mature serving and on-device deployment.
Autograd: you write the forward pass and the framework computes gradients, so backpropagation (adjusting weights to reduce error) is free. Transfer learning: start from a model pretrained on a huge dataset, then adapt it to your task.
import torch, torch.nn as nn
model = nn.Sequential(nn.Linear(30, 64), nn.ReLU(), nn.Linear(64, 2))
opt, loss_fn = torch.optim.Adam(model.parameters(), lr=1e-3), nn.CrossEntropyLoss()
x, y = torch.randn(16, 30), torch.randint(0, 2, (16,)) # 16 rows, 30 features
for _ in range(50):
opt.zero_grad()
loss = loss_fn(model(x), y)
loss.backward() # autograd computes gradients
opt.step() # update weights
print("final loss:", round(loss.item(), 4))
Use deep learning for unstructured data; for tabular data, scikit-learn.
NLP — Hugging Face Transformers and spaCy
Transformers gives a uniform API to pretrained transformer models across text, vision, and audio — including encoder (BERT-family) models. spaCy is an industrial-strength NLP pipeline for tokenisation, POS tagging, parsing, and NER.
Embeddings are dense vectors representing meaning; similar texts land near each other, enabling search and classification. NER (Named Entity Recognition) locates spans like people, orgs, dates. A fine-tuned encoder beats an LLM on cost, latency, and consistency at high volume.
from transformers import pipeline
clf = pipeline("sentiment-analysis") # pretrained, no training
print(clf("The shipment arrived late and damaged."))
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("Acme Corp shipped 500 units from Berlin on March 3rd.")
print([(ent.text, ent.label_) for ent in doc.ents])
Computer vision — Ultralytics/YOLO and OpenCV
Ultralytics provides the YOLO real-time family — object detection, instance segmentation, pose estimation, and classification with minimal code, deployable to edge and cloud. OpenCV is the classic computer vision library for image/video I/O, transforms, and operators.
Object detection returns boxes + class labels per object (classification labels a whole image). YOLO fine-tunes from pretrained weights for QC.
from ultralytics import YOLO
model = YOLO("yolo11n.pt") # small pretrained detector
for box in model("factory_line.jpg")[0].boxes:
print(model.names[int(box.cls)], round(float(box.conf), 3))
import cv2 # OpenCV: classical preprocessing
gray = cv2.cvtColor(cv2.imread("part.jpg"), cv2.COLOR_BGR2GRAY)
cv2.imwrite("edges.jpg", cv2.Canny(gray, 100, 200))
On edge, for a defined task, a task-specific detector beats a multimodal LLM.
Time series — statsmodels, Prophet, NeuralForecast
statsmodels estimates classical statistical models — ARIMA, SARIMAX, exponential smoothing — with confidence intervals and diagnostics. Prophet (from Meta) fits an additive model of trend + seasonality + holidays, robust to missing data. NeuralForecast offers scalable neural models (NBEATS, NHITS, transformers) with a scikit-learn-style API for many series.
A time series is data indexed by time, where order matters and observations correlate — what general ML and LLMs handle poorly. Seasonality = a repeating cycle.
import pandas as pd
from prophet import Prophet
df = pd.DataFrame({"ds": pd.date_range("2025-01-01", periods=180), "y": range(180)})
m = Prophet().fit(df) # trend + seasonality + holidays
future = m.make_future_dataframe(periods=30)
print(m.predict(future)[["ds", "yhat", "yhat_lower", "yhat_upper"]].tail())
LLMs are poor at numerical forecasting; these models give calibrated intervals and respect seasonality. See troubleshooting for drift and leakage.
MLOps — MLflow and Weights & Biases
MLflow is an open-source platform for the ML lifecycle: experiment tracking, a model registry, packaging, and deployment. Weights & Biases is a hosted platform for tracking, dashboards, run comparison, and versioning.
Experiment tracking records each run's parameters, metrics, and artifacts so results are reproducible — the gap between "worked once on my laptop" and a retrainable system.
import mlflow
from sklearn.linear_model import LogisticRegression
with mlflow.start_run():
mlflow.log_param("C", 0.5)
m = LogisticRegression(C=0.5).fit(X_tr, y_tr)
mlflow.log_metric("accuracy", m.score(X_te, y_te))
mlflow.sklearn.log_model(m, "model") # params, metric, artifact logged
The default question
For any new problem, ask: is the input structured and the task well-defined? If yes, a specialised model above is almost always faster, cheaper, and more auditable than an LLM. Reserve LLMs for open-ended reasoning over unstructured input — and even then, a classical model often makes a good pre-filter.
Related: rules, cheat sheet, decision patterns, troubleshooting, course home.