AI/ML Tools & Frameworks Reference

Reference intermediate

A categorised map of the tools that do the work beyond LLMs — what the rest of ML runs on. Key terms are defined inline.

The landscape at a glance

Domain Go-to tool(s) Reach for it when...
Tabular ML scikit-learn Rows-and-columns data; fast, interpretable
Deep learning PyTorch (research), TensorFlow/Keras (serving) Neural nets on images, audio, text
NLP Transformers, spaCy Encoders, NER, classification, embeddings
Computer vision Ultralytics/YOLO, OpenCV Detection/segmentation; image transforms
Time series statsmodels, Prophet, NeuralForecast ARIMA, seasonal forecasting, many series
MLOps MLflow, Weights & Biases Log runs, compare metrics, version models

General ML — scikit-learn

The standard Python library for classical ML, built on NumPy and SciPy: classification, regression, clustering, dimensionality reduction, model selection, preprocessing; no deep learning or GPU.

Cross-validation (rotate which fold is held out) estimates unseen performance honestly and catches overfitting. Precision = of items flagged positive, how many really were; recall = of truly positive ones, how many caught.

from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score, train_test_split
from sklearn.metrics import classification_report

X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
clf = RandomForestClassifier(n_estimators=200, random_state=0)
print("CV accuracy:", cross_val_score(clf, X_tr, y_tr, cv=5).mean())
clf.fit(X_tr, y_tr)
print(classification_report(y_te, clf.predict(X_te)))  # precision / recall / F1

Reach for it first on any tabular problem: fast, interpretable, and gradient-boosted trees usually beat a neural net here. An LLM would be slower and hard to audit.

Deep learning — PyTorch and TensorFlow

PyTorch is a tensor library with GPU acceleration and automatic differentiation (autograd) for neural networks — the dominant framework for research and fine-tuning. TensorFlow is an end-to-end ML platform whose Keras API and TensorFlow Lite give mature serving and on-device deployment.

Autograd: you write the forward pass and the framework computes gradients, so backpropagation (adjusting weights to reduce error) is free. Transfer learning: start from a model pretrained on a huge dataset, then adapt it to your task.

import torch, torch.nn as nn

model = nn.Sequential(nn.Linear(30, 64), nn.ReLU(), nn.Linear(64, 2))
opt, loss_fn = torch.optim.Adam(model.parameters(), lr=1e-3), nn.CrossEntropyLoss()
x, y = torch.randn(16, 30), torch.randint(0, 2, (16,))   # 16 rows, 30 features
for _ in range(50):
    opt.zero_grad()
    loss = loss_fn(model(x), y)
    loss.backward()              # autograd computes gradients
    opt.step()                   # update weights
print("final loss:", round(loss.item(), 4))

Use deep learning for unstructured data; for tabular data, scikit-learn.

NLP — Hugging Face Transformers and spaCy

Transformers gives a uniform API to pretrained transformer models across text, vision, and audio — including encoder (BERT-family) models. spaCy is an industrial-strength NLP pipeline for tokenisation, POS tagging, parsing, and NER.

Embeddings are dense vectors representing meaning; similar texts land near each other, enabling search and classification. NER (Named Entity Recognition) locates spans like people, orgs, dates. A fine-tuned encoder beats an LLM on cost, latency, and consistency at high volume.

from transformers import pipeline
clf = pipeline("sentiment-analysis")          # pretrained, no training
print(clf("The shipment arrived late and damaged."))

import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("Acme Corp shipped 500 units from Berlin on March 3rd.")
print([(ent.text, ent.label_) for ent in doc.ents])

Computer vision — Ultralytics/YOLO and OpenCV

Ultralytics provides the YOLO real-time family — object detection, instance segmentation, pose estimation, and classification with minimal code, deployable to edge and cloud. OpenCV is the classic computer vision library for image/video I/O, transforms, and operators.

Object detection returns boxes + class labels per object (classification labels a whole image). YOLO fine-tunes from pretrained weights for QC.

from ultralytics import YOLO
model = YOLO("yolo11n.pt")            # small pretrained detector
for box in model("factory_line.jpg")[0].boxes:
    print(model.names[int(box.cls)], round(float(box.conf), 3))

import cv2                            # OpenCV: classical preprocessing
gray = cv2.cvtColor(cv2.imread("part.jpg"), cv2.COLOR_BGR2GRAY)
cv2.imwrite("edges.jpg", cv2.Canny(gray, 100, 200))

On edge, for a defined task, a task-specific detector beats a multimodal LLM.

Time series — statsmodels, Prophet, NeuralForecast

statsmodels estimates classical statistical models — ARIMA, SARIMAX, exponential smoothing — with confidence intervals and diagnostics. Prophet (from Meta) fits an additive model of trend + seasonality + holidays, robust to missing data. NeuralForecast offers scalable neural models (NBEATS, NHITS, transformers) with a scikit-learn-style API for many series.

A time series is data indexed by time, where order matters and observations correlate — what general ML and LLMs handle poorly. Seasonality = a repeating cycle.

import pandas as pd
from prophet import Prophet

df = pd.DataFrame({"ds": pd.date_range("2025-01-01", periods=180), "y": range(180)})
m = Prophet().fit(df)                              # trend + seasonality + holidays
future = m.make_future_dataframe(periods=30)
print(m.predict(future)[["ds", "yhat", "yhat_lower", "yhat_upper"]].tail())

LLMs are poor at numerical forecasting; these models give calibrated intervals and respect seasonality. See troubleshooting for drift and leakage.

MLOps — MLflow and Weights & Biases

MLflow is an open-source platform for the ML lifecycle: experiment tracking, a model registry, packaging, and deployment. Weights & Biases is a hosted platform for tracking, dashboards, run comparison, and versioning.

Experiment tracking records each run's parameters, metrics, and artifacts so results are reproducible — the gap between "worked once on my laptop" and a retrainable system.

import mlflow
from sklearn.linear_model import LogisticRegression

with mlflow.start_run():
    mlflow.log_param("C", 0.5)
    m = LogisticRegression(C=0.5).fit(X_tr, y_tr)
    mlflow.log_metric("accuracy", m.score(X_te, y_te))
    mlflow.sklearn.log_model(m, "model")    # params, metric, artifact logged

The default question

For any new problem, ask: is the input structured and the task well-defined? If yes, a specialised model above is almost always faster, cheaper, and more auditable than an LLM. Reserve LLMs for open-ended reasoning over unstructured input — and even then, a classical model often makes a good pre-filter.

Related: rules, cheat sheet, decision patterns, troubleshooting, course home.