MLOps — Models to Production
Learning Outcomes
- Track every experiment — parameters, metrics, and artifacts — so any result is reproducible with MLflow or Weights & Biases.
- Version and register models with a signature and an alias, then load the exact production version by name.
- Serve a model two ways — batch scoring and a real-time HTTP endpoint — and choose between them by latency and cost.
- Detect data drift in production with a statistical test instead of waiting for accuracy to collapse.
- Automate retraining behind guardrail metrics so a worse model can never be promoted by accident.
A model that scores well in a notebook has done about 20% of its job. The rest — making it reproducible, deployable, observable, and self-healing — is MLOps: the discipline that gets a trained model from a laptop to a system other people depend on. If DevOps is "ship code reliably," MLOps is "ship code and data and a model that silently degrades over time, reliably." This lesson assumes you can already train a model (see Lesson 2: Machine Learning Fundamentals); now we make it production-grade.
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | The MLOps lifecycle and why notebooks don't ship |
| Step 1–2 | 11 min | Experiment tracking with MLflow |
| Step 3 | 7 min | Experiment tracking with Weights & Biases |
| Step 4 | 9 min | Model versioning and the registry |
| Step 5 | 10 min | Serving: batch vs real-time |
| Step 6 | 10 min | Drift monitoring in production |
| Step 7 | 10 min | Automated retraining with guardrails |
Before You Begin
Pre-work:
- Complete Lesson 2: Machine Learning Fundamentals — be comfortable with train/test splits and a metric like F1.
- Skim Lesson 1: The AI Landscape so the classical-ML vs LLM contrast lands.
- Have a Python 3.10+ environment and a terminal you can run
uvicornfrom.
Shopping List:
- A virtual environment (
python -m venv .venv, then activate it). - The libraries below — all cross-platform via pip:
pip install mlflow scikit-learn pandas scipy fastapi uvicorn wandb joblib pyarrow
- A free Weights & Biases account for Step 3 (optional — skip the cloud parts and still follow along).
Before tools, fix the mental model. A classical-ML system has three moving parts DevOps never had to coordinate: code, data, and the model (the file produced by training). Change any one and the output changes.
Two failure modes dominate production ML:
- Reproducibility gaps — the good result happened, but nobody recorded which data, library versions, or hyperparameters (knobs you set before training, like tree depth) produced it.
- Training-serving skew — the model sees one thing in training and a subtly different thing in production (different scaling, missing columns, stale data), so it quietly underperforms.
| Stage | What happens | Failure mode it prevents |
|---|---|---|
| Track | Log params, metrics, artifacts of every run | Lost, irreproducible results |
| Version | Register the model file with a signature | Deploying the wrong/old model |
| Serve | Expose the model (batch or real-time) | Training-serving skew |
| Monitor | Watch inputs and outputs in production | Silent accuracy decay |
| Retrain | Refresh on new data, gated by metrics | Promoting a worse model |
An experiment is a named group of training attempts; each attempt is a run. Tracking records, per run: hyperparameters, the resulting metrics (single numbers measuring quality), and artifacts (files — the model, plots, a dataset hash).
A quick metric primer. For a fraud classifier, precision = of the transactions we flagged, how many were really fraud; recall = of all real fraud, how much we caught. F1 is their harmonic mean — one number that punishes ignoring either. ROC-AUC measures how well the model ranks fraud above non-fraud across thresholds.
import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.metrics import f1_score, roc_auc_score
X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
mlflow.set_experiment("fraud-baseline")
with mlflow.start_run(run_name="rf-100-trees"):
params = {"n_estimators": 100, "max_depth": 8, "random_state": 42}
model = RandomForestClassifier(**params).fit(X_train, y_train)
proba = model.predict_proba(X_test)[:, 1]
preds = model.predict(X_test)
mlflow.log_params(params)
mlflow.log_metric("f1", f1_score(y_test, preds))
mlflow.log_metric("roc_auc", roc_auc_score(y_test, proba))
mlflow.sklearn.log_model(model, artifact_path="model")
Run mlflow ui in the same directory: every run is a row you can sort by F1, compare side by side, and reopen months later with the exact params intact.
mlruns/ folder. For a team, point it at a shared tracking server and an artifact store via the tracking URI. The model-logging keyword has evolved across versions — if artifact_path warns as deprecated, switch to the name argument; behaviour is the same.Weights & Biases (W&B) is a hosted alternative with the same idea — runs, configs, metrics, artifacts — but a stronger emphasis on live dashboards, sweeps, and team collaboration. The workflow: create an Artifact, add files to it, then log it to the run.
import wandb, joblib
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import f1_score
run = wandb.init(project="fraud-detection",
config={"learning_rate": 0.1, "n_estimators": 200})
cfg = run.config
model = GradientBoostingClassifier(
learning_rate=cfg.learning_rate, n_estimators=cfg.n_estimators
).fit(X_train, y_train)
run.log({"f1": f1_score(y_test, model.predict(X_test))})
joblib.dump(model, "model.joblib")
artifact = wandb.Artifact("fraud-model", type="model")
artifact.add_file("model.joblib")
run.log_artifact(artifact)
run.finish()
| Concern | MLflow | Weights & Biases |
|---|---|---|
| Hosting | Self-host or managed | Primarily managed (SaaS) |
| Strength | Built-in registry, open standard | Live dashboards, sweeps |
| Offline | Works fully offline by default | Cloud-first (offline mode exists) |
| Lock-in | Low — open formats | Higher — tied to the platform |
A logged run is not yet a deployable artifact. The model registry is a versioned catalogue: each registered model has numbered versions, and you attach a moving label — an alias like production — to whichever version is live. (Older MLflow used named stages such as Staging/Production; those are deprecated, and aliases plus tags are now the recommended approach.)
You also want a model signature: a recorded schema of expected input columns and output type. It guards against training-serving skew — if production sends the wrong columns, it fails loudly instead of returning garbage.
import mlflow
from mlflow.models import infer_signature
from mlflow import MlflowClient
signature = infer_signature(X_train, model.predict(X_train))
with mlflow.start_run():
info = mlflow.sklearn.log_model(
model, artifact_path="model",
signature=signature,
registered_model_name="fraud-classifier",
)
client = MlflowClient()
client.set_registered_model_alias(
name="fraud-classifier", alias="production",
version=info.registered_model_version,
)
# Serving loads by alias, never by a hard-coded version:
prod_model = mlflow.pyfunc.load_model("models:/fraud-classifier@production")
Inference is using the trained model to make predictions. Two delivery shapes:
- Batch inference — score a large set of rows on a schedule. Optimises throughput (rows per hour) and cost. Good for "email tomorrow's churn-risk list."
- Real-time inference — score one request synchronously over HTTP. Optimises latency (milliseconds per call). Required for "block this card swipe now."
Real-time endpoint with FastAPI, loading the registered model by alias:
# serve.py — run with: uvicorn serve:app --host 0.0.0.0 --port 8000
from fastapi import FastAPI
from pydantic import BaseModel
import mlflow.pyfunc
import pandas as pd
app = FastAPI()
model = mlflow.pyfunc.load_model("models:/fraud-classifier@production")
class Transaction(BaseModel):
amount: float
merchant_risk: float
hour_of_day: int
@app.post("/predict")
def predict(tx: Transaction):
row = pd.DataFrame([tx.model_dump()])
score = float(model.predict(row)[0])
return {"fraud_score": score, "flag": score > 0.5}
Batch scoring is just a scheduled script:
# score_batch.py — run nightly via cron or an orchestrator
import pandas as pd
import mlflow.pyfunc
FEATURES = ["amount", "merchant_risk", "hour_of_day"]
model = mlflow.pyfunc.load_model("models:/fraud-classifier@production")
df = pd.read_parquet("data/transactions_2026-05-31.parquet")
df["fraud_score"] = model.predict(df[FEATURES])
df.to_parquet("data/scored_2026-05-31.parquet")
| Dimension | Batch | Real-time |
|---|---|---|
| Optimises | Throughput, cost | Latency |
| Trigger | Schedule | Per request |
| Infra | A job runner | A live, scaled service |
| Use case | Nightly risk list | Block a card swipe |
Models do not crash when they go wrong — they quietly get less accurate as the world moves on. Two distinct things shift:
- Data drift — the inputs change shape. Average transaction amount jumps after a pricing change; the model still runs but sees unfamiliar values.
- Concept drift — the relationship between inputs and the answer changes. Fraudsters invent a new pattern, so the same features now mean something different.
You catch data drift fast — before you even have labels — with a statistical test. The Kolmogorov–Smirnov (KS) test compares two samples of a numeric feature and returns a p-value; a tiny p-value means the live distribution differs from the training one.
import numpy as np
from scipy.stats import ks_2samp
def drift_report(reference: np.ndarray, current: np.ndarray, alpha=0.01):
stat, p_value = ks_2samp(reference, current)
return {
"ks_statistic": round(float(stat), 4),
"p_value": round(float(p_value), 6),
"drift_detected": bool(p_value < alpha),
}
reference = X_train["mean radius"].to_numpy()
live = X_test["mean radius"].to_numpy() + 3.0 # simulate a shift
print(drift_report(reference, live))
# -> {'ks_statistic': ..., 'p_value': ..., 'drift_detected': True}
Concept drift needs ground truth — the eventual real outcome (did the flagged transaction turn out to be fraud?). Once labels arrive, recompute live F1 on a rolling window and alert when it falls below a threshold.
scipy is enough to start.The payoff of everything above: a model that refreshes itself safely. Retraining means training a new version on newer data. The danger is promoting a worse model, so retraining must be gated by a guardrail metric — a candidate is promoted only if it beats a hard floor on a held-out set.
# retrain_if_needed.py — trains, evaluates, promotes only if it passes
from mlflow import MlflowClient
DRIFT_THRESHOLD = 0.2 # KS statistic
F1_GUARDRAIL = 0.85 # never promote below this
def should_retrain(drift_stat: float, live_f1: float) -> bool:
return drift_stat > DRIFT_THRESHOLD or live_f1 < F1_GUARDRAIL
def promote(version: str):
MlflowClient().set_registered_model_alias(
"fraud-classifier", "production", version)
if should_retrain(drift_stat=0.27, live_f1=0.82):
new_version = train_and_register() # logs params, metrics, signature
candidate_f1 = evaluate(new_version) # on a fresh held-out set
if candidate_f1 >= F1_GUARDRAIL:
promote(new_version) # alias flips atomically
# else: keep current production model, page a human
Schedule it with cron (or an orchestrator like Airflow / Prefect for retries and lineage):
# /etc/cron.d/ml-retrain — 02:00 daily
0 2 * * * mluser python /opt/ml/retrain_if_needed.py >> /var/log/retrain.log 2>&1
For CI/CD the same idea moves into your pipeline: on a merge or schedule, a job trains, runs the guardrail evaluation, and the alias flip is the deploy. Because serving loads models:/fraud-classifier@production, promotion needs no service redeploy — it picks up the new version on next load. (Express the pipeline in your CI tool's own config rather than hand-templating it here.)
production alias back at the previous version. No rebuild, no redeploy. This one indirection removes most of the fear from shipping models.Questions & Answers
Key Takeaways
- A notebook result is not a product. The MLOps lifecycle — track, version, serve, monitor, retrain — turns a one-off score into a system people can rely on.
- Track everything, automatically. Logging params, metrics, and artifacts per run with MLflow or W&B makes results reproducible and comparable; an untracked good result is a lucky accident you can't repeat.
- The registry decouples "which model is live" from your code. A signature catches training-serving skew, and an alias makes promotion and rollback an instant flag flip rather than a redeploy.
- Choose serving by latency, default to batch. Batch is cheaper and simpler; pay for real-time only when the decision must happen inside a live request.
- Models decay silently — monitor inputs and outputs. Catch data drift now with a cheap statistical test, and concept drift later with ground-truth labels, before accuracy quietly collapses.
- Automate retraining behind guardrails. Retrain on a signal, not a clock, and never promote a candidate that fails a hard metric floor — so the pipeline can never make the model worse.
Next Steps: Lesson 8: When to Use Classical ML vs GenAI