MLOps — Models to Production

60 min advanced Lesson 7

Learning Outcomes

  • Track every experiment — parameters, metrics, and artifacts — so any result is reproducible with MLflow or Weights & Biases.
  • Version and register models with a signature and an alias, then load the exact production version by name.
  • Serve a model two ways — batch scoring and a real-time HTTP endpoint — and choose between them by latency and cost.
  • Detect data drift in production with a statistical test instead of waiting for accuracy to collapse.
  • Automate retraining behind guardrail metrics so a worse model can never be promoted by accident.

A model that scores well in a notebook has done about 20% of its job. The rest — making it reproducible, deployable, observable, and self-healing — is MLOps: the discipline that gets a trained model from a laptop to a system other people depend on. If DevOps is "ship code reliably," MLOps is "ship code and data and a model that silently degrades over time, reliably." This lesson assumes you can already train a model (see Lesson 2: Machine Learning Fundamentals); now we make it production-grade.

Lesson Plan

Segment Duration Topic
Intro 3 min The MLOps lifecycle and why notebooks don't ship
Step 1–2 11 min Experiment tracking with MLflow
Step 3 7 min Experiment tracking with Weights & Biases
Step 4 9 min Model versioning and the registry
Step 5 10 min Serving: batch vs real-time
Step 6 10 min Drift monitoring in production
Step 7 10 min Automated retraining with guardrails

Before You Begin

Pre-work:

Shopping List:

  • A virtual environment (python -m venv .venv, then activate it).
  • The libraries below — all cross-platform via pip:
pip install mlflow scikit-learn pandas scipy fastapi uvicorn wandb joblib pyarrow
  • A free Weights & Biases account for Step 3 (optional — skip the cloud parts and still follow along).

1 Understand the MLOps lifecycle

Before tools, fix the mental model. A classical-ML system has three moving parts DevOps never had to coordinate: code, data, and the model (the file produced by training). Change any one and the output changes.

Two failure modes dominate production ML:

  • Reproducibility gaps — the good result happened, but nobody recorded which data, library versions, or hyperparameters (knobs you set before training, like tree depth) produced it.
  • Training-serving skew — the model sees one thing in training and a subtly different thing in production (different scaling, missing columns, stale data), so it quietly underperforms.
Stage What happens Failure mode it prevents
Track Log params, metrics, artifacts of every run Lost, irreproducible results
Version Register the model file with a signature Deploying the wrong/old model
Serve Expose the model (batch or real-time) Training-serving skew
Monitor Watch inputs and outputs in production Silent accuracy decay
Retrain Refresh on new data, gated by metrics Promoting a worse model
NOTE
This applies to LLMs too — differently
Classical models drift on numeric feature distributions and are retrained on a cadence. LLMs are rarely retrained by you; instead you version the prompt and model ID and monitor outputs and cost. Same five stages, different levers. We return to this in Lesson 8.

2 Track experiments with MLflow

An experiment is a named group of training attempts; each attempt is a run. Tracking records, per run: hyperparameters, the resulting metrics (single numbers measuring quality), and artifacts (files — the model, plots, a dataset hash).

A quick metric primer. For a fraud classifier, precision = of the transactions we flagged, how many were really fraud; recall = of all real fraud, how much we caught. F1 is their harmonic mean — one number that punishes ignoring either. ROC-AUC measures how well the model ranks fraud above non-fraud across thresholds.

import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.metrics import f1_score, roc_auc_score

X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

mlflow.set_experiment("fraud-baseline")

with mlflow.start_run(run_name="rf-100-trees"):
    params = {"n_estimators": 100, "max_depth": 8, "random_state": 42}
    model = RandomForestClassifier(**params).fit(X_train, y_train)

    proba = model.predict_proba(X_test)[:, 1]
    preds = model.predict(X_test)

    mlflow.log_params(params)
    mlflow.log_metric("f1", f1_score(y_test, preds))
    mlflow.log_metric("roc_auc", roc_auc_score(y_test, proba))
    mlflow.sklearn.log_model(model, artifact_path="model")

Run mlflow ui in the same directory: every run is a row you can sort by F1, compare side by side, and reopen months later with the exact params intact.

TIP
Use a shared tracking server, not local files
By default MLflow writes to a local mlruns/ folder. For a team, point it at a shared tracking server and an artifact store via the tracking URI. The model-logging keyword has evolved across versions — if artifact_path warns as deprecated, switch to the name argument; behaviour is the same.

3 Track experiments with Weights & Biases

Weights & Biases (W&B) is a hosted alternative with the same idea — runs, configs, metrics, artifacts — but a stronger emphasis on live dashboards, sweeps, and team collaboration. The workflow: create an Artifact, add files to it, then log it to the run.

import wandb, joblib
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import f1_score

run = wandb.init(project="fraud-detection",
                 config={"learning_rate": 0.1, "n_estimators": 200})
cfg = run.config

model = GradientBoostingClassifier(
    learning_rate=cfg.learning_rate, n_estimators=cfg.n_estimators
).fit(X_train, y_train)

run.log({"f1": f1_score(y_test, model.predict(X_test))})

joblib.dump(model, "model.joblib")
artifact = wandb.Artifact("fraud-model", type="model")
artifact.add_file("model.joblib")
run.log_artifact(artifact)
run.finish()
Concern MLflow Weights & Biases
Hosting Self-host or managed Primarily managed (SaaS)
Strength Built-in registry, open standard Live dashboards, sweeps
Offline Works fully offline by default Cloud-first (offline mode exists)
Lock-in Low — open formats Higher — tied to the platform
NOTE
Pick one and standardise
The worst outcome is half the team on MLflow and half on W&B. Both solve the problem well. Choose by deployment model — self-hosted and open leans MLflow; a hosted dashboard with minimal ops leans W&B — then make it the default everywhere.

4 Version and register the model

A logged run is not yet a deployable artifact. The model registry is a versioned catalogue: each registered model has numbered versions, and you attach a moving label — an alias like production — to whichever version is live. (Older MLflow used named stages such as Staging/Production; those are deprecated, and aliases plus tags are now the recommended approach.)

You also want a model signature: a recorded schema of expected input columns and output type. It guards against training-serving skew — if production sends the wrong columns, it fails loudly instead of returning garbage.

import mlflow
from mlflow.models import infer_signature
from mlflow import MlflowClient

signature = infer_signature(X_train, model.predict(X_train))

with mlflow.start_run():
    info = mlflow.sklearn.log_model(
        model, artifact_path="model",
        signature=signature,
        registered_model_name="fraud-classifier",
    )

client = MlflowClient()
client.set_registered_model_alias(
    name="fraud-classifier", alias="production",
    version=info.registered_model_version,
)

# Serving loads by alias, never by a hard-coded version:
prod_model = mlflow.pyfunc.load_model("models:/fraud-classifier@production")
WARNING
Version the data, not just the model
A model version is only reproducible if you pin the data that produced it. Log a dataset hash or snapshot path as a run artifact. Otherwise "version 3" means nothing six months later when the training table has changed underneath you.

5 Serve the model — batch vs real-time

Inference is using the trained model to make predictions. Two delivery shapes:

  • Batch inference — score a large set of rows on a schedule. Optimises throughput (rows per hour) and cost. Good for "email tomorrow's churn-risk list."
  • Real-time inference — score one request synchronously over HTTP. Optimises latency (milliseconds per call). Required for "block this card swipe now."

Real-time endpoint with FastAPI, loading the registered model by alias:

# serve.py  — run with: uvicorn serve:app --host 0.0.0.0 --port 8000
from fastapi import FastAPI
from pydantic import BaseModel
import mlflow.pyfunc
import pandas as pd

app = FastAPI()
model = mlflow.pyfunc.load_model("models:/fraud-classifier@production")

class Transaction(BaseModel):
    amount: float
    merchant_risk: float
    hour_of_day: int

@app.post("/predict")
def predict(tx: Transaction):
    row = pd.DataFrame([tx.model_dump()])
    score = float(model.predict(row)[0])
    return {"fraud_score": score, "flag": score > 0.5}

Batch scoring is just a scheduled script:

# score_batch.py — run nightly via cron or an orchestrator
import pandas as pd
import mlflow.pyfunc

FEATURES = ["amount", "merchant_risk", "hour_of_day"]
model = mlflow.pyfunc.load_model("models:/fraud-classifier@production")

df = pd.read_parquet("data/transactions_2026-05-31.parquet")
df["fraud_score"] = model.predict(df[FEATURES])
df.to_parquet("data/scored_2026-05-31.parquet")
Dimension Batch Real-time
Optimises Throughput, cost Latency
Trigger Schedule Per request
Infra A job runner A live, scaled service
Use case Nightly risk list Block a card swipe
TIP
Default to batch
Real-time serving is more code, infrastructure, and failure modes. If the business can act on yesterday's scores, batch is cheaper and far simpler to operate. Pay for real-time only when the decision must happen inside a user-facing request.

6 Monitor for drift

Models do not crash when they go wrong — they quietly get less accurate as the world moves on. Two distinct things shift:

  • Data drift — the inputs change shape. Average transaction amount jumps after a pricing change; the model still runs but sees unfamiliar values.
  • Concept drift — the relationship between inputs and the answer changes. Fraudsters invent a new pattern, so the same features now mean something different.

You catch data drift fast — before you even have labels — with a statistical test. The Kolmogorov–Smirnov (KS) test compares two samples of a numeric feature and returns a p-value; a tiny p-value means the live distribution differs from the training one.

import numpy as np
from scipy.stats import ks_2samp

def drift_report(reference: np.ndarray, current: np.ndarray, alpha=0.01):
    stat, p_value = ks_2samp(reference, current)
    return {
        "ks_statistic": round(float(stat), 4),
        "p_value": round(float(p_value), 6),
        "drift_detected": bool(p_value < alpha),
    }

reference = X_train["mean radius"].to_numpy()
live = X_test["mean radius"].to_numpy() + 3.0   # simulate a shift
print(drift_report(reference, live))
# -> {'ks_statistic': ..., 'p_value': ..., 'drift_detected': True}

Concept drift needs ground truth — the eventual real outcome (did the flagged transaction turn out to be fraud?). Once labels arrive, recompute live F1 on a rolling window and alert when it falls below a threshold.

WARNING
Ground truth arrives late — design for it
For fraud, the real label can lag by weeks (chargebacks). You cannot measure concept drift in real time, so monitor data drift now as an early warning and concept drift later as confirmation. Build a feedback loop that joins predictions to outcomes when they finally land.
NOTE
Multiple features, one number
Population Stability Index (PSI) gives a single binned-feature drift score; KS is per-feature and assumption-light. Tools like Evidently bundle both into a dashboard, but a few lines of scipy is enough to start.

7 Automate retraining with guardrails

The payoff of everything above: a model that refreshes itself safely. Retraining means training a new version on newer data. The danger is promoting a worse model, so retraining must be gated by a guardrail metric — a candidate is promoted only if it beats a hard floor on a held-out set.

# retrain_if_needed.py — trains, evaluates, promotes only if it passes
from mlflow import MlflowClient

DRIFT_THRESHOLD = 0.2   # KS statistic
F1_GUARDRAIL = 0.85     # never promote below this

def should_retrain(drift_stat: float, live_f1: float) -> bool:
    return drift_stat > DRIFT_THRESHOLD or live_f1 < F1_GUARDRAIL

def promote(version: str):
    MlflowClient().set_registered_model_alias(
        "fraud-classifier", "production", version)

if should_retrain(drift_stat=0.27, live_f1=0.82):
    new_version = train_and_register()    # logs params, metrics, signature
    candidate_f1 = evaluate(new_version)  # on a fresh held-out set
    if candidate_f1 >= F1_GUARDRAIL:
        promote(new_version)              # alias flips atomically
    # else: keep current production model, page a human

Schedule it with cron (or an orchestrator like Airflow / Prefect for retries and lineage):

# /etc/cron.d/ml-retrain  — 02:00 daily
0 2 * * *  mluser  python /opt/ml/retrain_if_needed.py >> /var/log/retrain.log 2>&1

For CI/CD the same idea moves into your pipeline: on a merge or schedule, a job trains, runs the guardrail evaluation, and the alias flip is the deploy. Because serving loads models:/fraud-classifier@production, promotion needs no service redeploy — it picks up the new version on next load. (Express the pipeline in your CI tool's own config rather than hand-templating it here.)

TIP
Promotion is a flag flip, never a copy
Because the alias decouples "which version is live" from the serving code, rollback is instant: point the production alias back at the previous version. No rebuild, no redeploy. This one indirection removes most of the fear from shipping models.

Questions & Answers

Q: This is a lot of infrastructure for one model. Where do I start without over-engineering?
Start with two stages only: track (MLflow logging) and serve batch (a scheduled script loading the model). Those give you reproducibility and a deployed model with almost no infra. Add a registry, drift monitoring, and retraining once the model actually matters and stays live. MLOps is incremental — you do not need the full pipeline on day one.
Q: How often should I retrain — daily, weekly, monthly?
Don't retrain on a fixed clock; retrain on a signal. A daily cron that checks drift and guardrail metrics and only trains when they're breached beats blindly retraining every night (which wastes compute and can promote noise). Cadence depends on how fast your world moves: fraud and pricing shift in days; a demand forecast may be stable for months.
Q: Production accuracy dropped but the KS test shows no data drift. What happened?
That's the classic signature of concept drift: the inputs look the same, but the input-to-label relationship changed. Data-drift tests can't see it because they only look at feature distributions. You need ground-truth labels and a rolling live-metric monitor to catch it — which is exactly why the feedback loop in Step 6 matters.
Q: Should I serve with FastAPI myself, or use a managed serving platform?
For a handful of low-traffic models, the FastAPI service in Step 5 plus a container is perfectly adequate and fully under your control. Reach for a managed platform (cloud endpoints, KServe, BentoML, or the serving built into your tracking tool) when you need autoscaling, canary rollouts, GPU batching, or dozens of models.
Q: Does any of this apply if I'm only calling an LLM API?
Yes, in spirit. You won't train or register a model file, but you should still version prompts and the model ID, track which combination produced which outputs, monitor latency, cost, and quality, and set guardrails before swapping models. The five-stage lifecycle holds; "retraining" becomes "prompt and model upgrades." Lessons 8 and 9 dig into the hybrid case.

Key Takeaways

  1. A notebook result is not a product. The MLOps lifecycle — track, version, serve, monitor, retrain — turns a one-off score into a system people can rely on.
  2. Track everything, automatically. Logging params, metrics, and artifacts per run with MLflow or W&B makes results reproducible and comparable; an untracked good result is a lucky accident you can't repeat.
  3. The registry decouples "which model is live" from your code. A signature catches training-serving skew, and an alias makes promotion and rollback an instant flag flip rather than a redeploy.
  4. Choose serving by latency, default to batch. Batch is cheaper and simpler; pay for real-time only when the decision must happen inside a live request.
  5. Models decay silently — monitor inputs and outputs. Catch data drift now with a cheap statistical test, and concept drift later with ground-truth labels, before accuracy quietly collapses.
  6. Automate retraining behind guardrails. Retrain on a signal, not a clock, and never promote a candidate that fails a hard metric floor — so the pipeline can never make the model worse.

Next Steps: Lesson 8: When to Use Classical ML vs GenAI