Production Deployment

60 min advanced Lesson 10

Learning Outcomes

  • Serve a model behind an OpenAI-compatible HTTP API using Ollama, vLLM, or TGI
  • Containerise an inference service with Docker for reproducible deployments
  • Configure a reverse proxy to load-balance requests across multiple model instances
  • Instrument a deployment with metrics, logging, and health checks
  • Put access controls on an endpoint and model the cost of self-hosting versus cloud inference

Lesson Plan

Segment Duration Topic
Intro 3 min From laptop experiment to a service others call
Pick a server 8 min Ollama server mode vs vLLM vs TGI
Serve an API 12 min Bring up an OpenAI-compatible endpoint
Containerise 10 min Docker and docker-compose for reproducibility
Load balance 9 min Nginx in front of multiple replicas
Observe + control access 9 min Metrics, health checks, auth, rate limits
Cost model 6 min Self-host vs cloud maths
Wrap-up 3 min Checklist and next steps

Before You Begin

Pre-work:

Shopping List:

  • Docker Engine installed and running (docker --version)
  • A GPU for vLLM/TGI, or any machine for Ollama CPU/Metal serving
  • curl and Python 3 with the openai package for testing

1 Choosing a Serving Engine

"Serving" means exposing a model over HTTP so other processes — your app, a teammate's laptop, a CI job — can send a prompt and get tokens back. Three engines cover almost every case:

Engine Best for Hardware API shape
Ollama (server mode) Small teams, mixed hardware, fastest setup CPU, Metal, NVIDIA, AMD Native /api/* + OpenAI /v1/*
vLLM High-throughput GPU serving, concurrency NVIDIA/AMD GPU OpenAI-compatible /v1/*
TGI Hugging Face models, production GPU serving NVIDIA/AMD/Intel GPU /generate + OpenAI Messages API

The key idea: all three can speak the OpenAI Chat Completions format, so client code stays identical no matter which engine sits behind it. Prototype on Ollama and swap in vLLM for production without rewriting the calling app. Ollama's OpenAI-compatible layer covers /v1/chat/completions, /v1/completions, /v1/embeddings, and /v1/models.

NOTE
Why throughput differs
vLLM and TGI use continuous batching — they pack many in-flight requests into one GPU pass instead of serving them one at a time, often several times the tokens/sec of a naive server. Ollama is simpler and lighter; vLLM/TGI shine when many users hit the endpoint at once.

2 Standing Up an OpenAI-Compatible Endpoint

Ollama server mode. Ollama already runs a server; you just need it to bind beyond localhost. It binds 127.0.0.1:11434 by default — change the bind address with OLLAMA_HOST.

On macOS and Linux, export the variable and start the server (on Linux under systemd, set it with systemctl edit ollama.service):

export OLLAMA_HOST=0.0.0.0:11434
ollama serve
ollama pull llama3.1:8b   # in another terminal

Quit the Ollama tray app, set OLLAMA_HOST=0.0.0.0:11434 under "Edit environment variables for your account," then relaunch Ollama — it reads the variable on start.

Smoke-test it with curl http://localhost:11434/v1/chat/completions -d '{"model":"llama3.1:8b","messages":[{"role":"user","content":"say ready"}]}' — the exact OpenAI shape your apps will use.

vLLM (GPU). One command launches an OpenAI-compatible server on port 8000 by default. Use --tensor-parallel-size to split a large model across multiple GPUs.

pip install vllm
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --port 8000 --tensor-parallel-size 1 --max-model-len 8192

TGI (GPU, Docker). Hugging Face ships TGI as a container. Map host port 8080 to the container's port 80 and pass a Hugging Face model ID.

docker run --gpus all --shm-size 1g -p 8080:80 -v $PWD/data:/data \
  ghcr.io/huggingface/text-generation-inference --model-id teknium/OpenHermes-2.5-Mistral-7B

Because the API is OpenAI-shaped, the official openai SDK works against any of the three engines — point base_url at your server (use a non-empty key; local engines ignore it). The switch from a cloud provider to your own box is literally one line:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="local-not-checked")
resp = client.chat.completions.create(
    model="llama3.1:8b",
    messages=[{"role": "user", "content": "Summarise RAG in one sentence."}],
    stream=True,
)
for chunk in resp:
    print(chunk.choices[0].delta.content or "", end="", flush=True)
WARNING
0.0.0.0 means the network can reach you
Binding to 0.0.0.0 exposes the endpoint to every machine that can route to this host — on shared Wi-Fi that is the whole café. Only do this on a trusted network, or put it behind the firewall and reverse proxy (Step 4) before anyone outside touches it.
TIP
Always stream interactive UIs
Streaming does not speed up generation, but the user sees the first token in a fraction of a second instead of waiting for the whole answer. Set stream=True for any interactive surface.

3 Containerising for Reproducible Deployments

A bare ollama serve on someone's laptop is not a deployment — it dies when the laptop sleeps and nobody can reproduce its state. Docker fixes both. Use the official image and a docker-compose.yml so the whole stack comes up with one command.

# docker-compose.yml
services:
  ollama:
    image: ollama/ollama:latest
    ports:
      - "11434:11434"
    volumes:
      - ollama-models:/root/.ollama   # persist downloaded weights
    environment:
      - OLLAMA_HOST=0.0.0.0:11434
      - OLLAMA_NUM_PARALLEL=4          # concurrent requests per model
      - OLLAMA_MAX_LOADED_MODELS=2     # models resident at once
    restart: unless-stopped
    # NVIDIA GPUs: add a deploy.resources block reserving the nvidia driver
volumes:
  ollama-models:

OLLAMA_MAX_LOADED_MODELS defaults to 3 times the GPU count (3 for CPU), and OLLAMA_NUM_PARALLEL defaults to 1 parallel request per model. Tune both to your Lesson 9 numbers, then docker compose up -d and docker compose exec ollama ollama pull llama3.1:8b to warm a model so the first real request is not slow.

WARNING
Persist your model volume
Without the named volume, every docker compose down deletes your weights and the next start re-downloads tens of gigabytes. The ollama-models volume keeps them between restarts.
NOTE
GPU passthrough needs the toolkit
NVIDIA GPUs need the NVIDIA Container Toolkit installed before Docker can mount the GPU into the container. On Apple Silicon, Docker cannot pass through the Metal GPU — run Ollama natively and containerise only the surrounding services.

4 Load Balancing Across Replicas

A single instance is a single point of contention. When concurrent requests exceed what one instance handles smoothly, run several and put a reverse proxy in front. Nginx round-robins (or least_conn-balances) requests across replicas.

# nginx.conf
upstream llm_backends {
    least_conn;                  # route to the instance with fewest active conns
    server ollama-1:11434 max_fails=3 fail_timeout=30s;
    server ollama-2:11434 max_fails=3 fail_timeout=30s;
}
server {
    listen 80;
    location / {
        proxy_pass http://llm_backends;
        proxy_read_timeout 300s; # long generations must not time out
        proxy_buffering off;     # required so streamed tokens flow immediately
    }
}

Run the model containers (ollama-1, ollama-2) and the proxy on one Docker network:

docker network create llm-net
docker run -d --name ollama-1 --network llm-net -v ollama-models:/root/.ollama ollama/ollama
# repeat the line above for ollama-2, then start the load balancer:
docker run -d --name lb --network llm-net -p 80:80 \
  -v "$PWD/nginx.conf:/etc/nginx/conf.d/default.conf:ro" nginx:alpine
WARNING
Turn off proxy buffering for streaming
With proxy_buffering on, Nginx holds the whole response before forwarding, so streamed tokens arrive in one lump at the end. proxy_buffering off is non-negotiable for streamed LLM responses.
TIP
Scale up before scaling out
vLLM and TGI already batch many requests inside one instance, so a single well-configured GPU instance often beats several under-batched Ollama copies. Add replicas only after one instance is saturated.

5 Observability and Endpoint Access Control

You cannot operate what you cannot see. Three layers matter: alive (health), doing OK (metrics), and what happened (logs). Add a healthcheck to the compose service that curls http://localhost:11434/api/version so Docker restarts a wedged instance automatically.

vLLM and TGI both expose a Prometheus /metrics endpoint out of the box (scrape metrics_path: /metrics). Track the numbers that actually predict a bad user experience:

Metric Why it matters
Time to first token (TTFT) Perceived responsiveness; rises under queueing
Requests in queue / running Early warning of saturation
GPU memory used Distance from OOM crashes — see Troubleshooting
Error rate (5xx) Crashes, OOMs, timeouts

For logs, use structured logging shipped to a central store (Loki, CloudWatch) so you can search across replicas rather than tailing each container by hand.

NOTE
Watch TTFT, not just throughput
A server can post great average tokens/sec while individual users wait seconds for the first token because requests are queued. Time-to-first-token under load is the metric that most closely tracks how slow your service feels.

Now control access. Local engines ship with no auth — anyone who reaches the port runs your GPU for free and reads every prompt. Extend the Step 4 server block with TLS, a bearer-token check, and a per-client rate limit:

limit_req_zone $binary_remote_addr zone=llm:10m rate=10r/s;  # ~10 req/s per IP

# inside server { listen 443 ssl; ... } location / { ... }
if ($http_authorization != "Bearer CHANGE_ME") { return 401; }  # reject others
limit_req zone=llm burst=20 nodelay;

Access checklist for a self-hosted endpoint:

Control What it stops
TLS (HTTPS) Prompts and outputs sniffed in transit
API key / token Unauthorised use of your compute
Rate limiting A single client exhausting capacity or running up costs
Private network + firewall Public internet discovering an open port
Input size caps (max_model_len) Memory exhaustion via giant prompts
WARNING
Self-hosting does not mean unaudited
Running models locally is great for privacy — data never leaves your hardware — but an unauthenticated endpoint on a public IP is a worse leak than any cloud API. Treat the endpoint like any other production service: TLS, auth, and least-privilege network access. See the Rules & Best Practices supplemental.

6 Cost Modelling — Self-Host vs Cloud

The decision is rarely about raw price per token; it is about utilisation. Self-hosting is a fixed cost (you pay for the GPU busy or idle); cloud APIs are variable (you pay per token). The crossover depends on how busy the hardware stays. Estimate it from your Lesson 9 numbers:

monthly_tokens = 500_000_000               # tokens generated per month
self_host = 2000 / 24 + 30                  # amortised GPU + power per month
cloud = (monthly_tokens / 1_000_000) * 0.50 # per 1M tokens, your provider's rate
print("Self-host wins" if self_host < cloud else "Cloud wins")
Factor Favours self-hosting Favours cloud
Volume / utilisation High, steady, GPU busy most hours Low, spiky, occasional bursts
Data sensitivity Must never leave premises No hard requirement
Ops appetite Can run and patch servers Want zero ops
Quality ceiling Open weights good enough Need the very best model
TIP
Idle GPUs are the hidden cost
A GPU idle 90% of the day still costs its full amortised price. For bursty workloads, cloud APIs — or a cloud GPU spun up only when needed (Lesson 2) — often win despite a higher per-token rate. Plug in your provider's current rate: the calculation is durable, the figures are not.

Questions & Answers

Q: My single GPU is at 100% but requests still queue. More GPUs, or different software?
Switch engines before buying hardware. If you are on Ollama, move to vLLM or TGI — their continuous batching can multiply throughput on the same GPU. Only after a properly batched single instance is saturated does adding GPUs (--tensor-parallel-size for big models, or more replicas for many small requests) pay off.
Q: How do I do zero-downtime model upgrades?
Run at least two replicas. Drain one from the upstream, upgrade it, warm it with a real prompt, return it, then repeat for the other. Because the OpenAI API contract is stable, clients never notice. Always warm before returning — a cold model's first request can take many seconds while weights load.
Q: Is a local model behind my own API really more private than a cloud provider?
Only if you secure it. Data staying on your hardware is the win, but an unauthenticated endpoint bound to 0.0.0.0 on a routable IP is a bigger exposure than a reputable cloud API under a data-processing agreement. Privacy comes from TLS, auth, and network isolation — not from the word "local" alone.
Q: What happens when a prompt exceeds the model's context length in production?
The server returns a 4xx error rather than silently truncating — correct but ugly. Cap input size at the proxy and validate length client-side. Set --max-model-len (vLLM) deliberately, not to the model maximum: a smaller context uses less GPU memory per request and lets you batch more. See the Troubleshooting guide.

Key Takeaways

  1. Standardise on the OpenAI API shape — Ollama, vLLM, and TGI all speak /v1/chat/completions, so client code is portable and you can swap engines without rewrites.
  2. Match the engine to the constraint — Ollama for easy/mixed hardware, vLLM for high-throughput GPU concurrency, TGI for Hugging Face-native GPU serving.
  3. Containerise with a persistent volume — Docker makes deployments reproducible; the named model volume saves you re-downloading tens of gigabytes.
  4. Scale and observe deliberately — load-balance with Nginx (buffering off for streaming), and watch time-to-first-token and queue depth, not just throughput.
  5. Access control is your job now — add TLS, API keys, rate limits, and network isolation before any endpoint is reachable; "local" alone is not private.
  6. Cost is about utilisation — self-hosting wins on high, steady volume; bursty workloads usually favour cloud or on-demand cloud GPUs.

Next Steps: Back to all Local LLMs lessons