Production Deployment
Learning Outcomes
- Serve a model behind an OpenAI-compatible HTTP API using Ollama, vLLM, or TGI
- Containerise an inference service with Docker for reproducible deployments
- Configure a reverse proxy to load-balance requests across multiple model instances
- Instrument a deployment with metrics, logging, and health checks
- Put access controls on an endpoint and model the cost of self-hosting versus cloud inference
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | From laptop experiment to a service others call |
| Pick a server | 8 min | Ollama server mode vs vLLM vs TGI |
| Serve an API | 12 min | Bring up an OpenAI-compatible endpoint |
| Containerise | 10 min | Docker and docker-compose for reproducibility |
| Load balance | 9 min | Nginx in front of multiple replicas |
| Observe + control access | 9 min | Metrics, health checks, auth, rate limits |
| Cost model | 6 min | Self-host vs cloud maths |
| Wrap-up | 3 min | Checklist and next steps |
Before You Begin
Pre-work:
- Complete Lesson 9 — Performance Tuning so you know your tokens/sec and memory headroom
- Have a model running locally from Lesson 3 — Ollama
- Skim Lesson 6 — the OpenAI-compatible API is the contract we build on here
- Keep the Troubleshooting supplemental open
Shopping List:
- Docker Engine installed and running (
docker --version) - A GPU for vLLM/TGI, or any machine for Ollama CPU/Metal serving
curland Python 3 with theopenaipackage for testing
"Serving" means exposing a model over HTTP so other processes — your app, a teammate's laptop, a CI job — can send a prompt and get tokens back. Three engines cover almost every case:
| Engine | Best for | Hardware | API shape |
|---|---|---|---|
| Ollama (server mode) | Small teams, mixed hardware, fastest setup | CPU, Metal, NVIDIA, AMD | Native /api/* + OpenAI /v1/* |
| vLLM | High-throughput GPU serving, concurrency | NVIDIA/AMD GPU | OpenAI-compatible /v1/* |
| TGI | Hugging Face models, production GPU serving | NVIDIA/AMD/Intel GPU | /generate + OpenAI Messages API |
The key idea: all three can speak the OpenAI Chat Completions format, so client code stays identical no matter which engine sits behind it. Prototype on Ollama and swap in vLLM for production without rewriting the calling app. Ollama's OpenAI-compatible layer covers /v1/chat/completions, /v1/completions, /v1/embeddings, and /v1/models.
Ollama server mode. Ollama already runs a server; you just need it to bind beyond localhost. It binds 127.0.0.1:11434 by default — change the bind address with OLLAMA_HOST.
On macOS and Linux, export the variable and start the server (on Linux under systemd, set it with systemctl edit ollama.service):
export OLLAMA_HOST=0.0.0.0:11434
ollama serve
ollama pull llama3.1:8b # in another terminal
Quit the Ollama tray app, set OLLAMA_HOST=0.0.0.0:11434 under "Edit environment variables for your account," then relaunch Ollama — it reads the variable on start.
Smoke-test it with curl http://localhost:11434/v1/chat/completions -d '{"model":"llama3.1:8b","messages":[{"role":"user","content":"say ready"}]}' — the exact OpenAI shape your apps will use.
vLLM (GPU). One command launches an OpenAI-compatible server on port 8000 by default. Use --tensor-parallel-size to split a large model across multiple GPUs.
pip install vllm
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--port 8000 --tensor-parallel-size 1 --max-model-len 8192
TGI (GPU, Docker). Hugging Face ships TGI as a container. Map host port 8080 to the container's port 80 and pass a Hugging Face model ID.
docker run --gpus all --shm-size 1g -p 8080:80 -v $PWD/data:/data \
ghcr.io/huggingface/text-generation-inference --model-id teknium/OpenHermes-2.5-Mistral-7B
Because the API is OpenAI-shaped, the official openai SDK works against any of the three engines — point base_url at your server (use a non-empty key; local engines ignore it). The switch from a cloud provider to your own box is literally one line:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="local-not-checked")
resp = client.chat.completions.create(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Summarise RAG in one sentence."}],
stream=True,
)
for chunk in resp:
print(chunk.choices[0].delta.content or "", end="", flush=True)
0.0.0.0 exposes the endpoint to every machine that can route to this host — on shared Wi-Fi that is the whole café. Only do this on a trusted network, or put it behind the firewall and reverse proxy (Step 4) before anyone outside touches it.stream=True for any interactive surface.A bare ollama serve on someone's laptop is not a deployment — it dies when the laptop sleeps and nobody can reproduce its state. Docker fixes both. Use the official image and a docker-compose.yml so the whole stack comes up with one command.
# docker-compose.yml
services:
ollama:
image: ollama/ollama:latest
ports:
- "11434:11434"
volumes:
- ollama-models:/root/.ollama # persist downloaded weights
environment:
- OLLAMA_HOST=0.0.0.0:11434
- OLLAMA_NUM_PARALLEL=4 # concurrent requests per model
- OLLAMA_MAX_LOADED_MODELS=2 # models resident at once
restart: unless-stopped
# NVIDIA GPUs: add a deploy.resources block reserving the nvidia driver
volumes:
ollama-models:
OLLAMA_MAX_LOADED_MODELS defaults to 3 times the GPU count (3 for CPU), and OLLAMA_NUM_PARALLEL defaults to 1 parallel request per model. Tune both to your Lesson 9 numbers, then docker compose up -d and docker compose exec ollama ollama pull llama3.1:8b to warm a model so the first real request is not slow.
docker compose down deletes your weights and the next start re-downloads tens of gigabytes. The ollama-models volume keeps them between restarts.A single instance is a single point of contention. When concurrent requests exceed what one instance handles smoothly, run several and put a reverse proxy in front. Nginx round-robins (or least_conn-balances) requests across replicas.
# nginx.conf
upstream llm_backends {
least_conn; # route to the instance with fewest active conns
server ollama-1:11434 max_fails=3 fail_timeout=30s;
server ollama-2:11434 max_fails=3 fail_timeout=30s;
}
server {
listen 80;
location / {
proxy_pass http://llm_backends;
proxy_read_timeout 300s; # long generations must not time out
proxy_buffering off; # required so streamed tokens flow immediately
}
}
Run the model containers (ollama-1, ollama-2) and the proxy on one Docker network:
docker network create llm-net
docker run -d --name ollama-1 --network llm-net -v ollama-models:/root/.ollama ollama/ollama
# repeat the line above for ollama-2, then start the load balancer:
docker run -d --name lb --network llm-net -p 80:80 \
-v "$PWD/nginx.conf:/etc/nginx/conf.d/default.conf:ro" nginx:alpine
proxy_buffering on, Nginx holds the whole response before forwarding, so streamed tokens arrive in one lump at the end. proxy_buffering off is non-negotiable for streamed LLM responses.You cannot operate what you cannot see. Three layers matter: alive (health), doing OK (metrics), and what happened (logs). Add a healthcheck to the compose service that curls http://localhost:11434/api/version so Docker restarts a wedged instance automatically.
vLLM and TGI both expose a Prometheus /metrics endpoint out of the box (scrape metrics_path: /metrics). Track the numbers that actually predict a bad user experience:
| Metric | Why it matters |
|---|---|
| Time to first token (TTFT) | Perceived responsiveness; rises under queueing |
| Requests in queue / running | Early warning of saturation |
| GPU memory used | Distance from OOM crashes — see Troubleshooting |
| Error rate (5xx) | Crashes, OOMs, timeouts |
For logs, use structured logging shipped to a central store (Loki, CloudWatch) so you can search across replicas rather than tailing each container by hand.
Now control access. Local engines ship with no auth — anyone who reaches the port runs your GPU for free and reads every prompt. Extend the Step 4 server block with TLS, a bearer-token check, and a per-client rate limit:
limit_req_zone $binary_remote_addr zone=llm:10m rate=10r/s; # ~10 req/s per IP
# inside server { listen 443 ssl; ... } location / { ... }
if ($http_authorization != "Bearer CHANGE_ME") { return 401; } # reject others
limit_req zone=llm burst=20 nodelay;
Access checklist for a self-hosted endpoint:
| Control | What it stops |
|---|---|
| TLS (HTTPS) | Prompts and outputs sniffed in transit |
| API key / token | Unauthorised use of your compute |
| Rate limiting | A single client exhausting capacity or running up costs |
| Private network + firewall | Public internet discovering an open port |
Input size caps (max_model_len) |
Memory exhaustion via giant prompts |
The decision is rarely about raw price per token; it is about utilisation. Self-hosting is a fixed cost (you pay for the GPU busy or idle); cloud APIs are variable (you pay per token). The crossover depends on how busy the hardware stays. Estimate it from your Lesson 9 numbers:
monthly_tokens = 500_000_000 # tokens generated per month
self_host = 2000 / 24 + 30 # amortised GPU + power per month
cloud = (monthly_tokens / 1_000_000) * 0.50 # per 1M tokens, your provider's rate
print("Self-host wins" if self_host < cloud else "Cloud wins")
| Factor | Favours self-hosting | Favours cloud |
|---|---|---|
| Volume / utilisation | High, steady, GPU busy most hours | Low, spiky, occasional bursts |
| Data sensitivity | Must never leave premises | No hard requirement |
| Ops appetite | Can run and patch servers | Want zero ops |
| Quality ceiling | Open weights good enough | Need the very best model |
Questions & Answers
--tensor-parallel-size for big models, or more replicas for many small requests) pay off.0.0.0.0 on a routable IP is a bigger exposure than a reputable cloud API under a data-processing agreement. Privacy comes from TLS, auth, and network isolation — not from the word "local" alone.--max-model-len (vLLM) deliberately, not to the model maximum: a smaller context uses less GPU memory per request and lets you batch more. See the Troubleshooting guide.Key Takeaways
- Standardise on the OpenAI API shape — Ollama, vLLM, and TGI all speak
/v1/chat/completions, so client code is portable and you can swap engines without rewrites. - Match the engine to the constraint — Ollama for easy/mixed hardware, vLLM for high-throughput GPU concurrency, TGI for Hugging Face-native GPU serving.
- Containerise with a persistent volume — Docker makes deployments reproducible; the named model volume saves you re-downloading tens of gigabytes.
- Scale and observe deliberately — load-balance with Nginx (buffering off for streaming), and watch time-to-first-token and queue depth, not just throughput.
- Access control is your job now — add TLS, API keys, rate limits, and network isolation before any endpoint is reachable; "local" alone is not private.
- Cost is about utilisation — self-hosting wins on high, steady volume; bursty workloads usually favour cloud or on-demand cloud GPUs.
Next Steps: Back to all Local LLMs lessons