Connecting Local Models to Tools
Learning Outcomes
- Explain why the OpenAI-compatible API is the universal adapter for local models
- Deploy Open WebUI in Docker as a ChatGPT-style front end for Ollama
- Configure LM Studio as a local OpenAI-compatible inference server
- Wire a local model into VS Code as a coding assistant with Continue
- Point any OpenAI SDK or tool at a local endpoint by overriding the base URL
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why a running model is not yet a useful tool |
| Explain | 6 min | The OpenAI-compatible API as universal adapter |
| Demo | 9 min | Open WebUI in Docker |
| Demo | 8 min | LM Studio as a local server |
| Demo | 9 min | A VS Code assistant with Continue |
| Demo | 6 min | Driving the endpoint from code |
| Wrap-up | 4 min | Trade-offs and takeaways |
Before You Begin
Pre-work:
- Complete Lesson 3: Ollama — The Easiest Way to Start — Ollama installed, a model pulled
- Complete Lesson 5: Running Models from the CLI
- Have a 7B-8B class model handy, e.g.
llama3.1:8borqwen2.5-coder:7b
Shopping List:
- Ollama running (
ollama --version) and Docker (docker --version) for the Open WebUI demo - VS Code, plus Python 3.9+ with
pip - Enough free RAM/VRAM to keep one model resident (see Hardware Reference)
The most important fact in this lesson: Ollama already exposes an OpenAI-compatible HTTP API at http://localhost:11434/v1/ when running, with endpoints mirroring OpenAI's (/v1/chat/completions, /v1/models, /v1/embeddings).
That matters because hundreds of tools — chat UIs, IDE extensions, agent frameworks, SDKs — already speak the OpenAI wire format. Any of them connects to your local model by changing two things: the base URL and a throwaway API key:
curl -X POST http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "llama3.1:8b", "messages": [{ "role": "user", "content": "Say this is a test" }] }'
You get back a JSON object shaped exactly like OpenAI's — choices[0].message.content holds the reply. Every integration below is the same move: point the tool's base URL at your machine.
model field must match a model Ollama has — run ollama list for exact tags. Asking for llama3.1 when you pulled llama3.1:8b returns an error.A browser chat UI is what makes a local model usable day to day. Open WebUI is the de-facto standard: a self-hosted interface that runs fully offline, supports Ollama and any OpenAI-compatible API, and adds history, model switching, and accounts. The cleanest install is Docker; if Ollama is on the same host, use host.docker.internal:
docker run -d -p 3000:8080 \
--add-host=host.docker.internal:host-gateway \
-v open-webui:/app/backend/data \
--name open-webui \
--restart always \
ghcr.io/open-webui/open-webui:main
The mapping -p 3000:8080 exposes the container's port 8080 on host port 3000, so you browse to http://localhost:3000; the command works the same on macOS, Linux, and Windows with Docker Desktop. Two alternatives: image tag :ollama bundles Ollama in one container, and pip install open-webui then open-webui serve skips Docker (port 8080).
On first load you create a local admin account (stored in the volume); the model dropdown auto-populates from ollama list. The --add-host flag lets the container reach Ollama on port 11434 — on Linux, use --network=host with -e OLLAMA_BASE_URL=http://127.0.0.1:11434 instead.
docker exec open-webui curl -s http://host.docker.internal:11434/api/tags — if that fails, it is networking, not Open WebUI.LM Studio is a desktop app for browsing, downloading, and running models. Where Ollama is CLI-first, it excels at exploring repositories and experimenting with quantization levels (the technique that shrinks weights to lower precision — Q4, Q5, Q8 — for memory savings). It exposes its own OpenAI-compatible server, the same pattern as Ollama.
To enable it: download a model from the Discover tab, open the Developer tab, load the model, and toggle Start. It listens on port 1234 by default, with OpenAI-compatible routes under /v1. Test it like Ollama in Step 1, on a different port:
curl http://localhost:1234/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "local-model", "messages": [{ "role": "user", "content": "Say hi" }] }'
The payoff for developers: an in-editor assistant — chat, edit, autocomplete — backed by a local model, no code leaving your machine. Continue is an open-source VS Code (and JetBrains) extension supporting Ollama as a provider. Install it and edit config.yaml with two models — one for chat, one for completion:
name: Local Assistant
version: 0.0.1
schema: v1
models:
- name: Llama 3.1 8B Chat
provider: ollama
model: llama3.1:8b
apiBase: http://localhost:11434
roles:
- chat
- edit
- name: Qwen Coder Autocomplete
provider: ollama
model: qwen2.5-coder:7b
apiBase: http://localhost:11434
roles:
- autocomplete
The provider, model, and apiBase fields are the documented way to connect Continue to Ollama; capabilities like tool use are auto-detected. Pull both models, reload VS Code, then highlight a function and ask Continue to refactor it.
autocomplete and a larger model (7B-8B) for chat. Without a GPU, autocomplete can lag by seconds; if CPU-only, drop the autocomplete role.The integrations above are just clients. With the base-URL trick, you can wire a local model into your own scripts with the OpenAI SDK (pip install openai):
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1/",
api_key="ollama", # required by the SDK, ignored by Ollama
)
resp = client.chat.completions.create(
model="llama3.1:8b",
messages=[{"role": "user", "content": "One-liner to find files over 100MB?"}],
)
print(resp.choices[0].message.content)
The placeholder api_key="ollama" is exactly what Ollama's docs prescribe. Streaming works as in the cloud API (stream=True), and the same code points at LM Studio by switching the base URL to port 1234.
ollama show llama3.1:8b.By default Ollama binds to localhost, reachable only from your own machine — keep it that way unless you have a reason not to. If another device on your LAN needs access, change the bind address with OLLAMA_HOST:
OLLAMA_HOST=0.0.0.0:11434 ollama serve # bind to all interfaces
Do this only behind a trusted network. Understand the exposure first:
| Scenario | Exposure |
|---|---|
localhost only (default) |
None beyond your machine — keep for solo dev |
LAN bind (0.0.0.0) |
Anyone on the network can prompt your model — firewall it |
| Public internet | Anyone can run inference on your GPU — never expose |
Critically, Ollama's OpenAI layer ignores the API key — it accepts any string, so there is no built-in authentication. Before any network exposure, put a reverse proxy (Caddy or nginx) enforcing a bearer token in front of 127.0.0.1:11434, or just expose Open WebUI (port 3000) and keep Ollama on localhost.
Questions & Answers
--add-host=host.docker.internal:host-gateway (or --network=host on Linux) and that Ollama is running. Curl the host's /api/tags from inside the container to confirm.Key Takeaways
- The OpenAI-compatible API is the universal adapter — Ollama serves it at
localhost:11434/v1; any OpenAI-aware tool connects by changing the base URL and a throwaway key. - Open WebUI gives you a private ChatGPT — one Docker command on port 3000, auto-detecting your Ollama models, with accounts and history.
- LM Studio is the GUI for model hunting — its own OpenAI server on port 1234, best for browsing models and tuning quantization.
- Continue turns VS Code into a local copilot —
provider: ollamaandapiBaseinconfig.yaml, splitting a small completion model from a larger chat model. - Your own code is just another client — the OpenAI SDK with an overridden
base_urlmigrates any existing OpenAI app to local. - Local means localhost — keep it that way — there is no built-in auth, so never expose the raw endpoint without a proxy in front.
Next Steps: Lesson 7: Fine-Tuning Basics