Connecting Local Models to Tools

45 min intermediate Lesson 6

Learning Outcomes

  • Explain why the OpenAI-compatible API is the universal adapter for local models
  • Deploy Open WebUI in Docker as a ChatGPT-style front end for Ollama
  • Configure LM Studio as a local OpenAI-compatible inference server
  • Wire a local model into VS Code as a coding assistant with Continue
  • Point any OpenAI SDK or tool at a local endpoint by overriding the base URL

Lesson Plan

Segment Duration Topic
Intro 3 min Why a running model is not yet a useful tool
Explain 6 min The OpenAI-compatible API as universal adapter
Demo 9 min Open WebUI in Docker
Demo 8 min LM Studio as a local server
Demo 9 min A VS Code assistant with Continue
Demo 6 min Driving the endpoint from code
Wrap-up 4 min Trade-offs and takeaways

Before You Begin

Pre-work:

Shopping List:

  • Ollama running (ollama --version) and Docker (docker --version) for the Open WebUI demo
  • VS Code, plus Python 3.9+ with pip
  • Enough free RAM/VRAM to keep one model resident (see Hardware Reference)

1 The OpenAI-Compatible API — One Adapter to Rule Them All

The most important fact in this lesson: Ollama already exposes an OpenAI-compatible HTTP API at http://localhost:11434/v1/ when running, with endpoints mirroring OpenAI's (/v1/chat/completions, /v1/models, /v1/embeddings).

That matters because hundreds of tools — chat UIs, IDE extensions, agent frameworks, SDKs — already speak the OpenAI wire format. Any of them connects to your local model by changing two things: the base URL and a throwaway API key:

curl -X POST http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{ "model": "llama3.1:8b", "messages": [{ "role": "user", "content": "Say this is a test" }] }'

You get back a JSON object shaped exactly like OpenAI's — choices[0].message.content holds the reply. Every integration below is the same move: point the tool's base URL at your machine.

WARNING
The model name must match exactly
The model field must match a model Ollama has — run ollama list for exact tags. Asking for llama3.1 when you pulled llama3.1:8b returns an error.

2 Open WebUI — A ChatGPT-Style Front End

A browser chat UI is what makes a local model usable day to day. Open WebUI is the de-facto standard: a self-hosted interface that runs fully offline, supports Ollama and any OpenAI-compatible API, and adds history, model switching, and accounts. The cleanest install is Docker; if Ollama is on the same host, use host.docker.internal:

docker run -d -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui \
  --restart always \
  ghcr.io/open-webui/open-webui:main

The mapping -p 3000:8080 exposes the container's port 8080 on host port 3000, so you browse to http://localhost:3000; the command works the same on macOS, Linux, and Windows with Docker Desktop. Two alternatives: image tag :ollama bundles Ollama in one container, and pip install open-webui then open-webui serve skips Docker (port 8080).

On first load you create a local admin account (stored in the volume); the model dropdown auto-populates from ollama list. The --add-host flag lets the container reach Ollama on port 11434 — on Linux, use --network=host with -e OLLAMA_BASE_URL=http://127.0.0.1:11434 instead.

TIP
Empty dropdown? Check networking
It means the container cannot see Ollama. Run docker exec open-webui curl -s http://host.docker.internal:11434/api/tags — if that fails, it is networking, not Open WebUI.

3 LM Studio — A GUI Server for Model Hunters

LM Studio is a desktop app for browsing, downloading, and running models. Where Ollama is CLI-first, it excels at exploring repositories and experimenting with quantization levels (the technique that shrinks weights to lower precision — Q4, Q5, Q8 — for memory savings). It exposes its own OpenAI-compatible server, the same pattern as Ollama.

To enable it: download a model from the Discover tab, open the Developer tab, load the model, and toggle Start. It listens on port 1234 by default, with OpenAI-compatible routes under /v1. Test it like Ollama in Step 1, on a different port:

curl http://localhost:1234/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{ "model": "local-model", "messages": [{ "role": "user", "content": "Say hi" }] }'
NOTE
Ollama vs LM Studio, and a VRAM warning
Ollama wins for scripting and reproducible Modelfiles; LM Studio for model shopping and GPU-offload sliders. Both speak OpenAI, so you are never locked in — but running both at once means two runtimes competing for memory. On a 16GB GPU, an 8B model in each will likely OOM.

4 A VS Code Coding Assistant with Continue

The payoff for developers: an in-editor assistant — chat, edit, autocomplete — backed by a local model, no code leaving your machine. Continue is an open-source VS Code (and JetBrains) extension supporting Ollama as a provider. Install it and edit config.yaml with two models — one for chat, one for completion:

name: Local Assistant
version: 0.0.1
schema: v1

models:
  - name: Llama 3.1 8B Chat
    provider: ollama
    model: llama3.1:8b
    apiBase: http://localhost:11434
    roles:
      - chat
      - edit
  - name: Qwen Coder Autocomplete
    provider: ollama
    model: qwen2.5-coder:7b
    apiBase: http://localhost:11434
    roles:
      - autocomplete

The provider, model, and apiBase fields are the documented way to connect Continue to Ollama; capabilities like tool use are auto-detected. Pull both models, reload VS Code, then highlight a function and ask Continue to refactor it.

TIP
Split chat and autocomplete models
Autocomplete must be fast — use a small code model (1B-3B) for autocomplete and a larger model (7B-8B) for chat. Without a GPU, autocomplete can lag by seconds; if CPU-only, drop the autocomplete role.

5 Driving the Endpoint from Your Own Code

The integrations above are just clients. With the base-URL trick, you can wire a local model into your own scripts with the OpenAI SDK (pip install openai):

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1/",
    api_key="ollama",  # required by the SDK, ignored by Ollama
)
resp = client.chat.completions.create(
    model="llama3.1:8b",
    messages=[{"role": "user", "content": "One-liner to find files over 100MB?"}],
)
print(resp.choices[0].message.content)

The placeholder api_key="ollama" is exactly what Ollama's docs prescribe. Streaming works as in the cloud API (stream=True), and the same code points at LM Studio by switching the base URL to port 1234.

WARNING
Compatibility covers the protocol, not behaviour
Local models have smaller context windows and different prompting quirks. Code assuming a 128K context may need tuning — check the limit with ollama show llama3.1:8b.

6 Exposing the Endpoint Safely (and Keeping It Local)

By default Ollama binds to localhost, reachable only from your own machine — keep it that way unless you have a reason not to. If another device on your LAN needs access, change the bind address with OLLAMA_HOST:

OLLAMA_HOST=0.0.0.0:11434 ollama serve   # bind to all interfaces

Do this only behind a trusted network. Understand the exposure first:

Scenario Exposure
localhost only (default) None beyond your machine — keep for solo dev
LAN bind (0.0.0.0) Anyone on the network can prompt your model — firewall it
Public internet Anyone can run inference on your GPU — never expose

Critically, Ollama's OpenAI layer ignores the API key — it accepts any string, so there is no built-in authentication. Before any network exposure, put a reverse proxy (Caddy or nginx) enforcing a bearer token in front of 127.0.0.1:11434, or just expose Open WebUI (port 3000) and keep Ollama on localhost.

WARNING
Do not port-forward Ollama to the internet
An open endpoint is a free GPU for whoever finds it — and a data-leak channel if you have RAG or files wired in. Real auth, TLS, and rate limiting come in Lesson 10. Until then, localhost.

Questions & Answers

Q: If everything speaks the OpenAI API, why bother with Ollama vs LM Studio vs vLLM at all?
The wire protocol is the same; the runtime is not. Ollama optimises for simplicity and reproducible Modelfiles, LM Studio for GUI experimentation, vLLM (Lesson 10) for high-throughput serving. Pick by workload — and since they all speak OpenAI, switching later is a config change, not a rewrite.
Q: My VS Code autocomplete is painfully slow. Is Continue broken?
Almost certainly not. Autocomplete fires on nearly every keystroke pause — the most latency-sensitive integration there is — and a CPU-only 8B model cannot keep up. Switch the autocomplete role to a sub-2B model, or disable it and keep chat/edit. Lesson 9 covers more speed.
Q: Is any of my data leaving my machine when I use these tools?
Inference stays local — prompts go to localhost and never touch a vendor. Caveats: tools may phone home for update checks or telemetry, and models download from a remote registry. Check each app's telemetry settings for zero outbound traffic.
Q: Open WebUI's model dropdown is empty. What did I get wrong?
Networking, 95% of the time — the container cannot reach Ollama. Confirm you launched it with --add-host=host.docker.internal:host-gateway (or --network=host on Linux) and that Ollama is running. Curl the host's /api/tags from inside the container to confirm.

Key Takeaways

  1. The OpenAI-compatible API is the universal adapter — Ollama serves it at localhost:11434/v1; any OpenAI-aware tool connects by changing the base URL and a throwaway key.
  2. Open WebUI gives you a private ChatGPT — one Docker command on port 3000, auto-detecting your Ollama models, with accounts and history.
  3. LM Studio is the GUI for model hunting — its own OpenAI server on port 1234, best for browsing models and tuning quantization.
  4. Continue turns VS Code into a local copilot — provider: ollama and apiBase in config.yaml, splitting a small completion model from a larger chat model.
  5. Your own code is just another client — the OpenAI SDK with an overridden base_url migrates any existing OpenAI app to local.
  6. Local means localhost — keep it that way — there is no built-in auth, so never expose the raw endpoint without a proxy in front.

Next Steps: Lesson 7: Fine-Tuning Basics