Running Models from the CLI
Learning Outcomes
- Drive the Ollama CLI beyond chat — runs, pulls, one-shot prompts, and model management
- Pipe text in and out of a local model and compose it with standard Unix tools
- Script reproducible inference using the local HTTP API and structured JSON output
- Batch-process a folder of documents through a model and write structured results
- Run llama.cpp directly for control Ollama hides, and manage several models without exhausting memory
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why the CLI beats the chat box for real work |
| Demo 1 | 6 min | Ollama CLI in depth — run, pull, show, ps |
| Demo 2 | 6 min | Piping and one-shot prompts |
| Demo 3 | 9 min | Scripting the API: options and structured JSON |
| Demo 4 | 7 min | Batch processing a folder |
| Demo 5 | 5 min | llama.cpp direct for maximum control |
| Explain | 2 min | Managing multiple models and memory |
| Wrap-up | 2 min | Key takeaways, preview Lesson 6 |
Before You Begin
Pre-work:
- Complete Lesson 3: Ollama — The Easiest Way to Start so Ollama is installed with at least one model pulled
- Skim Lesson 4: Model Selection for parameter counts (7B, 13B) and quantization levels (Q4, Q5, Q8)
Shopping List:
- Ollama running locally (
ollama --versionprints a version) - A small instruct model pulled, e.g.
llama3.2:3borqwen2.5:7b jqfor parsing JSON (brew install jqon macOS,sudo apt install jqon Debian/Ubuntu)- A folder of plain-text or markdown files to practise batch processing on
Most people only ever type ollama run. The CLI does far more, and knowing it turns Ollama into a scriptable engine:
ollama list # models on disk, with size and date
ollama ps # models loaded into RAM/VRAM right now
ollama pull qwen2.5:7b # download (or update) a model
ollama show llama3.2:3b # parameters, context length, template
ollama rm old-model # delete to reclaim disk
ollama cp llama3.2:3b mine # clone under a new name
The two you will reach for constantly are ps and show. ollama ps reports what is resident and whether it landed on GPU or CPU:
NAME SIZE PROCESSOR UNTIL
qwen2.5:7b 5.5 GB 100% GPU 4 minutes from now
The UNTIL column matters: Ollama unloads a model 5 minutes after the last request by default. ollama show reveals the model's real context window and chat template — check it before scripting against the model.
qwen2.5:7b vs qwen2.5:7b-instruct-q8_0. A bare name resolves to a default tag — always pin an explicit tag in scripts so results are reproducible across machines.ollama run accepts a prompt as an argument, runs it once, prints the answer, and exits — no interactive session. That single fact unlocks the entire Unix toolbox.
# One-shot, then pipe a file in via stdin
ollama run llama3.2:3b "Summarise relativity in one sentence."
cat report.txt | ollama run llama3.2:3b "Summarise the text above in 3 bullets:"
# Turn git commit subjects into changelog entries
git log --since=yesterday --pretty=%s \
| ollama run llama3.2:3b "Rewrite each line as a changelog entry:"
# Explain a build error and page through the result
make 2>&1 | ollama run qwen2.5:7b "Explain this build error and fix:" | less
Anything that produces text on stdout can feed the model, and the model's stdout can feed anything else. Inside an interactive session, /set parameter temperature 0 makes output deterministic, /show info prints the model's parameters and template, and /bye exits.
ollama run mixes spinners and text on stdout, awkward to parse. For automation, talk to the HTTP server Ollama runs on localhost:11434 (started by ollama serve). The cleanest endpoint is /api/generate: set stream false for one JSON object instead of a token stream, and put the controls the chat box hides under options:
curl -s http://localhost:11434/api/generate -d '{
"model": "qwen2.5:7b",
"prompt": "List two facts about the moon.",
"stream": false,
"options": { "temperature": 0, "num_ctx": 8192, "seed": 42 }
}' | jq -r '.response'
The key options: temperature (0 is deterministic, 0.7 for prose), num_ctx (context window in tokens — larger uses more memory), seed (fixes sampling for reproducibility), and num_predict (max tokens, -1 for unlimited). For multi-turn or system-prompt work, use /api/chat, which takes a messages array with role and content fields exactly like the OpenAI format — the compatibility that underpins Lesson 6.
temperature to 0 and pin a seed whenever a script needs the same input to yield the same output. This is the single biggest reliability win when wrapping a local model in automation.Forcing structured JSON. Free-form text is hard to consume downstream. Ollama can constrain the model to a JSON schema so it fills your exact shape — the difference between a fragile regex and a clean jq pipeline. Pass the schema in format:
curl -s http://localhost:11434/api/generate -d '{
"model": "qwen2.5:7b",
"prompt": "Extract details from: Dr. Ada Lovelace, age 36, mathematician.",
"stream": false,
"format": {
"type": "object",
"properties": {
"name": { "type": "string" },
"age": { "type": "integer" },
"role": { "type": "string" }
},
"required": ["name", "age", "role"]
}
}' | jq '.response | fromjson'
The .response is a JSON string, so fromjson parses it into a real object. Two rules make this reliable: tell the model in the prompt to extract or produce data (the schema enforces shape, the prompt supplies intent), and mark fields required so you never get a partial object.
Now combine everything: walk a directory, run each file through the model, write one result per file. A cloud API would bill and rate-limit this; local, it is free and runs as fast as your hardware allows.
#!/usr/bin/env bash
set -euo pipefail
MODEL="qwen2.5:7b"; OUT="./summaries"; mkdir -p "$OUT"
for file in ./documents/*.txt; do
name="$(basename "$file" .txt)"
# jq builds the request body and safely escapes the file contents
body="$(jq -n --arg m "$MODEL" --arg text "$(cat "$file")" '{
model: $m, stream: false, options: { temperature: 0 },
prompt: ("Summarise the document below in two sentences:\n\n" + $text)
}')"
curl -s http://localhost:11434/api/generate -d "$body" \
| jq -r '.response' > "$OUT/$name.summary.txt"
done
The key detail is building the body with jq -n. Document text is full of quotes, newlines, and backslashes that break a hand-built JSON string; letting jq escape it is the difference between a script that works on your sample and one that survives real data. To include other file types, swap the glob for a find ... -name '*.md' -o -name '*.py' loop.
On macOS and Linux the script is portable as written — bash, curl, and find behave identically.
Run it inside WSL for a real bash, curl, and jq. If Ollama runs on the Windows host (not inside WSL), set export OLLAMA_HOST=http://host.docker.internal:11434 and use $OLLAMA_HOST/api/generate. If Ollama runs inside WSL, localhost:11434 works as on macOS.
OLLAMA_KEEP_ALIVE=30m so the model stays resident across the whole batch instead of reloading between files. For a 7B model on a laptop, reloading per file can dominate total runtime.Under the hood Ollama uses llama.cpp. When you need control it does not expose — a GGUF file you downloaded yourself, exact GPU layer offloading, or a custom sampler — run llama.cpp directly. GGUF is its single-file model format; you download one .gguf per model and quantization from a model hub.
# Build once (Metal on macOS, CUDA on NVIDIA), then run a local GGUF
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp && cmake -B build && cmake --build build --config Release
./build/bin/llama-cli -m ./models/qwen2.5-7b-q4_k_m.gguf \
-p "Explain how a hash map works." -n 256 -c 8192 -ngl 99 --temp 0
The flags map onto the API options: -m model path, -p prompt, -n tokens, -c context size, --temp temperature. The headline flag is -ngl, the number of transformer layers to offload to the GPU — set it as high as VRAM allows; -ngl 99 offloads everything and llama.cpp caps it at the model's layer count. llama.cpp also ships an OpenAI-compatible server — llama-server -m model.gguf -ngl 99 -c 8192 --port 8080 — that you POST to at /v1/chat/completions, a lightweight API without the full Ollama daemon.
-ngl tuning for a tight VRAM budget, or bleeding-edge features before they land in Ollama.Memory is the constraint once you juggle several models. Each occupies RAM (or VRAM) roughly equal to its file size plus the KV cache for the context window — two 7B Q4 models will not coexist on an 8GB GPU. A rough per-model budget:
| Model size | Quantization | Approx. memory to load |
|---|---|---|
| 3B | Q4 | ~2.5 GB |
| 7B | Q4 | ~5 GB |
| 13B | Q4 | ~8 GB |
| 70B | Q4 | ~40 GB |
Add headroom for the context window — a large num_ctx adds gigabytes of KV cache on top. Control loading with environment variables before starting ollama serve, and evict on demand:
export OLLAMA_KEEP_ALIVE=1h # keep models loaded for an hour
export OLLAMA_MAX_LOADED_MODELS=2 # at most 2 resident at once
ollama serve
ollama stop qwen2.5:7b # unload now, free memory
For a two-model pipeline, process the whole batch with model A, ollama stop it, then run model B — per-phase swapping keeps each model resident; per-item swapping thrashes memory and reloads constantly. On Apple Silicon, CPU and GPU share one unified-memory pool, so the RAM-vs-VRAM split does not apply — a 32GB Mac can load models a 32GB-RAM PC with an 8GB GPU cannot.
ollama ps shows a model on 100% CPU when you expected GPU, it did not fit in VRAM and silently fell back to RAM — generation will be far slower. Reduce num_ctx, drop to a smaller quantization, or pick a smaller model. Lesson 9 covers diagnosing this in depth.Questions & Answers
seed, GPU floating-point non-determinism, or a different num_ctx truncating long inputs all vary output. Pin seed and num_ctx and keep the same model tag and quantization. Even then GPU kernels add tiny nondeterminism — for bit-exact output, run on CPU.ollama run from my program, or hit the API?/api/generate and /api/chat endpoints give JSON, real HTTP status codes, and the full options object. Reserve ollama run for interactive use and shell one-liners.OLLAMA_KEEP_ALIVE is the default 5 minutes, or you stop/start between files, the model reloads from disk each time — seconds of overhead per file. Set OLLAMA_KEEP_ALIVE high and confirm with ollama ps that it stays resident.num_ctx are dropped. Raise num_ctx (at extra memory cost), or chunk the document and summarise the chunks, the problem RAG addresses in Lesson 8. Check the real limit with ollama show before assuming a huge context.Key Takeaways
- The CLI is the interface —
ps,show,pull, andstopcontrol what runs and what is loaded. - One-shot plus pipes —
ollama run model "prompt"composes with any tool that speaks stdin/stdout. - Script the API, not the chat —
/api/generatewithstream: falsegives parseable JSON and reproducible runs. - Constrain output with a schema — the
formatfield forces valid JSON, but still validate the values. - Keep models resident across a batch —
OLLAMA_KEEP_ALIVEplusjq -nbody construction makes a reliable pipeline. - Drop to llama.cpp for control — direct GGUF loading and
-ngltuning matter when memory is tight.
Next Steps: Lesson 6: Connecting Local Models to Tools