Running Models from the CLI

40 min intermediate Lesson 5

Learning Outcomes

  • Drive the Ollama CLI beyond chat — runs, pulls, one-shot prompts, and model management
  • Pipe text in and out of a local model and compose it with standard Unix tools
  • Script reproducible inference using the local HTTP API and structured JSON output
  • Batch-process a folder of documents through a model and write structured results
  • Run llama.cpp directly for control Ollama hides, and manage several models without exhausting memory

Lesson Plan

Segment Duration Topic
Intro 3 min Why the CLI beats the chat box for real work
Demo 1 6 min Ollama CLI in depth — run, pull, show, ps
Demo 2 6 min Piping and one-shot prompts
Demo 3 9 min Scripting the API: options and structured JSON
Demo 4 7 min Batch processing a folder
Demo 5 5 min llama.cpp direct for maximum control
Explain 2 min Managing multiple models and memory
Wrap-up 2 min Key takeaways, preview Lesson 6

Before You Begin

Pre-work:

Shopping List:

  • Ollama running locally (ollama --version prints a version)
  • A small instruct model pulled, e.g. llama3.2:3b or qwen2.5:7b
  • jq for parsing JSON (brew install jq on macOS, sudo apt install jq on Debian/Ubuntu)
  • A folder of plain-text or markdown files to practise batch processing on

1 The Ollama CLI, Properly

Most people only ever type ollama run. The CLI does far more, and knowing it turns Ollama into a scriptable engine:

ollama list          # models on disk, with size and date
ollama ps            # models loaded into RAM/VRAM right now
ollama pull qwen2.5:7b      # download (or update) a model
ollama show llama3.2:3b     # parameters, context length, template
ollama rm old-model         # delete to reclaim disk
ollama cp llama3.2:3b mine  # clone under a new name

The two you will reach for constantly are ps and show. ollama ps reports what is resident and whether it landed on GPU or CPU:

NAME           SIZE     PROCESSOR    UNTIL
qwen2.5:7b     5.5 GB   100% GPU     4 minutes from now

The UNTIL column matters: Ollama unloads a model 5 minutes after the last request by default. ollama show reveals the model's real context window and chat template — check it before scripting against the model.

TIP
Tip
Tags after the colon select size and quantization, e.g. qwen2.5:7b vs qwen2.5:7b-instruct-q8_0. A bare name resolves to a default tag — always pin an explicit tag in scripts so results are reproducible across machines.
NOTE
What is quantization
Quantization stores weights at lower precision (4-bit, 5-bit, 8-bit) instead of 16-bit: a smaller file and faster inference at a small quality cost. Q4 is the common sweet spot; Q8 is near-lossless but roughly twice the size. See Lesson 4.

2 One-Shot Prompts and Piping

ollama run accepts a prompt as an argument, runs it once, prints the answer, and exits — no interactive session. That single fact unlocks the entire Unix toolbox.

# One-shot, then pipe a file in via stdin
ollama run llama3.2:3b "Summarise relativity in one sentence."
cat report.txt | ollama run llama3.2:3b "Summarise the text above in 3 bullets:"

# Turn git commit subjects into changelog entries
git log --since=yesterday --pretty=%s \
  | ollama run llama3.2:3b "Rewrite each line as a changelog entry:"

# Explain a build error and page through the result
make 2>&1 | ollama run qwen2.5:7b "Explain this build error and fix:" | less

Anything that produces text on stdout can feed the model, and the model's stdout can feed anything else. Inside an interactive session, /set parameter temperature 0 makes output deterministic, /show info prints the model's parameters and template, and /bye exits.

WARNING
Watch Out
Local models still hallucinate. When you pipe output into a script that takes action — deleting files, running commands — never let raw output drive a destructive step without a human or validator in between. Treat generated text as a suggestion, not a command.

3 Scripting the API: Options and Structured Output

ollama run mixes spinners and text on stdout, awkward to parse. For automation, talk to the HTTP server Ollama runs on localhost:11434 (started by ollama serve). The cleanest endpoint is /api/generate: set stream false for one JSON object instead of a token stream, and put the controls the chat box hides under options:

curl -s http://localhost:11434/api/generate -d '{
  "model": "qwen2.5:7b",
  "prompt": "List two facts about the moon.",
  "stream": false,
  "options": { "temperature": 0, "num_ctx": 8192, "seed": 42 }
}' | jq -r '.response'

The key options: temperature (0 is deterministic, 0.7 for prose), num_ctx (context window in tokens — larger uses more memory), seed (fixes sampling for reproducibility), and num_predict (max tokens, -1 for unlimited). For multi-turn or system-prompt work, use /api/chat, which takes a messages array with role and content fields exactly like the OpenAI format — the compatibility that underpins Lesson 6.

TIP
Tip
Set temperature to 0 and pin a seed whenever a script needs the same input to yield the same output. This is the single biggest reliability win when wrapping a local model in automation.

Forcing structured JSON. Free-form text is hard to consume downstream. Ollama can constrain the model to a JSON schema so it fills your exact shape — the difference between a fragile regex and a clean jq pipeline. Pass the schema in format:

curl -s http://localhost:11434/api/generate -d '{
  "model": "qwen2.5:7b",
  "prompt": "Extract details from: Dr. Ada Lovelace, age 36, mathematician.",
  "stream": false,
  "format": {
    "type": "object",
    "properties": {
      "name": { "type": "string" },
      "age":  { "type": "integer" },
      "role": { "type": "string" }
    },
    "required": ["name", "age", "role"]
  }
}' | jq '.response | fromjson'

The .response is a JSON string, so fromjson parses it into a real object. Two rules make this reliable: tell the model in the prompt to extract or produce data (the schema enforces shape, the prompt supplies intent), and mark fields required so you never get a partial object.

NOTE
How it works
Ollama applies grammar-constrained sampling: at each step it only allows tokens that keep the output valid against your schema, so the model literally cannot emit a malformed brace. But a small model (1B-3B) can satisfy the schema while inventing values — valid JSON is not correct JSON, so validate the contents before trusting extracted data.

4 Batch Processing a Folder

Now combine everything: walk a directory, run each file through the model, write one result per file. A cloud API would bill and rate-limit this; local, it is free and runs as fast as your hardware allows.

#!/usr/bin/env bash
set -euo pipefail
MODEL="qwen2.5:7b"; OUT="./summaries"; mkdir -p "$OUT"

for file in ./documents/*.txt; do
  name="$(basename "$file" .txt)"

  # jq builds the request body and safely escapes the file contents
  body="$(jq -n --arg m "$MODEL" --arg text "$(cat "$file")" '{
    model: $m, stream: false, options: { temperature: 0 },
    prompt: ("Summarise the document below in two sentences:\n\n" + $text)
  }')"

  curl -s http://localhost:11434/api/generate -d "$body" \
    | jq -r '.response' > "$OUT/$name.summary.txt"
done

The key detail is building the body with jq -n. Document text is full of quotes, newlines, and backslashes that break a hand-built JSON string; letting jq escape it is the difference between a script that works on your sample and one that survives real data. To include other file types, swap the glob for a find ... -name '*.md' -o -name '*.py' loop.

On macOS and Linux the script is portable as written — bash, curl, and find behave identically.

Run it inside WSL for a real bash, curl, and jq. If Ollama runs on the Windows host (not inside WSL), set export OLLAMA_HOST=http://host.docker.internal:11434 and use $OLLAMA_HOST/api/generate. If Ollama runs inside WSL, localhost:11434 works as on macOS.

TIP
Tip
Set OLLAMA_KEEP_ALIVE=30m so the model stays resident across the whole batch instead of reloading between files. For a 7B model on a laptop, reloading per file can dominate total runtime.

5 Going Direct with llama.cpp

Under the hood Ollama uses llama.cpp. When you need control it does not expose — a GGUF file you downloaded yourself, exact GPU layer offloading, or a custom sampler — run llama.cpp directly. GGUF is its single-file model format; you download one .gguf per model and quantization from a model hub.

# Build once (Metal on macOS, CUDA on NVIDIA), then run a local GGUF
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp && cmake -B build && cmake --build build --config Release

./build/bin/llama-cli -m ./models/qwen2.5-7b-q4_k_m.gguf \
  -p "Explain how a hash map works." -n 256 -c 8192 -ngl 99 --temp 0

The flags map onto the API options: -m model path, -p prompt, -n tokens, -c context size, --temp temperature. The headline flag is -ngl, the number of transformer layers to offload to the GPU — set it as high as VRAM allows; -ngl 99 offloads everything and llama.cpp caps it at the model's layer count. llama.cpp also ships an OpenAI-compatible server — llama-server -m model.gguf -ngl 99 -c 8192 --port 8080 — that you POST to at /v1/chat/completions, a lightweight API without the full Ollama daemon.

NOTE
When to drop down to llama.cpp
Stay on Ollama for everyday work — it manages downloads, memory, and templates. Reach for raw llama.cpp when you need a model not in the Ollama library, precise -ngl tuning for a tight VRAM budget, or bleeding-edge features before they land in Ollama.

6 Managing Multiple Models and Memory

Memory is the constraint once you juggle several models. Each occupies RAM (or VRAM) roughly equal to its file size plus the KV cache for the context window — two 7B Q4 models will not coexist on an 8GB GPU. A rough per-model budget:

Model size Quantization Approx. memory to load
3B Q4 ~2.5 GB
7B Q4 ~5 GB
13B Q4 ~8 GB
70B Q4 ~40 GB

Add headroom for the context window — a large num_ctx adds gigabytes of KV cache on top. Control loading with environment variables before starting ollama serve, and evict on demand:

export OLLAMA_KEEP_ALIVE=1h            # keep models loaded for an hour
export OLLAMA_MAX_LOADED_MODELS=2      # at most 2 resident at once
ollama serve
ollama stop qwen2.5:7b                  # unload now, free memory

For a two-model pipeline, process the whole batch with model A, ollama stop it, then run model B — per-phase swapping keeps each model resident; per-item swapping thrashes memory and reloads constantly. On Apple Silicon, CPU and GPU share one unified-memory pool, so the RAM-vs-VRAM split does not apply — a 32GB Mac can load models a 32GB-RAM PC with an 8GB GPU cannot.

WARNING
Watch Out
If ollama ps shows a model on 100% CPU when you expected GPU, it did not fit in VRAM and silently fell back to RAM — generation will be far slower. Reduce num_ctx, drop to a smaller quantization, or pick a smaller model. Lesson 9 covers diagnosing this in depth.

Questions & Answers

Q: My batch script gives different results every run even at temperature 0. Why?
Temperature 0 fixes sampling, but a non-fixed seed, GPU floating-point non-determinism, or a different num_ctx truncating long inputs all vary output. Pin seed and num_ctx and keep the same model tag and quantization. Even then GPU kernels add tiny nondeterminism — for bit-exact output, run on CPU.
Q: Should I shell out to ollama run from my program, or hit the API?
Hit the API. Shelling out means parsing text mixed with spinner output, no clean error codes, and a process spawn per call. The /api/generate and /api/chat endpoints give JSON, real HTTP status codes, and the full options object. Reserve ollama run for interactive use and shell one-liners.
Q: My batch job is far slower than a single prompt felt. What is wrong?
Almost always model reloading. If OLLAMA_KEEP_ALIVE is the default 5 minutes, or you stop/start between files, the model reloads from disk each time — seconds of overhead per file. Set OLLAMA_KEEP_ALIVE high and confirm with ollama ps that it stays resident.
Q: A long document gets truncated or the answer ignores its end. Why?
You exceeded the context window — tokens past num_ctx are dropped. Raise num_ctx (at extra memory cost), or chunk the document and summarise the chunks, the problem RAG addresses in Lesson 8. Check the real limit with ollama show before assuming a huge context.

Key Takeaways

  1. The CLI is the interface — ps, show, pull, and stop control what runs and what is loaded.
  2. One-shot plus pipes — ollama run model "prompt" composes with any tool that speaks stdin/stdout.
  3. Script the API, not the chat — /api/generate with stream: false gives parseable JSON and reproducible runs.
  4. Constrain output with a schema — the format field forces valid JSON, but still validate the values.
  5. Keep models resident across a batch — OLLAMA_KEEP_ALIVE plus jq -n body construction makes a reliable pipeline.
  6. Drop to llama.cpp for control — direct GGUF loading and -ngl tuning matter when memory is tight.

Next Steps: Lesson 6: Connecting Local Models to Tools