Ollama & CLI Cheat Sheet

Reference intermediate

A terminal-first quick reference for the commands you reach for daily when running models locally. For deeper context see the model guide, the hardware reference, and the troubleshooting guide.

Ollama: Core Commands

Ollama wraps llama.cpp with a model registry and a daemon. Run ollama [command] --help for the full flag list on your version.

Command What it does Example
ollama pull Download a model from the registry ollama pull llama3.2
ollama run Pull (if needed) then start an interactive chat ollama run qwen2.5:7b
ollama list (ollama ls) List models on disk with sizes ollama ls
ollama ps Show models currently loaded in memory ollama ps
ollama stop Unload a running model from memory ollama stop llama3.2
ollama show Show a model's details and Modelfile ollama show llama3.2
ollama cp Copy / alias a model ollama cp llama3.2 my-llama
ollama rm Delete a model from disk ollama rm llama3.2
ollama create Build a model from a Modelfile ollama create my-bot -f ./Modelfile
ollama push Upload a model to a registry namespace ollama push user/my-bot
ollama serve Start the API server in the foreground ollama serve

Pulling & Running

A model tag is name:tag, where the tag usually encodes size and quantization (lossy compression that shrinks weights to fit memory). Omit it for the default.

# Pull a specific size/quant, then run it
ollama pull llama3.2:3b
ollama run llama3.2:3b

# Pull an explicit quant tag (Q4_K_M = 4-bit, the common default)
ollama pull qwen2.5:7b-instruct-q4_K_M

# One-shot prompt, no interactive session (good for scripts)
ollama run llama3.2 "Summarise in one sentence: $(cat notes.txt)"

# Pipe stdin into the model
cat error.log | ollama run llama3.2 "Explain the root cause"

Inside an interactive ollama run session: /set system "..." sets a session system prompt, /set parameter num_ctx 8192 changes a runtime parameter live, /show info prints the model's parameters, /clear resets context, """ opens a multi-line block, and /bye exits.

Managing Memory & Disk

# What is loaded now, and how much VRAM/RAM it uses
ollama ps

# Free memory now instead of waiting for the idle timeout
ollama stop qwen2.5:7b

# See disk usage, then reclaim space
ollama ls
ollama rm old-model:13b

Use ollama ps to confirm whether a model landed on GPU or fell back to CPU before blaming slow tokens on the model. See the troubleshooting guide for OOM and GPU fixes.

Useful Environment Variables

Set these in your shell profile or before ollama serve (ollama serve --help prints the full list).

Variable Purpose Example
OLLAMA_HOST Bind address/port for the server and client OLLAMA_HOST=0.0.0.0:11434
OLLAMA_MODELS Directory where model blobs are stored OLLAMA_MODELS=/data/ollama
OLLAMA_KEEP_ALIVE How long an idle model stays in memory OLLAMA_KEEP_ALIVE=30m
# Serve on the LAN with the model store on a big disk
OLLAMA_HOST=0.0.0.0:11434 OLLAMA_MODELS=/data/models ollama serve

Modelfile Template

A Modelfile customises a base model's system prompt, sampling parameters, template, and optional adapters. FROM is the only required instruction.

# Modelfile
FROM llama3.2

# Sampling and context parameters
PARAMETER temperature 0.6
PARAMETER num_ctx 8192
PARAMETER top_p 0.9
PARAMETER stop "<|eot_id|>"

# Baked-in personality
SYSTEM You are a terse senior engineer. Answer in code first, prose second.

# Optional: apply a fine-tuned LoRA adapter
# ADAPTER ./my-lora

Build and run:

ollama create senior-eng -f ./Modelfile
ollama run senior-eng

Instructions: FROM (required) sets the base model or GGUF path; PARAMETER sets a runtime value (below); SYSTEM bakes in a system message; TEMPLATE defines the Go-template prompt format; ADAPTER applies a (Q)LoRA adapter; LICENSE embeds license text; MESSAGE seeds canned history.

Common PARAMETER values:

Parameter Meaning Default
num_ctx Context window size in tokens 2048
temperature Higher = more creative/random 0.8
top_k Limits sampling to the K likeliest tokens 40
top_p Nucleus sampling probability mass 0.9
repeat_penalty Penalty for repeating tokens 1.1
num_predict Max tokens to generate (-1 = unlimited) -1
seed Fixed seed for reproducible output 0
stop Stop sequence (repeat for multiple) —

llama.cpp Invocations

For lower-level control than Ollama, drive llama.cpp directly. It runs GGUF files: -m points at a local model, -hf downloads from Hugging Face, and -ngl (GPU layers) is the biggest performance lever.

# One-shot generation from a local GGUF
llama-cli -m ./model.gguf -p "Write a haiku about cold compute" -n 100

# Conversation mode with a system prompt and bigger context
llama-cli -m ./model.gguf -cnv -sys "You are a helpful assistant" -c 8192

# Download from Hugging Face, offload all layers to GPU
llama-cli -hf ggml-org/gemma-3-1b-it-GGUF -ngl 99 -cnv

Most-used llama-cli flags:

Flag Meaning
-m, --model Path to a local GGUF model file
-hf, --hf-repo Download model from a Hugging Face repo (user/model:quant)
-p, --prompt Prompt to start generation
-sys, --system-prompt System prompt
-cnv, --conversation Interactive chat mode (-no-cnv to disable)
-c, --ctx-size Context window in tokens (0 = use model default)
-n, --n-predict Tokens to generate (-1 = unlimited)
-ngl, --n-gpu-layers Layers to offload to GPU/VRAM
-t, --threads CPU threads for generation
--temp Sampling temperature (default 0.80)

llama.cpp Server Mode

llama-server exposes an OpenAI-compatible API and web UI on 127.0.0.1:8080 by default.

# Serve a local model with 2K context
llama-server -m ./models/ggml-model.gguf -c 2048

# Serve on the LAN, all layers on GPU, from a HF repo
llama-server -hf ggml-org/gemma-3-1b-it-GGUF --host 0.0.0.0 --n-gpu-layers 99

Quantization Tags at a Glance

The quant suffix trades quality for memory and speed. Q4_K_M is the usual sweet spot; drop to Q3/Q2 only when memory-starved, reach for Q8/F16 with headroom.

Tag Bits/weight Trade-off
Q2_K ~2-bit Smallest, noticeable quality loss
Q4_K_M ~4-bit Recommended default balance
Q5_K_M ~5-bit Better quality, more memory
Q8_0 8-bit Near-lossless, larger footprint
F16 16-bit Full half-precision weights

For matching quant levels to hardware, see the model guide and hardware reference.