Ollama & CLI Cheat Sheet
A terminal-first quick reference for the commands you reach for daily when running models locally. For deeper context see the model guide, the hardware reference, and the troubleshooting guide.
Ollama: Core Commands
Ollama wraps llama.cpp with a model registry and a daemon. Run ollama [command] --help for the full flag list on your version.
| Command | What it does | Example |
|---|---|---|
ollama pull |
Download a model from the registry | ollama pull llama3.2 |
ollama run |
Pull (if needed) then start an interactive chat | ollama run qwen2.5:7b |
ollama list (ollama ls) |
List models on disk with sizes | ollama ls |
ollama ps |
Show models currently loaded in memory | ollama ps |
ollama stop |
Unload a running model from memory | ollama stop llama3.2 |
ollama show |
Show a model's details and Modelfile | ollama show llama3.2 |
ollama cp |
Copy / alias a model | ollama cp llama3.2 my-llama |
ollama rm |
Delete a model from disk | ollama rm llama3.2 |
ollama create |
Build a model from a Modelfile | ollama create my-bot -f ./Modelfile |
ollama push |
Upload a model to a registry namespace | ollama push user/my-bot |
ollama serve |
Start the API server in the foreground | ollama serve |
Pulling & Running
A model tag is name:tag, where the tag usually encodes size and quantization (lossy compression that shrinks weights to fit memory). Omit it for the default.
# Pull a specific size/quant, then run it
ollama pull llama3.2:3b
ollama run llama3.2:3b
# Pull an explicit quant tag (Q4_K_M = 4-bit, the common default)
ollama pull qwen2.5:7b-instruct-q4_K_M
# One-shot prompt, no interactive session (good for scripts)
ollama run llama3.2 "Summarise in one sentence: $(cat notes.txt)"
# Pipe stdin into the model
cat error.log | ollama run llama3.2 "Explain the root cause"
Inside an interactive ollama run session: /set system "..." sets a session system prompt, /set parameter num_ctx 8192 changes a runtime parameter live, /show info prints the model's parameters, /clear resets context, """ opens a multi-line block, and /bye exits.
Managing Memory & Disk
# What is loaded now, and how much VRAM/RAM it uses
ollama ps
# Free memory now instead of waiting for the idle timeout
ollama stop qwen2.5:7b
# See disk usage, then reclaim space
ollama ls
ollama rm old-model:13b
Use ollama ps to confirm whether a model landed on GPU or fell back to CPU before blaming slow tokens on the model. See the troubleshooting guide for OOM and GPU fixes.
Useful Environment Variables
Set these in your shell profile or before ollama serve (ollama serve --help prints the full list).
| Variable | Purpose | Example |
|---|---|---|
OLLAMA_HOST |
Bind address/port for the server and client | OLLAMA_HOST=0.0.0.0:11434 |
OLLAMA_MODELS |
Directory where model blobs are stored | OLLAMA_MODELS=/data/ollama |
OLLAMA_KEEP_ALIVE |
How long an idle model stays in memory | OLLAMA_KEEP_ALIVE=30m |
# Serve on the LAN with the model store on a big disk
OLLAMA_HOST=0.0.0.0:11434 OLLAMA_MODELS=/data/models ollama serve
Modelfile Template
A Modelfile customises a base model's system prompt, sampling parameters, template, and optional adapters. FROM is the only required instruction.
# Modelfile
FROM llama3.2
# Sampling and context parameters
PARAMETER temperature 0.6
PARAMETER num_ctx 8192
PARAMETER top_p 0.9
PARAMETER stop "<|eot_id|>"
# Baked-in personality
SYSTEM You are a terse senior engineer. Answer in code first, prose second.
# Optional: apply a fine-tuned LoRA adapter
# ADAPTER ./my-lora
Build and run:
ollama create senior-eng -f ./Modelfile
ollama run senior-eng
Instructions: FROM (required) sets the base model or GGUF path; PARAMETER sets a runtime value (below); SYSTEM bakes in a system message; TEMPLATE defines the Go-template prompt format; ADAPTER applies a (Q)LoRA adapter; LICENSE embeds license text; MESSAGE seeds canned history.
Common PARAMETER values:
| Parameter | Meaning | Default |
|---|---|---|
num_ctx |
Context window size in tokens | 2048 |
temperature |
Higher = more creative/random | 0.8 |
top_k |
Limits sampling to the K likeliest tokens | 40 |
top_p |
Nucleus sampling probability mass | 0.9 |
repeat_penalty |
Penalty for repeating tokens | 1.1 |
num_predict |
Max tokens to generate (-1 = unlimited) | -1 |
seed |
Fixed seed for reproducible output | 0 |
stop |
Stop sequence (repeat for multiple) | — |
llama.cpp Invocations
For lower-level control than Ollama, drive llama.cpp directly. It runs GGUF files: -m points at a local model, -hf downloads from Hugging Face, and -ngl (GPU layers) is the biggest performance lever.
# One-shot generation from a local GGUF
llama-cli -m ./model.gguf -p "Write a haiku about cold compute" -n 100
# Conversation mode with a system prompt and bigger context
llama-cli -m ./model.gguf -cnv -sys "You are a helpful assistant" -c 8192
# Download from Hugging Face, offload all layers to GPU
llama-cli -hf ggml-org/gemma-3-1b-it-GGUF -ngl 99 -cnv
Most-used llama-cli flags:
| Flag | Meaning |
|---|---|
-m, --model |
Path to a local GGUF model file |
-hf, --hf-repo |
Download model from a Hugging Face repo (user/model:quant) |
-p, --prompt |
Prompt to start generation |
-sys, --system-prompt |
System prompt |
-cnv, --conversation |
Interactive chat mode (-no-cnv to disable) |
-c, --ctx-size |
Context window in tokens (0 = use model default) |
-n, --n-predict |
Tokens to generate (-1 = unlimited) |
-ngl, --n-gpu-layers |
Layers to offload to GPU/VRAM |
-t, --threads |
CPU threads for generation |
--temp |
Sampling temperature (default 0.80) |
llama.cpp Server Mode
llama-server exposes an OpenAI-compatible API and web UI on 127.0.0.1:8080 by default.
# Serve a local model with 2K context
llama-server -m ./models/ggml-model.gguf -c 2048
# Serve on the LAN, all layers on GPU, from a HF repo
llama-server -hf ggml-org/gemma-3-1b-it-GGUF --host 0.0.0.0 --n-gpu-layers 99
Quantization Tags at a Glance
The quant suffix trades quality for memory and speed. Q4_K_M is the usual sweet spot; drop to Q3/Q2 only when memory-starved, reach for Q8/F16 with headroom.
| Tag | Bits/weight | Trade-off |
|---|---|---|
Q2_K |
~2-bit | Smallest, noticeable quality loss |
Q4_K_M |
~4-bit | Recommended default balance |
Q5_K_M |
~5-bit | Better quality, more memory |
Q8_0 |
8-bit | Near-lossless, larger footprint |
F16 |
16-bit | Full half-precision weights |
For matching quant levels to hardware, see the model guide and hardware reference.