Ollama — The Easiest Way to Start
Learning Outcomes
- Install Ollama on macOS, Linux, or Windows and confirm the server is running
- Pull and run your first model with a single command
- Drive the interactive chat REPL using its built-in slash commands
- Inspect what Ollama stores on disk and how it allocates memory
- Customise a model with a Modelfile and serve it over the local HTTP API
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | What Ollama is and why we start here |
| Install | 6 min | Installing on macOS, Linux, Windows |
| Demo 1 | 6 min | Pulling and running your first model |
| Explain | 6 min | What happens behind the scenes |
| Demo 2 | 6 min | The interactive REPL and slash commands |
| Demo 3 | 6 min | Managing models on disk |
| Demo 4 | 5 min | Customising with a Modelfile + the HTTP API |
| Wrap-up | 2 min | Key takeaways and next lesson |
Before You Begin
Pre-work:
- Read Lesson 1: Why Run Local? for the trade-offs.
- Complete Lesson 2: Hardware Requirements and confirm at least 16GB of RAM (or unified memory on Apple Silicon).
- Open a terminal you are comfortable in.
Shopping List:
- macOS 14 Sonoma or later, a modern Linux distro, or Windows 10/11.
- At least 8GB free disk space (models range from ~2GB to 40GB+).
- An internet connection for the initial model download.
- Roughly 5GB of free RAM/VRAM for the small model we pull first.
Ollama is a single binary bundling a model runtime (built on llama.cpp), a model library, and a local HTTP server.
Download the app from the website, or install from the command line. The one-line script also covers Linux:
curl -fsSL https://ollama.com/install.sh | sh
macOS requires version 14 Sonoma or later; Ollama runs as a menu-bar app that starts the background server. Linux: the same command registers a systemd service — manage it with systemctl status ollama and journalctl -u ollama -f.
Download OllamaSetup.exe from the download page and run it. Ollama runs natively on Windows 10/11 with a background service and tray icon — no WSL required. Use the same ollama commands from PowerShell or Windows Terminal.
Confirm the install from any terminal:
ollama --version
Then confirm the local server is reachable. Ollama listens on port 11434 by default:
curl http://localhost:11434/api/version
curl ... | sh pattern runs remote code as your user. To inspect it first, download install.sh, read it, then run it — good hygiene for any one-line installer.The fastest path is ollama run. If the model is not on disk, Ollama downloads it, then drops you into a chat prompt. Start with a small, capable model:
ollama run llama3.2
A progress bar shows the download, then a prompt appears. Type a message and press Enter:
>>> Explain what a context window is in one sentence.
To download without chatting (useful in scripts or CI), use pull:
ollama pull qwen3:8b
Names follow a name:tag convention; the tag usually encodes size and/or quantization — storing weights at lower precision (e.g. 4-bit vs 16-bit) to shrink memory use, at a small accuracy cost. Choosing tags is covered in Lesson 4: Model Selection.
| Action | Command | When to use |
|---|---|---|
| Download + chat | ollama run NAME |
Interactive use |
| Download only | ollama pull NAME |
Scripts, pre-caching |
| Specific tag | ollama run NAME:TAG |
Pin a size/quant |
8b (8 billion parameter) model at 4-bit quantization is roughly 4.5GB; 70b tags can exceed 40GB. Check the size on the library page before pulling on a metered connection. The first prompt after a pull is slower because weights load from disk; later prompts reuse the loaded model.When you run a model, three things happen:
- Resolve and download. Ollama checks its store for the model+tag. If missing, it downloads the GGUF weights (the format llama.cpp uses) plus a manifest with the template and defaults.
- Load into memory. The weights map into RAM. With a supported GPU — NVIDIA CUDA, AMD ROCm, or Apple's Metal on Apple Silicon — Ollama offloads as many layers as fit into VRAM/unified memory.
- Serve. A background process (
ollama serve) holds the loaded model and answers HTTP requests on port 11434.
See what is loaded and where (CPU vs GPU):
ollama ps
NAME ID SIZE PROCESSOR UNTIL
llama3.2:latest a1b2c3d4e5f6 6.7 GB 100% GPU 4 minutes from now
A loaded model is unloaded automatically after a few minutes of inactivity to free memory. The UNTIL column shows when.
SIZE in ollama ps is the in-memory footprint, close to the on-disk size for the chosen quantization. If a model will not fit fully in GPU memory, Ollama splits it: some layers run on the GPU, the rest on CPU. Partial offload still works — it is just slower.Inside an ollama run session you get a REPL with built-in slash commands for managing the session without quitting:
| Command | What it does |
|---|---|
/? or /help |
Show available commands |
/set |
Set session parameters (e.g. /set parameter temperature 0.2) |
/show |
Show model details (/show info, /show modelfile) |
/load MODEL |
Switch to another model in the same session |
/save MODEL |
Save the current session as a new named model |
/clear |
Clear the conversation context |
/bye |
Exit the session |
Keyboard shortcuts mirror common shells:
| Shortcut | Action |
|---|---|
Ctrl+C |
Stop the model mid-response |
Ctrl+D |
Exit the session (same as /bye) |
Ctrl+L |
Clear the screen |
For multi-line prompts (pasting a code block, say), wrap them in triple quotes ("""). You can also tune behaviour live — lowering temperature makes output more deterministic:
>>> /set parameter temperature 0.1
>>> /show info
0 the model picks the most likely next token nearly every time (good for code and extraction); near 1 it samples more freely (good for brainstorming).Pull a second model so you have at least two locally, then list what you have:
ollama pull gemma3:4b
ollama list
NAME ID SIZE MODIFIED
llama3.2:latest a1b2c3d4e5f6 2.0 GB 10 minutes ago
qwen3:8b f6e5d4c3b2a1 5.2 GB 8 minutes ago
gemma3:4b 001122334455 3.3 GB 1 minute ago
Models, manifests, and blobs live in one store. On macOS and a Linux user-install that is ~/.ollama/models (a Linux service install uses /usr/share/ollama/.ollama/models); on Windows, C:\Users\<you>\.ollama\models. Inspect its size, then free space by removing unused models:
du -sh ~/.ollama/models
ollama rm gemma3:4b
Relocate the store by setting the OLLAMA_MODELS environment variable before starting the server — handy when your home volume is small but you have a large secondary drive.
llama3.2 and llama3.2:1b downloads two separate model files. They share nothing on disk. Run ollama list periodically — it is easy to accumulate tens of gigabytes of forgotten variants.A Modelfile is a small text file that layers your defaults onto a base model — a system prompt and sampling parameters. It is the local equivalent of a "custom GPT," but the result is a real, named model on your machine. Create a file named Modelfile (no extension):
FROM llama3.2
# Lower temperature for consistent, factual answers
PARAMETER temperature 0.2
PARAMETER num_ctx 8192
SYSTEM """
You are a terse senior engineer. Answer in at most three sentences.
Prefer concrete commands and code over prose. If unsure, say so.
"""
The instructions: FROM (required base model), PARAMETER (runtime settings like temperature, num_ctx for context size, top_p, repeat_penalty), SYSTEM (persistent system prompt), plus TEMPLATE, ADAPTER (for LoRA fine-tunes — see Lesson 7), and MESSAGE for few-shot examples.
Build and run your custom model:
ollama create terse-eng -f ./Modelfile
ollama run terse-eng "How do I list the largest files in a directory?"
To see how an existing model is configured, dump its Modelfile:
ollama show --modelfile llama3.2
Finally, the same model is callable over the HTTP API — the bridge that lets local models plug into editors and tools later. The default base URL is http://localhost:11434:
curl http://localhost:11434/api/chat -d '{
"model": "terse-eng",
"messages": [{ "role": "user", "content": "One-liner to count lines in all .py files?" }],
"stream": false
}'
ollama create.Questions & Answers
OLLAMA_HOST) to serve other devices, you must firewall it yourself — the API has no authentication by default. Exposing the server safely is covered in [Lesson 10: Production Deployment](/courses/06-local-llms/lesson-10/).ollama ps and check the PROCESSOR column. 100% CPU on a GPU machine usually means the GPU driver or CUDA/ROCm runtime is not detected, or the model is too large for VRAM and fell back to CPU. Try a smaller model to confirm, then work through the [Troubleshooting](/courses/06-local-llms/supplemental/troubleshooting/) guide.Key Takeaways
- One install, one command to run. The
curl ... | shinstaller (macOS/Linux) or Windows installer gives you a binary;ollama run NAMEdownloads and chats in one step. - The server is always there. Ollama runs a background server on
http://localhost:11434;ollama psshows what is loaded and whether it is on CPU or GPU. - The REPL is a workspace. Slash commands like
/set parameter,/show,/clear, and/byetune and manage a session without restarting. - Models are files you manage.
ollama list,ollama rm, and the~/.ollama/modelsstore keep disk usage in check — each tag is a separate download. - Modelfiles make models yours. A few lines of
FROM,PARAMETER, andSYSTEMproduce a reusable, named, version-controllable assistant. - The HTTP API is the integration point. The model you chat with is callable over
/api/chat— how every tool in later lessons connects.
Next Steps: Lesson 4: Model Selection