Ollama — The Easiest Way to Start

40 min beginner Lesson 3

Learning Outcomes

  • Install Ollama on macOS, Linux, or Windows and confirm the server is running
  • Pull and run your first model with a single command
  • Drive the interactive chat REPL using its built-in slash commands
  • Inspect what Ollama stores on disk and how it allocates memory
  • Customise a model with a Modelfile and serve it over the local HTTP API

Lesson Plan

Segment Duration Topic
Intro 3 min What Ollama is and why we start here
Install 6 min Installing on macOS, Linux, Windows
Demo 1 6 min Pulling and running your first model
Explain 6 min What happens behind the scenes
Demo 2 6 min The interactive REPL and slash commands
Demo 3 6 min Managing models on disk
Demo 4 5 min Customising with a Modelfile + the HTTP API
Wrap-up 2 min Key takeaways and next lesson

Before You Begin

Pre-work:

Shopping List:

  • macOS 14 Sonoma or later, a modern Linux distro, or Windows 10/11.
  • At least 8GB free disk space (models range from ~2GB to 40GB+).
  • An internet connection for the initial model download.
  • Roughly 5GB of free RAM/VRAM for the small model we pull first.

1 Install Ollama

Ollama is a single binary bundling a model runtime (built on llama.cpp), a model library, and a local HTTP server.

Download the app from the website, or install from the command line. The one-line script also covers Linux:

curl -fsSL https://ollama.com/install.sh | sh

macOS requires version 14 Sonoma or later; Ollama runs as a menu-bar app that starts the background server. Linux: the same command registers a systemd service — manage it with systemctl status ollama and journalctl -u ollama -f.

Download OllamaSetup.exe from the download page and run it. Ollama runs natively on Windows 10/11 with a background service and tray icon — no WSL required. Use the same ollama commands from PowerShell or Windows Terminal.

Confirm the install from any terminal:

ollama --version

Then confirm the local server is reachable. Ollama listens on port 11434 by default:

curl http://localhost:11434/api/version
TIP
Inspect before you pipe
The curl ... | sh pattern runs remote code as your user. To inspect it first, download install.sh, read it, then run it — good hygiene for any one-line installer.

2 Pull and Run Your First Model

The fastest path is ollama run. If the model is not on disk, Ollama downloads it, then drops you into a chat prompt. Start with a small, capable model:

ollama run llama3.2

A progress bar shows the download, then a prompt appears. Type a message and press Enter:

>>> Explain what a context window is in one sentence.

To download without chatting (useful in scripts or CI), use pull:

ollama pull qwen3:8b

Names follow a name:tag convention; the tag usually encodes size and/or quantization — storing weights at lower precision (e.g. 4-bit vs 16-bit) to shrink memory use, at a small accuracy cost. Choosing tags is covered in Lesson 4: Model Selection.

Action Command When to use
Download + chat ollama run NAME Interactive use
Download only ollama pull NAME Scripts, pre-caching
Specific tag ollama run NAME:TAG Pin a size/quant
WARNING
Disk and bandwidth
An 8b (8 billion parameter) model at 4-bit quantization is roughly 4.5GB; 70b tags can exceed 40GB. Check the size on the library page before pulling on a metered connection. The first prompt after a pull is slower because weights load from disk; later prompts reuse the loaded model.

3 What Happens Behind the Scenes

When you run a model, three things happen:

  1. Resolve and download. Ollama checks its store for the model+tag. If missing, it downloads the GGUF weights (the format llama.cpp uses) plus a manifest with the template and defaults.
  2. Load into memory. The weights map into RAM. With a supported GPU — NVIDIA CUDA, AMD ROCm, or Apple's Metal on Apple Silicon — Ollama offloads as many layers as fit into VRAM/unified memory.
  3. Serve. A background process (ollama serve) holds the loaded model and answers HTTP requests on port 11434.

See what is loaded and where (CPU vs GPU):

ollama ps
NAME           ID            SIZE      PROCESSOR    UNTIL
llama3.2:latest a1b2c3d4e5f6  6.7 GB    100% GPU     4 minutes from now

A loaded model is unloaded automatically after a few minutes of inactivity to free memory. The UNTIL column shows when.

NOTE
Quantization in one line
The SIZE in ollama ps is the in-memory footprint, close to the on-disk size for the chosen quantization. If a model will not fit fully in GPU memory, Ollama splits it: some layers run on the GPU, the rest on CPU. Partial offload still works — it is just slower.

4 Driving the Interactive REPL

Inside an ollama run session you get a REPL with built-in slash commands for managing the session without quitting:

Command What it does
/? or /help Show available commands
/set Set session parameters (e.g. /set parameter temperature 0.2)
/show Show model details (/show info, /show modelfile)
/load MODEL Switch to another model in the same session
/save MODEL Save the current session as a new named model
/clear Clear the conversation context
/bye Exit the session

Keyboard shortcuts mirror common shells:

Shortcut Action
Ctrl+C Stop the model mid-response
Ctrl+D Exit the session (same as /bye)
Ctrl+L Clear the screen

For multi-line prompts (pasting a code block, say), wrap them in triple quotes ("""). You can also tune behaviour live — lowering temperature makes output more deterministic:

>>> /set parameter temperature 0.1
>>> /show info
TIP
Temperature, briefly
Temperature controls randomness in token selection. Near 0 the model picks the most likely next token nearly every time (good for code and extraction); near 1 it samples more freely (good for brainstorming).

5 Managing Models on Disk

Pull a second model so you have at least two locally, then list what you have:

ollama pull gemma3:4b
ollama list
NAME            ID            SIZE      MODIFIED
llama3.2:latest a1b2c3d4e5f6  2.0 GB    10 minutes ago
qwen3:8b        f6e5d4c3b2a1  5.2 GB    8 minutes ago
gemma3:4b       001122334455  3.3 GB    1 minute ago

Models, manifests, and blobs live in one store. On macOS and a Linux user-install that is ~/.ollama/models (a Linux service install uses /usr/share/ollama/.ollama/models); on Windows, C:\Users\<you>\.ollama\models. Inspect its size, then free space by removing unused models:

du -sh ~/.ollama/models
ollama rm gemma3:4b

Relocate the store by setting the OLLAMA_MODELS environment variable before starting the server — handy when your home volume is small but you have a large secondary drive.

WARNING
Quantization is per-tag, not per-name
Pulling llama3.2 and llama3.2:1b downloads two separate model files. They share nothing on disk. Run ollama list periodically — it is easy to accumulate tens of gigabytes of forgotten variants.

6 Customise a Model with a Modelfile

A Modelfile is a small text file that layers your defaults onto a base model — a system prompt and sampling parameters. It is the local equivalent of a "custom GPT," but the result is a real, named model on your machine. Create a file named Modelfile (no extension):

FROM llama3.2

# Lower temperature for consistent, factual answers
PARAMETER temperature 0.2
PARAMETER num_ctx 8192

SYSTEM """
You are a terse senior engineer. Answer in at most three sentences.
Prefer concrete commands and code over prose. If unsure, say so.
"""

The instructions: FROM (required base model), PARAMETER (runtime settings like temperature, num_ctx for context size, top_p, repeat_penalty), SYSTEM (persistent system prompt), plus TEMPLATE, ADAPTER (for LoRA fine-tunes — see Lesson 7), and MESSAGE for few-shot examples.

Build and run your custom model:

ollama create terse-eng -f ./Modelfile
ollama run terse-eng "How do I list the largest files in a directory?"

To see how an existing model is configured, dump its Modelfile:

ollama show --modelfile llama3.2

Finally, the same model is callable over the HTTP API — the bridge that lets local models plug into editors and tools later. The default base URL is http://localhost:11434:

curl http://localhost:11434/api/chat -d '{
  "model": "terse-eng",
  "messages": [{ "role": "user", "content": "One-liner to count lines in all .py files?" }],
  "stream": false
}'
NOTE
This API is the bridge
Every GUI and integration in Lesson 6 talks to this same endpoint. Ollama also exposes an OpenAI-compatible path, so tools that expect the OpenAI format can point at your local server with no code changes. Tip: check Modelfiles into git so a teammate can reproduce your exact local assistant with one ollama create.

Questions & Answers

Q: I have 16GB of RAM. Will pulling a 70B model just crash my machine?
It will not run usefully. A 70B model at 4-bit quantization needs ~40GB+; on a 16GB machine Ollama falls back to disk-backed swap, making generation painfully slow or triggering an out-of-memory failure. Stick to 1B-8B models on 16GB. See the [Hardware Reference](/courses/06-local-llms/supplemental/hardware-reference/) for a sizing table.
Q: Does Ollama send my prompts anywhere?
Inference runs entirely on your machine — prompts and responses do not leave it. The only network traffic is downloading weights when you pull, plus optional update checks. That privacy guarantee is the whole point of running local ([Lesson 1](/courses/06-local-llms/lesson-01/)).
Q: Is port 11434 exposed to my network or the internet?
By default the server binds to localhost only, so other machines cannot reach it. If you change the bind address (via OLLAMA_HOST) to serve other devices, you must firewall it yourself — the API has no authentication by default. Exposing the server safely is covered in [Lesson 10: Production Deployment](/courses/06-local-llms/lesson-10/).
Q: My GPU is not being used — generation is slow. What now?
Run ollama ps and check the PROCESSOR column. 100% CPU on a GPU machine usually means the GPU driver or CUDA/ROCm runtime is not detected, or the model is too large for VRAM and fell back to CPU. Try a smaller model to confirm, then work through the [Troubleshooting](/courses/06-local-llms/supplemental/troubleshooting/) guide.
Q: How is Ollama different from running llama.cpp directly?
Ollama is a wrapper around the same engine — it manages downloads, storage, templates, and the server for you. Running llama.cpp directly gives finer control over flags and quantization, which we use in [Lesson 5](/courses/06-local-llms/lesson-05/) and [Lesson 9](/courses/06-local-llms/lesson-09/). For getting started, Ollama is the lower-friction choice.

Key Takeaways

  1. One install, one command to run. The curl ... | sh installer (macOS/Linux) or Windows installer gives you a binary; ollama run NAME downloads and chats in one step.
  2. The server is always there. Ollama runs a background server on http://localhost:11434; ollama ps shows what is loaded and whether it is on CPU or GPU.
  3. The REPL is a workspace. Slash commands like /set parameter, /show, /clear, and /bye tune and manage a session without restarting.
  4. Models are files you manage. ollama list, ollama rm, and the ~/.ollama/models store keep disk usage in check — each tag is a separate download.
  5. Modelfiles make models yours. A few lines of FROM, PARAMETER, and SYSTEM produce a reusable, named, version-controllable assistant.
  6. The HTTP API is the integration point. The model you chat with is callable over /api/chat — how every tool in later lessons connects.

Next Steps: Lesson 4: Model Selection