Model Selection
Learning Outcomes
- Identify the major open-weight model families and who maintains them
- Decode parameter counts (7B, 13B, 70B) and what they cost in memory
- Choose a quantization level (Q4_K_M, Q5_K_M, Q8_0) for your hardware
- Match model families to tasks — coding, reasoning, multilingual, vision
- Benchmark candidate models on standardised prompts and pick a default
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why model choice beats tuning |
| Explain | 7 min | The open-weight families |
| Explain | 6 min | Parameter counts and their cost |
| Demo | 8 min | Quantization — reading GGUF tags |
| Demo | 6 min | Pulling and tagging variants |
| Explain | 5 min | Task fit — picking by job |
| Demo | 4 min | A benchmark harness |
| Wrap-up | 1 min | Takeaways, next lesson |
Before You Begin
Pre-work:
- Finish Lesson 2: Hardware Requirements so you know your RAM and VRAM ceiling
- Finish Lesson 3: Ollama and confirm
ollama --versionworks - Optionally skim the Model Selection Cheat Sheet
Shopping List:
- Ollama installed and running
- 30-50 GB of free disk for pulling a few model variants
- A terminal with
curlandjq(for the benchmark step)
"Open-weight" means the trained parameters are published for download — you run them offline, no API. A handful of families dominate what you run locally:
| Family | Maintainer | Strong at |
|---|---|---|
| Llama | Meta | General chat, huge fine-tune ecosystem |
| Mistral / Mixtral | Mistral AI | Efficiency; Mixtral is Mixture-of-Experts |
| Phi | Microsoft | Reasoning above its parameter count |
| Qwen | Alibaba | Coding, multilingual, agentic |
| Gemma | Google DeepMind | Single-GPU efficiency, vision |
| DeepSeek | DeepSeek | Reasoning and math (the "R1" line) |
| gpt-oss | OpenAI | Tool use and reasoning |
Run ollama list to see what you have, then browse the catalogue at ollama.com/library.
The number before the B is the parameter count in billions — a 7B model has roughly 7 billion weights. More parameters means more capability and more memory: footprint is roughly parameters x bytes-per-parameter, set by quantization (Step 3).
Sizing for dense models at a common 4-bit quant:
| Size | Approx. RAM/VRAM at Q4 | Typical home |
|---|---|---|
| 1B-3B | 1-3 GB | Phones, laptops, CPU-only |
| 7B-8B | 5-6 GB | 8 GB GPU, 16 GB Mac — the everyday default |
| 13B-14B | 9-11 GB | 12-16 GB GPU, 32 GB Mac |
| 30B-34B | 20-24 GB | 24 GB GPU, 32-64 GB Mac |
| 70B+ | 40 GB+ | Multi-GPU, 64-128 GB Mac |
These are floors, not targets — leave headroom for the KV cache (the conversation's running memory, which grows with context length).
ollama show llama3.1:8b # what a pulled model weighs
ollama ps # live: resident size + CPU/GPU split
Quantization stores each weight in fewer bits. Full precision is 16-bit (FP16); 4-bit cuts the footprint roughly 4x at a small, usually acceptable quality cost. The GGUF format used by llama.cpp and Ollama encodes this in the tag, e.g. Q4_K_M.
How to read a GGUF quant tag:
- The number is bits per weight:
Q4= 4-bit,Q5= 5-bit,Q8= 8-bit. _Kis a "k-quant" — mixed-precision that spends more bits on the tensors that matter most.- The trailing letter is the size class:
_S(small),_M(medium),_L(large).
Approximate size for a Llama-3.1-8B model:
| Quant | Size | When to use |
|---|---|---|
| Q8_0 | ~8.0 GB | Near-lossless; max fidelity when memory allows |
| Q6_K | ~6.1 GB | Very high quality with real savings |
| Q5_K_M | ~5.3 GB | High quality; safe step down |
| Q4_K_M | ~4.6 GB | The recommended default sweet spot |
| Q3_K_M | ~3.5 GB | Visible degradation; tight memory only |
| Q2_K | ~3.0 GB | Last resort to make it run at all |
Heuristic: start at Q4_K_M; step up to Q5_K_M or Q6_K with headroom; go below Q4 only when forced.
ollama pull qwen3:8b # default quant (typically a Q4 k-quant)
ollama pull qwen3:8b-q5_K_M # request a higher-fidelity build
To compare models fairly, pull several side by side. Ollama keeps each tag as a separate entry, so you can hold several sizes and quants at once.
ollama pull llama3.1:8b
ollama pull qwen3:8b
ollama pull phi4:14b
ollama list # confirm what's resident, with sizes
Storage lives in one directory you can relocate via the OLLAMA_MODELS env var — handy when models are 4-8 GB each.
Models live under ~/.ollama/models. To move them to an external drive, set the variable, then restart Ollama (Linux is identical):
export OLLAMA_MODELS=/Volumes/FastSSD/ollama-models
Models default to %USERPROFILE%\.ollama\models. Set a user environment variable named OLLAMA_MODELS, then restart Ollama from the tray icon.
A comparison sweep eats disk fast — clean up with ollama rm <tag> as you rule variants out.
qwen3:8b can point at a refreshed build later. For anything you benchmark or ship, record the exact tag and digest from ollama show.Families have personalities — durable tendencies, not guarantees. Verify on your prompts, but this is a sane starting map:
| Task | First choice |
|---|---|
| Code generation / agentic coding | Qwen (Coder variants) |
| Reasoning at small size | Phi |
| Math / step-by-step reasoning | DeepSeek (R1 line) |
| Multilingual | Qwen, Gemma |
| Vision / image + text | Gemma (multimodal builds) |
| Broad, well-supported default | Llama |
| Efficiency on modest hardware | Mistral |
Two more axes matter:
- Instruct vs base. A
-instruct(or-it) build follows instructions; a base build just continues text and frustrates you in a chat loop. Always pick instruct for interactive use. - License. Open weights still ship a license — some permit commercial use freely, some restrict it. Read the model card before building on it.
# Reasoning models "think" before answering — give them room
ollama run deepseek-r1:8b "A bat and ball cost 1.10; the bat costs 1.00 more. How much is the ball?"
Reputation gets you a shortlist; your own prompts pick the winner. Put your real prompts in prompts.txt and run every candidate against every one. The --verbose flag prints the eval rate (tokens/second), so you read quality and speed at once:
for m in llama3.1:8b qwen3:8b phi4:14b; do
echo "===== $m ====="
while IFS= read -r p; do
ollama run "$m" "$p" --verbose
done < prompts.txt
done
For structured scoring, hit Ollama's local API (port 11434) for the metrics:
curl -s http://localhost:11434/api/generate -d '{
"model": "qwen3:8b",
"prompt": "Write a haiku about local inference.",
"stream": false
}' | jq '{eval_count, tps: (.eval_count / (.eval_duration/1e9))}'
Score on correctness, format, and tokens/second; re-run when a new release lands.
Questions & Answers
-instruct or -it variant for interactive or tool-driven work. Default Ollama tags are usually instruct-tuned, but check with ollama show.Key Takeaways
- Pick by family, then slot. Family by task (Qwen for code, Phi for small-size reasoning, Gemma for vision, Llama as a default), then a size that fits, then a quant for headroom.
- Parameters are a memory budget. 7B-8B is the everyday sweet spot; fit in VRAM with KV-cache room before going bigger.
- Q4_K_M is the default quant. Step up to Q5_K_M or Q6_K with memory to spare; go below Q4 only when nothing else fits.
- Instruct, not base. Use the instruction-tuned variant for chat and tools; check the license first.
- Benchmark on your own prompts. A short harness over real tasks beats any leaderboard.
- Pin and record. Lock the exact tag and digest so fast-moving releases never break reproducibility.
Next Steps: Lesson 5: Running Models from the CLI