Model Selection Guide
A quick-reference grid for choosing an open-weight model and quantization level. Names and sizes move fast — always confirm the exact tag against ollama.com/library before pulling. Pair this with the Hardware Reference and the Troubleshooting guide.
Model Families at a Glance
These families dominate the local landscape (the same set Lesson 4 introduces). "Parameters" (the learned weights, counted in billions, e.g. 8B) drive both quality and memory. "MoE" means Mixture-of-Experts — only a fraction of the total weights activate per token, so a large model runs cheaper than its headline size suggests.
| Family | Vendor | Typical sizes (Ollama) | Strong at | Notes |
|---|---|---|---|---|
| Llama | Meta | 1B, 3B, 8B, 70B, 405B; Llama 4 is MoE | General chat, instruction following, broad ecosystem | Largest fine-tune/tooling community; small variants good for edge |
| Mistral | Mistral AI | 7B, Nemo 12B, Small 22–24B, Large 123B; Ministral 3–14B for edge | Efficiency per parameter, coding, long context (Nemo ~128k) | Apache-2.0 base 7B; strong quality-to-size ratio |
| Phi | Microsoft | Mini ~3.8B, Medium/Phi-4 ~14B | Reasoning and math at small sizes, multilingual (Phi-4-mini) | Punches above its weight; great for 8–16GB machines |
| Qwen | Alibaba | 0.5B–72B (Qwen2.5); 0.6B–235B with MoE (Qwen3) | Coding, agentic workflows, multilingual, long context | Wide size ladder; coder-tuned variants for repo-level work |
| Gemma | Google DeepMind | 270M–27B (Gemma 3); newer gens add MoE + multimodal | Multimodal (vision), efficient single-GPU use, tool calling | "Runs on a single GPU" design point; vision in recent gens |
| DeepSeek | DeepSeek | R1 1.5B–70B distills; V3 (large MoE) | Step-by-step reasoning and math (the R1 line) | R1 "distill" sizes bring reasoning to consumer HW and show their working |
| gpt-oss | OpenAI | 20B, 120B (MoE) | Tool use and agentic reasoning | OpenAI's open-weight release; strong tool-calling |
Sizes above reflect what is published in the Ollama library; vendors release new generations frequently, so treat the ladder (small / mid / large) as the durable part and the exact version as perishable.
Pick a Family by Task
| Task | First reach for | Why |
|---|---|---|
| General assistant / chat | Llama (8B) or Qwen (7–14B) | Balanced, well-supported, plentiful fine-tunes |
| Coding / repo work | Qwen coder variant, Mistral, Llama | Coder-tuned Qwen and Mistral excel at code generation |
| Reasoning / math | DeepSeek (R1 line), Phi (mini/medium) | DeepSeek R1 distills show their working; Phi reasons above its size at 4–14B |
| Multilingual | Qwen, Phi-4-mini | Broad language coverage |
| Vision / multimodal | Gemma, Llama (vision variants) | Native image input |
| Edge / very low RAM | Ministral, Llama 1–3B, Gemma small | Built for constrained devices |
Parameter Size vs. What It's For
| Size band | Examples | Good for | Rough RAM (Q4) |
|---|---|---|---|
| ~1–3B | Llama 3B, Gemma small, Ministral 3B | Autocomplete, classification, edge, fast drafts | ~2–3 GB |
| ~7–9B | Llama 8B, Mistral 7B, Qwen 7B, Gemma 9B | The everyday sweet spot for laptops/single GPU | ~5–6 GB |
| ~12–27B | Nemo 12B, Phi-4 14B, Mistral Small 24B, Gemma 27B | Noticeably better reasoning; needs more VRAM/RAM | ~8–18 GB |
| ~70B+ | Llama 70B, Mistral Large 123B | Near-frontier quality; workstation/multi-GPU territory | ~40 GB+ |
Rule of thumb: a dense model at 4-bit needs roughly (params in billions) × 0.6 GB of memory for weights, plus headroom for the KV cache (grows with context length).
Quantization: Quality vs. Memory
Quantization stores weights at lower precision to shrink the model. The _K_M suffix is a llama.cpp "K-quant" that keeps important layers at higher precision. Figures below are measured on Llama-3.1-8B per the llama.cpp quantize docs.
| Quant | Bits/weight | Size (8B model) | Quality | When to use |
|---|---|---|---|---|
| Q4_K_M | ~4.89 | ~4.6 GiB | Good; small perplexity rise | Default — best size/quality balance |
| Q5_K_M | ~5.70 | ~5.3 GiB | Very close to full precision | When you have spare memory |
| Q8_0 | ~8.50 | ~8.0 GiB | Essentially lossless | Quality-critical work, small models |
Lower than Q4 (Q3, Q2) saves memory but degrades quality faster — usable for very large models you otherwise couldn't fit, less so for small ones. Scale the size estimate linearly for other parameter counts (a 70B at Q4 ≈ 70/8 × 4.6 ≈ 40 GB).
# Inspect a model's quant/size in Ollama
ollama show llama3.1 --modelfile | grep -i quant
ollama list # shows on-disk size per tag
# Pull a specific quantization tag
ollama pull llama3.1:8b-instruct-q4_K_M
ollama pull qwen3:8b
# Quantize your own GGUF with llama.cpp
./llama-quantize ./model-f16.gguf ./model-q4_k_m.gguf Q4_K_M
./llama-quantize ./model-f16.gguf ./model-q8_0.gguf Q8_0
What Fits in N GB?
Unified-memory machines (Apple Silicon) and dedicated VRAM behave similarly here — both must hold weights plus KV cache. Leave a few GB free for the OS. These are conservative Q4_K_M targets.
| Memory | Comfortable | Tight / short context only |
|---|---|---|
| 8 GB | 1–3B models; one 7B at a time | 7–8B with reduced context |
| 16 GB | 7–8B at Q4/Q5 with room to spare | 12–14B (Phi-4, Nemo) |
| 24 GB | 12–14B comfortably; 24B at Q4 | ~27–32B at low quant |
| 32 GB | 24–27B at Q4/Q5 | ~70B only at aggressive quant |
| 48–64 GB | 70B-class at Q4 | larger MoE models |
| 96 GB+ | 70B at Q5/Q8; large MoE | 100B+ dense |
# Quick capability check before pulling
ollama ps # what's loaded and its memory footprint
free -h # Linux: available RAM
vm_stat | head # macOS: page stats (unified memory)
Context length matters: a long context (32k+) can add several GB of KV cache on top of the weights. If a model loads but OOMs mid-chat, lower
num_ctx(Ollama) or-c(llama.cpp) before dropping to a smaller model.
Choosing Quickly — A Decision Path
- Fit first. Find your memory band above; that caps the parameter size.
- Pick the family for your task from the task table.
- Default to Q4_K_M. Move up to Q5/Q8 only if memory allows and quality matters.
- Test on your own prompts — benchmarks rarely match your workload.
# Side-by-side smoke test on a fixed prompt
for m in llama3.1:8b qwen3:8b mistral:7b phi4:14b; do
echo "=== $m ==="
ollama run "$m" "Summarise the trade-offs of LoRA vs full fine-tuning in 3 bullets."
done
See Also
- Hardware Reference — GPU/Apple Silicon sizing
- CLI & Config Quick Reference — Ollama, llama.cpp, vLLM commands
- Troubleshooting — OOM, slow generation, downloads
- Rules & Best Practices — model management and responsible use