Model Selection Guide

Reference intermediate

A quick-reference grid for choosing an open-weight model and quantization level. Names and sizes move fast — always confirm the exact tag against ollama.com/library before pulling. Pair this with the Hardware Reference and the Troubleshooting guide.

Model Families at a Glance

These families dominate the local landscape (the same set Lesson 4 introduces). "Parameters" (the learned weights, counted in billions, e.g. 8B) drive both quality and memory. "MoE" means Mixture-of-Experts — only a fraction of the total weights activate per token, so a large model runs cheaper than its headline size suggests.

Family Vendor Typical sizes (Ollama) Strong at Notes
Llama Meta 1B, 3B, 8B, 70B, 405B; Llama 4 is MoE General chat, instruction following, broad ecosystem Largest fine-tune/tooling community; small variants good for edge
Mistral Mistral AI 7B, Nemo 12B, Small 22–24B, Large 123B; Ministral 3–14B for edge Efficiency per parameter, coding, long context (Nemo ~128k) Apache-2.0 base 7B; strong quality-to-size ratio
Phi Microsoft Mini ~3.8B, Medium/Phi-4 ~14B Reasoning and math at small sizes, multilingual (Phi-4-mini) Punches above its weight; great for 8–16GB machines
Qwen Alibaba 0.5B–72B (Qwen2.5); 0.6B–235B with MoE (Qwen3) Coding, agentic workflows, multilingual, long context Wide size ladder; coder-tuned variants for repo-level work
Gemma Google DeepMind 270M–27B (Gemma 3); newer gens add MoE + multimodal Multimodal (vision), efficient single-GPU use, tool calling "Runs on a single GPU" design point; vision in recent gens
DeepSeek DeepSeek R1 1.5B–70B distills; V3 (large MoE) Step-by-step reasoning and math (the R1 line) R1 "distill" sizes bring reasoning to consumer HW and show their working
gpt-oss OpenAI 20B, 120B (MoE) Tool use and agentic reasoning OpenAI's open-weight release; strong tool-calling

Sizes above reflect what is published in the Ollama library; vendors release new generations frequently, so treat the ladder (small / mid / large) as the durable part and the exact version as perishable.

Pick a Family by Task

Task First reach for Why
General assistant / chat Llama (8B) or Qwen (7–14B) Balanced, well-supported, plentiful fine-tunes
Coding / repo work Qwen coder variant, Mistral, Llama Coder-tuned Qwen and Mistral excel at code generation
Reasoning / math DeepSeek (R1 line), Phi (mini/medium) DeepSeek R1 distills show their working; Phi reasons above its size at 4–14B
Multilingual Qwen, Phi-4-mini Broad language coverage
Vision / multimodal Gemma, Llama (vision variants) Native image input
Edge / very low RAM Ministral, Llama 1–3B, Gemma small Built for constrained devices

Parameter Size vs. What It's For

Size band Examples Good for Rough RAM (Q4)
~1–3B Llama 3B, Gemma small, Ministral 3B Autocomplete, classification, edge, fast drafts ~2–3 GB
~7–9B Llama 8B, Mistral 7B, Qwen 7B, Gemma 9B The everyday sweet spot for laptops/single GPU ~5–6 GB
~12–27B Nemo 12B, Phi-4 14B, Mistral Small 24B, Gemma 27B Noticeably better reasoning; needs more VRAM/RAM ~8–18 GB
~70B+ Llama 70B, Mistral Large 123B Near-frontier quality; workstation/multi-GPU territory ~40 GB+

Rule of thumb: a dense model at 4-bit needs roughly (params in billions) × 0.6 GB of memory for weights, plus headroom for the KV cache (grows with context length).

Quantization: Quality vs. Memory

Quantization stores weights at lower precision to shrink the model. The _K_M suffix is a llama.cpp "K-quant" that keeps important layers at higher precision. Figures below are measured on Llama-3.1-8B per the llama.cpp quantize docs.

Quant Bits/weight Size (8B model) Quality When to use
Q4_K_M ~4.89 ~4.6 GiB Good; small perplexity rise Default — best size/quality balance
Q5_K_M ~5.70 ~5.3 GiB Very close to full precision When you have spare memory
Q8_0 ~8.50 ~8.0 GiB Essentially lossless Quality-critical work, small models

Lower than Q4 (Q3, Q2) saves memory but degrades quality faster — usable for very large models you otherwise couldn't fit, less so for small ones. Scale the size estimate linearly for other parameter counts (a 70B at Q4 ≈ 70/8 × 4.6 ≈ 40 GB).

# Inspect a model's quant/size in Ollama
ollama show llama3.1 --modelfile | grep -i quant
ollama list                 # shows on-disk size per tag

# Pull a specific quantization tag
ollama pull llama3.1:8b-instruct-q4_K_M
ollama pull qwen3:8b
# Quantize your own GGUF with llama.cpp
./llama-quantize ./model-f16.gguf ./model-q4_k_m.gguf Q4_K_M
./llama-quantize ./model-f16.gguf ./model-q8_0.gguf   Q8_0

What Fits in N GB?

Unified-memory machines (Apple Silicon) and dedicated VRAM behave similarly here — both must hold weights plus KV cache. Leave a few GB free for the OS. These are conservative Q4_K_M targets.

Memory Comfortable Tight / short context only
8 GB 1–3B models; one 7B at a time 7–8B with reduced context
16 GB 7–8B at Q4/Q5 with room to spare 12–14B (Phi-4, Nemo)
24 GB 12–14B comfortably; 24B at Q4 ~27–32B at low quant
32 GB 24–27B at Q4/Q5 ~70B only at aggressive quant
48–64 GB 70B-class at Q4 larger MoE models
96 GB+ 70B at Q5/Q8; large MoE 100B+ dense
# Quick capability check before pulling
ollama ps                   # what's loaded and its memory footprint
free -h                     # Linux: available RAM
vm_stat | head              # macOS: page stats (unified memory)

Context length matters: a long context (32k+) can add several GB of KV cache on top of the weights. If a model loads but OOMs mid-chat, lower num_ctx (Ollama) or -c (llama.cpp) before dropping to a smaller model.

Choosing Quickly — A Decision Path

  1. Fit first. Find your memory band above; that caps the parameter size.
  2. Pick the family for your task from the task table.
  3. Default to Q4_K_M. Move up to Q5/Q8 only if memory allows and quality matters.
  4. Test on your own prompts — benchmarks rarely match your workload.
# Side-by-side smoke test on a fixed prompt
for m in llama3.1:8b qwen3:8b mistral:7b phi4:14b; do
  echo "=== $m ==="
  ollama run "$m" "Summarise the trade-offs of LoRA vs full fine-tuning in 3 bullets."
done

See Also