Model Selection

40 min intermediate Lesson 4

Learning Outcomes

  • Identify the major open-weight model families and who maintains them
  • Decode parameter counts (7B, 13B, 70B) and what they cost in memory
  • Choose a quantization level (Q4_K_M, Q5_K_M, Q8_0) for your hardware
  • Match model families to tasks — coding, reasoning, multilingual, vision
  • Benchmark candidate models on standardised prompts and pick a default

Lesson Plan

Segment Duration Topic
Intro 3 min Why model choice beats tuning
Explain 7 min The open-weight families
Explain 6 min Parameter counts and their cost
Demo 8 min Quantization — reading GGUF tags
Demo 6 min Pulling and tagging variants
Explain 5 min Task fit — picking by job
Demo 4 min A benchmark harness
Wrap-up 1 min Takeaways, next lesson

Before You Begin

Pre-work:

Shopping List:

  • Ollama installed and running
  • 30-50 GB of free disk for pulling a few model variants
  • A terminal with curl and jq (for the benchmark step)

1 The Open-Weight Families

"Open-weight" means the trained parameters are published for download — you run them offline, no API. A handful of families dominate what you run locally:

Family Maintainer Strong at
Llama Meta General chat, huge fine-tune ecosystem
Mistral / Mixtral Mistral AI Efficiency; Mixtral is Mixture-of-Experts
Phi Microsoft Reasoning above its parameter count
Qwen Alibaba Coding, multilingual, agentic
Gemma Google DeepMind Single-GPU efficiency, vision
DeepSeek DeepSeek Reasoning and math (the "R1" line)
gpt-oss OpenAI Tool use and reasoning

Run ollama list to see what you have, then browse the catalogue at ollama.com/library.

NOTE
Versions move fast, principles don't
Version numbers change every few months. Pick by family reputation and size class, then grab the newest release in that slot. The family table is durable; any version number you see online is a snapshot.
TIP
MoE in one sentence
A Mixture-of-Experts model has many parameters but routes each token through only a small subset, so a large MoE runs closer to a small dense model's speed — while still needing memory for all the weights.

2 Reading Parameter Counts

The number before the B is the parameter count in billions — a 7B model has roughly 7 billion weights. More parameters means more capability and more memory: footprint is roughly parameters x bytes-per-parameter, set by quantization (Step 3).

Sizing for dense models at a common 4-bit quant:

Size Approx. RAM/VRAM at Q4 Typical home
1B-3B 1-3 GB Phones, laptops, CPU-only
7B-8B 5-6 GB 8 GB GPU, 16 GB Mac — the everyday default
13B-14B 9-11 GB 12-16 GB GPU, 32 GB Mac
30B-34B 20-24 GB 24 GB GPU, 32-64 GB Mac
70B+ 40 GB+ Multi-GPU, 64-128 GB Mac

These are floors, not targets — leave headroom for the KV cache (the conversation's running memory, which grows with context length).

ollama show llama3.1:8b    # what a pulled model weighs
ollama ps                  # live: resident size + CPU/GPU split
WARNING
A bigger model you can't fit is slower, not smarter
If a model spills past VRAM it offloads layers to system RAM and tokens-per-second can drop 5-10x. An 8B that fits entirely in VRAM beats a 70B that thrashes. Size to fit first.

3 Quantization — Q4, Q5, Q8 Decoded

Quantization stores each weight in fewer bits. Full precision is 16-bit (FP16); 4-bit cuts the footprint roughly 4x at a small, usually acceptable quality cost. The GGUF format used by llama.cpp and Ollama encodes this in the tag, e.g. Q4_K_M.

How to read a GGUF quant tag:

  • The number is bits per weight: Q4 = 4-bit, Q5 = 5-bit, Q8 = 8-bit.
  • _K is a "k-quant" — mixed-precision that spends more bits on the tensors that matter most.
  • The trailing letter is the size class: _S (small), _M (medium), _L (large).

Approximate size for a Llama-3.1-8B model:

Quant Size When to use
Q8_0 ~8.0 GB Near-lossless; max fidelity when memory allows
Q6_K ~6.1 GB Very high quality with real savings
Q5_K_M ~5.3 GB High quality; safe step down
Q4_K_M ~4.6 GB The recommended default sweet spot
Q3_K_M ~3.5 GB Visible degradation; tight memory only
Q2_K ~3.0 GB Last resort to make it run at all

Heuristic: start at Q4_K_M; step up to Q5_K_M or Q6_K with headroom; go below Q4 only when forced.

ollama pull qwen3:8b              # default quant (typically a Q4 k-quant)
ollama pull qwen3:8b-q5_K_M       # request a higher-fidelity build
TIP
Quant beats size for fit decisions
For most coding and chat work, a smaller model at higher quant gives steadier output than a bigger one at Q3. Beyond GGUF you'll meet AWQ, GPTQ, and the IQ-series — Lesson 9 covers those; GGUF k-quants are all you need for Ollama today.

4 Pulling and Tagging Variants

To compare models fairly, pull several side by side. Ollama keeps each tag as a separate entry, so you can hold several sizes and quants at once.

ollama pull llama3.1:8b
ollama pull qwen3:8b
ollama pull phi4:14b
ollama list                 # confirm what's resident, with sizes

Storage lives in one directory you can relocate via the OLLAMA_MODELS env var — handy when models are 4-8 GB each.

Models live under ~/.ollama/models. To move them to an external drive, set the variable, then restart Ollama (Linux is identical):

export OLLAMA_MODELS=/Volumes/FastSSD/ollama-models

Models default to %USERPROFILE%\.ollama\models. Set a user environment variable named OLLAMA_MODELS, then restart Ollama from the tray icon.

A comparison sweep eats disk fast — clean up with ollama rm <tag> as you rule variants out.

WARNING
Pin tags for reproducible work
A bare tag like qwen3:8b can point at a refreshed build later. For anything you benchmark or ship, record the exact tag and digest from ollama show.

5 Matching Models to Tasks

Families have personalities — durable tendencies, not guarantees. Verify on your prompts, but this is a sane starting map:

Task First choice
Code generation / agentic coding Qwen (Coder variants)
Reasoning at small size Phi
Math / step-by-step reasoning DeepSeek (R1 line)
Multilingual Qwen, Gemma
Vision / image + text Gemma (multimodal builds)
Broad, well-supported default Llama
Efficiency on modest hardware Mistral

Two more axes matter:

  • Instruct vs base. A -instruct (or -it) build follows instructions; a base build just continues text and frustrates you in a chat loop. Always pick instruct for interactive use.
  • License. Open weights still ship a license — some permit commercial use freely, some restrict it. Read the model card before building on it.
# Reasoning models "think" before answering — give them room
ollama run deepseek-r1:8b "A bat and ball cost 1.10; the bat costs 1.00 more. How much is the ball?"
TIP
Pick a family, then a slot
Family by task (above), size class by what fits (Step 2), quant by headroom (Step 3). Three decisions in that order gets you a defensible model nearly every time.

6 A Repeatable Benchmark Harness

Reputation gets you a shortlist; your own prompts pick the winner. Put your real prompts in prompts.txt and run every candidate against every one. The --verbose flag prints the eval rate (tokens/second), so you read quality and speed at once:

for m in llama3.1:8b qwen3:8b phi4:14b; do
  echo "===== $m ====="
  while IFS= read -r p; do
    ollama run "$m" "$p" --verbose
  done < prompts.txt
done

For structured scoring, hit Ollama's local API (port 11434) for the metrics:

curl -s http://localhost:11434/api/generate -d '{
  "model": "qwen3:8b",
  "prompt": "Write a haiku about local inference.",
  "stream": false
}' | jq '{eval_count, tps: (.eval_count / (.eval_duration/1e9))}'

Score on correctness, format, and tokens/second; re-run when a new release lands.

WARNING
Benchmark on your tasks, not a leaderboard
Public leaderboards are gamed and rarely reflect your prompts or hardware. Hold temperature and context length fixed across models, or you're measuring your settings rather than the models — Lesson 9 covers those knobs.

Questions & Answers

Q: Is a 70B model always better than a 7B for my use case?
No. For focused tasks like completion, classification, or extraction, a well-chosen 7B-8B at Q5/Q8 often matches a 70B at Q4 while running far faster. Bigger helps most on open-ended reasoning and broad world knowledge — and a 70B that doesn't fit in memory will crawl. Benchmark before assuming you need it.
Q: How much quality do I actually lose dropping from Q8 to Q4?
For k-quants like Q4_K_M the loss is small and often imperceptible in everyday chat and coding — that's why it's the default. The drop becomes visible below Q4 (Q3, and especially Q2), where reasoning and formatting degrade. Stay at Q4_K_M or above unless memory forces it.
Q: Can I run a model whose listed size is larger than my VRAM?
Yes, but overflow layers offload to system RAM (on Apple Silicon, unified memory blurs this line), with a large speed penalty once you spill. The clean answer is a size and quant that fit in VRAM with room for the KV cache — CPU offload is a fallback, not a plan. Lesson 9 covers it.
Q: Why did the model ignore my instructions and just ramble?
You almost certainly pulled a base/pretrained build instead of an instruct one — base models continue text rather than follow commands. Use the -instruct or -it variant for interactive or tool-driven work. Default Ollama tags are usually instruct-tuned, but check with ollama show.

Key Takeaways

  1. Pick by family, then slot. Family by task (Qwen for code, Phi for small-size reasoning, Gemma for vision, Llama as a default), then a size that fits, then a quant for headroom.
  2. Parameters are a memory budget. 7B-8B is the everyday sweet spot; fit in VRAM with KV-cache room before going bigger.
  3. Q4_K_M is the default quant. Step up to Q5_K_M or Q6_K with memory to spare; go below Q4 only when nothing else fits.
  4. Instruct, not base. Use the instruction-tuned variant for chat and tools; check the license first.
  5. Benchmark on your own prompts. A short harness over real tasks beats any leaderboard.
  6. Pin and record. Lock the exact tag and digest so fast-moving releases never break reproducibility.

Next Steps: Lesson 5: Running Models from the CLI