Hardware Reference
Memory Requirements by Model Size and Quantization
The single number that decides whether a model will run is memory — VRAM on a discrete GPU, or unified memory on Apple Silicon. The figures below are for model weights only. Add 1-4 GB on top for the KV cache (the key/value cache holds the running context window and grows with conversation length).
B = billion parameters. Q4/Q5/Q8 = quantization levels — the bits used per weight; lower means smaller and faster but slightly less accurate.
| Model size | FP16 (full) | Q8_0 (~8-bit) | Q5_K_M (~5-bit) | Q4_K_M (~4-bit) | Practical minimum |
|---|---|---|---|---|---|
| 3B | ~6 GB | ~3.5 GB | ~2.5 GB | ~2 GB | 8 GB RAM, CPU OK |
| 7-8B | ~15 GB | ~8 GB | ~5.5 GB | ~4.5 GB | 8 GB VRAM / 16 GB RAM |
| 13-14B | ~28 GB | ~14 GB | ~10 GB | ~8 GB | 12 GB VRAM / 16 GB RAM |
| 30-34B | ~68 GB | ~36 GB | ~24 GB | ~20 GB | 24 GB VRAM / 32 GB RAM |
| 70B | ~140 GB | ~75 GB | ~50 GB | ~40 GB | 48 GB VRAM / 64 GB RAM |
Q4_K_M is the default sweet spot for most local use: roughly a 4x size reduction versus FP16 with minor quality loss. Move up to Q5_K_M or Q8_0 only if the weights still fit comfortably in memory. The full ladder of GGUF quant types (from IQ1_S up through Q8_0 and F16) is documented in the llama.cpp quantize tool.
Estimate weight size yourself:
weight_size_GB ≈ (params_in_billions × bits_per_weight) / 8
# e.g. 7B at Q4 ≈ (7 × 4.5) / 8 ≈ 4 GB
NVIDIA vs AMD vs Apple Silicon
| Vendor | Runner support | Notes |
|---|---|---|
| NVIDIA | CUDA — broadest, most mature | Compute capability 5.0+ and driver 531+; cap 5.0-6.2 needs driver 570+ |
| AMD | ROCm v7 (Linux + Windows) and Vulkan | Radeon RX/PRO, Ryzen AI, Instinct supported; less mature than CUDA |
| Apple | Metal — built in | Unified memory; no discrete VRAM cap to fight |
These support requirements come from Ollama's official GPU documentation. Vulkan provides additional GPU support on Windows and Linux, including Intel and select AMD cards.
Check your NVIDIA card's compute capability before buying:
nvidia-smi --query-gpu=name,compute_cap,memory.total --format=csv
NVIDIA consumer cards: VRAM is what matters
Pick by VRAM, not by model-year branding. A card with more VRAM runs bigger models; a faster card with less VRAM just runs small models faster. As a durable rule:
| VRAM | Comfortably runs (Q4_K_M) |
|---|---|
| 8 GB | up to 8B |
| 12 GB | up to 13-14B |
| 16 GB | 14B comfortably, 30B tight |
| 24 GB | up to 34B; 70B with heavy offload |
| 2x 24 GB | 70B at Q4 across two GPUs |
When a model exceeds VRAM, Ollama offloads layers to system RAM automatically — it avoids a crash, but throughput drops sharply. Treat RAM offload as a fallback, not a plan.
Apple Silicon: unified memory is the advantage
On M-series Macs the GPU and CPU share one memory pool, so "VRAM" is just a slice of your total RAM. A 64 GB Mac can run models that would need a multi-GPU rig on a PC. Token generation is memory-bandwidth-bound, and Apple's high-bandwidth unified memory is why these chips punch above their wattage.
By default macOS reserves a chunk of RAM for the system: roughly two-thirds is GPU-addressable on machines with 36 GB or less, and about three-quarters above that. You can raise the GPU allocation on macOS Sonoma and later:
# Allow 24 GB (24576 MB) for the GPU; takes effect immediately, resets on reboot
sudo sysctl iogpu.wired_limit_mb=24576
Leave at least 4-8 GB for macOS itself — pushing the limit too high causes beachballs or a hard lock. Watch Memory Pressure in Activity Monitor; if it goes red, dial the limit back down.
Buy-for-headroom table (Apple, Q4_K_M):
| Unified memory | Comfortable ceiling |
|---|---|
| 16 GB | up to 8B |
| 32 GB | up to 13-14B, small 30B |
| 64 GB | up to 34B comfortably |
| 128 GB+ | 70B and beyond |
Storage Planning
Models live on disk as GGUF files and are loaded into memory on demand. Plan generously — collections grow fast.
| Item | Typical disk footprint |
|---|---|
| One 7-8B model (Q4) | 4-5 GB |
| One 13-14B model (Q4) | 8 GB |
| One 70B model (Q4) | ~40 GB |
| A working library of 4-6 models | 50-150 GB |
Use an NVMe SSD if you can — load time scales with read speed, and a model that takes seconds to load from NVMe can take a minute from a spinning disk. Inspect and prune what Ollama has pulled:
ollama list # show installed models and sizes
du -sh ~/.ollama/models # total disk used by the model store
ollama rm llama3.1:8b # remove a model you no longer need
Cloud GPU Options
When local hardware falls short — or before you commit to buying a card — rent a GPU by the hour. You keep local-style control (your own runner, your own models) without the capital outlay. Common providers include RunPod, Vast.ai, and Lambda. Prices move constantly, so treat these as relative tiers rather than fixed quotes; check each provider's live pricing page before launching.
| Provider | Model | Best for |
|---|---|---|
| RunPod | On-demand pods, per-second billing | Quick experiments, spin-up/spin-down workflows |
| Vast.ai | Marketplace of third-party hosts | Lowest prices; reliability varies by host |
| Lambda | First-party data-centre instances | Consistent, support-backed multi-GPU jobs |
Rough GPU-to-workload mapping (rentable cards):
| Cloud GPU | VRAM | Runs comfortably |
|---|---|---|
| RTX 4090 | 24 GB | up to 34B (Q4) |
| A100 80GB | 80 GB | 70B at higher quant |
| H100 | 80 GB | 70B fast; fine-tuning |
Cost discipline: cloud GPUs bill while running, so terminate idle instances and store models/datasets on persistent volumes so you do not re-download tens of gigabytes each session.
# Point any local tool at a remote Ollama host over the network
export OLLAMA_HOST=http://<remote-ip>:11434
ollama list
Related References
- Match models to your hardware: Model Selection Cheat Sheet
- Recommended models by task: Model Guide
- Fix out-of-memory and GPU-not-detected errors: Troubleshooting
- Management and security guidelines: Rules and Best Practices
- Course home: Local LLMs