Local LLM Best Practices
A durable set of rules for running open-weight models on your own hardware. These hold regardless of which model family is hot this month. For per-model recommendations see the model selection cheat sheet, for hardware specifics the hardware reference, and for fixes the troubleshooting guide.
Rule 1 — Match Model Size to Your Memory
A model's parameters (the 7B, 13B, 70B numbers) determine how much memory the weights occupy. At a given quantization, the file size is roughly your floor for RAM (CPU) or VRAM (GPU) — plus headroom for the KV cache (the running memory of the conversation) and OS. As a planning rule, leave 2-4 GB free on top of the model file.
| Model size | Q4 footprint | Min memory (comfortable) | Typical home |
|---|---|---|---|
| 1-3B | ~1-2 GB | 8 GB | Any laptop, CPU fine |
| 7-8B | ~4-5 GB | 16 GB | Mainstream sweet spot |
| 13-14B | ~8-9 GB | 24 GB | 12 GB+ GPU or 32 GB Mac |
| 30-34B | ~18-20 GB | 32 GB+ | 24 GB GPU or 48 GB Mac |
| 70B | ~38-42 GB | 64 GB+ | Dual GPU or 64-96 GB Mac |
If the model does not fit, it either refuses to load or spills to disk and crawls. Size down before you size up.
Rule 2 — Start Quantized, Then Decide if You Need More
Quantization shrinks weights from 16-bit floats to 4/5/8-bit integers, trading a little quality for a lot of memory and speed. Q4 is the default starting point — it runs almost everything and the quality loss is small. Only move to Q5/Q6/Q8 if you can measure a deficiency that matters for your task.
# Ollama tags carry the quantization; pull an explicit one
ollama pull llama3.1:8b-instruct-q4_K_M
ollama pull qwen2.5-coder:7b-instruct-q5_K_M
# Compare two quants on the same prompt before committing memory
ollama run llama3.1:8b-instruct-q4_K_M "Summarise the CAP theorem in 3 bullets"
ollama run llama3.1:8b-instruct-q8_0 "Summarise the CAP theorem in 3 bullets"
| Quant | Quality | Use when |
|---|---|---|
| Q4_K_M | Good | Default for chat, coding, drafting |
| Q5_K_M | Slightly better | You see Q4 errors and have spare memory |
| Q6_K / Q8_0 | Near-full | Eval, structured output, sensitive tasks |
| FP16 | Reference | Benchmarking only; rarely worth the RAM |
Rule 3 — Measure Tokens per Second, Not Vibes
Throughput (tokens/sec) is the honest metric. "Feels slow" is not actionable; "9 tok/s on a 13B at Q4" is. Establish a baseline number for each model so you can tell whether a config change actually helped.
# Ollama prints eval rate with --verbose
ollama run llama3.1:8b --verbose "Write a haiku about latency"
# -> look for "eval rate: NN.N tokens/s"
# llama.cpp ships a dedicated benchmark
llama-bench -m models/llama-3.1-8b-q4_k_m.gguf -p 512 -n 128
| Reading | Interpretation |
|---|---|
| Prompt eval rate | How fast it ingests your input (matters for long context) |
| Eval/gen rate | How fast it writes the answer (what you feel) |
| < 5 tok/s | Likely spilling to CPU/disk — size down or quantize harder |
| Sudden drop mid-session | KV cache filled RAM — see Rule 4 |
Rule 4 — Keep Context Tight
The context window (how many tokens the model holds at once) is not free. Every token in context consumes KV-cache memory and slows prompt processing. Set the window to what the task needs, not the maximum the model allows.
# Ollama: cap context with num_ctx (default is often 2k-4k)
ollama run llama3.1:8b
>>> /set parameter num_ctx 8192
# llama.cpp: -c sets context length explicitly
llama-cli -m model.gguf -c 8192 -p "..."
- Trim system prompts and few-shot examples to the minimum that works.
- Summarise or drop old turns instead of letting a chat grow unbounded.
- For document work, retrieve the relevant chunks (RAG) rather than pasting whole files.
- A 32k window you don't use still costs memory at load time — request it only when needed.
Rule 5 — Prefer GGUF on CPU and Apple Silicon
GGUF is the file format used by llama.cpp (and Ollama, which wraps it). It is built for CPU and Apple Silicon's unified memory, supports memory-mapped loading, and runs the widest range of quantizations. On a Mac or a GPU-less box, GGUF is almost always the right choice.
| Format | Best on | Notes |
|---|---|---|
| GGUF | CPU, Apple Silicon (Metal), mixed | Default for local desktop use |
| AWQ / GPTQ | NVIDIA GPU (vLLM, TGI) | Faster GPU serving, less CPU-friendly |
| Safetensors (FP16) | Fine-tuning, conversion source | Convert to GGUF/AWQ before serving |
# Apple Silicon: Ollama uses Metal automatically. Confirm GPU is engaged:
ollama run llama3.1:8b --verbose "hi" # eval rate should be well above CPU-only
# Convert a HF safetensors model to GGUF with llama.cpp tooling
python convert_hf_to_gguf.py ./my-model --outfile my-model-f16.gguf
llama-quantize my-model-f16.gguf my-model-q4_k_m.gguf Q4_K_M
Rule 6 — Isolate Serving from Your Main Machine
The moment a model is on a port, treat it like any network service. Ollama and llama.cpp servers ship with no authentication — anyone who can reach the port can use your model and read your prompts. Bind to localhost for personal use; put it behind a reverse proxy with auth before exposing it.
# Default: bind to loopback only (safe for a single machine)
OLLAMA_HOST=127.0.0.1:11434 ollama serve
# Exposing to a LAN/team: do NOT do 0.0.0.0 naked.
# Run it in a container and front it with auth + TLS.
docker run -d --name ollama -p 127.0.0.1:11434:11434 \
-v ollama:/root/.ollama ollama/ollama
# Then reverse-proxy (nginx/Caddy) with an API key or basic auth in front.
- Keep inference off your daily-driver if you value its uptime — a runaway 70B load can lock a workstation. A dedicated box, container, or cloud GPU instance isolates the blast radius.
- Pin model versions and verify downloads; a swapped weight file is a supply-chain risk.
- Local models are unrestricted — there is no provider filter. You own the outputs and the responsibility for them.
Rule 7 — Treat Quality as Task-Specific
There is no single "best" local model; there is the best model for this task at this size on this hardware. Keep a tiny eval set of 5-10 prompts that represent your real work and run candidates through it.
# Loop a fixed prompt set across models for a quick head-to-head
for m in qwen2.5-coder:7b llama3.1:8b mistral:7b; do
echo "=== $m ==="
ollama run "$m" "Refactor this for readability:\n$(cat sample.py)"
done
| Task | Lean toward |
|---|---|
| Coding | Code-tuned models (e.g. Qwen Coder family) |
| Reasoning / math | Larger or reasoning-tuned instruct models |
| Fast drafting / chat | 7-8B instruct at Q4 |
| Tight RAM | 1-3B models; accept narrower competence |
When prompt engineering and retrieval stop closing the gap, that is your signal to consider fine-tuning — not before.
Quick Self-Check
| Question | If "no"... |
|---|---|
| Does the model + headroom fit in memory? | Size down or quantize harder (Rules 1-2) |
| Do I know my tokens/sec baseline? | Run --verbose / llama-bench (Rule 3) |
Is num_ctx sized to the task? |
Lower it; trim the prompt (Rule 4) |
| GGUF on CPU/Mac, AWQ/GPTQ on NVIDIA? | Convert to the right format (Rule 5) |
| Is the server bound to localhost or authed? | Lock it down before exposing (Rule 6) |
| Did I eval on my own prompts? | Build a 5-10 prompt set (Rule 7) |