Performance Tuning
Learning Outcomes
- Choose the right quantization format (GGUF, AWQ, GPTQ) for your hardware and goals
- Benchmark tokens-per-second and time-to-first-token to find your real bottleneck
- Manage context length and KV-cache memory so you do not silently overflow VRAM
- Configure GPU layer offloading, CPU fallback, and Apple Silicon Metal settings
- Enable batching and parallel requests to raise throughput on a single GPU
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | The three speed levers and how to measure them |
| Benchmark | 8 min | Establishing a baseline: tok/s and TTFT |
| Quantization | 10 min | GGUF, AWQ, GPTQ — what each one is for |
| Memory | 9 min | GPU layers, KV cache, context length |
| Throughput | 8 min | Flash attention, batching, parallel requests |
| Platform | 7 min | Apple Silicon Metal and multi-GPU NVIDIA |
| Wrap-up | 5 min | Tuning checklist and key takeaways |
Before You Begin
Pre-work:
- Complete Lesson 3: Ollama and Lesson 5: Running Models from the CLI
- Skim the Hardware Reference to know your VRAM and memory ceiling
- Have a model pulled in Ollama and, ideally, a
llama.cppbuild available
Shopping List:
- Ollama running (
ollama servereachable on127.0.0.1:11434) - A
llama.cppbuild with GPU support, or the prebuilt binaries - A model in two quant levels to compare (e.g. a Q4_K_M and a Q8_0 GGUF)
nvidia-smi(NVIDIA) or Activity Monitor /asitop(Apple Silicon) to watch memory
You cannot tune what you do not measure. Two numbers matter: tokens per second (throughput once generation starts) and time to first token (TTFT) — the lag before the first word, set by prompt processing. Ollama prints both with --verbose:
ollama run llama3.1:8b --verbose "Write a haiku about garbage collection."
Key fields: prompt eval rate (drives TTFT), eval rate (generation tok/s), load duration. For a repeatable benchmark, use the API — eval_count and eval_duration (nanoseconds) give exact tok/s:
curl -s http://localhost:11434/api/generate -d '{
"model": "llama3.1:8b",
"prompt": "Explain mutexes in two sentences.",
"stream": false
}' | python3 -c 'import sys,json; d=json.load(sys.stdin); print(round(d["eval_count"]/(d["eval_duration"]/1e9),1),"tok/s")'
Run it three times and take the middle value — the first run pays model-load cost.
num_predict to the request options to cap generated tokens so every run does the same amount of work.Quantization stores weights at lower precision (4/5/8-bit) instead of 16-bit floats — less memory and faster math at some cost to accuracy. Three formats dominate; which you want depends on your runtime.
| Format | Runtime | Hardware | Best for |
|---|---|---|---|
| GGUF | llama.cpp, Ollama, LM Studio | CPU, Apple Silicon, any GPU | Desktop, mixed CPU/GPU, Macs |
| AWQ | vLLM, TGI | NVIDIA GPU | High-throughput serving |
| GPTQ | vLLM, TGI, ExLlama | NVIDIA GPU | GPU serving, widely available |
GGUF is what llama.cpp and Ollama use. Its "K-quant" variants (like Q4_K_M) do importance-weighted bit allocation: critical attention and output tensors keep higher precision while feed-forward layers are squeezed harder. The name reads as bits + method + size tier (M is medium; S/L also exist). Types range from Q2_K through Q4_K_S/M, Q5_K_S/M, Q6_K, Q8_0, plus newer I-quants (IQ2_XXS to IQ4_NL) that pack smaller via an importance matrix. Produce one yourself:
# Convert an F16 GGUF down to a 4-bit K-quant
./llama-quantize model-f16.gguf model-q4_k_m.gguf Q4_K_M
AWQ (Activation-aware Weight Quantization) drops FP16 to INT4 while protecting the weights activations care most about, cutting memory and latency on NVIDIA GPUs. GPTQ is a similar 4-bit GPU scheme. Both run through vLLM (Lesson 10), set with --quantization awq or detected from config.
Q4_K_M is the everyday sweet spot — ~4x smaller with little perceptible quality loss. Drop to Q3 or I-quants only to fit a bigger model; step up to Q5_K_M or Q8_0 with memory to spare.A model is a stack of transformer layers. If they all fit in VRAM, generation runs at full GPU speed; if not, the runtime spills layers to CPU RAM, and every CPU layer is far slower. Getting as many layers onto the GPU as fit is the biggest speed lever. In llama.cpp that is the -ngl (GPU layers) flag:
# Offload all layers to the GPU; drop the number if you hit OOM
./llama-cli -m model-q4_k_m.gguf -ngl 999 -p "Summarize the CAP theorem."
The load log prints how many layers landed on GPU versus CPU. If you hit an OOM error, lower -ngl until it fits. Ollama offloads automatically, but you can force the layer count per request with "options": {"num_gpu": 33} in the API body.
Watch what happens while a request runs. On NVIDIA (including Linux, and Windows inside WSL2 for GPU passthrough), run watch -n1 nvidia-smi to see VRAM and utilization climb. On Apple Silicon, memory is unified, so "offloading" just decides how much of the shared pool the GPU uses; watch it with Activity Monitor or asitop.
Every token the model attends to is stored in the KV cache (keys and values of past tokens). It grows linearly with context length and is allocated up front from num_ctx, so a window larger than you need wastes memory that could have held more weight layers.
Set context deliberately in Ollama — globally with OLLAMA_CONTEXT_LENGTH or per request with num_ctx. The default is small; raising it allocates more KV cache.
OLLAMA_CONTEXT_LENGTH=8192 ollama serve # or "options": {"num_ctx": 8192} per request
You can also shrink the cache by quantizing it. Ollama's OLLAMA_KV_CACHE_TYPE defaults to f16; q8_0 uses about half the memory and q4_0 a quarter — it requires Flash Attention on (Step 5).
| Setting | KV-cache memory | Quality impact |
|---|---|---|
f16 (default) |
Baseline (1x) | None |
q8_0 |
~1/2 | Negligible |
q4_0 |
~1/4 | Small but measurable |
So far we tuned a single request. Now raise throughput — total tokens per second.
Flash attention is a memory-efficient attention algorithm. Enabling it in Ollama can significantly reduce memory usage as context grows, and is the prerequisite for Step 4's KV-cache quantization.
OLLAMA_FLASH_ATTENTION=1 ollama serve
Batching / parallel requests is the big win when you serve multiple users or a pipeline: the runtime interleaves requests to fill GPU compute instead of handling one at a time. OLLAMA_NUM_PARALLEL sets max concurrent requests per model (default 1), OLLAMA_MAX_LOADED_MODELS caps loaded models, and OLLAMA_KEEP_ALIVE keeps a model resident to avoid reload costs.
Combine everything into one tuned server start:
OLLAMA_FLASH_ATTENTION=1 \
OLLAMA_KV_CACHE_TYPE=q8_0 \
OLLAMA_NUM_PARALLEL=4 \
OLLAMA_CONTEXT_LENGTH=8192 \
OLLAMA_KEEP_ALIVE=30m \
ollama serve
OLLAMA_NUM_PARALLEL roughly doubles KV-cache memory. Raise it only while watching VRAM. For heavy multi-user serving, vLLM (Lesson 10) batches more efficiently.The last gains are platform-specific.
Apple Silicon (Metal + unified memory). llama.cpp and Ollama use the Metal backend automatically on M-series chips, so layers run on the GPU with no PCIe copy — CPU and GPU share one pool. Keep -ngl 999. For large models, raise the GPU's share of unified memory with the wired-memory limit:
# Raise GPU's unified-memory share (MB). Resets on reboot; leave the OS headroom.
sudo sysctl iogpu.wired_limit_mb=57344
On Linux with an NVIDIA card, use the multi-GPU split on the Windows tab — the flags are identical.
Multi-GPU NVIDIA (inside WSL2). With two or more cards, split a model across them. In llama.cpp, --tensor-split sets the per-GPU share:
# Split a 70B model across two GPUs, 60/40 by VRAM
./llama-cli -m model-70b-q4_k_m.gguf -ngl 999 \
--split-mode layer --tensor-split 60,40 -p "Explain Raft consensus."
For serving, vLLM uses tensor parallelism: --tensor-parallel-size must be a power of two dividing the attention-head count (Lesson 10).
--tensor-split or mismatched cards lets the slowest GPU bottleneck the whole model. Match cards where you can, and verify each one's utilization in nvidia-smi — idle cards mean a bad split.Tuning is iterative: change one setting, restart ollama serve, re-benchmark, keep what helps. Work top-down — the biggest wins come first.
| Order | Lever | Typical impact |
|---|---|---|
| 1 | Fit all layers in VRAM (-ngl / num_gpu) |
Largest — CPU layers are 5-20x slower |
| 2 | Pick the right quant (Q4_K_M baseline) |
Large — frees memory for more layers |
| 3 | Right-size context (num_ctx) |
Medium — recovers VRAM for layers/batch |
| 4 | Flash attention + KV quant | Medium — more context in the same memory |
| 5 | Parallel requests (OLLAMA_NUM_PARALLEL) |
Throughput only, costs KV memory |
| 6 | Platform flags (Metal limit / tensor split) | Situational |
Questions & Answers
num_ctx, enable flash attention, and make sure prompt-eval runs on the GPU (check the load log for CPU layers). On a cold model, part of the lag is load time — keep it resident with OLLAMA_KEEP_ALIVE.nvidia-smi, and reach for vLLM if you genuinely need many concurrent users.Key Takeaways
- Measure first. Track tokens/sec and time-to-first-token with a fixed prompt and output length; discard cold-start runs.
- Fit the model in VRAM first. Getting every layer onto the GPU (
-ngl/num_gpu) is the largest speedup — offloaded CPU layers are far slower. - Pick quant by runtime. GGUF (Q4_K_M default) for llama.cpp, Ollama, Apple Silicon; AWQ/GPTQ for NVIDIA serving with vLLM/TGI.
- Context costs memory. KV cache grows with
num_ctx; right-size the window and quantize it (q8_0halves it) with flash attention on. - Batch for throughput.
OLLAMA_NUM_PARALLELraises total tokens served but costs KV cache per slot — scale while watching VRAM. - Exploit your platform. Metal and the wired-memory limit on Apple Silicon;
--tensor-splitacross NVIDIA GPUs.
Next Steps: Lesson 10: Production Deployment