Performance Tuning

50 min advanced Lesson 9

Learning Outcomes

  • Choose the right quantization format (GGUF, AWQ, GPTQ) for your hardware and goals
  • Benchmark tokens-per-second and time-to-first-token to find your real bottleneck
  • Manage context length and KV-cache memory so you do not silently overflow VRAM
  • Configure GPU layer offloading, CPU fallback, and Apple Silicon Metal settings
  • Enable batching and parallel requests to raise throughput on a single GPU

Lesson Plan

Segment Duration Topic
Intro 3 min The three speed levers and how to measure them
Benchmark 8 min Establishing a baseline: tok/s and TTFT
Quantization 10 min GGUF, AWQ, GPTQ — what each one is for
Memory 9 min GPU layers, KV cache, context length
Throughput 8 min Flash attention, batching, parallel requests
Platform 7 min Apple Silicon Metal and multi-GPU NVIDIA
Wrap-up 5 min Tuning checklist and key takeaways

Before You Begin

Pre-work:

Shopping List:

  • Ollama running (ollama serve reachable on 127.0.0.1:11434)
  • A llama.cpp build with GPU support, or the prebuilt binaries
  • A model in two quant levels to compare (e.g. a Q4_K_M and a Q8_0 GGUF)
  • nvidia-smi (NVIDIA) or Activity Monitor / asitop (Apple Silicon) to watch memory

1 Measure First — Establish a Baseline

You cannot tune what you do not measure. Two numbers matter: tokens per second (throughput once generation starts) and time to first token (TTFT) — the lag before the first word, set by prompt processing. Ollama prints both with --verbose:

ollama run llama3.1:8b --verbose "Write a haiku about garbage collection."

Key fields: prompt eval rate (drives TTFT), eval rate (generation tok/s), load duration. For a repeatable benchmark, use the API — eval_count and eval_duration (nanoseconds) give exact tok/s:

curl -s http://localhost:11434/api/generate -d '{
  "model": "llama3.1:8b",
  "prompt": "Explain mutexes in two sentences.",
  "stream": false
}' | python3 -c 'import sys,json; d=json.load(sys.stdin); print(round(d["eval_count"]/(d["eval_duration"]/1e9),1),"tok/s")'

Run it three times and take the middle value — the first run pays model-load cost.

TIP
Title
Benchmark with a fixed prompt and output length. Add num_predict to the request options to cap generated tokens so every run does the same amount of work.

2 Quantization Formats — GGUF, AWQ, and GPTQ

Quantization stores weights at lower precision (4/5/8-bit) instead of 16-bit floats — less memory and faster math at some cost to accuracy. Three formats dominate; which you want depends on your runtime.

Format Runtime Hardware Best for
GGUF llama.cpp, Ollama, LM Studio CPU, Apple Silicon, any GPU Desktop, mixed CPU/GPU, Macs
AWQ vLLM, TGI NVIDIA GPU High-throughput serving
GPTQ vLLM, TGI, ExLlama NVIDIA GPU GPU serving, widely available

GGUF is what llama.cpp and Ollama use. Its "K-quant" variants (like Q4_K_M) do importance-weighted bit allocation: critical attention and output tensors keep higher precision while feed-forward layers are squeezed harder. The name reads as bits + method + size tier (M is medium; S/L also exist). Types range from Q2_K through Q4_K_S/M, Q5_K_S/M, Q6_K, Q8_0, plus newer I-quants (IQ2_XXS to IQ4_NL) that pack smaller via an importance matrix. Produce one yourself:

# Convert an F16 GGUF down to a 4-bit K-quant
./llama-quantize model-f16.gguf model-q4_k_m.gguf Q4_K_M

AWQ (Activation-aware Weight Quantization) drops FP16 to INT4 while protecting the weights activations care most about, cutting memory and latency on NVIDIA GPUs. GPTQ is a similar 4-bit GPU scheme. Both run through vLLM (Lesson 10), set with --quantization awq or detected from config.

NOTE
Title
Rule of thumb: Q4_K_M is the everyday sweet spot — ~4x smaller with little perceptible quality loss. Drop to Q3 or I-quants only to fit a bigger model; step up to Q5_K_M or Q8_0 with memory to spare.

3 GPU Layer Offloading and CPU Fallback

A model is a stack of transformer layers. If they all fit in VRAM, generation runs at full GPU speed; if not, the runtime spills layers to CPU RAM, and every CPU layer is far slower. Getting as many layers onto the GPU as fit is the biggest speed lever. In llama.cpp that is the -ngl (GPU layers) flag:

# Offload all layers to the GPU; drop the number if you hit OOM
./llama-cli -m model-q4_k_m.gguf -ngl 999 -p "Summarize the CAP theorem."

The load log prints how many layers landed on GPU versus CPU. If you hit an OOM error, lower -ngl until it fits. Ollama offloads automatically, but you can force the layer count per request with "options": {"num_gpu": 33} in the API body.

Watch what happens while a request runs. On NVIDIA (including Linux, and Windows inside WSL2 for GPU passthrough), run watch -n1 nvidia-smi to see VRAM and utilization climb. On Apple Silicon, memory is unified, so "offloading" just decides how much of the shared pool the GPU uses; watch it with Activity Monitor or asitop.

WARNING
Leave VRAM headroom
The KV cache (Step 4) is allocated on top of the weights, so filling VRAM with weights lets a long prompt trigger an OOM crash mid-generation. See the Troubleshooting guide for recovery.

4 Context Length and KV-Cache Memory

Every token the model attends to is stored in the KV cache (keys and values of past tokens). It grows linearly with context length and is allocated up front from num_ctx, so a window larger than you need wastes memory that could have held more weight layers.

Set context deliberately in Ollama — globally with OLLAMA_CONTEXT_LENGTH or per request with num_ctx. The default is small; raising it allocates more KV cache.

OLLAMA_CONTEXT_LENGTH=8192 ollama serve   # or "options": {"num_ctx": 8192} per request

You can also shrink the cache by quantizing it. Ollama's OLLAMA_KV_CACHE_TYPE defaults to f16; q8_0 uses about half the memory and q4_0 a quarter — it requires Flash Attention on (Step 5).

Setting KV-cache memory Quality impact
f16 (default) Baseline (1x) None
q8_0 ~1/2 Negligible
q4_0 ~1/4 Small but measurable
TIP
Title
Right-size the window. If prompts are 2K tokens and replies are short, 8K context is plenty. Reserving 128K on an 8B model can eat more memory than the weights themselves — for no benefit.

5 Throughput — Flash Attention and Parallel Requests

So far we tuned a single request. Now raise throughput — total tokens per second.

Flash attention is a memory-efficient attention algorithm. Enabling it in Ollama can significantly reduce memory usage as context grows, and is the prerequisite for Step 4's KV-cache quantization.

OLLAMA_FLASH_ATTENTION=1 ollama serve

Batching / parallel requests is the big win when you serve multiple users or a pipeline: the runtime interleaves requests to fill GPU compute instead of handling one at a time. OLLAMA_NUM_PARALLEL sets max concurrent requests per model (default 1), OLLAMA_MAX_LOADED_MODELS caps loaded models, and OLLAMA_KEEP_ALIVE keeps a model resident to avoid reload costs.

Combine everything into one tuned server start:

OLLAMA_FLASH_ATTENTION=1 \
OLLAMA_KV_CACHE_TYPE=q8_0 \
OLLAMA_NUM_PARALLEL=4 \
OLLAMA_CONTEXT_LENGTH=8192 \
OLLAMA_KEEP_ALIVE=30m \
ollama serve
NOTE
Title
Parallelism is not free: each slot gets its own KV cache, so doubling OLLAMA_NUM_PARALLEL roughly doubles KV-cache memory. Raise it only while watching VRAM. For heavy multi-user serving, vLLM (Lesson 10) batches more efficiently.

6 Platform Tuning — Apple Silicon and Multi-GPU NVIDIA

The last gains are platform-specific.

Apple Silicon (Metal + unified memory). llama.cpp and Ollama use the Metal backend automatically on M-series chips, so layers run on the GPU with no PCIe copy — CPU and GPU share one pool. Keep -ngl 999. For large models, raise the GPU's share of unified memory with the wired-memory limit:

# Raise GPU's unified-memory share (MB). Resets on reboot; leave the OS headroom.
sudo sysctl iogpu.wired_limit_mb=57344

On Linux with an NVIDIA card, use the multi-GPU split on the Windows tab — the flags are identical.

Multi-GPU NVIDIA (inside WSL2). With two or more cards, split a model across them. In llama.cpp, --tensor-split sets the per-GPU share:

# Split a 70B model across two GPUs, 60/40 by VRAM
./llama-cli -m model-70b-q4_k_m.gguf -ngl 999 \
  --split-mode layer --tensor-split 60,40 -p "Explain Raft consensus."

For serving, vLLM uses tensor parallelism: --tensor-parallel-size must be a power of two dividing the attention-head count (Lesson 10).

WARNING
Title
On multi-GPU rigs, an uneven --tensor-split or mismatched cards lets the slowest GPU bottleneck the whole model. Match cards where you can, and verify each one's utilization in nvidia-smi — idle cards mean a bad split.

7 Put It Together — A Tuning Loop

Tuning is iterative: change one setting, restart ollama serve, re-benchmark, keep what helps. Work top-down — the biggest wins come first.

Order Lever Typical impact
1 Fit all layers in VRAM (-ngl / num_gpu) Largest — CPU layers are 5-20x slower
2 Pick the right quant (Q4_K_M baseline) Large — frees memory for more layers
3 Right-size context (num_ctx) Medium — recovers VRAM for layers/batch
4 Flash attention + KV quant Medium — more context in the same memory
5 Parallel requests (OLLAMA_NUM_PARALLEL) Throughput only, costs KV memory
6 Platform flags (Metal limit / tensor split) Situational
TIP
Title
Stop when you hit your target, not when you run out of knobs. If 40 tok/s is fine for chat, do not chase 45 — save aggressive tuning for batch and multi-user serving.

Questions & Answers

Q: My tokens/sec looks fine but the first response takes forever. What is wrong?
That is high time-to-first-token — prompt processing, dominated by long prompts and large contexts. Shrink num_ctx, enable flash attention, and make sure prompt-eval runs on the GPU (check the load log for CPU layers). On a cold model, part of the lag is load time — keep it resident with OLLAMA_KEEP_ALIVE.
Q: Will a more aggressive quant make the model noticeably dumber?
Q4_K_M is hard to distinguish from full precision for most tasks. Degradation becomes real below ~4 bits (Q3 and the smaller I-quants) and shows up first on reasoning, code, and long-context work. Benchmark on YOUR prompts — compare outputs across two quant levels, not just speed.
Q: I cranked OLLAMA_NUM_PARALLEL to 8 and it got slower, not faster. Why?
Two reasons. First, each slot allocates its own KV cache, so you may have pushed weights or cache out of VRAM into slow CPU memory. Second, a single consumer GPU saturates quickly — past a point you are time-slicing the same compute, adding latency without throughput. Raise it gradually while watching nvidia-smi, and reach for vLLM if you genuinely need many concurrent users.
Q: Should I use GGUF or AWQ/GPTQ?
Your runtime decides, not preference. Ollama, llama.cpp, LM Studio, Macs, and mixed CPU/GPU all want GGUF. NVIDIA serving with vLLM or TGI wants AWQ or GPTQ. You cannot mix them — a GGUF file will not load in vLLM's AWQ path, and vice versa.

Key Takeaways

  1. Measure first. Track tokens/sec and time-to-first-token with a fixed prompt and output length; discard cold-start runs.
  2. Fit the model in VRAM first. Getting every layer onto the GPU (-ngl / num_gpu) is the largest speedup — offloaded CPU layers are far slower.
  3. Pick quant by runtime. GGUF (Q4_K_M default) for llama.cpp, Ollama, Apple Silicon; AWQ/GPTQ for NVIDIA serving with vLLM/TGI.
  4. Context costs memory. KV cache grows with num_ctx; right-size the window and quantize it (q8_0 halves it) with flash attention on.
  5. Batch for throughput. OLLAMA_NUM_PARALLEL raises total tokens served but costs KV cache per slot — scale while watching VRAM.
  6. Exploit your platform. Metal and the wired-memory limit on Apple Silicon; --tensor-split across NVIDIA GPUs.

Next Steps: Lesson 10: Production Deployment