Troubleshooting
Problem/cause/fix reference for the issues you hit most running models locally, assuming Ollama or llama.cpp at the terminal. For background, see the model guide and hardware reference.
Vocabulary used throughout:
- VRAM — dedicated GPU memory. Models must fit here for fast inference.
- Quantization — compressing weights to fewer bits (Q4, Q5, Q8). Lower bits = smaller, less memory, lower quality.
- Context window — tokens (roughly word-pieces) the model reads at once, prompt plus reply.
- Offloading — splitting a model between GPU and CPU/RAM when it won't fully fit in VRAM.
Out-of-Memory Errors
Symptoms: CUDA out of memory, ggml_metal_init: failed to allocate, a Killed process, or model requires more system memory than is available.
| Cause | Fix |
|---|---|
| Model too large for VRAM/RAM | Drop to a smaller quant (q4_K_M not q8_0) or fewer params (13B → 7B). |
| Context set too high | Lower num_ctx. A 32K context can cost several GB of KV cache atop the weights. |
| Other apps holding VRAM | Check nvidia-smi or Activity Monitor, then close browsers / GPU containers. |
| All layers forced onto GPU | Offload fewer layers; let CPU handle the rest (slower but it runs). |
| Multiple models loaded | Ollama keeps models warm; set OLLAMA_MAX_LOADED_MODELS=1. |
Estimate before pulling: weights ≈ (billions of params) × (bits ÷ 8) GB, plus 1-3 GB KV cache. A 7B model at Q4 needs ~4-5 GB.
# Smaller context cuts KV-cache memory (set in an Ollama session)
ollama run llama3.1:8b
/set parameter num_ctx 4096
# llama.cpp: offload only 20 layers to GPU, rest on CPU
./llama-cli -m model-q4_k_m.gguf -ngl 20 -c 4096 -p "Hello"
Very Slow Inference
Symptoms: a few tokens per second or worse, long pauses before the first token, fans roaring.
| Cause | Fix |
|---|---|
| Running on CPU only | Confirm the GPU is used (see GPU section). On CPU, use a smaller/more-quantized model. |
| Model spilling to RAM | Reduce quant or context so the model fits in VRAM. Partial offload is much slower. |
| Context too large | Long prompts re-process every turn. Trim system prompts; keep num_ctx minimal. |
| Thermal throttling | Improve cooling; on laptops plug in to mains so the GPU runs at full clocks. |
| Swapping to disk | If RAM is exhausted the OS swaps — check free -h and use a smaller model. |
# Measure real throughput (tokens/sec) and load time
ollama run llama3.1:8b --verbose "Write one sentence about the sea."
Rule of thumb: once any layers offload to CPU, speed can fall 5-20x. Getting the whole model into VRAM is the biggest lever.
GPU Not Detected
Symptoms: inference uses 100% CPU, nvidia-smi shows no process, or logs say no compatible GPUs were discovered / no usable GPU found.
| Cause | Fix |
|---|---|
| Driver missing/outdated (NVIDIA) | Install/upgrade the driver; verify nvidia-smi lists the card. |
| CUDA toolkit mismatch | Match the CUDA version your build expects; reinstall Ollama after a driver upgrade. |
| GPU build not compiled | Rebuild llama.cpp with the right backend (CUDA, Metal, ROCm); a CPU-only build never touches the GPU. |
| Docker can't see GPU | Install the NVIDIA Container Toolkit and run with --gpus all. |
| AMD card on Linux | Use a ROCm build; confirm the card is supported. |
| Apple Silicon | Metal is automatic. If slow, you're on an x86 build under Rosetta — install the arm64 binary. |
# Is the NVIDIA GPU visible, and which driver/CUDA version?
nvidia-smi
# Build llama.cpp with CUDA (rebuild clean if you flip backends)
cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release
# Run Ollama in Docker with GPU access
docker run --gpus all -d -v ollama:/root/.ollama -p 11434:11434 ollama/ollama
Model Download Failures
Symptoms: pull stalls, Error: max retries exceeded, checksum/digest mismatch, or EOF partway through.
| Cause | Fix |
|---|---|
| Flaky network / large file | Re-run the same ollama pull — it resumes rather than restarting. |
| Disk full | Models are multi-GB. Check df -h; free space or change the storage dir. |
| Corrupted partial download | Remove the model and pull again for a clean fetch. |
| Behind a proxy/firewall | Export HTTPS_PROXY/HTTP_PROXY before pulling. |
| Wrong tag name | Browse the library; tags are case-sensitive and version-specific. |
| Custom GGUF won't import | Verify it's a valid GGUF (not a partial download) before ollama create. |
# Resume a stalled pull (re-run); or remove + re-pull clean
ollama pull mistral:7b-instruct
ollama rm mistral:7b-instruct && ollama pull mistral:7b-instruct
# Move model storage to a bigger disk, then restart
export OLLAMA_MODELS=/mnt/big-disk/ollama-models
ollama serve
Garbled or Repetitive Output
Symptoms: nonsense tokens, infinite repetition, broken formatting, ignored template, or mixed languages.
| Cause | Fix |
|---|---|
| Wrong prompt/chat template | Use an -instruct/-chat tag and let Ollama apply its template; don't hand-format markup. |
| Quantization too aggressive | Q2/Q3 can degrade badly. Step up to Q4_K_M or higher. |
| Corrupted weights file | Re-pull the model; a bad download produces garbage. |
| Repetition with no penalty | Add a repeat penalty and a sensible temperature. |
| Base model used as chat | Base models complete text, not follow instructions — switch to the instruct variant. |
| Context overflow truncating prompt | See context-length below; a truncated prompt gives off-topic replies. |
# Tame repetition and randomness (Ollama, then llama.cpp)
/set parameter temperature 0.7
/set parameter repeat_penalty 1.15
./llama-cli -m model-q4_k_m.gguf --temp 0.7 --repeat-penalty 1.15 -p "..."
If output is fine at short prompts but degrades on long ones, it's almost always context length.
Context-Length Errors
Symptoms: context length exceeded, input is too long, the model "forgets" the start of a long document, or answers an earlier turn.
| Cause | Fix |
|---|---|
Prompt longer than num_ctx |
Raise num_ctx (costs memory) or shorten the input. |
| Default context too small | Models often default low even when they support more — set it explicitly. |
| KV cache blew up memory | Raising context raises memory use; balance against the OOM section. |
| RAG stuffing too many chunks | Retrieve fewer, more relevant chunks rather than maxing the window. |
| Long multi-turn chat | Summarise or /clear periodically; old turns count against the limit. |
# Set context window explicitly (Ollama, then llama.cpp -c)
ollama run qwen2.5:7b
/set parameter num_ctx 8192
./llama-cli -m model-q4_k_m.gguf -c 8192 -p "Summarise: ..."
Token budgeting: roughly 0.75 words per token in English. A 4,000-word document is ~5,300 tokens — an 8K window has room for the reply, a 4K window truncates it.
Still Stuck?
| Step | What to do |
|---|---|
| Read the logs | ollama serve in a foreground terminal prints the real error; llama.cpp prints offload info at startup. |
| Reproduce minimally | Try a small known-good model (1-3B instruct) to isolate model vs. setup. |
| Confirm the basics | ollama --version, nvidia-smi, df -h, free -h — most issues are memory, disk, or driver. |
| Check the guides | Compare against the hardware reference and the rules. |
Back to the Local LLMs course.