Troubleshooting

Reference intermediate

Problem/cause/fix reference for the issues you hit most running models locally, assuming Ollama or llama.cpp at the terminal. For background, see the model guide and hardware reference.

Vocabulary used throughout:

  • VRAM — dedicated GPU memory. Models must fit here for fast inference.
  • Quantization — compressing weights to fewer bits (Q4, Q5, Q8). Lower bits = smaller, less memory, lower quality.
  • Context window — tokens (roughly word-pieces) the model reads at once, prompt plus reply.
  • Offloading — splitting a model between GPU and CPU/RAM when it won't fully fit in VRAM.

Out-of-Memory Errors

Symptoms: CUDA out of memory, ggml_metal_init: failed to allocate, a Killed process, or model requires more system memory than is available.

Cause Fix
Model too large for VRAM/RAM Drop to a smaller quant (q4_K_M not q8_0) or fewer params (13B → 7B).
Context set too high Lower num_ctx. A 32K context can cost several GB of KV cache atop the weights.
Other apps holding VRAM Check nvidia-smi or Activity Monitor, then close browsers / GPU containers.
All layers forced onto GPU Offload fewer layers; let CPU handle the rest (slower but it runs).
Multiple models loaded Ollama keeps models warm; set OLLAMA_MAX_LOADED_MODELS=1.

Estimate before pulling: weights ≈ (billions of params) × (bits ÷ 8) GB, plus 1-3 GB KV cache. A 7B model at Q4 needs ~4-5 GB.

# Smaller context cuts KV-cache memory (set in an Ollama session)
ollama run llama3.1:8b
/set parameter num_ctx 4096

# llama.cpp: offload only 20 layers to GPU, rest on CPU
./llama-cli -m model-q4_k_m.gguf -ngl 20 -c 4096 -p "Hello"

Very Slow Inference

Symptoms: a few tokens per second or worse, long pauses before the first token, fans roaring.

Cause Fix
Running on CPU only Confirm the GPU is used (see GPU section). On CPU, use a smaller/more-quantized model.
Model spilling to RAM Reduce quant or context so the model fits in VRAM. Partial offload is much slower.
Context too large Long prompts re-process every turn. Trim system prompts; keep num_ctx minimal.
Thermal throttling Improve cooling; on laptops plug in to mains so the GPU runs at full clocks.
Swapping to disk If RAM is exhausted the OS swaps — check free -h and use a smaller model.
# Measure real throughput (tokens/sec) and load time
ollama run llama3.1:8b --verbose "Write one sentence about the sea."

Rule of thumb: once any layers offload to CPU, speed can fall 5-20x. Getting the whole model into VRAM is the biggest lever.

GPU Not Detected

Symptoms: inference uses 100% CPU, nvidia-smi shows no process, or logs say no compatible GPUs were discovered / no usable GPU found.

Cause Fix
Driver missing/outdated (NVIDIA) Install/upgrade the driver; verify nvidia-smi lists the card.
CUDA toolkit mismatch Match the CUDA version your build expects; reinstall Ollama after a driver upgrade.
GPU build not compiled Rebuild llama.cpp with the right backend (CUDA, Metal, ROCm); a CPU-only build never touches the GPU.
Docker can't see GPU Install the NVIDIA Container Toolkit and run with --gpus all.
AMD card on Linux Use a ROCm build; confirm the card is supported.
Apple Silicon Metal is automatic. If slow, you're on an x86 build under Rosetta — install the arm64 binary.
# Is the NVIDIA GPU visible, and which driver/CUDA version?
nvidia-smi

# Build llama.cpp with CUDA (rebuild clean if you flip backends)
cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release

# Run Ollama in Docker with GPU access
docker run --gpus all -d -v ollama:/root/.ollama -p 11434:11434 ollama/ollama

Model Download Failures

Symptoms: pull stalls, Error: max retries exceeded, checksum/digest mismatch, or EOF partway through.

Cause Fix
Flaky network / large file Re-run the same ollama pull — it resumes rather than restarting.
Disk full Models are multi-GB. Check df -h; free space or change the storage dir.
Corrupted partial download Remove the model and pull again for a clean fetch.
Behind a proxy/firewall Export HTTPS_PROXY/HTTP_PROXY before pulling.
Wrong tag name Browse the library; tags are case-sensitive and version-specific.
Custom GGUF won't import Verify it's a valid GGUF (not a partial download) before ollama create.
# Resume a stalled pull (re-run); or remove + re-pull clean
ollama pull mistral:7b-instruct
ollama rm mistral:7b-instruct && ollama pull mistral:7b-instruct

# Move model storage to a bigger disk, then restart
export OLLAMA_MODELS=/mnt/big-disk/ollama-models
ollama serve

Garbled or Repetitive Output

Symptoms: nonsense tokens, infinite repetition, broken formatting, ignored template, or mixed languages.

Cause Fix
Wrong prompt/chat template Use an -instruct/-chat tag and let Ollama apply its template; don't hand-format markup.
Quantization too aggressive Q2/Q3 can degrade badly. Step up to Q4_K_M or higher.
Corrupted weights file Re-pull the model; a bad download produces garbage.
Repetition with no penalty Add a repeat penalty and a sensible temperature.
Base model used as chat Base models complete text, not follow instructions — switch to the instruct variant.
Context overflow truncating prompt See context-length below; a truncated prompt gives off-topic replies.
# Tame repetition and randomness (Ollama, then llama.cpp)
/set parameter temperature 0.7
/set parameter repeat_penalty 1.15
./llama-cli -m model-q4_k_m.gguf --temp 0.7 --repeat-penalty 1.15 -p "..."

If output is fine at short prompts but degrades on long ones, it's almost always context length.

Context-Length Errors

Symptoms: context length exceeded, input is too long, the model "forgets" the start of a long document, or answers an earlier turn.

Cause Fix
Prompt longer than num_ctx Raise num_ctx (costs memory) or shorten the input.
Default context too small Models often default low even when they support more — set it explicitly.
KV cache blew up memory Raising context raises memory use; balance against the OOM section.
RAG stuffing too many chunks Retrieve fewer, more relevant chunks rather than maxing the window.
Long multi-turn chat Summarise or /clear periodically; old turns count against the limit.
# Set context window explicitly (Ollama, then llama.cpp -c)
ollama run qwen2.5:7b
/set parameter num_ctx 8192
./llama-cli -m model-q4_k_m.gguf -c 8192 -p "Summarise: ..."

Token budgeting: roughly 0.75 words per token in English. A 4,000-word document is ~5,300 tokens — an 8K window has room for the reply, a 4K window truncates it.

Still Stuck?

Step What to do
Read the logs ollama serve in a foreground terminal prints the real error; llama.cpp prints offload info at startup.
Reproduce minimally Try a small known-good model (1-3B instruct) to isolate model vs. setup.
Confirm the basics ollama --version, nvidia-smi, df -h, free -h — most issues are memory, disk, or driver.
Check the guides Compare against the hardware reference and the rules.

Back to the Local LLMs course.