06
No Hype AI: Local LLMs
Can you really run AI on your own hardware? Yes. Learn to deploy, fine-tune, and use local language models.
The case for running LLMs on your own hardware — privacy, cost, offline access, and freedom — weighed honestly against the real trade-offs.
What you actually need — GPUs (NVIDIA, AMD), Apple Silicon, RAM by model size, storage, and cloud GPU options.
Install Ollama on macOS, Linux, or Windows, pull your first model, chat with it, and customise it with a Modelfile.
The open-weight landscape — Llama, Mistral, Phi, Qwen, Gemma — model sizes, quantization levels (Q4/Q5/Q8), and task fit.
A full command-line workflow — scripting, piping, batch processing, and managing multiple models with Ollama and llama.cpp.
Make local models useful — Open WebUI, LM Studio, a VS Code coding assistant, and the OpenAI-compatible API.
Adapt a model to your domain — when to fine-tune vs RAG, LoRA/QLoRA, data prep, and training with Unsloth and Axolotl.
Build a private retrieval-augmented system locally — ingestion, chunking, local embeddings, and vector stores (ChromaDB, Qdrant).
Squeeze out speed — quantization formats (GGUF, AWQ, GPTQ), context management, batching, GPU memory, and CPU offloading.
Serve models in production — vLLM, TGI, Ollama server mode, Docker, load balancing, monitoring, access control, and cost modelling.