06

No Hype AI: Local LLMs

Can you really run AI on your own hardware? Yes. Learn to deploy, fine-tune, and use local language models.

10 lessons • 5 supplemental references

01
Why Run Local?

The case for running LLMs on your own hardware — privacy, cost, offline access, and freedom — weighed honestly against the real trade-offs.

30 min • beginner
02
Hardware Requirements

What you actually need — GPUs (NVIDIA, AMD), Apple Silicon, RAM by model size, storage, and cloud GPU options.

35 min • beginner
03
Ollama — The Easiest Way to Start

Install Ollama on macOS, Linux, or Windows, pull your first model, chat with it, and customise it with a Modelfile.

40 min • beginner
04
Model Selection

The open-weight landscape — Llama, Mistral, Phi, Qwen, Gemma — model sizes, quantization levels (Q4/Q5/Q8), and task fit.

40 min • intermediate
05
Running Models from the CLI

A full command-line workflow — scripting, piping, batch processing, and managing multiple models with Ollama and llama.cpp.

40 min • intermediate
06
Connecting Local Models to Tools

Make local models useful — Open WebUI, LM Studio, a VS Code coding assistant, and the OpenAI-compatible API.

45 min • intermediate
07
Fine-Tuning Basics

Adapt a model to your domain — when to fine-tune vs RAG, LoRA/QLoRA, data prep, and training with Unsloth and Axolotl.

50 min • intermediate
08
Local RAG — Your Own Knowledge Base

Build a private retrieval-augmented system locally — ingestion, chunking, local embeddings, and vector stores (ChromaDB, Qdrant).

55 min • advanced
09
Performance Tuning

Squeeze out speed — quantization formats (GGUF, AWQ, GPTQ), context management, batching, GPU memory, and CPU offloading.

50 min • advanced
10
Production Deployment

Serve models in production — vLLM, TGI, Ollama server mode, Docker, load balancing, monitoring, access control, and cost modelling.

60 min • advanced