Why Run Local?
Learning Outcomes
- Articulate the four core reasons to run an LLM locally: privacy, cost, availability, and freedom
- Weigh those benefits honestly against quality, hardware, and maintenance trade-offs
- Understand where open-weight models stand today versus the proprietary frontier
- Apply a simple decision framework to your own use cases
- Decide whether local, cloud, or a hybrid setup is the right call for you
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | What "running local" actually means |
| The case | 10 min | Privacy, cost, availability, and freedom |
| The trade-offs | 8 min | Quality, hardware, and maintenance, told honestly |
| State of open weights | 5 min | How close are open models to the frontier? |
| Decision framework | 3 min | Scoring your own use cases |
| Wrap-up | 1 min | When to go local, and what's next |
Before You Begin
Pre-work:
- None — this is the opening lesson of the course. You only need an open mind and a use case or two in your head.
- If you have a specific task you wish you could run privately or offline, hold it in mind; we will score it at the end.
Shopping List:
- A notepad (paper or digital) for the decision-framework exercise
- No software, no hardware, no purchases — this lesson is about whether local makes sense before you spend anything
When you use ChatGPT, Claude, or Gemini, your prompt travels over the internet to a company's data centre, a model on their hardware generates a response, and the text comes back. Running local means the model weights live on your machine and inference happens on your CPU, GPU, or Apple Silicon — nothing leaves the device.
The thing that makes this possible is open-weight models: trained models whose parameters are published for anyone to download and run. Meta's Llama, Mistral, Microsoft's Phi, Alibaba's Qwen, and Google's Gemma are all open-weight families you can pull down and run today. Tools like Ollama (Lesson 3) have collapsed setup from a weekend of compiling into a single command:
# The whole pitch in two lines — full setup is Lesson 3.
# Download a small open-weight model and chat with it, entirely on your machine:
ollama run llama3.2
That one command downloads the weights, loads them, and drops you into a chat — no API key, no account, no network round-trip once the download finishes.
So the question is not can you run AI on your own hardware — you can. The real question this lesson answers is should you, and for which jobs. The rest of the course is the how; this lesson is the why.
This is the headline reason, and the strongest. When inference runs on your machine, your prompts and the model's responses never leave the device. There is no API provider logging requests, no terms-of-service clause about using your data for training, no third party in the loop at all.
That matters concretely for:
- Regulated data — health records, legal documents, or financial data where sending text to a third-party API may breach HIPAA, GDPR, or client confidentiality
- Proprietary code and trade secrets — internal source code, unreleased product specs, or research you cannot risk leaking
- Personal journaling and sensitive drafting — anything you would simply rather no company ever sees
With a local model, the only network traffic is the one-time download of the weights. After that, you can pull the network cable and the model keeps working. That is a privacy guarantee no cloud service can match — not because cloud providers are careless, but because the data physically never moves.
Privacy is the biggest draw, but three more reasons round out the case.
Cost. Cloud APIs bill per token, and heavy use adds up fast — a developer running an AI coding assistant all day, or a pipeline processing thousands of documents, can run a meaningful monthly bill. A local model has a hardware cost up front and then near-zero marginal cost: once the weights are on disk, you can run a million tokens or ten million for the price of electricity. For high-volume, repetitive workloads, local crosses over and becomes cheaper.
Availability. A local model works on a plane, in a basement, on a boat, or during an internet outage. No rate limits, no "the service is experiencing high demand," no API key expiring mid-task. The model is a file on your disk and it answers every time.
Freedom. You control the model. You can pick exactly which one runs, pin a version so it never changes underneath you, adjust how it behaves, and fine-tune it on your own data (Lesson 7). Cloud providers deprecate models, change behaviour silently, and apply content filters you cannot turn off. Local inference hands all of that back to you.
| Reason | What you gain | Who it matters most for |
|---|---|---|
| Privacy | Data never leaves the device | Regulated industries, sensitive work |
| Cost | Near-zero marginal cost after hardware | High-volume, repetitive workloads |
| Availability | Works offline, no rate limits | Travellers, unreliable connections |
| Freedom | Full control over model and behaviour | Tinkerers, fine-tuners, version-pinners |
A balanced case has to admit the costs. Local is not strictly better — it trades convenience for control.
Model quality. The very best proprietary models (the frontier from OpenAI, Anthropic, and Google) still lead the best open-weight models on the hardest reasoning, coding, and long-context tasks. The gap has narrowed dramatically and a good open model is more than enough for most everyday work — but if you need absolute top-tier capability on a hard problem, the cloud frontier is still ahead.
Hardware cost and limits. A capable local setup needs real memory. You can run small models on a laptop, but the larger, more capable models want a 16-24 GB GPU or a high-memory Mac (the whole of Lesson 2). That is a real up-front cost, and your hardware caps which models you can run at all.
Maintenance burden. With the cloud, the provider handles updates, scaling, and uptime. Run local and you are the operations team: installing tools, downloading and managing model files, troubleshooting drivers, and keeping things current. It is very manageable — this course teaches exactly that — but it is not zero effort.
It helps to know where the open ecosystem actually stands, because it is the single biggest factor in whether local is "good enough" for you.
The short version: open-weight models have closed most of the gap for everyday tasks. A modern open model in the 8B-32B range handles summarisation, drafting, Q&A, classification, and a great deal of coding capably — work that a year or two ago would have demanded a frontier API. For routine assistant duties, the difference is often imperceptible.
Where the frontier still leads is the hard end: multi-step reasoning, very long context, the trickiest coding problems, and the broadest world knowledge. If your work lives mostly there, weigh that honestly.
Two practical notes carried through the rest of the course:
- Think in size classes, not model names. The landscape moves monthly. Anchor on tiers — a "7B everyday model" or a "32B coding model" — and refill each tier with whatever is currently best. Lesson 4 covers model selection in depth.
- Quantization makes it feasible. Compressing weights from 16-bit to roughly 4-bit (Lesson 2) shrinks a model to about a quarter of its size with only a small quality cost, which is what lets capable models fit on consumer hardware at all.
The honest way to settle the "is it good enough?" question is to stop reading benchmarks and run a current model on your own work:
# Pull a current general-purpose model and test it on YOUR real prompts (full workflow in Lessons 3-4):
ollama pull qwen2.5:7b
ollama run qwen2.5:7b
Pull out your notepad. For each task you are considering, score these five questions — roughly, on a 0-3 scale where higher means "more reason to go local":
- How sensitive is the data? (0 = public, 3 = confidential / regulated)
- How high is the volume? (0 = a few prompts a day, 3 = constant or batch processing)
- How much do you need offline access? (0 = always online, 3 = frequently disconnected)
- How important is control and customisation? (0 = defaults are fine, 3 = you want to fine-tune and pin versions)
- How hard is the task? (0 = needs the frontier's best reasoning, 3 = routine, well within open-model range)
Add up the score. As a rough guide:
| Total score | Recommendation |
|---|---|
| 0-5 | Stick with cloud APIs — local is unlikely to pay off |
| 6-10 | A hybrid setup makes sense — local for sensitive or high-volume work, cloud for the hardest tasks |
| 11-15 | Local is a strong fit — this course is for you |
Most people land in the hybrid band, and that is a perfectly good answer: run a local model for private, routine, high-volume work, and reach for a cloud frontier model on the occasional hard problem. Local and cloud are not an either/or — the smart move is using each where it wins.
Questions & Answers
Key Takeaways
- You absolutely can run AI on your own hardware. Open-weight models plus tools like Ollama make it a single command — the real question is whether it fits your use case, not whether it is possible.
- Four reasons to go local: privacy (data never leaves the device), cost (near-zero marginal cost at high volume), availability (offline, no rate limits), and freedom (full control over the model).
- Privacy is the strongest case. If you would not type it into a stranger's computer, that is a candidate for local inference — the data physically never moves.
- Be honest about the trade-offs. The proprietary frontier still leads on the hardest tasks, capable models need real hardware, and you become the operations team.
- Open models are good enough for most everyday work. Think in size classes rather than model names, and judge quality by running current models on your own prompts.
- Hybrid is the realistic default. Score your use cases on data sensitivity, volume, offline need, control, and task difficulty — most people land on local-for-most, cloud-for-the-hardest.
Next Steps: Lesson 2: Hardware Requirements