Hardware Requirements
Learning Outcomes
- Calculate a model's memory footprint from its parameter count and quantization
- Compare NVIDIA, AMD, and Apple Silicon and pick the right path for local inference
- Map model sizes (7B, 13B, 70B) to realistic RAM and VRAM budgets
- Plan storage and bandwidth for downloading and caching model weights
- Decide when to rent a cloud GPU instead of buying hardware
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | What hardware gates local inference |
| Memory math | 8 min | The VRAM formula, quantization, a sizing map |
| GPUs | 11 min | NVIDIA, AMD, and "compute capability" |
| Apple Silicon | 5 min | Unified memory and why Macs punch above their weight |
| Storage & cloud | 7 min | Disk, bandwidth, cache, renting GPUs |
| Wrap-up | 1 min | Benchmark your box, preview Ollama |
Before You Begin
Pre-work:
- Read Lesson 1: Why Run Local? so you know which use cases justify the spend.
- Find your specs — total RAM, GPU model, free disk. On macOS/Linux,
system_profiler SPHardwareDataType,free -h, andnvidia-smicover most of it.
Shopping List:
- A terminal you are comfortable in
- Your machine's RAM, GPU, and VRAM figures
- No purchases required — this lesson is about deciding what to run before you spend
Local inference is mostly a memory problem. The weights must live in fast memory the compute unit can reach — VRAM on a GPU, or unified memory on Apple Silicon. If they fit you get good speed; spill into system RAM and generation crawls.
Footprint is three things: parameters (trained weights in billions — the "7B" in a name), bytes per parameter (set by quantization — compressing 16-bit weights to 4-bit), and overhead (KV cache plus activations). A practical estimate:
memory_GB ≈ parameters_in_billions × bytes_per_param × 1.2
The 1.2 is rough 20% headroom for KV cache and activations. Bytes per parameter run ~2.0 at FP16 (full), ~1.0 at Q8 (near-lossless), and ~0.55 at Q4_K_M (the sweet spot). So a 7B at Q4_K_M needs roughly 7 × 0.55 × 1.2 ≈ 4.6 GB, versus ~17 GB at FP16. Quantization is the single biggest lever you have.
That gives a sizing map — Q4-class footprints including context, the practical minimum memory (VRAM, or usable unified memory):
| Model size | Memory (Q4) | Example models |
|---|---|---|
| 1B-4B | 1-4 GB | small Gemma / Phi / Llama |
| 7B-8B | ~5-6 GB | Llama 8B, Mistral 7B, Qwen 7B |
| 13B-14B | ~9-10 GB | 13B chat / coding models |
| 30B-34B | ~20-22 GB | 32B coding / reasoning models |
| 70B | ~42-48 GB | 70B flagship open-weight |
Anchor on the size class, not a release — a "32B coding model" is a target you re-fill with the current best. With Ollama you pick a quant by tag (llama3.1:8b defaults to Q4_K_M; append -q8_0 for Q8). Q4_K_M is the default sweet spot, stepped up only with spare memory.
Context also claims memory: the KV cache grows with conversation length, adding ~1-3 GB at 32K tokens and several GB at 128K — for large documents, size up a tier.
NVIDIA is the path of least resistance: broadest tool support, mature CUDA drivers, most docs. What matters is VRAM (how big a model fits) and compute capability (whether the card is supported). Ollama needs compute capability 5.0+ and driver 531+ (legacy 5.0-6.2 cards want 570+). Run nvidia-smi to check driver and VRAM; most cards from the GTX 10-series on clear the 5.0 bar. What consumer cards run at Q4:
| GPU (VRAM) | Comfortable at Q4 |
|---|---|
| 8 GB (3060 Ti / 4060) | 7B-8B, short contexts |
| 12-16 GB (3060, 4060 Ti) | 13B — best value |
| 24 GB (3090 / 4090) | 32B class — "serious local" |
| 2× 24 GB | 70B, layers split across cards |
Data-centre cards (A100, H100) suit fine-tuning and multi-user serving — overkill for one person, usually rented.
AMD works, but the road is bumpier. Acceleration goes through ROCm, AMD's CUDA equivalent: Ollama needs the ROCm v7 driver stack on Linux (or ROCm v7 / HIP7 on Windows), covering recent Radeon RX, Radeon PRO, Ryzen AI, and Instinct cards. On Linux, confirm the GPU is visible:
rocminfo | grep -i "Marketing Name"
VRAM math matches NVIDIA — a 24 GB Radeon runs the same 32B Q4 model; the difference is driver maturity, not capacity. Acceleration paths, fastest first: native ROCm (supported cards), then Vulkan (unsupported AMD and Intel GPUs), then CPU.
Apple Silicon (M1 through M4, including Pro/Max/Ultra) is unusually good at local LLMs thanks to unified memory: the CPU and GPU share one high-bandwidth pool, so there is no separate VRAM and no PCIe copy. A 64 GB MacBook can give most of that to a model that needs a stack of GPUs on a PC.
Speed depends on unified memory and bandwidth, which climbs steeply up the line — the M4 Max supports up to 128 GB at up to 546 GB/s. Best of all, Ollama enables Metal acceleration automatically — no drivers or CUDA setup:
ollama run llama3.1:8b # uses the Mac GPU, zero config
Mac sizing tracks unified RAM at Q4: 16 GB runs 7B-8B, 32 GB up to ~32B, 64 GB a 70B, 128 GB a 70B at higher quant or long context. The OS reserves some, so not all is usable.
Memory decides what runs; disk decides how many models you keep and how fast they load. A 7B Q4 model is ~4-5 GB, a 70B is ~40 GB — plan for 100 GB+ free and prefer an NVMe SSD (near-instant loads versus tens of seconds on disk). Manage the cache:
ollama list # downloaded models and sizes
ollama rm llama3.1:70b # remove a model to reclaim disk
du -sh ~/.ollama/models # total cache size
To move the cache to a bigger drive, set OLLAMA_MODELS before starting Ollama (default differs by platform):
macOS and Linux default to ~/.ollama/models. Override with export OLLAMA_MODELS=/Volumes/BigSSD/ollama-models, then start Ollama.
Windows defaults to %USERPROFILE%\.ollama\models. Set OLLAMA_MODELS in System Properties → Environment Variables, then restart Ollama.
On a home connection a 40 GB model can take over an hour — pull large models before you need them.
Cannot run the models you want? You need not buy a GPU. Marketplaces rent instances by the hour — handy for trying a 70B, one-off fine-tuning, or serving.
| Provider | Character | Good for |
|---|---|---|
| RunPod | Convenient, container-friendly | Quick spin-ups, serving |
| Vast.ai | Marketplace, often cheapest | Cost-sensitive runs |
| Lambda | Managed, steadier availability | Longer training |
Pricing moves constantly, but the durable picture: a consumer RTX 4090 rents in the low hundreds of cents per hour and data-centre A100/H100 cards in the low single digits of dollars per hour, with marketplaces cheapest on consumer GPUs and managed providers charging a premium for reliability. A few dollars an hour buys months of experiments before buying breaks even — rent for occasional jobs, own for private work.
Questions & Answers
Key Takeaways
- Memory is the gate. Estimate footprint with
params_billions × bytes_per_param × 1.2and fit it in VRAM (or unified memory) with headroom — offloading to RAM costs a 5-20x slowdown. - Quantization is your biggest lever. Q4 cuts a model to ~a quarter of full-precision size with small quality cost — it makes a 7B fit in under 5 GB.
- Three hardware paths. NVIDIA is the default (compute capability 5.0+), AMD works via ROCm or Vulkan with caveats, and Apple Silicon's unified memory runs large models with no setup.
- Size class over model name. Plan around 7B / 13B / 32B / 70B tiers — a 24 GB GPU or 32-64 GB Mac is the sweet spot; 70B needs 64 GB+ or two GPUs.
- Budget disk and bandwidth. Keep ~100 GB free on NVMe, pull large models in advance, manage the cache with
ollama list/rm. - Rent when buying does not pay. Cloud GPUs rent data-centre VRAM by the hour — but renting moves data off-device, so own a box for private work.
Next Steps: Lesson 3: Ollama — The Easiest Way to Start