Hardware Requirements

35 min beginner Lesson 2

Learning Outcomes

  • Calculate a model's memory footprint from its parameter count and quantization
  • Compare NVIDIA, AMD, and Apple Silicon and pick the right path for local inference
  • Map model sizes (7B, 13B, 70B) to realistic RAM and VRAM budgets
  • Plan storage and bandwidth for downloading and caching model weights
  • Decide when to rent a cloud GPU instead of buying hardware

Lesson Plan

Segment Duration Topic
Intro 3 min What hardware gates local inference
Memory math 8 min The VRAM formula, quantization, a sizing map
GPUs 11 min NVIDIA, AMD, and "compute capability"
Apple Silicon 5 min Unified memory and why Macs punch above their weight
Storage & cloud 7 min Disk, bandwidth, cache, renting GPUs
Wrap-up 1 min Benchmark your box, preview Ollama

Before You Begin

Pre-work:

  • Read Lesson 1: Why Run Local? so you know which use cases justify the spend.
  • Find your specs — total RAM, GPU model, free disk. On macOS/Linux, system_profiler SPHardwareDataType, free -h, and nvidia-smi cover most of it.

Shopping List:

  • A terminal you are comfortable in
  • Your machine's RAM, GPU, and VRAM figures
  • No purchases required — this lesson is about deciding what to run before you spend

1 The One Number That Gates Everything: Memory

Local inference is mostly a memory problem. The weights must live in fast memory the compute unit can reach — VRAM on a GPU, or unified memory on Apple Silicon. If they fit you get good speed; spill into system RAM and generation crawls.

Footprint is three things: parameters (trained weights in billions — the "7B" in a name), bytes per parameter (set by quantization — compressing 16-bit weights to 4-bit), and overhead (KV cache plus activations). A practical estimate:

memory_GB ≈ parameters_in_billions × bytes_per_param × 1.2

The 1.2 is rough 20% headroom for KV cache and activations. Bytes per parameter run ~2.0 at FP16 (full), ~1.0 at Q8 (near-lossless), and ~0.55 at Q4_K_M (the sweet spot). So a 7B at Q4_K_M needs roughly 7 × 0.55 × 1.2 ≈ 4.6 GB, versus ~17 GB at FP16. Quantization is the single biggest lever you have.

That gives a sizing map — Q4-class footprints including context, the practical minimum memory (VRAM, or usable unified memory):

Model size Memory (Q4) Example models
1B-4B 1-4 GB small Gemma / Phi / Llama
7B-8B ~5-6 GB Llama 8B, Mistral 7B, Qwen 7B
13B-14B ~9-10 GB 13B chat / coding models
30B-34B ~20-22 GB 32B coding / reasoning models
70B ~42-48 GB 70B flagship open-weight

Anchor on the size class, not a release — a "32B coding model" is a target you re-fill with the current best. With Ollama you pick a quant by tag (llama3.1:8b defaults to Q4_K_M; append -q8_0 for Q8). Q4_K_M is the default sweet spot, stepped up only with spare memory.

Context also claims memory: the KV cache grows with conversation length, adding ~1-3 GB at 32K tokens and several GB at 128K — for large documents, size up a tier.

WARNING
Bigger is not automatically better
An 8B model that fits in VRAM and runs fast often beats a 70B that spills to RAM and crawls at a token a second. Fit and speed matter as much as parameter count. Model families are compared in Lesson 4.

2 NVIDIA GPUs — the Default Path

NVIDIA is the path of least resistance: broadest tool support, mature CUDA drivers, most docs. What matters is VRAM (how big a model fits) and compute capability (whether the card is supported). Ollama needs compute capability 5.0+ and driver 531+ (legacy 5.0-6.2 cards want 570+). Run nvidia-smi to check driver and VRAM; most cards from the GTX 10-series on clear the 5.0 bar. What consumer cards run at Q4:

GPU (VRAM) Comfortable at Q4
8 GB (3060 Ti / 4060) 7B-8B, short contexts
12-16 GB (3060, 4060 Ti) 13B — best value
24 GB (3090 / 4090) 32B class — "serious local"
2× 24 GB 70B, layers split across cards

Data-centre cards (A100, H100) suit fine-tuning and multi-user serving — overkill for one person, usually rented.

WARNING
Spilling into RAM kills speed
When a model exceeds your VRAM, Ollama offloads the overflow to system RAM. That prevents a crash but can slow generation 5-20x. Size the model to fit in VRAM with headroom, not barely.

3 AMD GPUs — Viable, With Caveats

AMD works, but the road is bumpier. Acceleration goes through ROCm, AMD's CUDA equivalent: Ollama needs the ROCm v7 driver stack on Linux (or ROCm v7 / HIP7 on Windows), covering recent Radeon RX, Radeon PRO, Ryzen AI, and Instinct cards. On Linux, confirm the GPU is visible:

rocminfo | grep -i "Marketing Name"

VRAM math matches NVIDIA — a 24 GB Radeon runs the same 32B Q4 model; the difference is driver maturity, not capacity. Acceleration paths, fastest first: native ROCm (supported cards), then Vulkan (unsupported AMD and Intel GPUs), then CPU.

TIP
Check support before you buy
Choosing an AMD card for local LLMs? Verify it is on the current ROCm support list — a supported 24 GB Radeon is great value; an unsupported card stuck on Vulkan or CPU is not.

4 Apple Silicon — Unified Memory Changes the Math

Apple Silicon (M1 through M4, including Pro/Max/Ultra) is unusually good at local LLMs thanks to unified memory: the CPU and GPU share one high-bandwidth pool, so there is no separate VRAM and no PCIe copy. A 64 GB MacBook can give most of that to a model that needs a stack of GPUs on a PC.

Speed depends on unified memory and bandwidth, which climbs steeply up the line — the M4 Max supports up to 128 GB at up to 546 GB/s. Best of all, Ollama enables Metal acceleration automatically — no drivers or CUDA setup:

ollama run llama3.1:8b   # uses the Mac GPU, zero config

Mac sizing tracks unified RAM at Q4: 16 GB runs 7B-8B, 32 GB up to ~32B, 64 GB a 70B, 128 GB a 70B at higher quant or long context. The OS reserves some, so not all is usable.

NOTE
The Mac trade-off
A single 64 GB Mac runs a 70B from one shared pool — quiet, low power, no multi-GPU plumbing — where a PC needs two 24 GB GPUs. The catch: peak throughput trails a high-end NVIDIA card. Macs win on capacity and convenience, not speed.

5 Storage and Download Bandwidth

Memory decides what runs; disk decides how many models you keep and how fast they load. A 7B Q4 model is ~4-5 GB, a 70B is ~40 GB — plan for 100 GB+ free and prefer an NVMe SSD (near-instant loads versus tens of seconds on disk). Manage the cache:

ollama list              # downloaded models and sizes
ollama rm llama3.1:70b   # remove a model to reclaim disk
du -sh ~/.ollama/models  # total cache size

To move the cache to a bigger drive, set OLLAMA_MODELS before starting Ollama (default differs by platform):

macOS and Linux default to ~/.ollama/models. Override with export OLLAMA_MODELS=/Volumes/BigSSD/ollama-models, then start Ollama.

Windows defaults to %USERPROFILE%\.ollama\models. Set OLLAMA_MODELS in System Properties → Environment Variables, then restart Ollama.

On a home connection a 40 GB model can take over an hour — pull large models before you need them.

NOTE
Quantization saves disk too
Q4 shrinks the download as well: a 70B that is ~140 GB at FP16 is ~40 GB at Q4 — a quarter of the bytes to fetch and store.

6 Cloud GPUs — Rent Instead of Buy

Cannot run the models you want? You need not buy a GPU. Marketplaces rent instances by the hour — handy for trying a 70B, one-off fine-tuning, or serving.

Provider Character Good for
RunPod Convenient, container-friendly Quick spin-ups, serving
Vast.ai Marketplace, often cheapest Cost-sensitive runs
Lambda Managed, steadier availability Longer training

Pricing moves constantly, but the durable picture: a consumer RTX 4090 rents in the low hundreds of cents per hour and data-centre A100/H100 cards in the low single digits of dollars per hour, with marketplaces cheapest on consumer GPUs and managed providers charging a premium for reliability. A few dollars an hour buys months of experiments before buying breaks even — rent for occasional jobs, own for private work.

WARNING
Rented does not mean private
Cloud GPUs run your data on someone else's machine, undercutting the privacy argument from Lesson 1 — rent for testing and training, own for confidential work. And hourly billing runs until you destroy the instance, not when you log out — tear it down when you finish.

Questions & Answers

Q: I have 16 GB of RAM and no dedicated GPU. Am I locked out?
No — small models run on CPU. A 7B Q4 model fits in ~5 GB and generates text, just slowly (a handful of tokens per second). Stick to 7B-and-under quantized models and short contexts. A cheap 8-12 GB GPU is the biggest single upgrade.
Q: Should I buy a Mac or build a PC with an NVIDIA card?
For broadest tool compatibility and peak throughput, NVIDIA wins — most projects target CUDA first. For the largest model on a single quiet, low-power box, a high-memory Mac wins: unified memory lets one machine hold a 70B. Want it to "just work"? Apple Silicon is hard to beat.
Q: Does more VRAM make generation faster, or just allow bigger models?
Primarily it lets bigger models fit without spilling to slow RAM. Raw token speed is driven more by bandwidth and compute than capacity — but the worst slowdowns come from overflowing VRAM, so enough headroom to avoid offload is the biggest speed factor in practice. (Lesson 9.)
Q: Can I split a model across two GPUs?
Yes — Ollama and llama.cpp split a model's layers across GPUs, which is how people run 70B models on two 24 GB cards. They need not be identical, but matched cards perform most predictably. (Lesson 9.)

Key Takeaways

  1. Memory is the gate. Estimate footprint with params_billions × bytes_per_param × 1.2 and fit it in VRAM (or unified memory) with headroom — offloading to RAM costs a 5-20x slowdown.
  2. Quantization is your biggest lever. Q4 cuts a model to ~a quarter of full-precision size with small quality cost — it makes a 7B fit in under 5 GB.
  3. Three hardware paths. NVIDIA is the default (compute capability 5.0+), AMD works via ROCm or Vulkan with caveats, and Apple Silicon's unified memory runs large models with no setup.
  4. Size class over model name. Plan around 7B / 13B / 32B / 70B tiers — a 24 GB GPU or 32-64 GB Mac is the sweet spot; 70B needs 64 GB+ or two GPUs.
  5. Budget disk and bandwidth. Keep ~100 GB free on NVMe, pull large models in advance, manage the cache with ollama list / rm.
  6. Rent when buying does not pay. Cloud GPUs rent data-centre VRAM by the hour — but renting moves data off-device, so own a box for private work.

Next Steps: Lesson 3: Ollama — The Easiest Way to Start