Fine-Tuning Basics
Learning Outcomes
- Decide when fine-tuning beats prompt engineering or RAG
- Explain how LoRA and QLoRA make training feasible on consumer GPUs
- Build a clean instruction dataset in JSONL chat format
- Run a QLoRA training job with Unsloth and with Axolotl
- Evaluate the tuned model and ship it to Ollama as a GGUF
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | What fine-tuning is and is not |
| Decision | 7 min | Fine-tune vs prompt vs RAG |
| Explain | 8 min | LoRA and QLoRA on consumer hardware |
| Demo | 8 min | Preparing a JSONL dataset |
| Demo | 10 min | QLoRA training with Unsloth |
| Demo | 6 min | The same job with Axolotl |
| Demo | 5 min | Evaluation and overfitting |
| Wrap-up | 3 min | Export to Ollama, takeaways |
Before You Begin
Pre-work:
- Finish Lesson 4: Model Selection for parameter counts and quantization
- Finish Lesson 6: Connecting Local Models to Tools to serve the result
- Skim the Hardware Reference to confirm VRAM headroom
Shopping List:
- An NVIDIA GPU with 8 GB+ VRAM, or a cloud GPU (RunPod / Vast.ai / Lambda)
- Python 3.10+ and a fresh virtual environment
- A Hugging Face account and token (
huggingface-cli login) - 30-60 GB free disk for weights, checkpoints, and GGUF exports
- Ollama installed for the export step
Fine-tuning changes the weights so the model behaves differently by default. Powerful, and often the wrong tool — most "it won't do what I want" problems are solved more cheaply by a better prompt or by retrieval. Fine-tune only to change behaviour, format, or style, not to add facts.
| Need | Best tool | Why |
|---|---|---|
| Answer about your private docs | RAG (Lesson 8) | Retrieval stays current without retraining |
| Always reply in a strict JSON shape | Fine-tune | Format baked into weights |
| Adopt a house tone or persona | Fine-tune | Style is hard to enforce by prompt |
| One-off task with clear instructions | Prompt engineering | Zero cost, instant |
| Domain jargon and reasoning patterns | Fine-tune + RAG | Tune the style, retrieve the facts |
Rule of thumb: if the answer changes when your data changes, use RAG; if the way it answers must change, fine-tune. Try a prompt first — it is free.
Full fine-tuning updates every weight — for a 7B model in 16-bit, optimizer state needs 80+ GB of VRAM, beyond any consumer card.
LoRA (Low-Rank Adaptation) freezes the original weights and trains a small pair of low-rank matrices added to specific layers — a few million parameters, not billions. The output is a small adapter (tens of MB). QLoRA loads the frozen base in 4-bit before training the adapter, letting a 7-9B model tune inside ~6-8 GB. Per Unsloth, QLoRA's 4-bit base cuts memory ~4x versus LoRA's 16-bit.
The knobs you will touch most, with Unsloth's recommended starts :
| Parameter | Controls | Start |
|---|---|---|
r (rank) |
Adapter capacity | 16 or 32 |
lora_alpha |
Effect scaling | r, or 2 * r |
target_modules |
Layers adapted | all attention + MLP projections |
learning_rate |
Step size | 2e-4 |
| epochs | Passes over data | 1-3 |
Data quality dominates — 200 clean examples beat 5,000 noisy ones. The format is JSONL (one JSON object per line); the model-agnostic shape is chat (messages):
{"messages": [{"role": "system", "content": "You are a terse SRE assistant; reply only with shell commands."}, {"role": "user", "content": "Disk usage by directory, largest first."}, {"role": "assistant", "content": "du -h --max-depth=1 . | sort -rh"}]}
Keep the system prompt identical to what you serve with. The older "alpaca" format (instruction/input/output) works too, but chat travels better across model families. Validate first — one bad line wastes a run: python -c "import json;[json.loads(l) for l in open('train.jsonl')]".
Data hygiene that matters most:
| Rule | Why |
|---|---|
| Deduplicate examples | Repeats cause memorization, not learning |
| Hold out 5-10% for eval | Cannot measure overfitting without unseen data |
| Keep responses consistent in style | The model learns the average of your data |
| Match the inference system prompt | Train/serve mismatch silently degrades quality |
val_set_size: 0.1), and spend your time here, not on hyperparameters.Unsloth patches the Hugging Face stack for faster, lower-memory single-GPU runs. Install, then run a QLoRA script:
python -m venv .venv && source .venv/bin/activate
pip install unsloth
from unsloth import FastLanguageModel
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Qwen2.5-3B-Instruct",
max_seq_length=2048, load_in_4bit=True,
)
model = FastLanguageModel.get_peft_model(
model, r=16, lora_alpha=16, use_gradient_checkpointing="unsloth",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
)
dataset = load_dataset("json", data_files="train.jsonl", split="train")
trainer = SFTTrainer(
model=model, tokenizer=tokenizer, train_dataset=dataset,
args=SFTConfig(
per_device_train_batch_size=2, gradient_accumulation_steps=4,
learning_rate=2e-4, num_train_epochs=2, output_dir="outputs",
),
)
trainer.train()
# Save the adapter, plus a merged 16-bit copy for export
model.save_pretrained("qwen-sre-lora")
model.save_pretrained_merged("qwen-sre-merged", tokenizer, save_method="merged_16bit")
Run it and watch the loss drop then flatten. On out-of-memory, lower max_seq_length and the batch size, raising gradient_accumulation_steps.
per_device_train_batch_size x gradient_accumulation_steps — here, batch 8 while holding only 2 sequences in memory. Once a run works, pip freeze > requirements.txt to keep it reproducible.Axolotl is config-driven: describe the run in YAML and launch with one command — better for reproducible recipes. Create qlora.yml:
base_model: Qwen/Qwen2.5-3B-Instruct
load_in_4bit: true
adapter: qlora
datasets:
- path: train.jsonl
type: chat_template
val_set_size: 0.1
output_dir: ./outputs/qlora-out
sequence_len: 2048
lora_r: 16
lora_alpha: 16
lora_target_modules: [q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj]
learning_rate: 0.0002
num_epochs: 2
micro_batch_size: 2
gradient_accumulation_steps: 4
After pip install axolotl, preprocess then train. Per the quickstart, load_in_4bit + adapter: qlora picks the method:
axolotl preprocess qlora.yml
axolotl train qlora.yml
A model that scores great on training data but worse in real use has overfit. Two failure modes:
- Overfitting: training loss keeps falling while eval loss rises — the gap is the tell.
- Catastrophic forgetting: good at your task, worse at everything else. LoRA reduces this, but heavy training on a narrow dataset still causes it.
Generate on held-out prompts and compare tuned vs base:
from unsloth import FastLanguageModel
model, tok = FastLanguageModel.from_pretrained("qwen-sre-merged", load_in_4bit=True)
FastLanguageModel.for_inference(model)
for line in open("eval_prompts.txt"):
msgs = [{"role": "user", "content": line.strip()}]
ids = tok.apply_chat_template(msgs, return_tensors="pt", add_generation_prompt=True).to("cuda")
print(line.strip(), "->", tok.decode(model.generate(ids, max_new_tokens=128)[0], skip_special_tokens=True))
Read four signals: eval loss diverging from train loss (overfitting), held-out accuracy (gains), general-knowledge checks (forgetting), and the base comparison. To fix overfitting, drop to 1 epoch, lower the LR, or add varied data; for forgetting, fewer epochs, mix in general examples, or lower the rank.
An adapter is useless until you can run it. The cleanest path is a merged GGUF in Ollama, reusing the Lesson 6 tooling. Export with Unsloth's save_pretrained_gguf and a quantization_method like q4_k_m:
model.save_pretrained_gguf("qwen-sre-gguf", tokenizer, quantization_method="q4_k_m")
Then write a Modelfile — FROM takes a GGUF directly and ollama create registers it:
FROM ./qwen-sre-gguf/unsloth.Q4_K_M.gguf
SYSTEM "You are a terse SRE assistant; reply only with shell commands."
PARAMETER temperature 0.3
ollama create qwen-sre -f Modelfile
ollama run qwen-sre "reload nginx without dropping connections"
Keep the Modelfile and GGUF in the same folder. To skip the merge instead, point FROM at the base and add ADAPTER ./qwen-sre-lora.
Apple Silicon cannot do the QLoRA step, so train on a rented NVIDIA GPU, scp the GGUF to your Mac, then serve over Metal. (Linux matches this usage.)
Run the export and ollama create inside WSL2 so paths and CUDA behave; reference the GGUF with a WSL path, not C:\.
q4_k_m for the best size/quality balance, or q8_0 with ample VRAM — see the [Model Guide](/courses/06-local-llms/supplemental/model-guide/).Questions & Answers
pip freeze > requirements.txt and commit it; use a clean venv per project. See [Troubleshooting](/courses/06-local-llms/supplemental/troubleshooting/).Key Takeaways
- Fine-tune behaviour, not facts. If the answer changes with your data, use RAG; if the way it answers must change, fine-tune — and try a prompt first.
- LoRA/QLoRA make it feasible. Freezing the base and training a small low-rank adapter (4-bit base for QLoRA) brings 3-9B fine-tuning into single-GPU memory.
- Defaults exist. Rank 16-32, alpha = rank or 2x rank, all attention + MLP projections, lr 2e-4, 1-3 epochs. Tune only if evaluation tells you to.
- Data quality is the whole game. Clean, consistent, deduplicated JSONL with a held-out split beats any hyperparameter trick.
- Evaluate against the base on held-out and out-of-distribution prompts to catch overfitting and forgetting, then export to GGUF for Ollama.
Next Steps: Lesson 8: Local RAG — Your Own Knowledge Base