Fine-Tuning Basics

50 min intermediate Lesson 7

Learning Outcomes

  • Decide when fine-tuning beats prompt engineering or RAG
  • Explain how LoRA and QLoRA make training feasible on consumer GPUs
  • Build a clean instruction dataset in JSONL chat format
  • Run a QLoRA training job with Unsloth and with Axolotl
  • Evaluate the tuned model and ship it to Ollama as a GGUF

Lesson Plan

Segment Duration Topic
Intro 3 min What fine-tuning is and is not
Decision 7 min Fine-tune vs prompt vs RAG
Explain 8 min LoRA and QLoRA on consumer hardware
Demo 8 min Preparing a JSONL dataset
Demo 10 min QLoRA training with Unsloth
Demo 6 min The same job with Axolotl
Demo 5 min Evaluation and overfitting
Wrap-up 3 min Export to Ollama, takeaways

Before You Begin

Pre-work:

Shopping List:

  • An NVIDIA GPU with 8 GB+ VRAM, or a cloud GPU (RunPod / Vast.ai / Lambda)
  • Python 3.10+ and a fresh virtual environment
  • A Hugging Face account and token (huggingface-cli login)
  • 30-60 GB free disk for weights, checkpoints, and GGUF exports
  • Ollama installed for the export step
WARNING
Apple Silicon note
QLoRA leans on 4-bit CUDA kernels (bitsandbytes) that do not run on Metal. On a Mac, do small LoRA runs with MLX; otherwise the flows below assume an NVIDIA GPU — or rent a cloud one by the hour.

1 Fine-Tune, Prompt, or RAG?

Fine-tuning changes the weights so the model behaves differently by default. Powerful, and often the wrong tool — most "it won't do what I want" problems are solved more cheaply by a better prompt or by retrieval. Fine-tune only to change behaviour, format, or style, not to add facts.

Need Best tool Why
Answer about your private docs RAG (Lesson 8) Retrieval stays current without retraining
Always reply in a strict JSON shape Fine-tune Format baked into weights
Adopt a house tone or persona Fine-tune Style is hard to enforce by prompt
One-off task with clear instructions Prompt engineering Zero cost, instant
Domain jargon and reasoning patterns Fine-tune + RAG Tune the style, retrieve the facts

Rule of thumb: if the answer changes when your data changes, use RAG; if the way it answers must change, fine-tune. Try a prompt first — it is free.

NOTE
Fine-tuning does not teach facts reliably
Facts pushed into weights produce confident hallucinations that go stale when your data updates. Treat fine-tuning as behaviour shaping; let RAG carry the knowledge.

2 How LoRA and QLoRA Make This Feasible

Full fine-tuning updates every weight — for a 7B model in 16-bit, optimizer state needs 80+ GB of VRAM, beyond any consumer card.

LoRA (Low-Rank Adaptation) freezes the original weights and trains a small pair of low-rank matrices added to specific layers — a few million parameters, not billions. The output is a small adapter (tens of MB). QLoRA loads the frozen base in 4-bit before training the adapter, letting a 7-9B model tune inside ~6-8 GB. Per Unsloth, QLoRA's 4-bit base cuts memory ~4x versus LoRA's 16-bit.

The knobs you will touch most, with Unsloth's recommended starts :

Parameter Controls Start
r (rank) Adapter capacity 16 or 32
lora_alpha Effect scaling r, or 2 * r
target_modules Layers adapted all attention + MLP projections
learning_rate Step size 2e-4
epochs Passes over data 1-3
NOTE
Why all the projections
Targeting both attention and MLP projections (not just attention) lets a LoRA run approach full fine-tuning quality, for a little memory.

3 Preparing Your Training Data

Data quality dominates — 200 clean examples beat 5,000 noisy ones. The format is JSONL (one JSON object per line); the model-agnostic shape is chat (messages):

{"messages": [{"role": "system", "content": "You are a terse SRE assistant; reply only with shell commands."}, {"role": "user", "content": "Disk usage by directory, largest first."}, {"role": "assistant", "content": "du -h --max-depth=1 . | sort -rh"}]}

Keep the system prompt identical to what you serve with. The older "alpaca" format (instruction/input/output) works too, but chat travels better across model families. Validate first — one bad line wastes a run: python -c "import json;[json.loads(l) for l in open('train.jsonl')]".

Data hygiene that matters most:

Rule Why
Deduplicate examples Repeats cause memorization, not learning
Hold out 5-10% for eval Cannot measure overfitting without unseen data
Keep responses consistent in style The model learns the average of your data
Match the inference system prompt Train/serve mismatch silently degrades quality
WARNING
Garbage in, confident garbage out
The model reproduces sloppy formatting, inconsistent tone, and wrong answers in your data. Carve off an eval set up front (Axolotl: val_set_size: 0.1), and spend your time here, not on hyperparameters.

4 Training a QLoRA Adapter with Unsloth

Unsloth patches the Hugging Face stack for faster, lower-memory single-GPU runs. Install, then run a QLoRA script:

python -m venv .venv && source .venv/bin/activate
pip install unsloth
from unsloth import FastLanguageModel
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/Qwen2.5-3B-Instruct",
    max_seq_length=2048, load_in_4bit=True,
)
model = FastLanguageModel.get_peft_model(
    model, r=16, lora_alpha=16, use_gradient_checkpointing="unsloth",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
)
dataset = load_dataset("json", data_files="train.jsonl", split="train")
trainer = SFTTrainer(
    model=model, tokenizer=tokenizer, train_dataset=dataset,
    args=SFTConfig(
        per_device_train_batch_size=2, gradient_accumulation_steps=4,
        learning_rate=2e-4, num_train_epochs=2, output_dir="outputs",
    ),
)
trainer.train()

# Save the adapter, plus a merged 16-bit copy for export
model.save_pretrained("qwen-sre-lora")
model.save_pretrained_merged("qwen-sre-merged", tokenizer, save_method="merged_16bit")

Run it and watch the loss drop then flatten. On out-of-memory, lower max_seq_length and the batch size, raising gradient_accumulation_steps.

TIP
Effective batch size
Effective batch size = per_device_train_batch_size x gradient_accumulation_steps — here, batch 8 while holding only 2 sequences in memory. Once a run works, pip freeze > requirements.txt to keep it reproducible.

5 The Same Job with Axolotl

Axolotl is config-driven: describe the run in YAML and launch with one command — better for reproducible recipes. Create qlora.yml:

base_model: Qwen/Qwen2.5-3B-Instruct
load_in_4bit: true
adapter: qlora
datasets:
  - path: train.jsonl
    type: chat_template
val_set_size: 0.1
output_dir: ./outputs/qlora-out
sequence_len: 2048
lora_r: 16
lora_alpha: 16
lora_target_modules: [q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj]
learning_rate: 0.0002
num_epochs: 2
micro_batch_size: 2
gradient_accumulation_steps: 4

After pip install axolotl, preprocess then train. Per the quickstart, load_in_4bit + adapter: qlora picks the method:

axolotl preprocess qlora.yml
axolotl train qlora.yml
NOTE
Same concepts, different ergonomics
Unsloth suits fast single-GPU runs and notebooks; Axolotl suits reproducible YAML recipes and multi-GPU. Both use identical knobs — rank 16, alpha 16, seven modules, lr 2e-4, 2 epochs. The technique is constant; only the expression differs.

6 Evaluating and Avoiding Overfitting

A model that scores great on training data but worse in real use has overfit. Two failure modes:

  • Overfitting: training loss keeps falling while eval loss rises — the gap is the tell.
  • Catastrophic forgetting: good at your task, worse at everything else. LoRA reduces this, but heavy training on a narrow dataset still causes it.

Generate on held-out prompts and compare tuned vs base:

from unsloth import FastLanguageModel
model, tok = FastLanguageModel.from_pretrained("qwen-sre-merged", load_in_4bit=True)
FastLanguageModel.for_inference(model)
for line in open("eval_prompts.txt"):
    msgs = [{"role": "user", "content": line.strip()}]
    ids = tok.apply_chat_template(msgs, return_tensors="pt", add_generation_prompt=True).to("cuda")
    print(line.strip(), "->", tok.decode(model.generate(ids, max_new_tokens=128)[0], skip_special_tokens=True))

Read four signals: eval loss diverging from train loss (overfitting), held-out accuracy (gains), general-knowledge checks (forgetting), and the base comparison. To fix overfitting, drop to 1 epoch, lower the LR, or add varied data; for forgetting, fewer epochs, mix in general examples, or lower the rank.

WARNING
Beware the demo trap
Testing only on prompts that resemble training data hides overfitting. Always include prompts deliberately *outside* the training distribution.

7 Shipping It Back to Ollama

An adapter is useless until you can run it. The cleanest path is a merged GGUF in Ollama, reusing the Lesson 6 tooling. Export with Unsloth's save_pretrained_gguf and a quantization_method like q4_k_m:

model.save_pretrained_gguf("qwen-sre-gguf", tokenizer, quantization_method="q4_k_m")

Then write a Modelfile — FROM takes a GGUF directly and ollama create registers it:

FROM ./qwen-sre-gguf/unsloth.Q4_K_M.gguf
SYSTEM "You are a terse SRE assistant; reply only with shell commands."
PARAMETER temperature 0.3
ollama create qwen-sre -f Modelfile
ollama run qwen-sre "reload nginx without dropping connections"

Keep the Modelfile and GGUF in the same folder. To skip the merge instead, point FROM at the base and add ADAPTER ./qwen-sre-lora.

Apple Silicon cannot do the QLoRA step, so train on a rented NVIDIA GPU, scp the GGUF to your Mac, then serve over Metal. (Linux matches this usage.)

Run the export and ollama create inside WSL2 so paths and CUDA behave; reference the GGUF with a WSL path, not C:\.

TIP
Quantize for the target machine
Use q4_k_m for the best size/quality balance, or q8_0 with ample VRAM — see the [Model Guide](/courses/06-local-llms/supplemental/model-guide/).

Questions & Answers

Q: I have a 6 GB laptop GPU. Can I fine-tune at all, or is this hopeless?
You can do QLoRA on a small (1-3B) model at short sequence length and batch 1, but slowly. For serious runs, rent a cloud GPU (RunPod, Vast.ai, Lambda) by the hour, then pull the GGUF down to serve locally — training and serving need not share hardware.
Q: How many examples do I actually need?
For style and format shaping, a few hundred consistent examples often suffice; for complex behaviour, low thousands. Quality matters far more than count — add data before epochs.
Q: My tuned model nails my test prompts but got dumber at general questions. What happened?
Catastrophic forgetting — you trained too hard on a narrow distribution. Drop to 1 epoch, lower the LR, mix in general examples, or reduce the rank. LoRA forgets less than full fine-tuning, but it still can.
Q: The toolchain broke after a pip upgrade. How do I keep runs reproducible?
Minor bumps in transformers, trl, or bitsandbytes regularly break things. Once a run works, pip freeze > requirements.txt and commit it; use a clean venv per project. See [Troubleshooting](/courses/06-local-llms/supplemental/troubleshooting/).

Key Takeaways

  1. Fine-tune behaviour, not facts. If the answer changes with your data, use RAG; if the way it answers must change, fine-tune — and try a prompt first.
  2. LoRA/QLoRA make it feasible. Freezing the base and training a small low-rank adapter (4-bit base for QLoRA) brings 3-9B fine-tuning into single-GPU memory.
  3. Defaults exist. Rank 16-32, alpha = rank or 2x rank, all attention + MLP projections, lr 2e-4, 1-3 epochs. Tune only if evaluation tells you to.
  4. Data quality is the whole game. Clean, consistent, deduplicated JSONL with a held-out split beats any hyperparameter trick.
  5. Evaluate against the base on held-out and out-of-distribution prompts to catch overfitting and forgetting, then export to GGUF for Ollama.

Next Steps: Lesson 8: Local RAG — Your Own Knowledge Base