Computer Vision in 2026

60 min intermediate Lesson 3

Learning Outcomes

  • Distinguish the four core vision tasks — classification, detection, segmentation, and pose — and pick the right one for a problem
  • Run a real object-detection pipeline with YOLO and read boxes, confidence, and class labels from the results
  • Apply transfer learning to fine-tune a detector on your own images instead of training from scratch
  • Extract text from images with OCR and segment objects pixel-precisely with SAM and Mask R-CNN
  • Decide when a specialised vision model beats a multimodal LLM, and deploy one to an edge device

Lesson Plan

Segment Duration Topic
Intro 3 min Why most production vision is not an LLM
Concepts 9 min The four vision tasks and how CNNs see
Build 12 min Object detection with YOLO, end to end
Build 10 min Transfer learning — fine-tuning on your data
Build 8 min Segmentation: SAM and Mask R-CNN
Build 7 min Pose estimation and OCR
Decide 7 min Specialised models vs multimodal LLMs
Wrap-up 4 min Edge deployment and key takeaways

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ and a fresh virtual environment
  • pip install ultralytics torch torchvision transformers pillow (a few GB; first run downloads model weights)
  • A GPU helps for training but is not required — every example here runs on CPU, just slower
  • A handful of your own images (a photo with people/objects, and one with printed or handwritten text)

1 The Four Vision Tasks (and Why CNNs Beat LLMs Here)

Before any code, build the mental map. Computer vision is not one problem but four, and the right architecture depends on which you have.

Task Question it answers Output Real-world use
Classification "What is in this image?" One label per image Defect / no-defect sorting
Detection "What is here, and where?" Boxes + labels + scores Counting cars, finding PPE on workers
Segmentation "Which exact pixels are the object?" Per-pixel mask Tumour outlining, background removal
Pose estimation "Where are the keypoints?" Joint coordinates Ergonomics, sports analytics, fall detection

The workhorse behind all four is the convolutional neural network (CNN) — early layers learn small reusable filters (edges, textures), deeper layers compose them into shapes and objects. Convolution just means sliding a filter across the image and recording where it matches, which is why a CNN finds a cat in any corner of the frame without being told where to look.

A fair 2026 question: why not just send the image to a multimodal LLM? For these structured tasks, specialised CNNs win on the axes that matter in production.

Dimension Specialised vision CNN Multimodal LLM
Latency Milliseconds, runs on a camera Hundreds of ms to seconds, needs a server
Cost One-time training, free inference Per-image API cost forever
Output Exact pixel coordinates / boxes Prose you must parse
Throughput Thousands of frames/sec on a GPU Rate-limited
Determinism Same input → same boxes Can vary run to run
NOTE
Key Insight
An assembly line inspecting 60 parts per second cannot wait on an API and cannot tolerate non-deterministic answers. That is why the overwhelming majority of deployed vision is task-specific CNNs, not chat models. The LLM is for open-ended visual reasoning, not high-throughput measurement.

2 Object Detection with YOLO, End to End

Object detection finds every object and draws a bounding box (a rectangle given as pixel corners) around each, with a class label and confidence score. The dominant real-time family is YOLO ("You Only Look Once") — it processes the whole image in a single forward pass, fast enough for video. Current Ultralytics releases are the YOLO11 family and the newer YOLO26 flagship, both covering detection, segmentation, pose, classification, and oriented boxes from one API.

Run a pretrained detector — the *.pt weights file downloads automatically on first use:

from ultralytics import YOLO

# Load a small pretrained detector (trained on the 80-class COCO dataset).
model = YOLO("yolo11n.pt")   # 'n' = nano, the fastest variant

# Run inference on an image (a path, URL, PIL image, or numpy array).
results = model.predict("street.jpg", conf=0.35)

r = results[0]               # results is a list, one entry per image
for box in r.boxes:
    cls_id = int(box.cls[0])           # class index
    label  = r.names[cls_id]           # human-readable class name
    score  = float(box.conf[0])        # confidence 0..1
    x1, y1, x2, y2 = box.xyxy[0].tolist()   # corner pixels
    print(f"{label:12s} {score:.2f}  box=({x1:.0f},{y1:.0f},{x2:.0f},{y2:.0f})")

r.save(filename="street_annotated.jpg")   # save a copy with boxes drawn

You will see lines like person 0.91 box=(412,88,503,377). The conf=0.35 argument is the confidence threshold — detections below it are discarded. Raise it to cut false positives; lower it to catch faint objects you are missing.

Three metrics define detector quality, and you must know all three:

  • Precision — of the boxes the model drew, how many were correct. Low precision = too many false alarms.
  • Recall — of the objects actually present, how many it found. Low recall = it missed things.
  • mAP (mean Average Precision) — the standard single-number summary balancing precision and recall across all classes and box-overlap thresholds. Higher is better; it is what you compare models on.
TIP
Pick the smallest model that clears your bar
Variants trade accuracy for speed: nano is built for edge/real-time, larger ones for maximum mAP on a server. Start with nano, measure precision and recall on YOUR images, scale up only if you miss the target. Bigger costs latency and power on every single frame.
WARNING
COCO classes are generic
Pretrained YOLO knows 80 everyday categories (person, car, bottle...). It does not know cracked_solder_joint or ripe_strawberry. For anything domain-specific you must fine-tune — that is the next step.

3 Transfer Learning — Fine-Tuning on Your Own Images

Transfer learning is the single most important idea in applied vision. Instead of training from random weights (millions of images, days of compute), you start from a model already trained on a huge dataset and continue training it on your small one. The early CNN layers — edges, textures, shapes — already generalise, so you only teach the model your specific objects. This turns "I need 1,000,000 labelled images" into "I need a few hundred."

To fine-tune YOLO, supply a tiny dataset config and labelled images — one .txt per image, each line class_id x_center y_center width height with coordinates normalised 0–1. A dataset YAML describes where the data lives and the class names:

# data.yaml
path: ./pallet-dataset      # dataset root
train: images/train         # images for training
val: images/val             # images held out to measure performance
names:
  0: pallet
  1: forklift
  2: damaged_pallet

Then fine-tune with a few lines of Python:

from ultralytics import YOLO

# Start from pretrained weights — this is the transfer-learning step.
model = YOLO("yolo11n.pt")

model.train(
    data="data.yaml",
    epochs=50,          # passes over the training set
    imgsz=640,          # input resolution
    batch=16,
    patience=10,        # early-stop if val metrics stop improving
)

# Ultralytics evaluates on the val split automatically.
metrics = model.val()
print(metrics.box.map)     # mAP averaged over overlap thresholds
print(metrics.box.mp)      # mean precision
print(metrics.box.mr)      # mean recall

Two terms from above, defined plainly:

  • Epoch — one full pass over your training images. Too few and the model underfits (hasn't learned); too many and it overfits — memorising the training set so well it fails on new images.
  • Validation split — images the model never trains on, used only to measure honest performance. The gap between training and validation accuracy is your early-warning light for overfitting (see Lesson 2).
TIP
Label quality beats label quantity
300 carefully, consistently labelled images out-train 3,000 sloppy ones. Inconsistent boxes (sometimes including the shadow, sometimes not) teach the model contradictions. Audit a random sample of your labels before you ever hit train.
WARNING
Watch the train/val gap
If training mAP keeps climbing while validation mAP plateaus or drops, you are overfitting. The fix is rarely a bigger model — it is more varied data, augmentation, or stopping earlier. The patience argument automates that last one.

4 Pixel-Precise Segmentation: SAM and Mask R-CNN

A bounding box says "a defect is somewhere in this rectangle." Segmentation goes further: it labels every pixel, producing a mask that traces the object's exact outline — needed for measuring area, removing backgrounds, or medical/satellite imagery where shape matters. Two flavours answer different questions.

Mask R-CNN is the classic instance segmentation model: like a detector, but each detection also carries a pixel mask, trained on labelled classes ("this blob is a car, that one a person"). It is pretrained in torchvision:

import torch
from torchvision.models.detection import (
    maskrcnn_resnet50_fpn, MaskRCNN_ResNet50_FPN_Weights,
)
from torchvision.io import read_image

weights = MaskRCNN_ResNet50_FPN_Weights.DEFAULT
model = maskrcnn_resnet50_fpn(weights=weights).eval()
preprocess = weights.transforms()

img = read_image("street.jpg")
with torch.no_grad():
    out = model([preprocess(img)])[0]

labels = weights.meta["categories"]
for label_id, score, mask in zip(out["labels"], out["scores"], out["masks"]):
    if score > 0.7:
        print(labels[label_id], float(score), "mask:", tuple(mask.shape))

SAM 2 (Segment Anything Model 2, from Meta) is different: a promptable, class-agnostic foundation model. Give it a point or box prompt and it returns the mask for whatever object is there — no training, across both images and video. It does not know what the thing is called; it just isolates it. Via Ultralytics:

from ultralytics import SAM

model = SAM("sam2.1_b.pt")          # downloads on first use
# Prompt with a point (x, y) on the object you want segmented.
results = model("dog.jpg", points=[[640, 360]], labels=[1])
results[0].save("dog_masked.jpg")
Use this When
Mask R-CNN You have fixed known classes and want a self-contained, deployable model
SAM 2 You need to segment arbitrary objects from a prompt, or to bootstrap labels
NOTE
SAM as a labelling accelerator
A powerful production pattern: use SAM 2 to auto-generate pixel masks from cheap point clicks, then train a small fast Mask R-CNN or YOLO-seg on those masks for real-time inference. The foundation model does the slow general work once; the specialised model does the fast repeated work in production.

5 Pose Estimation and OCR

Two more high-value tasks that almost never need an LLM. Pose estimation locates keypoints — joints like shoulders, elbows, wrists — and connects them into a skeleton. It powers ergonomics monitoring, fall detection, rep-counting, and sports analytics. YOLO ships pose models returning per-person keypoint coordinates:

from ultralytics import YOLO

pose = YOLO("yolo11n-pose.pt")
results = pose.predict("gym.jpg")

# Each person has a set of (x, y, confidence) keypoints.
for person in results[0].keypoints.data:   # persons x joints x 3
    nose = person[0]                       # joint 0 is the nose in COCO order
    print("nose at", float(nose[0]), float(nose[1]), "conf", float(nose[2]))

OCR (Optical Character Recognition) turns pixels of text into strings — invoices, receipts, license plates, scanned forms. Modern OCR pairs a vision encoder with a text decoder. Hugging Face's transformers exposes such models behind a simple pipeline:

from transformers import pipeline
from PIL import Image

# A transformer OCR model: a vision encoder reads the crop,
# a text decoder writes out the characters.
ocr = pipeline("image-to-text", model="microsoft/trocr-base-printed")

text = ocr(Image.open("receipt_line.png"))
print(text[0]["generated_text"])

TrOCR works best on a single line or word crop, not a full page — so the real pipeline is detect text regions, then OCR each crop. For full documents, dedicated OCR toolkits add layout analysis and multi-language support on top.

TIP
OCR then structure, don't OCR-and-pray
OCR gives you raw strings, not meaning. Pair specialised OCR (fast, cheap, offline) for the pixel-to-text step with light post-processing — or, for messy unstructured documents, an LLM downstream to turn that text into fields. More on this hybrid pattern in [Lesson 9](/courses/07-beyond-genai/lesson-09/).
WARNING
Keypoint order is a contract
Pose models output joints in a fixed index order (the COCO 17-keypoint convention is common). Hard-coding 'index 0 = nose' only holds for that convention — always check the model card before wiring downstream logic to joint indices.

6 Specialised Vision Model vs Multimodal LLM

You have both tools in hand. The judgement call — central to this course — is which to reach for. Three realistic briefs:

Brief A — "Count defective bottle caps on a line at 40 parts/sec." Specialised YOLO-seg, no contest: you need millisecond latency, exact counts, determinism, zero per-image cost. An LLM is too slow, too expensive at volume, and its prose ("a few caps might be misaligned") is unusable for a control system.

Brief B — "User uploads a fridge photo; suggest recipes from what's inside." Multimodal LLM, comfortably. Open-ended task, unbounded categories, seconds of latency tolerance, and you want natural-language reasoning, not pixel coordinates. Training a detector for every grocery item would be absurd.

Brief C — "Read 50,000 scanned invoices/night and extract totals." Hybrid. OCR + a text detector for the fast deterministic pixel-to-text pass; an LLM only on the messy minority where layout is ambiguous. Pure-LLM is slow and costly at volume; pure-OCR chokes on irregular layouts.

Decision sketch:
  high throughput / fixed classes / need coordinates / edge ........ specialised CNN
  open-ended / unbounded classes / natural-language output ......... multimodal LLM
  large volume + occasional ambiguity ............................. hybrid: CNN first, LLM on hard cases
Signal Lean specialised CNN Lean multimodal LLM
Throughput High (video, lines) Low / interactive
Output needed Boxes, masks, coordinates Description, reasoning
Classes Fixed, known Open-ended, novel
Cost model Train once, infer free Per-image forever
Deployment Edge / offline OK Needs a server / API
NOTE
The cost crossover
At a handful of images a day, an LLM API is cheaper than training and hosting a model. At thousands or millions, a one-time-trained CNN on hardware you own is dramatically cheaper and faster. Estimate your real volume first — it usually decides the architecture. Lesson 8 turns this instinct into a full decision framework.

7 Deploying Vision to the Edge

Specialised vision dominates industry because it runs where the camera is — a factory PC, a drone, a phone, a Raspberry Pi — with no cloud round-trip. Getting there means exporting your trained model from PyTorch into a portable, optimised format.

from ultralytics import YOLO

model = YOLO("runs/detect/train/weights/best.pt")

# Export to ONNX — an open, framework-neutral format many edge runtimes load.
model.export(format="onnx")

# Other common targets, depending on the hardware:
#   format="tflite"      -> phones / microcontrollers
#   format="engine"      -> NVIDIA TensorRT for Jetson / GPUs
#   format="coreml"      -> Apple devices

Two techniques shrink models for constrained hardware:

  • Quantization — store weights as 8-bit integers instead of 32-bit floats. Roughly 4x smaller and faster, usually with small accuracy loss. The first lever for edge.
  • Pruning — remove weights/channels that contribute little, shrinking the network.

After exporting, re-measure precision and recall on the edge format — quantization can nudge accuracy, and you want to know before deployment.

On macOS, format="coreml" targets the Apple Neural Engine; ONNX runs everywhere via onnxruntime. On Linux edge boxes (Jetson, Raspberry Pi), prefer ONNX or, on NVIDIA hardware, TensorRT engine for the biggest speedup. Install with pip install onnxruntime.

On Windows, ONNX with onnxruntime (optionally the DirectML build for GPU) is the most portable path. Install pip install onnxruntime-directml for GPU, or pip install onnxruntime for CPU.

WARNING
The cloud-vs-edge trade is real
Edge gives low latency, privacy (frames never leave the device), and offline operation — but limited compute, so you need a small/quantized model. The cloud gives big models and easy updates but adds latency, cost, and a network dependency. Most serious deployments run a small model on the edge and escalate only hard cases upstream.

Questions & Answers

Q: My detector gets ~95% accuracy in training but misses obvious objects in the real plant. What went wrong?
Almost certainly a train/test distribution gap, not a model flaw — your training images differ from production in lighting, angle, motion blur, or backgrounds. Accuracy is also the wrong metric for detection; check precision and recall (and mAP) on a validation set drawn from the actual deployment camera. The fix is collecting and labelling representative real-world images, far more often than a bigger model.
Q: Can't I skip all of this and just send every frame to a multimodal LLM?
For low volume and open-ended questions, yes. For high-throughput, latency-sensitive, or coordinate-precise tasks, no — too slow, too expensive per frame, non-deterministic, and it gives you prose instead of boxes. Do the volume math: at thousands of images a day, a one-time-trained CNN on your own hardware wins decisively on cost and speed.
Q: How many labelled images do I actually need to fine-tune a detector?
Thanks to transfer learning, far fewer than from scratch — often a few hundred per class for a working prototype, more for hard or rare classes. Diversity and label consistency matter more than raw count: cover the real lighting, angles, and edge cases you will see in production. Start small, measure, and add data targeting the specific failures you observe.
Q: SAM segments anything without training — why would I ever train Mask R-CNN?
SAM is promptable and class-agnostic: it isolates an object you point at but does not know what it is, and it is heavier to run. Mask R-CNN (or YOLO-seg) is trained on your fixed classes, runs fast and fully automatically, and deploys as one self-contained model. The common pattern is SAM to bootstrap masks cheaply, then a small trained model for fast production inference.
Q: Do I need a GPU to do any of this?
For inference with small models, no — nano YOLO runs fine on CPU, and edge devices run quantized models in real time. For training/fine-tuning a GPU is strongly recommended (often minutes vs hours per run); a free cloud notebook GPU is a fine starting point if you don't own one.

Key Takeaways

  1. Four tasks, not one — classification, detection, segmentation, and pose each need a different output and architecture; naming your task correctly is half the design.
  2. YOLO is the real-time detection workhorse — one API gives boxes, confidence, and labels fast enough for video, and you judge it on precision, recall, and mAP, never raw accuracy.
  3. Transfer learning is the unlock — start from pretrained weights and fine-tune on a few hundred well-labelled images instead of training from scratch on millions.
  4. Segmentation = pixels, pose = keypoints, OCR = text — Mask R-CNN and SAM 2 trace exact outlines, pose models locate joints, OCR turns pixels into strings; each is specialised, cheap, and deterministic.
  5. Specialised CNN vs multimodal LLM is a volume-and-output decision — high throughput, fixed classes, and coordinate outputs favour CNNs; open-ended reasoning favours LLMs; large messy pipelines favour a hybrid.
  6. Production vision lives at the edge — export to ONNX/TFLite, quantize for constrained hardware, and re-measure metrics after export so the camera, not the cloud, does the work.

Next Steps: Lesson 4: NLP Beyond Chat