Computer Vision in 2026
Learning Outcomes
- Distinguish the four core vision tasks — classification, detection, segmentation, and pose — and pick the right one for a problem
- Run a real object-detection pipeline with YOLO and read boxes, confidence, and class labels from the results
- Apply transfer learning to fine-tune a detector on your own images instead of training from scratch
- Extract text from images with OCR and segment objects pixel-precisely with SAM and Mask R-CNN
- Decide when a specialised vision model beats a multimodal LLM, and deploy one to an edge device
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why most production vision is not an LLM |
| Concepts | 9 min | The four vision tasks and how CNNs see |
| Build | 12 min | Object detection with YOLO, end to end |
| Build | 10 min | Transfer learning — fine-tuning on your data |
| Build | 8 min | Segmentation: SAM and Mask R-CNN |
| Build | 7 min | Pose estimation and OCR |
| Decide | 7 min | Specialised models vs multimodal LLMs |
| Wrap-up | 4 min | Edge deployment and key takeaways |
Before You Begin
Pre-work:
- Complete Lesson 2: Machine Learning Fundamentals — you need the train/validate/test and metrics vocabulary
- Skim Lesson 1: The AI Landscape for where vision sits in the wider field
- Be comfortable with Python, pip, and a virtual environment
Shopping List:
- Python 3.10+ and a fresh virtual environment
pip install ultralytics torch torchvision transformers pillow(a few GB; first run downloads model weights)- A GPU helps for training but is not required — every example here runs on CPU, just slower
- A handful of your own images (a photo with people/objects, and one with printed or handwritten text)
Before any code, build the mental map. Computer vision is not one problem but four, and the right architecture depends on which you have.
| Task | Question it answers | Output | Real-world use |
|---|---|---|---|
| Classification | "What is in this image?" | One label per image | Defect / no-defect sorting |
| Detection | "What is here, and where?" | Boxes + labels + scores | Counting cars, finding PPE on workers |
| Segmentation | "Which exact pixels are the object?" | Per-pixel mask | Tumour outlining, background removal |
| Pose estimation | "Where are the keypoints?" | Joint coordinates | Ergonomics, sports analytics, fall detection |
The workhorse behind all four is the convolutional neural network (CNN) — early layers learn small reusable filters (edges, textures), deeper layers compose them into shapes and objects. Convolution just means sliding a filter across the image and recording where it matches, which is why a CNN finds a cat in any corner of the frame without being told where to look.
A fair 2026 question: why not just send the image to a multimodal LLM? For these structured tasks, specialised CNNs win on the axes that matter in production.
| Dimension | Specialised vision CNN | Multimodal LLM |
|---|---|---|
| Latency | Milliseconds, runs on a camera | Hundreds of ms to seconds, needs a server |
| Cost | One-time training, free inference | Per-image API cost forever |
| Output | Exact pixel coordinates / boxes | Prose you must parse |
| Throughput | Thousands of frames/sec on a GPU | Rate-limited |
| Determinism | Same input → same boxes | Can vary run to run |
Object detection finds every object and draws a bounding box (a rectangle given as pixel corners) around each, with a class label and confidence score. The dominant real-time family is YOLO ("You Only Look Once") — it processes the whole image in a single forward pass, fast enough for video. Current Ultralytics releases are the YOLO11 family and the newer YOLO26 flagship, both covering detection, segmentation, pose, classification, and oriented boxes from one API.
Run a pretrained detector — the *.pt weights file downloads automatically on first use:
from ultralytics import YOLO
# Load a small pretrained detector (trained on the 80-class COCO dataset).
model = YOLO("yolo11n.pt") # 'n' = nano, the fastest variant
# Run inference on an image (a path, URL, PIL image, or numpy array).
results = model.predict("street.jpg", conf=0.35)
r = results[0] # results is a list, one entry per image
for box in r.boxes:
cls_id = int(box.cls[0]) # class index
label = r.names[cls_id] # human-readable class name
score = float(box.conf[0]) # confidence 0..1
x1, y1, x2, y2 = box.xyxy[0].tolist() # corner pixels
print(f"{label:12s} {score:.2f} box=({x1:.0f},{y1:.0f},{x2:.0f},{y2:.0f})")
r.save(filename="street_annotated.jpg") # save a copy with boxes drawn
You will see lines like person 0.91 box=(412,88,503,377). The conf=0.35 argument is the confidence threshold — detections below it are discarded. Raise it to cut false positives; lower it to catch faint objects you are missing.
Three metrics define detector quality, and you must know all three:
- Precision — of the boxes the model drew, how many were correct. Low precision = too many false alarms.
- Recall — of the objects actually present, how many it found. Low recall = it missed things.
- mAP (mean Average Precision) — the standard single-number summary balancing precision and recall across all classes and box-overlap thresholds. Higher is better; it is what you compare models on.
cracked_solder_joint or ripe_strawberry. For anything domain-specific you must fine-tune — that is the next step.Transfer learning is the single most important idea in applied vision. Instead of training from random weights (millions of images, days of compute), you start from a model already trained on a huge dataset and continue training it on your small one. The early CNN layers — edges, textures, shapes — already generalise, so you only teach the model your specific objects. This turns "I need 1,000,000 labelled images" into "I need a few hundred."
To fine-tune YOLO, supply a tiny dataset config and labelled images — one .txt per image, each line class_id x_center y_center width height with coordinates normalised 0–1. A dataset YAML describes where the data lives and the class names:
# data.yaml
path: ./pallet-dataset # dataset root
train: images/train # images for training
val: images/val # images held out to measure performance
names:
0: pallet
1: forklift
2: damaged_pallet
Then fine-tune with a few lines of Python:
from ultralytics import YOLO
# Start from pretrained weights — this is the transfer-learning step.
model = YOLO("yolo11n.pt")
model.train(
data="data.yaml",
epochs=50, # passes over the training set
imgsz=640, # input resolution
batch=16,
patience=10, # early-stop if val metrics stop improving
)
# Ultralytics evaluates on the val split automatically.
metrics = model.val()
print(metrics.box.map) # mAP averaged over overlap thresholds
print(metrics.box.mp) # mean precision
print(metrics.box.mr) # mean recall
Two terms from above, defined plainly:
- Epoch — one full pass over your training images. Too few and the model underfits (hasn't learned); too many and it overfits — memorising the training set so well it fails on new images.
- Validation split — images the model never trains on, used only to measure honest performance. The gap between training and validation accuracy is your early-warning light for overfitting (see Lesson 2).
patience argument automates that last one.A bounding box says "a defect is somewhere in this rectangle." Segmentation goes further: it labels every pixel, producing a mask that traces the object's exact outline — needed for measuring area, removing backgrounds, or medical/satellite imagery where shape matters. Two flavours answer different questions.
Mask R-CNN is the classic instance segmentation model: like a detector, but each detection also carries a pixel mask, trained on labelled classes ("this blob is a car, that one a person"). It is pretrained in torchvision:
import torch
from torchvision.models.detection import (
maskrcnn_resnet50_fpn, MaskRCNN_ResNet50_FPN_Weights,
)
from torchvision.io import read_image
weights = MaskRCNN_ResNet50_FPN_Weights.DEFAULT
model = maskrcnn_resnet50_fpn(weights=weights).eval()
preprocess = weights.transforms()
img = read_image("street.jpg")
with torch.no_grad():
out = model([preprocess(img)])[0]
labels = weights.meta["categories"]
for label_id, score, mask in zip(out["labels"], out["scores"], out["masks"]):
if score > 0.7:
print(labels[label_id], float(score), "mask:", tuple(mask.shape))
SAM 2 (Segment Anything Model 2, from Meta) is different: a promptable, class-agnostic foundation model. Give it a point or box prompt and it returns the mask for whatever object is there — no training, across both images and video. It does not know what the thing is called; it just isolates it. Via Ultralytics:
from ultralytics import SAM
model = SAM("sam2.1_b.pt") # downloads on first use
# Prompt with a point (x, y) on the object you want segmented.
results = model("dog.jpg", points=[[640, 360]], labels=[1])
results[0].save("dog_masked.jpg")
| Use this | When |
|---|---|
| Mask R-CNN | You have fixed known classes and want a self-contained, deployable model |
| SAM 2 | You need to segment arbitrary objects from a prompt, or to bootstrap labels |
Two more high-value tasks that almost never need an LLM. Pose estimation locates keypoints — joints like shoulders, elbows, wrists — and connects them into a skeleton. It powers ergonomics monitoring, fall detection, rep-counting, and sports analytics. YOLO ships pose models returning per-person keypoint coordinates:
from ultralytics import YOLO
pose = YOLO("yolo11n-pose.pt")
results = pose.predict("gym.jpg")
# Each person has a set of (x, y, confidence) keypoints.
for person in results[0].keypoints.data: # persons x joints x 3
nose = person[0] # joint 0 is the nose in COCO order
print("nose at", float(nose[0]), float(nose[1]), "conf", float(nose[2]))
OCR (Optical Character Recognition) turns pixels of text into strings — invoices, receipts, license plates, scanned forms. Modern OCR pairs a vision encoder with a text decoder. Hugging Face's transformers exposes such models behind a simple pipeline:
from transformers import pipeline
from PIL import Image
# A transformer OCR model: a vision encoder reads the crop,
# a text decoder writes out the characters.
ocr = pipeline("image-to-text", model="microsoft/trocr-base-printed")
text = ocr(Image.open("receipt_line.png"))
print(text[0]["generated_text"])
TrOCR works best on a single line or word crop, not a full page — so the real pipeline is detect text regions, then OCR each crop. For full documents, dedicated OCR toolkits add layout analysis and multi-language support on top.
You have both tools in hand. The judgement call — central to this course — is which to reach for. Three realistic briefs:
Brief A — "Count defective bottle caps on a line at 40 parts/sec." Specialised YOLO-seg, no contest: you need millisecond latency, exact counts, determinism, zero per-image cost. An LLM is too slow, too expensive at volume, and its prose ("a few caps might be misaligned") is unusable for a control system.
Brief B — "User uploads a fridge photo; suggest recipes from what's inside." Multimodal LLM, comfortably. Open-ended task, unbounded categories, seconds of latency tolerance, and you want natural-language reasoning, not pixel coordinates. Training a detector for every grocery item would be absurd.
Brief C — "Read 50,000 scanned invoices/night and extract totals." Hybrid. OCR + a text detector for the fast deterministic pixel-to-text pass; an LLM only on the messy minority where layout is ambiguous. Pure-LLM is slow and costly at volume; pure-OCR chokes on irregular layouts.
Decision sketch:
high throughput / fixed classes / need coordinates / edge ........ specialised CNN
open-ended / unbounded classes / natural-language output ......... multimodal LLM
large volume + occasional ambiguity ............................. hybrid: CNN first, LLM on hard cases
| Signal | Lean specialised CNN | Lean multimodal LLM |
|---|---|---|
| Throughput | High (video, lines) | Low / interactive |
| Output needed | Boxes, masks, coordinates | Description, reasoning |
| Classes | Fixed, known | Open-ended, novel |
| Cost model | Train once, infer free | Per-image forever |
| Deployment | Edge / offline OK | Needs a server / API |
Specialised vision dominates industry because it runs where the camera is — a factory PC, a drone, a phone, a Raspberry Pi — with no cloud round-trip. Getting there means exporting your trained model from PyTorch into a portable, optimised format.
from ultralytics import YOLO
model = YOLO("runs/detect/train/weights/best.pt")
# Export to ONNX — an open, framework-neutral format many edge runtimes load.
model.export(format="onnx")
# Other common targets, depending on the hardware:
# format="tflite" -> phones / microcontrollers
# format="engine" -> NVIDIA TensorRT for Jetson / GPUs
# format="coreml" -> Apple devices
Two techniques shrink models for constrained hardware:
- Quantization — store weights as 8-bit integers instead of 32-bit floats. Roughly 4x smaller and faster, usually with small accuracy loss. The first lever for edge.
- Pruning — remove weights/channels that contribute little, shrinking the network.
After exporting, re-measure precision and recall on the edge format — quantization can nudge accuracy, and you want to know before deployment.
On macOS, format="coreml" targets the Apple Neural Engine; ONNX runs everywhere via onnxruntime. On Linux edge boxes (Jetson, Raspberry Pi), prefer ONNX or, on NVIDIA hardware, TensorRT engine for the biggest speedup. Install with pip install onnxruntime.
On Windows, ONNX with onnxruntime (optionally the DirectML build for GPU) is the most portable path. Install pip install onnxruntime-directml for GPU, or pip install onnxruntime for CPU.
Questions & Answers
Key Takeaways
- Four tasks, not one — classification, detection, segmentation, and pose each need a different output and architecture; naming your task correctly is half the design.
- YOLO is the real-time detection workhorse — one API gives boxes, confidence, and labels fast enough for video, and you judge it on precision, recall, and mAP, never raw accuracy.
- Transfer learning is the unlock — start from pretrained weights and fine-tune on a few hundred well-labelled images instead of training from scratch on millions.
- Segmentation = pixels, pose = keypoints, OCR = text — Mask R-CNN and SAM 2 trace exact outlines, pose models locate joints, OCR turns pixels into strings; each is specialised, cheap, and deterministic.
- Specialised CNN vs multimodal LLM is a volume-and-output decision — high throughput, fixed classes, and coordinate outputs favour CNNs; open-ended reasoning favours LLMs; large messy pipelines favour a hybrid.
- Production vision lives at the edge — export to ONNX/TFLite, quantize for constrained hardware, and re-measure metrics after export so the camera, not the cloud, does the work.
Next Steps: Lesson 4: NLP Beyond Chat