Careers in AI
Learning Outcomes
- Distinguish the core AI roles of 2026 by what they actually ship, not their titles
- Map your existing Python and engineering skills onto role requirements to find your shortest path
- Build a small portfolio artefact that proves a skill instead of merely claiming it
- Assess which AI skills stay durable as the field automates parts of itself
- Design a concrete, time-boxed personal learning plan with measurable milestones
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why "prompt engineer" is not the centre |
| Explain | 8 min | The role map and what each role requires |
| Build | 9 min | Self-assessment script: score yourself |
| Build | 9 min | Build one portfolio artefact |
| Build | 8 min | Sample the adjacent domains (CV, NLP) |
| Explain | 5 min | Skill durability: what survives automation |
| Wrap-up | 3 min | Your 12-week plan + key takeaways |
Before You Begin
Pre-work:
- Have a working sense of the ML workflow: Lesson 2: ML Fundamentals
- Understand the classical-vs-GenAI trade-off: Lesson 8: Classical ML vs GenAI
- Have seen how models reach production: Lesson 7: MLOps
Shopping List:
- Python 3.10+ and a virtual environment (
python3 -m venv .venv) pip install scikit-learnfor the portfolio step- A GitHub account (your portfolio lives there)
- 30 minutes of honest reflection: this lesson is partly about you
The first myth to kill: "AI career" is not one job, and "prompt engineer" is not its centre. Most production AI is classical, non-generative work — fraud, forecasting, recommendations, predictive maintenance. Roles cluster by what they ship:
| Role | Ships | Gap that sinks candidates |
|---|---|---|
| Data Scientist | Insight + a validated model | Can't explain why a model fails |
| ML Engineer | Production training & serving code | Notebook code that can't deploy |
| MLOps Engineer | Pipelines, monitoring, infra | No production incident experience |
| CV Engineer | Detection/segmentation systems | Ignores data quality, labelling |
| NLP Engineer | Extraction, classification, search | Reaches for an LLM for everything |
| Research Scientist | New methods, papers | Can't ship anything usable |
Three terms. Inference runs a trained model to get predictions (the production hot path); training fits it to data offline; transfer learning adapts a model trained on a huge generic dataset to your smaller task — far cheaper than from scratch.
Notice the "gap" column: what sinks candidates is almost never knowing fewer algorithms — it's the inability to make code production-ready, evaluate honestly, or pick the right tool.
Vague aspiration produces vague effort. Rate yourself 0-3 on each competency; this script ranks your fit per role and lists your smallest gaps. Edit me (0 = none … 3 = ship solo):
ROLES = {
"Data Scientist": ["stats", "sklearn", "evaluation", "communication"],
"ML Engineer": ["python_scale", "pytorch", "software_eng", "deployment"],
"MLOps Engineer": ["docker_cicd", "monitoring", "cloud", "software_eng"],
"CV Engineer": ["transfer_learning", "sklearn", "deployment", "data_quality"],
"NLP Engineer": ["transformers", "evaluation", "retrieval", "python_scale"],
}
me = {"sklearn": 2, "communication": 2, "python_scale": 3, "software_eng": 3,
"deployment": 2, "docker_cicd": 2, "data_quality": 2} # omitted => 0
for role, sk in ROLES.items():
pct = round(100 * sum(me.get(s, 0) for s in sk) / (3 * len(sk)))
gaps = [s for s in sk if me.get(s, 0) < 2] # below 2 is a gap
print(f"{role:16}{pct:3d}% close:", ", ".join(gaps))
Running it ranks the roles — the top row is your shortest path, its "close" list your study queue. Re-run monthly.
Claims are cheap; artefacts hire you. The highest-leverage move is a small, finished GitHub project showing judgement. This one reports precision/recall, uses cross-validation (training/testing across folds so your score isn't a fluke of one split — it also catches overfitting, memorising training data and failing on new data), and prints a confusion matrix:
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.metrics import classification_report, confusion_matrix
X, y = load_breast_cancer(return_X_y=True) # ships with scikit-learn
X_tr, X_te, y_tr, y_te = train_test_split(
X, y, test_size=0.25, stratify=y, random_state=42)
clf = RandomForestClassifier(n_estimators=200, random_state=42)
# Cross-validate on TRAINING data only; never peek at the test set.
cv = cross_val_score(clf, X_tr, y_tr, cv=5, scoring="f1")
print(f"CV F1 (5-fold): {cv.mean():.3f} +/- {cv.std():.3f}")
clf.fit(X_tr, y_tr)
preds = clf.predict(X_te)
print(classification_report(y_te, preds, digits=3))
print(confusion_matrix(y_te, preds))
You'll see precision/recall per class and a matrix of where it confused malignant for benign — the framing a medical or fraud team cares about. Now write a five-line README.md (problem, data, metric, meaning, one limitation): that README gets the interview.
Pipeline does this automatically.Don't pick a domain blind. Touch CV and NLP (pip install ultralytics transformers torch); whichever makes you want to keep going is the signal. Both snippets run as-is.
Computer vision with Ultralytics YOLO — a pretrained model detects common object classes immediately, no training needed:
from ultralytics import YOLO
model = YOLO("yolo11n.pt") # nano: smallest/fastest; auto-downloads
results = model("https://ultralytics.com/images/bus.jpg")
results[0].show() # draws boxes; results[0].boxes has the data
NLP with Hugging Face Transformers — calling pipeline with no model downloads a sensible default plus a tokeniser (which splits text into the sub-word units the model expects):
from transformers import pipeline
classifier = pipeline("sentiment-analysis")
print(classifier("This course finally made ML make sense."))
# -> [{'label': 'POSITIVE', 'score': 0.99...}]
So little code stands between you and a working system that the scarce skill is no longer invoking a model — it's choosing the right one, evaluating it honestly, and knowing when a classical model beats an LLM. Go deeper: vision, text, forecasting, serving.
AI tools make raw implementation cheaper every quarter. Invest where the half-life is longest:
| Skill | Durability | Why |
|---|---|---|
| Problem framing & metric design | Very high | Deciding what to optimise can't be automated |
| Evaluation & error analysis | Very high | Someone must judge if a model is safe to ship |
| Data quality / labelling | High | Garbage in still means garbage out |
| Systems & deployment thinking | High | Reliability and cost are perennial problems |
| Library-specific syntax | Low | Assistants generate it; APIs churn anyway |
The pattern: judgement and framing are durable; syntax and recall are commoditising. Spend your time on the durable questions — why did this model fail? which metric matches the business cost? classical ML or an LLM (Lesson 8)? — and pick up syntax as you need it.
Turn the above into a dated plan: pick one target role (your top bar from Step 2), end each phase with a shippable artefact:
TARGET ROLE: [e.g. ML Engineer] WHY: [one honest sentence]
Weeks 1-4 Foundation: close your top gap (e.g. PyTorch)
-> 1 finished project: precision/recall + a README
Weeks 5-8 Domain depth -> deploy 1 behind an API (Lesson 7)
Weeks 9-12 Production: Docker + monitoring + a post-mortem
-> 3 GitHub projects, each with a stated limitation
REVIEW: re-run assess.py at weeks 4, 8, 12.
SUCCESS: I can explain each project's metric AND its failure mode.
Two rules keep a plan alive: make every milestone an artefact, not a feeling ("deployed behind an API" is checkable, "understand deployment" is not), and timebox study to ship — a finished B-grade project beats an unfinished A.
Questions & Answers
Key Takeaways
- "AI career" is a family of distinct roles — judged by what they ship; most of the work is classical ML, not GenAI.
- Engineers have the shortest path in — ML Engineer and MLOps reward your software discipline plus a layer of ML literacy; no PhD needed.
- Artefacts hire you, claims don't — three small, honestly-evaluated projects beat thirty unfinished notebooks.
- Invest in durable skills — problem framing, metric design, and evaluation outlast any library; syntax and recall are commoditising.
- Become the person who picks the right tool — matching classical ML, deep learning, LLMs, or hybrids to a problem is the most future-proof AI skill of 2026.
- Plan in dated, shippable milestones — one target role, gaps closed just-in-time, one artefact per phase.
Next Steps: Back to all Beyond GenAI lessons