Careers in AI

45 min advanced Lesson 10

Learning Outcomes

  • Distinguish the core AI roles of 2026 by what they actually ship, not their titles
  • Map your existing Python and engineering skills onto role requirements to find your shortest path
  • Build a small portfolio artefact that proves a skill instead of merely claiming it
  • Assess which AI skills stay durable as the field automates parts of itself
  • Design a concrete, time-boxed personal learning plan with measurable milestones

Lesson Plan

Segment Duration Topic
Intro 3 min Why "prompt engineer" is not the centre
Explain 8 min The role map and what each role requires
Build 9 min Self-assessment script: score yourself
Build 9 min Build one portfolio artefact
Build 8 min Sample the adjacent domains (CV, NLP)
Explain 5 min Skill durability: what survives automation
Wrap-up 3 min Your 12-week plan + key takeaways

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ and a virtual environment (python3 -m venv .venv)
  • pip install scikit-learn for the portfolio step
  • A GitHub account (your portfolio lives there)
  • 30 minutes of honest reflection: this lesson is partly about you

1 The Role Map and What Each Role Requires

The first myth to kill: "AI career" is not one job, and "prompt engineer" is not its centre. Most production AI is classical, non-generative work — fraud, forecasting, recommendations, predictive maintenance. Roles cluster by what they ship:

Role Ships Gap that sinks candidates
Data Scientist Insight + a validated model Can't explain why a model fails
ML Engineer Production training & serving code Notebook code that can't deploy
MLOps Engineer Pipelines, monitoring, infra No production incident experience
CV Engineer Detection/segmentation systems Ignores data quality, labelling
NLP Engineer Extraction, classification, search Reaches for an LLM for everything
Research Scientist New methods, papers Can't ship anything usable

Three terms. Inference runs a trained model to get predictions (the production hot path); training fits it to data offline; transfer learning adapts a model trained on a huge generic dataset to your smaller task — far cheaper than from scratch.

Notice the "gap" column: what sinks candidates is almost never knowing fewer algorithms — it's the inability to make code production-ready, evaluate honestly, or pick the right tool.

NOTE
The 80/20 of AI jobs
Most paid AI work is classical ML and data engineering, not generative AI. A developer who can ship a calibrated fraud classifier behind a low-latency API is more employable than one who only writes clever prompts.
TIP
Engineers have a shortcut in
A fluent Python developer is closer to ML Engineer and MLOps than to Research Scientist. Those roles reward what you have — software discipline, testing, deployment — plus a layer of ML literacy. Lead with engineering; add ML on top. A PhD is needed only for Research Scientist roles.

2 Score Yourself: A Self-Assessment Script

Vague aspiration produces vague effort. Rate yourself 0-3 on each competency; this script ranks your fit per role and lists your smallest gaps. Edit me (0 = none … 3 = ship solo):

ROLES = {
    "Data Scientist": ["stats", "sklearn", "evaluation", "communication"],
    "ML Engineer":    ["python_scale", "pytorch", "software_eng", "deployment"],
    "MLOps Engineer": ["docker_cicd", "monitoring", "cloud", "software_eng"],
    "CV Engineer":    ["transfer_learning", "sklearn", "deployment", "data_quality"],
    "NLP Engineer":   ["transformers", "evaluation", "retrieval", "python_scale"],
}
me = {"sklearn": 2, "communication": 2, "python_scale": 3, "software_eng": 3,
      "deployment": 2, "docker_cicd": 2, "data_quality": 2}  # omitted => 0

for role, sk in ROLES.items():
    pct = round(100 * sum(me.get(s, 0) for s in sk) / (3 * len(sk)))
    gaps = [s for s in sk if me.get(s, 0) < 2]   # below 2 is a gap
    print(f"{role:16}{pct:3d}%  close:", ", ".join(gaps))

Running it ranks the roles — the top row is your shortest path, its "close" list your study queue. Re-run monthly.

NOTE
Why a 2 is the threshold
Anything below 2 counts as a gap, deliberately. Employers hire for trajectory, not perfection: you need a cluster of 2s and 3s matching one role, plus evidence you learn fast.

3 Build One Artefact That Proves a Skill

Claims are cheap; artefacts hire you. The highest-leverage move is a small, finished GitHub project showing judgement. This one reports precision/recall, uses cross-validation (training/testing across folds so your score isn't a fluke of one split — it also catches overfitting, memorising training data and failing on new data), and prints a confusion matrix:

from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.metrics import classification_report, confusion_matrix

X, y = load_breast_cancer(return_X_y=True)   # ships with scikit-learn
X_tr, X_te, y_tr, y_te = train_test_split(
    X, y, test_size=0.25, stratify=y, random_state=42)

clf = RandomForestClassifier(n_estimators=200, random_state=42)

# Cross-validate on TRAINING data only; never peek at the test set.
cv = cross_val_score(clf, X_tr, y_tr, cv=5, scoring="f1")
print(f"CV F1 (5-fold): {cv.mean():.3f} +/- {cv.std():.3f}")

clf.fit(X_tr, y_tr)
preds = clf.predict(X_te)
print(classification_report(y_te, preds, digits=3))
print(confusion_matrix(y_te, preds))

You'll see precision/recall per class and a matrix of where it confused malignant for benign — the framing a medical or fraud team cares about. Now write a five-line README.md (problem, data, metric, meaning, one limitation): that README gets the interview.

WARNING
Data leakage will embarrass you in interviews
Computing scaling or oversampling on the full dataset before splitting leaks test info into training; your score becomes a lie that collapses in production. Split first, then fit transforms on the training fold only — a Pipeline does this automatically.

4 Sample the Adjacent Domains

Don't pick a domain blind. Touch CV and NLP (pip install ultralytics transformers torch); whichever makes you want to keep going is the signal. Both snippets run as-is.

Computer vision with Ultralytics YOLO — a pretrained model detects common object classes immediately, no training needed:

from ultralytics import YOLO

model = YOLO("yolo11n.pt")          # nano: smallest/fastest; auto-downloads
results = model("https://ultralytics.com/images/bus.jpg")
results[0].show()                   # draws boxes; results[0].boxes has the data

NLP with Hugging Face Transformers — calling pipeline with no model downloads a sensible default plus a tokeniser (which splits text into the sub-word units the model expects):

from transformers import pipeline

classifier = pipeline("sentiment-analysis")
print(classifier("This course finally made ML make sense."))
# -> [{'label': 'POSITIVE', 'score': 0.99...}]

So little code stands between you and a working system that the scarce skill is no longer invoking a model — it's choosing the right one, evaluating it honestly, and knowing when a classical model beats an LLM. Go deeper: vision, text, forecasting, serving.

NOTE
Pretrained-first is the default
Across CV and NLP, the professional starting point is a pretrained model plus transfer learning, not training from scratch. Adapting, evaluating, and serving existing models matters more day-to-day than building new ones.

5 Skill Durability: What Survives Automation

AI tools make raw implementation cheaper every quarter. Invest where the half-life is longest:

Skill Durability Why
Problem framing & metric design Very high Deciding what to optimise can't be automated
Evaluation & error analysis Very high Someone must judge if a model is safe to ship
Data quality / labelling High Garbage in still means garbage out
Systems & deployment thinking High Reliability and cost are perennial problems
Library-specific syntax Low Assistants generate it; APIs churn anyway

The pattern: judgement and framing are durable; syntax and recall are commoditising. Spend your time on the durable questions — why did this model fail? which metric matches the business cost? classical ML or an LLM (Lesson 8)? — and pick up syntax as you need it.

TIP
Become the person who picks the right tool
The most durable AI skill in 2026 is knowing when classical ML, deep learning, an LLM, or a hybrid (see Lesson 9) is correct. Don't anchor your identity to one library — anchor to a problem domain. Tools are seasonal; problems are permanent.

6 Write Your 12-Week Learning Plan

Turn the above into a dated plan: pick one target role (your top bar from Step 2), end each phase with a shippable artefact:

TARGET ROLE: [e.g. ML Engineer]   WHY: [one honest sentence]

Weeks 1-4   Foundation: close your top gap (e.g. PyTorch)
            -> 1 finished project: precision/recall + a README
Weeks 5-8   Domain depth -> deploy 1 behind an API (Lesson 7)
Weeks 9-12  Production: Docker + monitoring + a post-mortem
            -> 3 GitHub projects, each with a stated limitation

REVIEW: re-run assess.py at weeks 4, 8, 12.
SUCCESS: I can explain each project's metric AND its failure mode.

Two rules keep a plan alive: make every milestone an artefact, not a feeling ("deployed behind an API" is checkable, "understand deployment" is not), and timebox study to ship — a finished B-grade project beats an unfinished A.

NOTE
Use the supplements as a checklist
Pair this plan with the decision patterns, the cheat-sheet of each domain's algorithms, and the tools reference.

Questions & Answers

Q: I'm a senior backend developer with no ML. Do I need a master's or PhD to switch?
No, not for ML Engineer or MLOps. They reward engineering discipline plus working ML literacy — both self-teachable and provable with the Step 3 portfolio. A PhD is needed only for Research Scientist roles. Lead with engineering; add ML on top.
Q: Won't AI coding assistants automate these jobs away before I finish learning?
They automate the perishable parts — boilerplate, syntax, scaffolding — not problem framing, metric design, honest evaluation, or the judgement to pick the right technique. Those grow more valuable as implementation gets cheaper. Use the assistants; invest above them.
Q: Should I specialise in one domain or stay general?
Go T-shaped: broad literacy across the domains plus depth in one. Generalists struggle to prove a hire-able edge; pure specialists are fragile when their niche cools. Pick the Step 4 domain you enjoyed — sustained interest produces the depth employers pay for.
Q: Are big salaries only for GenAI roles? Is classical ML a dead end?
The opposite. Most paid AI work is classical ML and data engineering, stable and less hype-cyclical than GenAI fashion. Compensation tracks the value you create and your seniority more than the buzzword in your title. Ranges vary by region, stage, and level — anchor to local data, not headlines.
Q: How do I get experience when every job wants experience?
Manufacture it. The Step 3 artefacts, an internal project at work ("I added a churn model to our admin tool"), Kaggle, and open-source contributions all count. The bar is not "paid ML job" — it's "shipped and honestly evaluated a usable model." Three of those beat a year of tutorials.

Key Takeaways

  1. "AI career" is a family of distinct roles — judged by what they ship; most of the work is classical ML, not GenAI.
  2. Engineers have the shortest path in — ML Engineer and MLOps reward your software discipline plus a layer of ML literacy; no PhD needed.
  3. Artefacts hire you, claims don't — three small, honestly-evaluated projects beat thirty unfinished notebooks.
  4. Invest in durable skills — problem framing, metric design, and evaluation outlast any library; syntax and recall are commoditising.
  5. Become the person who picks the right tool — matching classical ML, deep learning, LLMs, or hybrids to a problem is the most future-proof AI skill of 2026.
  6. Plan in dated, shippable milestones — one target role, gaps closed just-in-time, one artefact per phase.

Next Steps: Back to all Beyond GenAI lessons