Comparing Tools
Learning Outcomes
- Design a fair head-to-head test that sends one identical prompt to Claude, ChatGPT, and Gemini
- Score each response across five dimensions: accuracy, instruction-following, format quality, creativity, and speed
- Identify which tool wins for each task category — research, writing, analysis, coding, and creative work
- Recognise the platform features (Projects, Artifacts, Canvas, Code Interpreter, Workspace integration) that change a verdict
- Build a personal tool-choice matrix you can re-run whenever the models update
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why "best AI" is the wrong question |
| Setup | 5 min | Building a fair test harness across three tabs |
| Demo | 6 min | The five scoring dimensions |
| Demo | 8 min | Running the bake-off across five task types |
| Explain | 6 min | When features (not models) decide the winner |
| Wrap-up | 7 min | Building and maintaining your decision matrix |
Before You Begin
Pre-work:
- Working accounts on at least two tools, set up from Lesson 1: Claude Desktop Setup, Lesson 3: ChatGPT App & GPTs, and Lesson 5: Gemini Across Workspace
- Know where each tool's model picker lives (the dropdown at the top of a new chat)
- Skim Lesson 4: ChatGPT Canvas & Code Interpreter so the feature comparisons land
Shopping List:
- Three browser tabs (or desktop apps) side by side
- A scoring sheet — notebook, spreadsheet, or the feature-evaluation reference
- One real task from your own work as the test prompt
There is no single best AI tool — only the best one for a specific task, with a specific feature, on the day you ask. Claude offers Free, Pro, and Max tiers and an Opus/Sonnet/Haiku family. ChatGPT spans Free, Plus, and Pro tiers with built-in Canvas, GPTs, and data analysis. Gemini ships Free, Pro, and Ultra tiers and embeds into Gmail and Docs.
Because each vendor ships new models on its own schedule, a verdict you write today expires in weeks. So this lesson teaches a repeatable method, not a frozen leaderboard — three parts you re-run in five minutes whenever a tool updates: a fair harness, a scoring rubric, and a personal decision matrix from your own results.
your work — which tab to open first. Benchmark scores online ran on different prompts; trust your own bake-off over any headline.A fair test means every tool gets the same chance. Open three tabs, start a fresh chat in each, and hold five things steady: the prompt (identical text, copy-pasted), the model (a comparable tier in each), the context (no memory carried over), the attachments (same file, or none), and the starting state (web browsing off everywhere, or on everywhere).
Pick the model deliberately: in each dropdown, choose the most capable model your plan offers — reasoning against reasoning, not flagship against lightweight. Claude exposes Opus, Sonnet, and Haiku here. Then make the prompt demanding enough to separate them:
> From the three customer-churn findings below, produce: a one-paragraph
> summary; a table ranking them by urgency with a reason for each; and
> two actions for the highest-urgency one. Under 250 words, British spelling.
That one prompt tests comprehension, format control, prioritisation, brevity, and a spelling rule at once.
"It just felt better" is not data. Score every response 1 to 5 on five dimensions so the feeling becomes a comparable number:
| Dimension | A 5 looks like |
|---|---|
| Accuracy | Nothing invented; figures check out |
| Instruction-following | Hit the word limit, table, and spelling |
| Format quality | No fixing before you paste |
| Creativity | An interesting angle |
| Speed | First tokens in a second |
Score the churn prompt this way and totals usually land close, but the shape differs: one tool nails instructions, one format, one speed. Weight the dimensions for your work first — a lawyer doubles accuracy, a marketer creativity.
One prompt is a data point; five categories is a pattern. Run one per category through all three tools (same spreadsheet for analysis):
> RESEARCH: The consensus on intermittent fasting for over-50s — three findings, each with a caveat.
> WRITING: A 120-word LinkedIn post announcing a new accessibility feature. Warm but professional. No hashtags.
> ANALYSIS: From the attached sales CSV, find the three slowest months, explain the cause, recommend a promotion.
> CODING: A Sheets formula flagging any row where column D is over 20% above column D's average. Explain in plain words.
> CREATIVE: Five unexpected metaphors for "learning AI tools" — no cars, journeys, or magic wands.
Each stresses a different muscle — synthesis, voice, reasoning over data, technical-but-explained output, open invention. Record a winner per category. Research usually favours whichever tool has live browsing and analysis whichever runs the file; writing and creative come down to taste.
Often the smarter model loses because the other tool has a feature that fits the job — the model writes the answer, but the feature decides whether you can use it. Check each picker before judging:
| Capability | Claude | ChatGPT | Gemini |
|---|---|---|---|
| Persistent context across chats | Projects | Projects / GPTs | Gems |
| Live document/code workspace | Artifacts | Canvas | Canvas |
| Runs code on your data | Limited | Code Interpreter | Within the app |
| Lives in your email and docs | No | No | Workspace integration |
A feature flips a verdict when it changes what you can do: a tool that runs Code Interpreter on your spreadsheet beats one that only describes it, and drafting inside Gmail beats a better answer you copy-paste in.
your plan with the [feature-evaluation reference](/courses/03-ai-tools/supplemental/feature-evaluation/).Turn your scores into a one-glance map. Combine your category winners (Step 4) with the feature advantages (Step 5) into one matrix:
| When I need to... | Open first | Because |
|---|---|---|
| Analyse a long document | Claude | Long-form reasoning; Projects holds context |
| Crunch a spreadsheet | ChatGPT | Code Interpreter runs the file |
| Draft inside email or Docs | Gemini | Already in the Workspace |
| Co-edit over many rounds | Live-workspace tool | Artifacts / Canvas beat copy-paste |
Pin it where you will see it. But it has a shelf life: new model versions ship every few weeks, and one can move a tool from third to first. So re-run the bake-off whenever a vendor ships a new flagship, and quarterly regardless. Because you kept the identical prompt, each re-run is comparable — revealing trend lines no review site can give you about your own work.
Questions & Answers
Key Takeaways
- There is no single best tool — only the best one for a specific task and feature on a specific day.
- A fair harness is everything — identical prompt, comparable models, fresh chats, same attachments. Change one variable and the comparison is worthless.
- Score five dimensions — accuracy, instruction-following, format quality, creativity, and speed turn "it felt better" into a weightable grid.
- Features flip verdicts — Projects, Artifacts, Canvas, Code Interpreter, and Workspace integration often beat a marginally smarter model.
- Re-run, don't memorise — the matrix is the deliverable, but keep the saved prompt and re-test quarterly; the method outlives any leaderboard.
Next Steps: Lesson 7: Advanced Prompting in Apps