Comparing Tools

35 min intermediate Lesson 6

Learning Outcomes

  • Design a fair head-to-head test that sends one identical prompt to Claude, ChatGPT, and Gemini
  • Score each response across five dimensions: accuracy, instruction-following, format quality, creativity, and speed
  • Identify which tool wins for each task category — research, writing, analysis, coding, and creative work
  • Recognise the platform features (Projects, Artifacts, Canvas, Code Interpreter, Workspace integration) that change a verdict
  • Build a personal tool-choice matrix you can re-run whenever the models update

Lesson Plan

Segment Duration Topic
Intro 3 min Why "best AI" is the wrong question
Setup 5 min Building a fair test harness across three tabs
Demo 6 min The five scoring dimensions
Demo 8 min Running the bake-off across five task types
Explain 6 min When features (not models) decide the winner
Wrap-up 7 min Building and maintaining your decision matrix

Before You Begin

Pre-work:

Shopping List:

  • Three browser tabs (or desktop apps) side by side
  • A scoring sheet — notebook, spreadsheet, or the feature-evaluation reference
  • One real task from your own work as the test prompt

1 Why "Which Tool Is Best?" Is the Wrong Question

There is no single best AI tool — only the best one for a specific task, with a specific feature, on the day you ask. Claude offers Free, Pro, and Max tiers and an Opus/Sonnet/Haiku family. ChatGPT spans Free, Plus, and Pro tiers with built-in Canvas, GPTs, and data analysis. Gemini ships Free, Pro, and Ultra tiers and embeds into Gmail and Docs.

Because each vendor ships new models on its own schedule, a verdict you write today expires in weeks. So this lesson teaches a repeatable method, not a frozen leaderboard — three parts you re-run in five minutes whenever a tool updates: a fair harness, a scoring rubric, and a personal decision matrix from your own results.

NOTE
Key Insight
The goal is not to crown a champion but to know — for your work — which tab to open first. Benchmark scores online ran on different prompts; trust your own bake-off over any headline.

2 Building a Fair Test Harness

A fair test means every tool gets the same chance. Open three tabs, start a fresh chat in each, and hold five things steady: the prompt (identical text, copy-pasted), the model (a comparable tier in each), the context (no memory carried over), the attachments (same file, or none), and the starting state (web browsing off everywhere, or on everywhere).

Pick the model deliberately: in each dropdown, choose the most capable model your plan offers — reasoning against reasoning, not flagship against lightweight. Claude exposes Opus, Sonnet, and Haiku here. Then make the prompt demanding enough to separate them:

> From the three customer-churn findings below, produce: a one-paragraph
> summary; a table ranking them by urgency with a reason for each; and
> two actions for the highest-urgency one. Under 250 words, British spelling.

That one prompt tests comprehension, format control, prioritisation, brevity, and a spelling rule at once.

TIP
Tip
Save your test prompt in a plain text file. Reusing the identical wording each quarter is what makes results comparable — a reworded prompt is a new experiment.

3 The Five Scoring Dimensions

"It just felt better" is not data. Score every response 1 to 5 on five dimensions so the feeling becomes a comparable number:

Dimension A 5 looks like
Accuracy Nothing invented; figures check out
Instruction-following Hit the word limit, table, and spelling
Format quality No fixing before you paste
Creativity An interesting angle
Speed First tokens in a second

Score the churn prompt this way and totals usually land close, but the shape differs: one tool nails instructions, one format, one speed. Weight the dimensions for your work first — a lawyer doubles accuracy, a marketer creativity.

WARNING
Watch Out
The most common scoring mistake is rewarding length — a longer answer is not a better one. If the prompt said under 250 words, a 600-word response failed instruction-following.

4 Running the Bake-Off Across Five Task Types

One prompt is a data point; five categories is a pattern. Run one per category through all three tools (same spreadsheet for analysis):

> RESEARCH: The consensus on intermittent fasting for over-50s — three findings, each with a caveat.
> WRITING: A 120-word LinkedIn post announcing a new accessibility feature. Warm but professional. No hashtags.
> ANALYSIS: From the attached sales CSV, find the three slowest months, explain the cause, recommend a promotion.
> CODING: A Sheets formula flagging any row where column D is over 20% above column D's average. Explain in plain words.
> CREATIVE: Five unexpected metaphors for "learning AI tools" — no cars, journeys, or magic wands.

Each stresses a different muscle — synthesis, voice, reasoning over data, technical-but-explained output, open invention. Record a winner per category. Research usually favours whichever tool has live browsing and analysis whichever runs the file; writing and creative come down to taste.

TIP
Tip
On the analysis prompt, a tool that executes code on your file beats one that only describes what it would do — a feature difference, not a model one. See Step 5.

5 When Features — Not Models — Decide the Winner

Often the smarter model loses because the other tool has a feature that fits the job — the model writes the answer, but the feature decides whether you can use it. Check each picker before judging:

Capability Claude ChatGPT Gemini
Persistent context across chats Projects Projects / GPTs Gems
Live document/code workspace Artifacts Canvas Canvas
Runs code on your data Limited Code Interpreter Within the app
Lives in your email and docs No No Workspace integration

A feature flips a verdict when it changes what you can do: a tool that runs Code Interpreter on your spreadsheet beats one that only describes it, and drafting inside Gmail beats a better answer you copy-paste in.

NOTE
Key Insight
Re-score with features in mind: a 4/5 model that lets you edit live in Artifacts or Canvas may beat a 5/5 model that hands you text to copy out. Feature names also move between tiers — confirm one is on your plan with the [feature-evaluation reference](/courses/03-ai-tools/supplemental/feature-evaluation/).

6 Building — and Maintaining — Your Decision Matrix

Turn your scores into a one-glance map. Combine your category winners (Step 4) with the feature advantages (Step 5) into one matrix:

When I need to... Open first Because
Analyse a long document Claude Long-form reasoning; Projects holds context
Crunch a spreadsheet ChatGPT Code Interpreter runs the file
Draft inside email or Docs Gemini Already in the Workspace
Co-edit over many rounds Live-workspace tool Artifacts / Canvas beat copy-paste

Pin it where you will see it. But it has a shelf life: new model versions ship every few weeks, and one can move a tool from third to first. So re-run the bake-off whenever a vendor ships a new flagship, and quarterly regardless. Because you kept the identical prompt, each re-run is comparable — revealing trend lines no review site can give you about your own work.

NOTE
Key Insight
The real skill is adaptability: a frozen opinion about the best tool is wrong within a quarter, but a repeatable test stays right forever. Keep the matrix to a few rows you remember; Lesson 10 builds the update-tracking system.

Questions & Answers

Q: Do I really need to pay for all three tools to compare them?
No. Run the bake-off on free tiers first; you still learn each tool's strengths. Then upgrade only the one or two that win categories you work in — the matrix justifies where your money goes, not that you spend more.
Q: Isn't my scoring just subjective? How is a number I made up "data"?
It is subjective — deliberately, because your judgement of what is useful is what matters. Writing a number per dimension stops "vibes" from hiding a tool that quietly ignored half your instructions. Applied the same way to all three, it beats a gut feeling.
Q: The models change so fast — won't my results be out of date by next month?
Yes, by design. You are not building a permanent verdict but a five-minute test you re-run when something changes — the saved prompt is the asset, not the score. Lesson 10 makes the re-run a habit.
Q: Should I send my real, confidential work as the test prompt?
Not for the comparison itself. Build a prompt that mirrors the shape of your work — same length, same format demands — using non-sensitive or fictional data. You get a fair benchmark without uploading anything private to three vendors at once. See Lesson 8: File Handling for privacy.

Key Takeaways

  1. There is no single best tool — only the best one for a specific task and feature on a specific day.
  2. A fair harness is everything — identical prompt, comparable models, fresh chats, same attachments. Change one variable and the comparison is worthless.
  3. Score five dimensions — accuracy, instruction-following, format quality, creativity, and speed turn "it felt better" into a weightable grid.
  4. Features flip verdicts — Projects, Artifacts, Canvas, Code Interpreter, and Workspace integration often beat a marginally smarter model.
  5. Re-run, don't memorise — the matrix is the deliverable, but keep the saved prompt and re-test quarterly; the method outlives any leaderboard.

Next Steps: Lesson 7: Advanced Prompting in Apps