When to Use AI (and When Not To)

40 min beginner Lesson 1

Learning Outcomes

  • Apply a four-question decision framework to judge whether any product feature genuinely needs AI
  • Identify the anti-patterns that lead teams to add AI where simpler solutions would serve users better
  • Distinguish tasks where imperfect AI output is acceptable from tasks where errors are genuinely costly
  • Evaluate a product proposal against the framework and arrive at a defensible build/don't-build recommendation
  • Recognise the conditions that make AI unlock entirely new value rather than just replicate existing functionality

Lesson Plan

Segment Duration Topic
Intro 3 min Why "add AI to it" is the wrong starting question
The framework 8 min Four diagnostic questions every product decision needs
Question 1–2 deep dive 6 min Ambiguity and personalization: what they really mean
Question 3–4 deep dive 6 min Tolerance for imperfection and fallback design
Anti-patterns 7 min Five ways AI gets added for the wrong reasons
Success stories 6 min When AI creates capabilities that didn't exist before
Applying the framework 2 min How to run the test on your own roadmap
Wrap-up 2 min Key takeaways and next lesson

Before You Begin

Pre-work:

  • No previous lessons are required — this is Lesson 1. If you want context on the course arc, skim the AI Product Design course landing page.
  • Think of one or two AI features on your current or recent product roadmap. You will apply the framework to them as we go.
  • Bring any open questions about a specific proposal — the Q&A section at the end addresses the most common objections teams raise.

Shopping List:

  • A web browser
  • A notes document alongside — you will want to write down your framework verdicts as you work through the steps
  • Optionally: the AI Glossary tab, open in the background, for any terms of art you want to check

1 Why "Add AI to It" Is the Wrong Starting Question

Somewhere in the last two years, "we should add AI to this" became the default response to every product challenge. A struggling retention metric? Add an AI recommendation engine. Users can't find features? Add an AI assistant. Churn is up? Add an AI-powered engagement nudge. The instinct is understandable — AI is genuinely powerful, and the competitive pressure to ship AI features is real.

But the instinct produces predictable failures. Teams build AI features that users ignore, features that erode trust when they get things wrong in small but visible ways, and features that cost ten times as much to maintain as the rule-based or search-based solution they replaced. The graveyard of abandoned AI features is already crowded, and almost every failure traces back to the same root: a team that started with the technology and reverse-engineered a use case, rather than starting with the user problem and asking whether AI was the right tool.

The right starting question is not "can we add AI here?" — you almost certainly can. The right question is "should we, and what specifically are users trying to do that AI uniquely helps with?"

NOTE
Why This Matters
Google's People + AI Guidebook (the PAIR team's practical design resource) makes this explicit in its first principle: start with the user need, not the capability. Microsoft's Guidelines for Human-AI Interaction (Amershi et al., 2019) similarly ground every design decision in what the user is trying to accomplish — not what the model can produce.

This lesson gives you a four-question framework for making that call with your team in a structured way. By the end, you will be able to look at any proposed AI feature and arrive at a defensible answer: yes, this genuinely needs AI — or no, something simpler will serve users better and cost you less.


2 The Four-Question Decision Framework

Before building any AI feature, run it through four diagnostic questions. A genuine "yes" to all four is a strong signal that AI belongs here. "No" answers reveal where a simpler solution would be better — or where the feature design needs to change before AI is appropriate.

Question What you're probing
1. Is the task inherently ambiguous? Would a rule or a search index handle this adequately?
2. Does it benefit from personalization? Does the right answer vary meaningfully across users or contexts?
3. Is there tolerance for imperfect output? Can errors be surfaced, corrected, and absorbed without real harm?
4. Is there a viable fallback? If the AI fails, does the user still get a usable experience?

No single question is disqualifying alone — context matters. A task might be ambiguous but have zero tolerance for imperfect output, which pushes you toward a hybrid approach rather than full AI automation. What the framework surfaces is the combination that matters: where the answers are aligned, AI can genuinely add value; where they conflict, you need a more careful design.

TIP
How to Run This With Your Team
Write the four questions on a whiteboard and evaluate your proposed feature together. Disagreement between team members on any question is signal — it means you haven't aligned on what the feature is actually doing or who it's for. Surface that conflict now, not after you've shipped.

The four questions are not a checklist you run once. Return to them whenever the scope of a feature changes. A feature that clears all four at the proposal stage may stop clearing them after user research reveals the real task is more constrained than you assumed.


3 Questions 1 and 2: Ambiguity and Personalization

Question 1: Is the task inherently ambiguous?

A task is ambiguous when there is no single correct answer — the right output depends on context, intent, style, or nuance that cannot be captured in a rule. Writing an email, generating a summary for a specific audience, suggesting the next step in a workflow that varies by user: these are genuinely ambiguous. Sorting a list of items alphabetically, displaying the correct price from a database, or routing a support ticket to the right queue: these are not. Rules, search, and deterministic logic handle the latter category better than AI — they are faster, cheaper to run, and easier to test.

A useful diagnostic: could you write a finite set of rules that would produce correct output for at least 90% of cases? If yes, write the rules. AI adds cost, latency, and unpredictability without returning proportional value.

A real product decision that went wrong: A fintech company added an AI chatbot to their FAQ section. Users' questions were mostly in three categories — account balance, fee schedule, and how to dispute a charge — which could have been answered with keyword matching and three lookup queries. The chatbot hallucinated fee amounts, broke user trust, and was quietly rolled back after three months. The task had no meaningful ambiguity; AI was the wrong tool.

Question 2: Does it benefit from personalization?

Personalization means the right answer is meaningfully different for different users or contexts. A recommendation engine that surfaces the same top-ten items for everyone is not personalized — it is ranked content. True personalization shifts based on individual signals: a user's past behaviour, stated preferences, expertise level, or current context.

AI is appropriate here when the space of possible variations is too large for manual segmentation, and when you have enough signal to personalize usefully rather than randomly. It is not appropriate when your user base is small and homogeneous, when you have no behavioural data yet, or when the right answer is genuinely the same for everyone.

WARNING
Personalization Is Not Always Better
Early in a product's life, you often lack the data to personalize well — and poor personalization actively harms the experience. A new user who gets a blank slate because the AI has no signal on them is worse off than a user who gets a thoughtful, curated default. Start with sensible defaults and layer personalization only when the signal is reliable.
Feature Ambiguous? Personalized? Verdict
Auto-fill an address field No No Rules are better
Suggest the right email reply tone Yes Yes Good AI candidate
Show the user their account balance No No (same data) Lookup, no AI needed
Recommend the next lesson in a course Yes Yes Good AI candidate
Sort search results by date No No Rules are better
Surface anomalies in a user's spend Yes Yes Good AI candidate

4 Questions 3 and 4: Tolerance for Imperfection and the Fallback

Question 3: Is there tolerance for imperfect output?

AI systems are probabilistic. They are right most of the time and wrong some of the time, and — crucially — they are often wrong without signalling that they are wrong. Any product decision to use AI must honestly answer: what happens to the user when the output is incorrect? And does the cost of that error fit within acceptable bounds?

This is a spectrum, not a binary. On one end: a writing assistant that suggests an awkward sentence, which the user immediately sees and ignores. The cost of error is nearly zero. On the other end: an AI feature that surfaces a medical drug interaction warning — or silently omits one. The cost of error can be severe. Between these poles, most product decisions sit in a middle zone where the answer depends on context.

The right analysis combines two factors: how often will it be wrong (which depends on the model and the task) and how expensive is a wrong answer (which depends on what the user does with it and what your fallback is).

> Framework: error tolerance grid

             | Low error cost      | High error cost
-------------|---------------------|------------------
Low frequency| Green: build with AI| Yellow: build with
of errors    | minimal guardrails  | strong transparency
             |                     | and verification
-------------|---------------------|------------------
High frequency| Yellow: improve    | Red: do not use
of errors    | model or reduce     | AI here without
             | AI's scope          | a human review step

The grid is a first-pass diagnostic, not a final answer. A red-zone feature may still be buildable if you add human review before the output reaches the user — but that changes the cost structure and the UX fundamentally.

Question 4: Is there a viable fallback?

Every AI feature should degrade gracefully. If the model is unavailable, if confidence falls below a threshold, or if the task is outside the model's reliable capability, what does the user see? "Something went wrong" is not a fallback — it is an abandonment. A real fallback serves the user's underlying goal through a different path: a search interface instead of an AI answer, a static recommendation instead of a personalized one, a prompt to contact support instead of an AI that fails silently.

Fallback design is not a nice-to-have added at the end of the sprint. It belongs in the product spec from day one. Designing the fallback forces clarity on what the feature is actually for — and teams that cannot describe a fallback often discover they have not clearly defined the user need.

TIP
Design the Fallback First
Some teams find it useful to design the fallback experience before designing the AI experience. If the fallback is good enough that users would be satisfied with it, the AI experience becomes an enhancement rather than a dependency. If the fallback is terrible, the product has a systemic reliability risk that needs to be addressed before launch.
WARNING
Invisible Failures Are Worse Than Visible Ones
An AI feature that fails silently — returning a plausible but wrong answer with no indication of uncertainty — erodes trust faster than one that says 'I'm not sure about this.' Lesson 3 (User Trust and Transparency) covers how to signal uncertainty honestly. Lesson 4 (Designing for Errors) covers the full error design vocabulary.

5 Five Anti-Patterns: How AI Gets Added for the Wrong Reasons

These patterns appear repeatedly in product decisions that result in AI features no one uses, trusts, or benefits from. Recognising them early saves months of build time and prevents real user harm.

Anti-pattern 1: AI as a UI shortcut. The product has a complexity problem — too many menus, too many settings, a learning curve that is driving churn. The proposed solution is a conversational AI interface so users can "just ask" instead of navigating. The underlying problem is real, but the solution is usually wrong. Conversational UI is a high-friction pattern that requires users to know what to ask for. A well-structured menu, a better onboarding flow, or progressive disclosure often solves the complexity problem at a tenth of the cost and with far higher reliability. AI is not a substitute for good information architecture.

Anti-pattern 2: The confidence patch. An existing feature has known quality problems — search results that miss the mark, recommendations that feel stale, content filtering that catches too many false positives. The proposal is to replace the rule-based system with an AI model. Sometimes this is the right call. But often, the rules are wrong because the underlying data is poor, the taxonomy is out of date, or the problem definition has drifted. Pouring AI on top of a data or process problem does not fix it — it hides it behind a more expensive and harder-to-debug layer.

Anti-pattern 3: Novelty as a feature. "We need an AI feature in our next release" is a business requirement, not a user need. Products built around the AI capability rather than the user job-to-be-done produce features that feel impressive in a demo and are abandoned within weeks of launch. Users do not care whether an outcome was produced by AI, rules, or a team of humans — they care whether it helped them accomplish something. The moment AI becomes the feature rather than the delivery mechanism, the product is optimising for the wrong thing.

Anti-pattern 4: Automation without agency. Teams sometimes deploy AI that makes decisions on behalf of users without giving users meaningful visibility or control. An email prioritization system that buries messages the user would have wanted to see. A content moderation system that removes posts without explanation or appeal. An autofill feature that submits forms without a confirmation step. Each of these represents AI taking action in a context where the user's ability to review and override is essential. Microsoft's Human-AI Interaction guidelines specifically call out the need for appropriate user agency — particularly for actions that are difficult to reverse.

Anti-pattern 5: Premature autonomy. There is a spectrum from "AI suggests, human decides" to "AI decides and acts." Moving too far toward the autonomous end before trust is established — both the user's trust and the team's own validated confidence in the model's reliability — creates fragile systems and users who feel blindsided. Google's PAIR Guidebook recommends calibrating autonomy to the stakes: higher stakes warrant more human oversight at every stage of a feature's maturity.

NOTE
A Useful Heuristic
Before committing to AI, ask: what is the minimum viable version of this feature without AI? If the minimum viable version is useful and shippable, start there. You can add AI as a quality improvement once you understand the real user behaviour — and you will make better AI design decisions with that data.

6 Success Stories: When AI Unlocks Capabilities That Did Not Exist Before

The anti-patterns above should make you more rigorous, not more risk-averse. AI genuinely enables product capabilities that were either impossible or economically unviable before, and understanding the pattern of those successes helps you recognise the same opportunities in your own roadmap.

Pattern A: Crossing the personalization threshold. A music or video streaming product can manually curate a hundred playlists — but it cannot curate a distinct, evolving playlist for each of a hundred million users based on their individual listening history, context, and taste graph. The task is genuinely ambiguous (taste cannot be reduced to rules), the right answer varies significantly across users, errors are low-cost (you skip a song), and the fallback is a curated playlist that still works. All four framework questions align. This is the clearest case for AI.

Pattern B: Processing at human-impossible scale. A platform generates tens of thousands of user-written posts per minute. Moderating for policy violations, surfacing the highest-quality content, identifying emerging trends — none of these are feasible for a human team at that velocity. The tasks are ambiguous (policy edge cases), context-dependent (what counts as harassment varies by community norms), and benefit from continuous adaptation as new patterns emerge. AI is not optional here; it is what makes the product viable.

Pattern C: Turning expertise into scale. A legal research product wants to let junior associates ask natural-language questions about a corpus of case law and receive structured, cited answers. Previously, that required a senior partner's time. The task is inherently ambiguous (legal interpretation), personalised to the case at hand, tolerant of imperfect output when clearly caveated (the associate reviews and validates), and has a clear fallback (manual search). This is the "AI as expert assistant" pattern — it distributes specialised capability more broadly without removing human judgement from consequential decisions.

Pattern D: Closing the feedback loop in real time. A writing tool monitors the user's draft and surfaces specific, relevant suggestions as they type — not generic grammar corrections, but observations like "this paragraph buries the key point" or "this sentence is ambiguous to a non-technical reader." A static rule set could catch grammar; it cannot assess rhetorical clarity in context. The task requires understanding the full document, the user's apparent intent, and the likely reader — genuinely ambiguous, genuinely personalised.

TIP
The New Capability Test
Ask yourself: could we have shipped a version of this feature five years ago that did roughly the same job, without AI? If yes, the AI version may be a quality improvement — valuable, but not transformative. If no, you may be in Pattern territory: AI is enabling something genuinely new. Those are the features worth the investment and the complexity.
Product type What AI unlocked Framework alignment
Music streaming Per-user taste graph at scale Ambiguous + Personalized + Low-cost error + Clear fallback
Writing assistant Real-time rhetorical feedback Ambiguous + Contextual + Visible errors + Fallback to no suggestion
Legal research Natural-language query over case law Ambiguous + Case-specific + Reviewed by human + Fallback to search
Content moderation Policy enforcement at platform scale Ambiguous + Context-dependent + Errors reviewed + Human escalation

7 Applying the Framework to Your Own Roadmap

The framework is most useful when you run it on a specific, concrete feature proposal — not on a category. "Should we add AI to our product?" is not the question. "Should the next-step recommendation in our onboarding flow use an AI model to adapt to user behaviour, or should we ship a static decision tree first?" is the question.

Here is a worked example to show what that analysis looks like in practice.

Scenario: A B2B project management tool is considering adding an AI feature that writes a first-draft weekly status report for project managers, based on the tasks completed, comments made, and milestones updated during the week.

> Applying the four-question framework:

1. Is the task inherently ambiguous?
   Yes — summarising project activity into a narrative requires
   judgement about what is significant, how to frame delays
   without alarmism, and what the audience cares about.
   Rules cannot reliably make those calls.

2. Does it benefit from personalization?
   Moderately — different project managers have different
   reporting styles, different stakeholders have different
   preferences. Personalization improves quality over time
   but is not essential for the first version.

3. Is there tolerance for imperfect output?
   Yes — the PM reviews the draft before sending it.
   An awkward sentence or mischaracterised milestone is
   caught before reaching the stakeholder. The stakes of
   an error in the draft are low.

4. Is there a viable fallback?
   Yes — if the AI produces nothing useful, the PM writes
   the report manually, as they did before. The feature
   degrades to the previous state without breaking anything.

Verdict: Strong AI candidate. Build with a review step
and a "generate new version" option. Monitor edit
distance to see how much PMs change the drafts.

Run this exercise on two or three features from your own roadmap. Disagreement between team members on any single question is worth exploring — it usually surfaces an assumption about the user need or the task definition that has not been made explicit.

NOTE
What Comes Next
Once you have decided AI belongs in a feature, the next question is how it should appear in the UI. That is the subject of Lesson 2, which covers the four major AI UX patterns — chat interfaces, inline suggestions, autonomous agents, and ambient intelligence — and when each is appropriate.

Questions & Answers

Q: Our stakeholders are demanding AI features for competitive reasons, regardless of whether users actually need them. How do I use this framework when the decision has already been made above my level?
This is one of the most common real-world situations, and the framework is still useful here — it just shifts from a yes/no gate to a design input. If the feature is shipping regardless, the framework tells you how to design it responsibly: where to add transparency, what the fallback must be, and which error scenarios to prioritise. It also gives you a clear language for raising the highest-risk aspects with stakeholders: not "we shouldn't build this" but "this feature has a high error cost in this specific scenario — here is what we need to add to the design to handle it."
Q: How do we know whether our tolerance for imperfect output is realistic before we've actually seen the model's error rate in production?
You often do not — and that is a legitimate design constraint, not a reason to skip the question. The honest approach is to state your assumption explicitly ("we are assuming an error rate below X% on this class of task") and build a measurement plan that will confirm or disconfirm it. Set a guardrail metric — a maximum acceptable error correction rate or user complaint rate — and a rollback plan. Lesson 8 covers how to set those measurement thresholds. The key is not false certainty; it is documenting the assumption so you know what signal to watch for.
Q: The framework points toward "no AI" for our proposed feature, but our engineers have already built a prototype and are excited about it. How do we handle that?
This is a real cost, and it deserves honest acknowledgement rather than a framework that dismisses sunk effort. The question to bring to the team is: given what the framework is telling us, can we redesign the scope of this feature so it lands in a part of the task where AI genuinely adds value? Sometimes a prototype built for the wrong job reveals a different, adjacent job where AI is genuinely useful. If the answer is genuinely no, the earlier you stop, the less expensive it is — and "we built a prototype that taught us this is not the right application" is a legitimate and valuable outcome from a prototyping stage.
Q: We are an early-stage product with very little user data. Does the personalization question even apply to us?
Mostly no — and that is the right answer. Personalization without sufficient signal produces random variation dressed up as personalisation, which is worse than a thoughtful default. For early-stage products, the better questions to focus on are ambiguity and fallback. Get the AI feature working reliably for the median user first. Design the feedback mechanisms that will let you collect the signal you need to personalise later. Lesson 5 (Feedback Loops) and Lesson 6 (Personalization with AI) cover the progression from sensible defaults to genuine personalization as data matures.
Q: How should we handle features where the same AI model fails the framework for some users but passes it for others?
This is more common than teams expect, and it points toward a segmented rollout strategy rather than a single feature design. If the feature genuinely serves power users with established histories but fails new users with no signal, ship it to the segment it serves and use a different design — a static recommendation, a wizard, or a prompted setup flow — for the segment it does not yet serve well. Progressive personalisation, covered in Lesson 6, is the pattern for managing this transition over time.

Key Takeaways

  1. Start with the user need, not the capability — "We can add AI here" is never sufficient justification. The right question is whether AI uniquely solves a user problem that simpler tools cannot address as well.
  2. Run the four-question test on every proposal — Ambiguity, personalization, error tolerance, and fallback viability are the four conditions that distinguish genuine AI use cases from expensive experiments.
  3. Rules and search are often the right answer — For tasks that are deterministic, narrowly scoped, or low-ambiguity, a rule-based or lookup approach is faster, cheaper, more testable, and more reliable than AI.
  4. Anti-patterns are expensive to reverse — AI-as-UI-shortcut, the confidence patch, and premature autonomy each erode user trust in ways that are hard to recover from once shipped.
  5. Design the fallback first — If you cannot describe a useful experience when the AI fails, you have not finished defining the feature. A well-designed fallback also reveals what the feature is truly for.
  6. AI's strongest cases create genuinely new value — Personalization at scale, real-time contextual feedback, and expert assistance at volume are the patterns where AI enables something that did not exist before — and that is where the investment is worth it.

Next Steps: Lesson 2: AI UX Patterns