When to Use AI (and When Not To)
Learning Outcomes
- Apply a four-question decision framework to judge whether any product feature genuinely needs AI
- Identify the anti-patterns that lead teams to add AI where simpler solutions would serve users better
- Distinguish tasks where imperfect AI output is acceptable from tasks where errors are genuinely costly
- Evaluate a product proposal against the framework and arrive at a defensible build/don't-build recommendation
- Recognise the conditions that make AI unlock entirely new value rather than just replicate existing functionality
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why "add AI to it" is the wrong starting question |
| The framework | 8 min | Four diagnostic questions every product decision needs |
| Question 1–2 deep dive | 6 min | Ambiguity and personalization: what they really mean |
| Question 3–4 deep dive | 6 min | Tolerance for imperfection and fallback design |
| Anti-patterns | 7 min | Five ways AI gets added for the wrong reasons |
| Success stories | 6 min | When AI creates capabilities that didn't exist before |
| Applying the framework | 2 min | How to run the test on your own roadmap |
| Wrap-up | 2 min | Key takeaways and next lesson |
Before You Begin
Pre-work:
- No previous lessons are required — this is Lesson 1. If you want context on the course arc, skim the AI Product Design course landing page.
- Think of one or two AI features on your current or recent product roadmap. You will apply the framework to them as we go.
- Bring any open questions about a specific proposal — the Q&A section at the end addresses the most common objections teams raise.
Shopping List:
- A web browser
- A notes document alongside — you will want to write down your framework verdicts as you work through the steps
- Optionally: the AI Glossary tab, open in the background, for any terms of art you want to check
Somewhere in the last two years, "we should add AI to this" became the default response to every product challenge. A struggling retention metric? Add an AI recommendation engine. Users can't find features? Add an AI assistant. Churn is up? Add an AI-powered engagement nudge. The instinct is understandable — AI is genuinely powerful, and the competitive pressure to ship AI features is real.
But the instinct produces predictable failures. Teams build AI features that users ignore, features that erode trust when they get things wrong in small but visible ways, and features that cost ten times as much to maintain as the rule-based or search-based solution they replaced. The graveyard of abandoned AI features is already crowded, and almost every failure traces back to the same root: a team that started with the technology and reverse-engineered a use case, rather than starting with the user problem and asking whether AI was the right tool.
The right starting question is not "can we add AI here?" — you almost certainly can. The right question is "should we, and what specifically are users trying to do that AI uniquely helps with?"
This lesson gives you a four-question framework for making that call with your team in a structured way. By the end, you will be able to look at any proposed AI feature and arrive at a defensible answer: yes, this genuinely needs AI — or no, something simpler will serve users better and cost you less.
Before building any AI feature, run it through four diagnostic questions. A genuine "yes" to all four is a strong signal that AI belongs here. "No" answers reveal where a simpler solution would be better — or where the feature design needs to change before AI is appropriate.
| Question | What you're probing |
|---|---|
| 1. Is the task inherently ambiguous? | Would a rule or a search index handle this adequately? |
| 2. Does it benefit from personalization? | Does the right answer vary meaningfully across users or contexts? |
| 3. Is there tolerance for imperfect output? | Can errors be surfaced, corrected, and absorbed without real harm? |
| 4. Is there a viable fallback? | If the AI fails, does the user still get a usable experience? |
No single question is disqualifying alone — context matters. A task might be ambiguous but have zero tolerance for imperfect output, which pushes you toward a hybrid approach rather than full AI automation. What the framework surfaces is the combination that matters: where the answers are aligned, AI can genuinely add value; where they conflict, you need a more careful design.
The four questions are not a checklist you run once. Return to them whenever the scope of a feature changes. A feature that clears all four at the proposal stage may stop clearing them after user research reveals the real task is more constrained than you assumed.
Question 1: Is the task inherently ambiguous?
A task is ambiguous when there is no single correct answer — the right output depends on context, intent, style, or nuance that cannot be captured in a rule. Writing an email, generating a summary for a specific audience, suggesting the next step in a workflow that varies by user: these are genuinely ambiguous. Sorting a list of items alphabetically, displaying the correct price from a database, or routing a support ticket to the right queue: these are not. Rules, search, and deterministic logic handle the latter category better than AI — they are faster, cheaper to run, and easier to test.
A useful diagnostic: could you write a finite set of rules that would produce correct output for at least 90% of cases? If yes, write the rules. AI adds cost, latency, and unpredictability without returning proportional value.
A real product decision that went wrong: A fintech company added an AI chatbot to their FAQ section. Users' questions were mostly in three categories — account balance, fee schedule, and how to dispute a charge — which could have been answered with keyword matching and three lookup queries. The chatbot hallucinated fee amounts, broke user trust, and was quietly rolled back after three months. The task had no meaningful ambiguity; AI was the wrong tool.
Question 2: Does it benefit from personalization?
Personalization means the right answer is meaningfully different for different users or contexts. A recommendation engine that surfaces the same top-ten items for everyone is not personalized — it is ranked content. True personalization shifts based on individual signals: a user's past behaviour, stated preferences, expertise level, or current context.
AI is appropriate here when the space of possible variations is too large for manual segmentation, and when you have enough signal to personalize usefully rather than randomly. It is not appropriate when your user base is small and homogeneous, when you have no behavioural data yet, or when the right answer is genuinely the same for everyone.
| Feature | Ambiguous? | Personalized? | Verdict |
|---|---|---|---|
| Auto-fill an address field | No | No | Rules are better |
| Suggest the right email reply tone | Yes | Yes | Good AI candidate |
| Show the user their account balance | No | No (same data) | Lookup, no AI needed |
| Recommend the next lesson in a course | Yes | Yes | Good AI candidate |
| Sort search results by date | No | No | Rules are better |
| Surface anomalies in a user's spend | Yes | Yes | Good AI candidate |
Question 3: Is there tolerance for imperfect output?
AI systems are probabilistic. They are right most of the time and wrong some of the time, and — crucially — they are often wrong without signalling that they are wrong. Any product decision to use AI must honestly answer: what happens to the user when the output is incorrect? And does the cost of that error fit within acceptable bounds?
This is a spectrum, not a binary. On one end: a writing assistant that suggests an awkward sentence, which the user immediately sees and ignores. The cost of error is nearly zero. On the other end: an AI feature that surfaces a medical drug interaction warning — or silently omits one. The cost of error can be severe. Between these poles, most product decisions sit in a middle zone where the answer depends on context.
The right analysis combines two factors: how often will it be wrong (which depends on the model and the task) and how expensive is a wrong answer (which depends on what the user does with it and what your fallback is).
> Framework: error tolerance grid
| Low error cost | High error cost
-------------|---------------------|------------------
Low frequency| Green: build with AI| Yellow: build with
of errors | minimal guardrails | strong transparency
| | and verification
-------------|---------------------|------------------
High frequency| Yellow: improve | Red: do not use
of errors | model or reduce | AI here without
| AI's scope | a human review step
The grid is a first-pass diagnostic, not a final answer. A red-zone feature may still be buildable if you add human review before the output reaches the user — but that changes the cost structure and the UX fundamentally.
Question 4: Is there a viable fallback?
Every AI feature should degrade gracefully. If the model is unavailable, if confidence falls below a threshold, or if the task is outside the model's reliable capability, what does the user see? "Something went wrong" is not a fallback — it is an abandonment. A real fallback serves the user's underlying goal through a different path: a search interface instead of an AI answer, a static recommendation instead of a personalized one, a prompt to contact support instead of an AI that fails silently.
Fallback design is not a nice-to-have added at the end of the sprint. It belongs in the product spec from day one. Designing the fallback forces clarity on what the feature is actually for — and teams that cannot describe a fallback often discover they have not clearly defined the user need.
These patterns appear repeatedly in product decisions that result in AI features no one uses, trusts, or benefits from. Recognising them early saves months of build time and prevents real user harm.
Anti-pattern 1: AI as a UI shortcut. The product has a complexity problem — too many menus, too many settings, a learning curve that is driving churn. The proposed solution is a conversational AI interface so users can "just ask" instead of navigating. The underlying problem is real, but the solution is usually wrong. Conversational UI is a high-friction pattern that requires users to know what to ask for. A well-structured menu, a better onboarding flow, or progressive disclosure often solves the complexity problem at a tenth of the cost and with far higher reliability. AI is not a substitute for good information architecture.
Anti-pattern 2: The confidence patch. An existing feature has known quality problems — search results that miss the mark, recommendations that feel stale, content filtering that catches too many false positives. The proposal is to replace the rule-based system with an AI model. Sometimes this is the right call. But often, the rules are wrong because the underlying data is poor, the taxonomy is out of date, or the problem definition has drifted. Pouring AI on top of a data or process problem does not fix it — it hides it behind a more expensive and harder-to-debug layer.
Anti-pattern 3: Novelty as a feature. "We need an AI feature in our next release" is a business requirement, not a user need. Products built around the AI capability rather than the user job-to-be-done produce features that feel impressive in a demo and are abandoned within weeks of launch. Users do not care whether an outcome was produced by AI, rules, or a team of humans — they care whether it helped them accomplish something. The moment AI becomes the feature rather than the delivery mechanism, the product is optimising for the wrong thing.
Anti-pattern 4: Automation without agency. Teams sometimes deploy AI that makes decisions on behalf of users without giving users meaningful visibility or control. An email prioritization system that buries messages the user would have wanted to see. A content moderation system that removes posts without explanation or appeal. An autofill feature that submits forms without a confirmation step. Each of these represents AI taking action in a context where the user's ability to review and override is essential. Microsoft's Human-AI Interaction guidelines specifically call out the need for appropriate user agency — particularly for actions that are difficult to reverse.
Anti-pattern 5: Premature autonomy. There is a spectrum from "AI suggests, human decides" to "AI decides and acts." Moving too far toward the autonomous end before trust is established — both the user's trust and the team's own validated confidence in the model's reliability — creates fragile systems and users who feel blindsided. Google's PAIR Guidebook recommends calibrating autonomy to the stakes: higher stakes warrant more human oversight at every stage of a feature's maturity.
The anti-patterns above should make you more rigorous, not more risk-averse. AI genuinely enables product capabilities that were either impossible or economically unviable before, and understanding the pattern of those successes helps you recognise the same opportunities in your own roadmap.
Pattern A: Crossing the personalization threshold. A music or video streaming product can manually curate a hundred playlists — but it cannot curate a distinct, evolving playlist for each of a hundred million users based on their individual listening history, context, and taste graph. The task is genuinely ambiguous (taste cannot be reduced to rules), the right answer varies significantly across users, errors are low-cost (you skip a song), and the fallback is a curated playlist that still works. All four framework questions align. This is the clearest case for AI.
Pattern B: Processing at human-impossible scale. A platform generates tens of thousands of user-written posts per minute. Moderating for policy violations, surfacing the highest-quality content, identifying emerging trends — none of these are feasible for a human team at that velocity. The tasks are ambiguous (policy edge cases), context-dependent (what counts as harassment varies by community norms), and benefit from continuous adaptation as new patterns emerge. AI is not optional here; it is what makes the product viable.
Pattern C: Turning expertise into scale. A legal research product wants to let junior associates ask natural-language questions about a corpus of case law and receive structured, cited answers. Previously, that required a senior partner's time. The task is inherently ambiguous (legal interpretation), personalised to the case at hand, tolerant of imperfect output when clearly caveated (the associate reviews and validates), and has a clear fallback (manual search). This is the "AI as expert assistant" pattern — it distributes specialised capability more broadly without removing human judgement from consequential decisions.
Pattern D: Closing the feedback loop in real time. A writing tool monitors the user's draft and surfaces specific, relevant suggestions as they type — not generic grammar corrections, but observations like "this paragraph buries the key point" or "this sentence is ambiguous to a non-technical reader." A static rule set could catch grammar; it cannot assess rhetorical clarity in context. The task requires understanding the full document, the user's apparent intent, and the likely reader — genuinely ambiguous, genuinely personalised.
| Product type | What AI unlocked | Framework alignment |
|---|---|---|
| Music streaming | Per-user taste graph at scale | Ambiguous + Personalized + Low-cost error + Clear fallback |
| Writing assistant | Real-time rhetorical feedback | Ambiguous + Contextual + Visible errors + Fallback to no suggestion |
| Legal research | Natural-language query over case law | Ambiguous + Case-specific + Reviewed by human + Fallback to search |
| Content moderation | Policy enforcement at platform scale | Ambiguous + Context-dependent + Errors reviewed + Human escalation |
The framework is most useful when you run it on a specific, concrete feature proposal — not on a category. "Should we add AI to our product?" is not the question. "Should the next-step recommendation in our onboarding flow use an AI model to adapt to user behaviour, or should we ship a static decision tree first?" is the question.
Here is a worked example to show what that analysis looks like in practice.
Scenario: A B2B project management tool is considering adding an AI feature that writes a first-draft weekly status report for project managers, based on the tasks completed, comments made, and milestones updated during the week.
> Applying the four-question framework:
1. Is the task inherently ambiguous?
Yes — summarising project activity into a narrative requires
judgement about what is significant, how to frame delays
without alarmism, and what the audience cares about.
Rules cannot reliably make those calls.
2. Does it benefit from personalization?
Moderately — different project managers have different
reporting styles, different stakeholders have different
preferences. Personalization improves quality over time
but is not essential for the first version.
3. Is there tolerance for imperfect output?
Yes — the PM reviews the draft before sending it.
An awkward sentence or mischaracterised milestone is
caught before reaching the stakeholder. The stakes of
an error in the draft are low.
4. Is there a viable fallback?
Yes — if the AI produces nothing useful, the PM writes
the report manually, as they did before. The feature
degrades to the previous state without breaking anything.
Verdict: Strong AI candidate. Build with a review step
and a "generate new version" option. Monitor edit
distance to see how much PMs change the drafts.
Run this exercise on two or three features from your own roadmap. Disagreement between team members on any single question is worth exploring — it usually surfaces an assumption about the user need or the task definition that has not been made explicit.
Questions & Answers
Key Takeaways
- Start with the user need, not the capability — "We can add AI here" is never sufficient justification. The right question is whether AI uniquely solves a user problem that simpler tools cannot address as well.
- Run the four-question test on every proposal — Ambiguity, personalization, error tolerance, and fallback viability are the four conditions that distinguish genuine AI use cases from expensive experiments.
- Rules and search are often the right answer — For tasks that are deterministic, narrowly scoped, or low-ambiguity, a rule-based or lookup approach is faster, cheaper, more testable, and more reliable than AI.
- Anti-patterns are expensive to reverse — AI-as-UI-shortcut, the confidence patch, and premature autonomy each erode user trust in ways that are hard to recover from once shipped.
- Design the fallback first — If you cannot describe a useful experience when the AI fails, you have not finished defining the feature. A well-designed fallback also reveals what the feature is truly for.
- AI's strongest cases create genuinely new value — Personalization at scale, real-time contextual feedback, and expert assistance at volume are the patterns where AI enables something that did not exist before — and that is where the investment is worth it.
Next Steps: Lesson 2: AI UX Patterns