Case Studies

55 min advanced Lesson 10

Learning Outcomes

  • Analyse real AI products against the trust, error, feedback, and personalization frameworks covered in this course
  • Identify which UX pattern each case study uses and why that choice fits the underlying task
  • Extract transferable design principles from diverse product categories — creative tools, productivity, consumer, and enterprise
  • Distinguish decisions that were uniquely right for a specific context from decisions that generalise broadly
  • Apply course frameworks to critique an AI product and propose targeted improvements

Lesson Plan

Segment Duration Topic
Intro 3 min Why studying shipped products teaches what theory cannot
Case Study 1 9 min Writing assistants: AI as inline creative partner
Case Study 2 8 min Email triage and composition: ambient AI meeting user expectations
Case Study 3 8 min Code completion: trust, transparency, and acceptance rate
Case Study 4 8 min Consumer recommendation engines: personalisation at scale
Case Study 5 8 min Enterprise analytics AI: explaining the unexplainable
Cross-cutting principles 8 min What every case shares — and what only some get right
Wrap-up 3 min Key takeaways and applying this to your own product

Before You Begin

Pre-work:

Shopping List:

  • A browser with access to at least one of the products discussed (free accounts are sufficient for observation — you are studying the UX, not the model quality)
  • Your notes or a blank document to capture the principles as they emerge
  • Thirty minutes of uninterrupted thinking time after the lesson to complete the personal analysis exercise

1 Why Case Studies Beat Theory

Every framework in this course was built backwards — by observing what worked and what failed in shipped products, then generalising. That means the most reliable way to internalise the frameworks is to run the process in reverse: start from a product, peel back the design decisions, and ask why each choice was made.

This lesson does that for five products across four categories. In each case we ask the same set of questions:

Analysis question What it reveals
What problem does the AI solve here? Whether AI is the right tool at all (Lesson 1 framework)
Which UX pattern is in use? The interaction model and user expectations it creates (Lesson 2)
How does the product build or maintain trust? Transparency and confidence signals (Lesson 3)
How does it handle errors and uncertainty? Degradation and recovery design (Lesson 4)
How does feedback flow back into improvement? Explicit vs. implicit collection (Lesson 5)
Where does personalisation appear? Adaptation strategy and privacy tradeoffs (Lesson 6)
What would you change? Applying your own judgment

You do not need to have used each product deeply. The design patterns are visible from even a single session — and many decisions are evident from published case studies, help documentation, and onboarding flows.

NOTE
A Note on Naming
This lesson describes design patterns as they appeared in well-known products at the time of writing. Specific features evolve quickly; the design decisions and tradeoffs are durable. Where a feature has changed, the underlying principle still applies.

The goal is not to produce verdicts about whether each product is good or bad. The goal is to see the decisions clearly enough that you can make better ones yourself.


2 Case Study 1 — Writing Assistants: AI as Inline Creative Partner

The product category: AI writing assistants embedded inside a document editor — the model that Notion AI, Google Docs' "Help me write," and similar tools follow.

The problem being solved. Writers face two distinct problems: the blank page (getting started) and the stuck page (continuing when momentum stops). Both are caused by the same thing — converting a fuzzy intent into specific words is cognitively expensive. AI reduces that cost by producing a draft the writer can react to, which is faster than producing one from nothing.

UX pattern chosen: Inline suggestions. The AI lives inside the writer's existing workflow, not in a separate chat window. This is the right pattern here because context is the product: the AI reads what has already been written and continues it, or responds to a command that refers to selected text. Moving to a chat window would break the writing flow and force the writer to copy text back and forth.

Trust and transparency. Writing assistants handle a trust problem that other categories largely avoid: the output becomes the user's own work. Readers of the final document will not know which sentences were AI-generated. This creates two sub-problems:

  1. The user needs to trust that the suggestion is accurate enough to use (low stakes for creative text, higher stakes for factual claims embedded in a business document).
  2. The user needs to feel that accepting a suggestion is a genuine choice, not a subtle pressure.

Well-designed writing assistants solve both with visual separation: the suggested text appears in a distinct colour or fade, and it stays on-screen only until the user explicitly accepts, rejects, or replaces it. The visual boundary makes clear that the AI is offering, not deciding.

TIP
The Ghost Text Principle
Showing AI suggestions in a visually distinct state — often called ghost text or a suggestion overlay — is one of the most replicable trust patterns in AI product design. It communicates uncertainty without requiring any explicit confidence score. Users understand that text they can see but have not accepted is provisional.

Error handling. Writing AI fails in two modes: factual hallucination (inventing a claim or figure) and stylistic mismatch (producing text that does not sound like the user). Factual hallucinations are managed by keeping human review responsibility explicit — the accept gesture signals that the writer has read and endorsed the text. Stylistic mismatches are managed by short suggestions (two to four sentences) rather than long autonomous passages, because a small mismatch is easy to fix and a page-length mismatch requires starting over.

Feedback loops. Acceptance rate — what percentage of generated suggestions the user accepts unedited — is the primary implicit signal. Products in this category also track edit distance after acceptance: how much does the user change an accepted suggestion? High edit distance means the suggestion was technically accepted but practically wrong. The healthy metric is: accepted and used without major revision.

What the best products get right that weaker ones miss. The strongest writing assistants show the suggestion in the right place at the right time — triggered by a pause in typing or an explicit command, never pushing text onto the screen uninvited. Uninvited text feels intrusive; invited text feels helpful. The same model capability, with different trigger design, produces completely different trust responses.

Design decision Weaker version Stronger version
When to suggest On every pause On explicit command or natural stopping point
Length of suggestion Full paragraph or more Two to four sentences max
Acceptance gesture Auto-accept on continued typing Deliberate accept key (Tab or dedicated button)
Factual content Generated freely Hedged ("you may want to verify...") or avoided
WARNING
Where Writing Assistants Fail
Products that surface AI suggestions for factual documents — legal memos, financial reports, medical notes — without additional verification prompts are taking on significant liability. The inline suggestion pattern assumes the writer will review. When the document stakes are high and writers are under time pressure, that assumption breaks. Consider adding an explicit review step or confidence indicators for fact-sensitive content types.

3 Case Study 2 — Email Triage and Composition: Ambient AI Meets User Expectations

The product category: AI features built into an email client — the pattern demonstrated by Gmail's Smart Reply, Smart Compose, and Priority Inbox, and later by tools like Superhuman's AI summaries and draft generation.

The problem being solved. Email volume exceeds most people's capacity to process it. Two tasks consume disproportionate time: reading and routing (deciding which messages need a response, which need to wait, and which can be ignored) and composing (writing responses for messages that follow a predictable pattern). Both are rule-following under noise, which is a strong fit for AI.

UX patterns in use. Email AI products typically deploy two patterns simultaneously:

  • Ambient intelligence for triage — the AI filters, prioritises, and surfaces messages without asking the user to do anything. The experience is that your inbox simply shows the right things first.
  • Inline suggestions for composition — suggested replies and sentence completions appear as the user is reading or typing, without switching contexts.

This dual pattern creates a design tension: ambient AI is invisible by design, while inline AI is visible by design. Products that handle this well use ambient AI for low-consequence routing (ordering messages) and visible AI for consequential composition (words that will go to another person).

Trust challenges unique to email. Email has higher interpersonal stakes than a private document. An AI-generated reply carries your name and reaches another human who may adjust their relationship with you based on it. The trust problem is: users need to feel in control of their voice in a channel that is deeply personal.

Gmail's Smart Reply handles this with very short suggestions — three options, typically five words or fewer. A five-word suggestion is fast to evaluate and easy to edit. A full-paragraph suggestion requires careful review and creates editing work that undermines the time-saving premise. Length management is trust management here.

NOTE
Progressive Disclosure of AI Involvement
Gmail's Smart Compose surfaced suggestions character by character as the user typed, using grey ghost text. Users could accept a full sentence with Tab or keep typing to override. This design gave AI a visible but non-intrusive presence: always ready to help, never in the way. The trigger was the user's own typing — the AI was responding, not initiating.

Error handling and degradation. Triage AI fails when it misprioritises — surfacing a low-urgency message prominently and burying an urgent one. The graceful degradation pattern here is to make the triage decision reversible and visible: users can see that messages have been sorted, they can move things manually, and the category labels ("Important," "Other") are always explicit. Invisible sorting that users cannot inspect or override destroys trust when it gets something wrong.

Priority Inbox's decision to show visible category labels rather than simply silently re-ordering was a significant trust design decision. It turned ambient AI into inspectable AI, which allowed users to calibrate their trust over time.

Feedback loops. Email triage products rely almost entirely on implicit feedback: opening a deprioritised message tells the system it got the importance wrong; deleting without opening a prioritised message tells it the same. No thumbs-up or thumbs-down required. The richness of email interaction data (open, reply, archive, delete, move) gives triage models an abundance of signal that most AI products do not have.

What transfers. The key lesson from email AI is that ambient intelligence only works if users can see enough to calibrate trust, even if they do not see everything. Full transparency and full invisibility are both worse than inspectable ambient AI.

Inspectability level User experience Trust outcome
Fully invisible "Something weird is happening to my inbox" Erodes trust when errors surface
Visible labels and reversible decisions "I can see what it's doing and fix mistakes" Appropriate, calibrated trust
Explicit opt-in required for every decision "This takes as long as doing it myself" User fatigue; feature abandoned

4 Case Study 3 — Code Completion: Trust, Transparency, and Acceptance Rate

The product category: AI code completion tools — the pattern established by GitHub Copilot and followed by Cursor, JetBrains AI, and similar tools embedded in development environments.

Why this case study matters for non-engineers. You do not need to write code to draw lessons from code completion AI. It is the most instrumented, most studied inline AI product category available, and the design decisions it made — particularly around trust, latency, and feedback — have become reference patterns for every other inline AI tool.

The problem being solved. Experienced engineers spend a surprising amount of time on predictable work: boilerplate structure, common patterns they have written dozens of times, and function signatures they can almost remember. This is not the thinking part of programming — it is the typing part. AI handles typing; engineers handle thinking.

UX pattern: Inline suggestions. Code completion is inline suggestions in their purest form. The suggestion appears at the cursor, in context, and accepts or declines with a single keystroke. No modal, no chat window, no output to copy.

Trust signals in code completion. Code carries real consequences — it executes. An incorrect writing suggestion is embarrassing; an incorrect code suggestion is a potential bug or security flaw. The trust design problem is acute.

GitHub Copilot's approach was to show the suggestion visually but never explain it, trusting engineers to read and verify code before accepting it. This works for an expert audience: engineers are trained to read code critically. The same approach would fail for an audience less equipped to verify the output.

The deeper trust signal is latency. Suggestions that appear within a few hundred milliseconds feel like a responsive tool; suggestions that take several seconds feel like waiting for an external service. Perceived responsiveness is a proxy for reliability. Products in this category invest heavily in latency optimisation precisely because speed affects trust independently of accuracy.

TIP
Latency as a Trust Signal
In inline AI tools, response latency affects trust independently of correctness. A fast, occasionally wrong suggestion often earns higher trust than a slow, mostly right one. When scoping AI feature performance requirements, treat latency as a trust metric, not only a technical one.

Error handling. Code completion fails in two ways that have driven significant design work:

  1. Plausible but wrong suggestions — code that looks correct and compiles but has a subtle logical error or uses a deprecated API. The graceful degradation pattern is to show multiple alternatives (Copilot shows up to ten options with a keyboard shortcut) so the user is thinking comparatively rather than accepting the first thing they see.
  2. Context-blind suggestions — the model does not understand the broader system the code lives in, so it proposes something that works in isolation but breaks something elsewhere. This is harder to handle with UX alone, and the honest answer is to set expectations: inline suggestions are line-level and function-level assists, not architectural ones.

Feedback loops and the acceptance rate metric. Acceptance rate — what percentage of shown suggestions the user accepts — became the primary health metric for code completion AI. But teams studying Copilot found that acceptance rate alone was misleading: a user who accepts every suggestion without review inflates the metric without receiving value. The more informative metric is acceptance plus retention: was the accepted code still present ten minutes later, or was it deleted?

Edit distance after acceptance, combined with whether accepted code survived to commit, gives a far richer picture of whether the AI is genuinely helping.

> Evaluate the AI feature in your product: if you could track only two implicit feedback signals, which two would most reliably indicate that the AI helped a user accomplish something real — not just that the user interacted with it?

What transfers. The code completion case study is a masterclass in metric design. The lesson is that engagement metrics (suggestions shown, suggestions accepted) measure how often users interact with AI; outcome metrics (accepted code retained, task completed faster) measure whether the AI produced value. Building toward outcome metrics from day one prevents teams from optimising for activity and calling it success.

WARNING
Vanity Metrics in AI Products
High acceptance rate, high suggestion volume, and high daily active usage are all valid signals — but none of them directly measure whether the AI made users more effective. Before shipping, define at least one outcome metric that can only improve if the AI genuinely helped. The measurement framework from Lesson 8 is the right tool here.

5 Case Study 4 — Consumer Recommendation Engines: Personalisation at Scale

The product category: AI-driven recommendation systems in consumer products — the pattern used by Spotify's Discover Weekly, Netflix's homepage recommendations, and similar features in streaming, e-commerce, and social platforms.

The problem being solved. Catalogue depth is simultaneously a product's greatest asset and its most difficult UX challenge. A catalogue with ten million songs cannot be meaningfully browsed. Recommendations replace browsing with surfacing: the AI presents a curated slice of the catalogue it predicts will match the individual user's taste.

UX pattern: Ambient intelligence. Recommendations appear without the user requesting them. The interface does not ask what you like; it observes what you do and adapts. This is ambient intelligence at its most complete — the AI is doing work continuously in the background, and the user experiences only the output.

Personalisation design. Spotify's Discover Weekly is worth examining in detail because it made an explicit design decision that most recommendation systems avoid: it presented personalised recommendations in a named, weekly playlist with a clear premise ("songs you haven't heard, matched to your taste"). This gave the ambient AI a visible identity.

The design consequences of this decision were significant:

  • Users knew what to expect and when. A recommendation feed that updates unpredictably feels noisy; a weekly playlist feels like a delivery.
  • Users could evaluate it as a unit. "Discover Weekly was good this week" or "it missed this time" was a natural response. This created social sharing behaviour and word-of-mouth that a generic feed cannot generate.
  • The clear premise set calibrated expectations: these are new-to-you songs, not your favourites. Users did not mistake it for a playlist of things they already knew they liked.
NOTE
Named AI Features Build Accountability
Giving an AI-driven feature a distinct identity — a name, a cadence, a stated purpose — makes it accountable to the user. The Google PAIR Guidebook calls this setting correct expectations: users who know what the AI is trying to do are better equipped to evaluate whether it succeeded. Compare the discoverability and trust of Discover Weekly against a generic 'Recommended' feed on any platform.

The filter bubble problem. Recommendation AI that only surfaces content similar to what a user has already consumed creates a narrowing effect over time. A user whose listening history is dominated by one genre receives increasingly similar suggestions, missing content they might enjoy but would never discover. This is a design problem as much as a model problem.

Well-designed recommendation systems introduce deliberate diversity — a percentage of suggestions that sit at the edges of the user's observed taste rather than at its centre. Spotify's Discover Weekly reportedly allocated a fraction of each playlist to songs just outside the user's confirmed taste, surfaced because they were highly popular with users who otherwise resembled the listener. This breadth-vs-depth tradeoff is a product design decision, not an algorithmic inevitability.

Trust and the "why am I seeing this" question. Consumer recommendation products were among the first to face regulatory and user pressure to explain their recommendations. Netflix's star rating system (later changed to thumbs up/down) was an early attempt to make preferences explicit so users felt agency over what the system learned about them. The move to thumbs up/down was itself a UX decision: a five-star scale implied more precision than users actually had, and most people clustered ratings at extremes anyway.

Trust mechanism Design choice Trade-off
Explicit ratings User-controlled, high signal Low participation, burden on user
Explanation text ("Because you watched X") Transparent, accountable Can feel surveillance-like if overdone
User-editable taste profile Full control, high trust Complex to surface; most users do not use it
No explanation Simple interface Erodes trust when recommendations are bad

What transfers. Recommendation systems demonstrate that ambient AI requires a visible identity to build sustained trust. Invisible helpfulness can feel magical at first and creepy or random later. Naming what the AI is doing, giving it a cadence, and explaining the premise once — without requiring explanation on every interaction — is the design pattern that scales.

TIP
Introduce Diversity Deliberately
If your AI feature learns from user behaviour, build in explicit diversity logic from the start. Pure exploitation of observed preferences narrows over time. The right moment to add diversity to the design is before launch, not after users have already formed a narrow pattern. Lesson 6 on personalisation covers the technical approaches in more detail.

6 Case Study 5 — Enterprise Analytics AI: Explaining the Unexplainable

The product category: AI features embedded in business intelligence and analytics tools that surface AI-generated observations about business data. The worked example here is Microsoft Power BI, which ships this pattern as a small family of features: Quick Insights scans a whole semantic model for patterns, the report Insights pane flags anomalies, trends and KPI outliers as you read a report, and Explain the increase / decrease explains a single change in a chart.

The problem being solved. Business analysts spend significant time on questions that have deterministic answers hiding in large datasets: Why did revenue drop in this region last quarter? Which customer segments are churning at an unusual rate? What is anomalous about this month's performance? These questions require pattern recognition across many variables simultaneously — a task AI can do quickly, but which produces outputs that are hard for non-technical users to trust or act on.

UX pattern: A hybrid. Enterprise analytics AI sits between ambient intelligence and a conversational interface. The most successful implementations surface a curated set of insights proactively (ambient), but let users drill in and ask follow-up questions (conversational). The Microsoft Human-AI Interaction guidelines describe this as "match relevant social norms": in a professional analytics context, the right tone is closer to a trusted analyst presenting findings than a chatbot answering questions.

The explainability problem. Enterprise analytics AI faces a trust obstacle more severe than any of the previous cases: the users are professional decision-makers who will be held accountable for the decisions they make based on AI output. "The AI said so" is not an acceptable justification for a strategic decision. Explainability — the ability to show the reasoning behind an AI-generated insight — is not a nice-to-have in this context; it is a prerequisite for adoption.

Power BI's Explain the increase / decrease shows the pattern at its simplest. The analyst right-clicks a data point, and Power BI compares it with the previous point, works out which categories changed their share of the total most, and returns a chart plus a short description of which categories most influenced the change. Microsoft's own example: sales rise 55% quarter on quarter, but Computers and Home Appliances grew 63% while TV and Audio grew only 23%, so those categories are what the explanation calls out. The analyst can switch the chart between waterfall, scatter, 100% stacked column and ribbon views, and add it to the report as an ordinary visual. Microsoft describes the method only "simplistically" and says the ranking uses "various heuristics" — so this is not deep transparency. What it does give the analyst is enough surface area to judge whether the explanation is plausible and check it against their own domain knowledge.

NOTE
Plausible vs. Proven Explanations
Enterprise AI explanations usually show association, not causal proof. Power BI labels the output of its report Insights 'Possible Explanations', found by looking for movements in other dimensions that correlate with the anomaly — yet its Explain the increase documentation opens by saying you can 'discover the cause with just a few clicks'. Same vendor, two framings. Responsible product design keeps the UI language on the correlation side, so analysts do not treat correlations as conclusions.

Error handling at enterprise scale. The consequences of AI errors in enterprise analytics are business decisions made on wrong data. The graceful degradation pattern here has three components:

  1. Confidence levels made explicit. Showing not just the insight but how much of the data pattern supports it. Weak patterns should surface as tentative observations, not conclusions.
  2. Drill-down to the underlying data. Every AI-generated insight should link directly to the raw data that produced it. This converts "trust me" into "verify me."
  3. Acknowledgment of limitations. Tools that clearly document what data the AI had access to, what time period it analysed, and what variables it did not consider give users the context to evaluate whether the insight might be missing something critical.

Feedback loops. Enterprise analytics AI has access to a valuable but rarely used feedback signal: whether users act on the insight. If a sales team receives an AI-generated alert about customer churn risk and then takes no action, is that because the alert was wrong, or because they were unable to act? Distinguishing between "irrelevant insight" and "insight we could not act on" requires a lightweight feedback collection step that most tools do not include.

> You are designing the feedback mechanism for an enterprise analytics AI feature. What would you ask a user who dismissed an AI-generated insight without acting on it — in a way that takes less than five seconds to answer and still gives you signal about whether to improve the model or the user experience?

Personalisation in enterprise analytics. The most successful analytics AI tools learn what questions a specific analyst or team asks repeatedly and surface those dimensions more prominently. A sales operations analyst who always investigates pipeline coverage will benefit from an AI that leads with pipeline coverage anomalies. This is professional context personalisation — adapting not to personality or aesthetic preference but to job function and analytical style.

The privacy and governance tradeoff here is different from consumer personalisation. Enterprise users generally expect the tool to learn their role, and organisational data governance controls what data the AI can see in the first place. The personalisation questions are less about consent and more about role-based access and model transparency within the organisation.

Enterprise analytics AI design decision Trade-off
Proactive insight surfacing vs. query-driven only Proactive is more valuable but creates noise if recall is poor
Show contributing variables vs. show only the conclusion More transparency but more cognitive load; calibrate to user expertise
Team-level models vs. individual models Team models are more robust with less data; individual models are more precise but drift with personnel changes
Editable insight prioritisation High trust, low friction; requires UI investment most analytics tools defer
WARNING
The Autonomy Trap in Enterprise AI
Analytics AI that makes recommendations without surfacing the data behind them creates a dependency that degrades the quality of human judgment over time. Analysts who rely on AI to tell them what to look for gradually lose the skill of exploratory analysis. Lesson 9 on ethics covers the autonomy degradation risk; the design principle here is: always provide a path from the AI's conclusion back to the user's own data.

7 Cross-Cutting Principles: What Every Case Shares

Having examined five products across four categories, it is worth naming the patterns that appear regardless of product type. These are not universally obvious when reading about AI product design in the abstract — they become clear only when you look across cases.

Principle 1: Every successful AI feature has a clear problem statement that justifies the imperfection. Writing assistants are tolerated when they occasionally suggest the wrong word because the alternative — the blank page — is worse. Code completion is tolerated when it suggests wrong code because engineers review before accepting. Recommendation engines are tolerated when they miss because a curated selection from ten million items has always been better than unguided browsing. AI features that are introduced where a deterministic solution already works well get rejected — not because the AI is worse, but because imperfection is intolerable when perfection was previously available.

Principle 2: Trust is managed through visible boundaries, not confidence scores. Across all five cases, the dominant trust mechanism is visual or temporal separation between what the AI suggested and what the user accepted — ghost text, suggested labels, weekly playlist framing — not a numeric confidence percentage. Confidence scores require users to calibrate what 73% confidence means in context. Visual boundaries require only that users understand "this is provisional until I decide." Simpler trust mechanisms are more durable.

Principle 3: Error recovery must be easier than the original task. In every case, the product survives AI errors because correcting them costs less than doing the original task without AI. Accepting a writing suggestion and deleting it is one keypress. Moving a misprioritised email is two clicks. Dismissing a wrong code suggestion is one keystroke. If error recovery costs approach the cost of the task itself, users stop using the feature. Error cost is a design variable, not an accident.

Principle 4: Engagement metrics and outcome metrics diverge — and diverge more over time. Every case study revealed the same pattern: early teams measured engagement (suggestions shown, suggestions accepted, insights surfaced), later discovered that high engagement did not reliably predict user value, and invested in outcome metrics (retained code, better inbox response times, acted-on insights) only after the gap became obvious. Starting with outcome metrics is harder — they require more instrumentation and longer observation windows — but they are the only metrics that tell you whether the AI feature is actually doing its job.

Principle 5: The ambient-visible spectrum is a choice, not a default. Every product in this lesson made an explicit decision about how visible to make its AI. Gmail chose visible labels for sorting but invisible ranking; Spotify chose a named, weekly playlist rather than a continuously updating feed; Power BI presents Quick Insights as up to 32 separate cards, each a chart plus a short description, that the analyst can pin to a dashboard or run further insights on — and when its report Insights feature raises notifications unprompted, it stops showing them for the session if the user keeps dismissing them. The pattern from Google PAIR and the Microsoft Human-AI Interaction guidelines is consistent: make the AI's involvement visible enough to be trusted, invisible enough not to create cognitive overload. Neither extreme — full invisibility or full explanation — performs well.

NOTE
The Google PAIR Principle: Show Uncertainty Without Paralysing
The People + AI Guidebook from Google PAIR describes the challenge as making AI limitations legible without making the product feel unreliable. The best implementations in this lesson communicated uncertainty through design — ghost text, tentative language, drill-down availability — rather than through explicit probability notation. Design your uncertainty communication system before you design the happy path.

Principle 6: Feedback loops that require effort have low participation; feedback loops embedded in natural behaviour are abundant. Explicit thumbs-up and thumbs-down collection works in some contexts (chat interfaces, where the interaction is already conversational) and fails in others (mid-task inline suggestions, where rating the suggestion interrupts the work). The cases that have the richest feedback signals are those where user behaviour naturally encodes preference: accepting or rejecting a suggestion, moving or replying to an email, acting or not acting on an insight. Designing for abundant implicit feedback from the first release is more valuable than investing in polished explicit feedback UI.

A final cross-cutting observation: all five products share a common launch sequencing that the course's prototyping and measurement lessons predict. They launched with narrower AI scope than the product eventually achieved, measured carefully, and expanded AI autonomy only after building a track record. The temptation to launch ambitious AI features is real; the products that built lasting user trust started more conservatively.

Launch conservatively Expand carefully Signs it is ready to expand
Suggestions with explicit accept Lower friction acceptance Error correction rate below threshold
Visible AI labels Quieter background operation User calibration evident in behaviour
Single use case Multiple use cases Original use case healthy in outcome metrics
TIP
Your Own Expansion Trigger
Before launch, define the outcome metric threshold that will tell you the AI feature is ready to expand in scope or autonomy. Writing down 'we will expand when X reaches Y' before shipping protects the team from both premature expansion (driven by feature excitement) and excessive caution (driven by anecdotal reports of errors). Lesson 8's measurement framework has the tools to define this precisely.

Questions & Answers

Q: All five of these case studies are from large companies with engineering teams I could never match. What can a small team actually take from this?
The design decisions in each case study are separable from the engineering scale. Ghost text acceptance, visible category labels, named features with a clear premise, drill-down from insight to data — none of these require a large team to implement. What large teams have is more data for model training, not a monopoly on good design decisions. The principle that error recovery must be cheaper than the original task, for instance, is a design constraint that a two-person team can honour on day one. Start by taking the framework questions from Step 1 and applying them to your own feature before you have written a line of code.
Q: These products have been shipping for years and are still getting things wrong. How do I know when a design decision is genuinely good versus one I just haven't seen fail yet?
The test is whether the decision survives contact with a range of failure modes, not just the happy path. A genuinely good design decision holds under the subtle error, the catastrophic failure, and the edge case user. Ghost text, for example, works in writing assistants because it costs essentially nothing to dismiss, which means it holds even when the suggestion is completely wrong. If your design decision depends on the AI being right most of the time to feel good, it is fragile. Ask: what does this experience look like when the AI is badly wrong? If the answer is "users can recover cheaply and quickly," the decision is likely robust.
Q: How do I convince my organisation to invest in outcome metrics when engagement metrics are what leadership asks about in weekly reviews?
Show the divergence rather than arguing against engagement metrics. Pull two data series: engagement (suggestions accepted, features used) and an outcome proxy (retained completions, task time, repeat usage of AI-assisted content). If they move together, engagement is a reasonable proxy for now. If they diverge — high engagement, flat outcomes — you have the evidence the conversation needs. The argument is not "engagement metrics are wrong" but "here is a case where engagement is telling us something different from whether users are succeeding." Most leaders respond to data showing the metrics are decoupled. The measurement framework from Lesson 8 includes a worked example of this analysis.
Q: The enterprise analytics case study describes making AI reasoning "visible enough to be trusted" — but the explainability is still partial. Isn't partial transparency worse than no transparency, because users think they understand something they don't?
This is a real and live debate in AI product design, and the honest answer is that it depends on user sophistication and stakes. Partial transparency is dangerous when users overinterpret it — when a list of contributing variables is read as a causal proof rather than a correlation structure. The mitigation is language design: "these factors are associated with this outcome" is meaningfully different from "these factors caused this outcome," and the distinction matters most when a decision-maker is going to use the insight to justify a business decision. Calibrated language is not a full solution, but it reduces the overinterpretation risk significantly. The alternative — no explanation — has a worse failure mode: users who do not understand any reasoning develop uncalibrated trust that fails catastrophically when the model is wrong.
Q: Every case study here ends up recommending starting conservative and expanding slowly. But the market moves fast — won't cautious launches let competitors take the ground?
The cases where launching too ambitiously destroyed user trust are more numerous than the cases where being cautious ceded ground permanently. Trust, once lost to an AI feature that failed visibly, is expensive and slow to recover — user memory of a bad AI experience persists longer than memory of a slow feature release. The more accurate framing is not "conservative versus fast" but "conservative in AI autonomy, fast in everything else." You can ship quickly with limited AI scope — a writing assistant that suggests short completions rather than full paragraphs, a recommendation engine that covers one content type before five. Speed-to-market and scope-of-AI-autonomy are independent variables. Ship quickly; start the AI narrow; earn the expansion.

Key Takeaways

  1. Every good AI feature has a clear problem that justifies imperfection — users tolerate AI errors when the alternative is a genuinely worse experience, not when AI replaces something that already worked.
  2. Trust is built through visible boundaries, not confidence scores — ghost text, category labels, named weekly playlists, and drill-down availability all communicate provisionality more effectively than a percentage.
  3. Error recovery must cost less than the original task — if fixing an AI mistake is harder than doing the task without AI, the feature will fail regardless of how accurate the model is.
  4. Engagement metrics and outcome metrics diverge — define at least one outcome metric before launch that can only improve if users genuinely benefited, and track it from day one.
  5. The ambient-visible spectrum is a deliberate choice — make AI involvement visible enough to be trusted and invisible enough not to create cognitive overload; both extremes underperform.
  6. Launch with narrow AI scope and expand after earning a track record — the products with the deepest user trust started conservative, measured carefully, and expanded AI autonomy only after outcome metrics validated the expansion.

Next Steps: Back to all AI Product Design lessons