AI UX Patterns

50 min beginner Lesson 2

Learning Outcomes

  • Identify the four core AI UX patterns and describe what distinguishes each one
  • Evaluate which pattern fits a given product feature using a structured decision table
  • Recognise the failure modes that each pattern is most prone to, and name a design response for each
  • Analyse real-world AI products and map them to the correct pattern or pattern combination
  • Apply a pattern-selection framework to three hypothetical feature proposals and justify the choice

Lesson Plan

Segment Duration Topic
Intro 3 min Why patterns matter and how to use this library
Pattern 1 9 min Chat interfaces — conversational AI for complex, open-ended tasks
Pattern 2 9 min Inline suggestions — AI inside the user's existing flow
Pattern 3 9 min Autonomous agents — AI acting independently on the user's behalf
Pattern 4 8 min Ambient intelligence — AI working silently in the background
Choosing a pattern 7 min Decision framework and combination patterns
Wrap-up 5 min Key takeaways and applying the library to your own product

Before You Begin

Pre-work:

  • Complete Lesson 1: When to Use AI (and When Not To) — this lesson builds on the decision framework introduced there
  • Spend five minutes as a user of at least one AI-powered product (a writing assistant, an email tool, or a recommendation surface) and jot down what the interaction felt like from the user's side
  • Have the AI Product Design course landing page open in a second tab for reference

Shopping List:

  • A web browser — no installs or accounts required for this lesson
  • A blank document or notepad for the pattern-mapping exercise at the end of each step
  • Access to one product you work on or know well, so you can apply the frameworks to something real

1 Why Patterns Matter Before You Design

Before any wireframe or prompt engineering decision, you need to answer a deceptively simple question: how should the AI show up in this product? The answer shapes everything downstream — the interface, the trust signals, the error handling, and the success metrics.

A pattern is a named, reusable solution to a recurring design problem. In traditional UX, you recognise patterns like the hamburger menu, the infinite scroll, or the modal confirmation dialog. Each carries implicit user expectations: if you see a hamburger icon, you expect hidden navigation. Deviate from that expectation without a strong reason and you create confusion.

AI UX patterns work the same way. When a user sees a chat input in your product, they arrive with expectations drawn from every AI chat interface they have used before — a conversational back-and-forth, the ability to ask follow-ups, and probably some tolerance for imperfect answers. When they see a small "accept" button beside a suggested phrase, they expect a quick, low-stakes interaction they can dismiss in a keystroke. These expectations are your starting constraints.

The four patterns covered in this lesson — chat, inline suggestions, autonomous agents, and ambient intelligence — represent the design space that covers the vast majority of AI features shipping in products today. They differ across four key dimensions:

Dimension What it asks
User control How much does the user direct the AI's actions?
Visibility How visible is the AI's work to the user?
Interruption How much does the AI break the user's existing flow?
Stakes How costly is a mistake to fix?

You will use these four dimensions throughout this lesson as a quick evaluation tool. By the end, you will have a decision table you can apply to any feature proposal.

NOTE
Where These Patterns Come From
The Google PAIR People+AI Guidebook and the Microsoft Human-AI Interaction guidelines (Amershi et al.) both converge on these interaction modes after studying dozens of shipped AI products. They were not invented theoretically — they were discovered by observing how users actually interact with AI in practice.
TIP
One Pattern Per Surface, Not Per Product
A single product can use multiple patterns at once. Gmail uses ambient intelligence (spam filtering), inline suggestions (Smart Compose), and a quasi-chat pattern (Gemini integration). The decision you're making is which pattern fits this specific surface, not the whole product.

2 Pattern 1: Chat Interfaces

Definition: The user types a natural-language request and receives a natural-language response. Interaction is conversational — the user can refine, follow up, push back, and redirect within a single session.

Chat is the pattern most people picture when they think "AI product." It became the dominant mental model when large language models went mainstream, but it is not the right choice for most features. Understanding when it earns its place — and when it wastes it — is a foundational product design skill.

When chat works

Chat is the right pattern when the task is genuinely open-ended and benefits from dialogue to clarify intent. Legal research, complex project planning, exploratory data analysis, and anything involving ambiguous requirements are natural fits. The conversational loop lets the AI ask clarifying questions and the user course-correct without losing context.

It also works well when the user's goal is not fully formed. A user who opens a chat interface expecting to "figure out a marketing strategy" benefits from a system that asks questions, proposes framings, and builds understanding incrementally. A user who knows exactly what they want — "resize this image to 800px wide" — does not.

> Example prompt a user might type in a well-designed chat interface:
> "I need to write a user research plan for a new checkout redesign.
> I've never written one before. Where should I start?"

The multi-turn nature of chat is its strength: the AI can ask "How many users do you plan to interview?" or "Is this primarily a usability or a conversion study?" and get progressively more useful.

When chat fails

Chat fails when it is used as a lazy default for tasks that have a well-defined answer. If a user asks for the weather, a one-question-one-answer interaction is not meaningfully conversational — it is just a slow search box. Wrapping it in a chat interface adds latency and cognitive overhead without adding value.

It also fails when the task requires precision inputs the user cannot reliably express in prose. Scheduling a meeting with three constraints is faster through a structured form than through a conversation. Tasks with exact, enumerable parameters are poor chat candidates.

Chat works well Chat fails or feels excessive
Exploratory, open-ended tasks Precise, bounded tasks with known parameters
Tasks that benefit from back-and-forth clarification One-shot lookups (weather, definitions, status checks)
Tasks where the user's goal may shift mid-session High-frequency, repetitive actions (apply filter, move file)
Synthesising across multiple sources or topics Tasks better served by a form, toggle, or dropdown

What users expect from chat

Users expect chat interfaces to maintain context within a session. If they mentioned a constraint three messages ago, the AI should still honour it. They expect the ability to course-correct without starting over, and they expect the system to be honest when it does not know something rather than guessing confidently.

The most damaging chat failure mode is overconfident wrongness — the AI answers a question it cannot reliably answer and does so in the same fluent tone as answers it can. Users who have not yet calibrated their trust may not notice. We cover the design responses to this in Lesson 3: User Trust and Transparency.

WARNING
The Chat Tax
Every chat interaction demands the user to translate their goal into a sentence. This is a hidden cost — especially for users who are not confident writers, whose first language differs from the interface language, or who are working under time pressure. Before choosing chat, ask whether the user's need could be met faster with a well-structured UI that requires no typing at all.
TIP
First Message Sets the Frame
The placeholder text in a chat input is a design decision, not filler. 'Ask me anything' trains users to ask trivially narrow questions. 'Describe the outcome you're working toward' trains them to give useful context. Seed the right behaviour from the first interaction.

3 Pattern 2: Inline Suggestions

Definition: The AI makes contextual suggestions within the user's existing workflow — text completions, code proposals, smart replies, or alternative phrasings — which the user can accept, modify, or dismiss without leaving their current context.

This is AI's least disruptive form. The user stays in their primary task; the AI surfaces assistance as an optional offer. Gmail's Smart Compose, GitHub Copilot's code completions, and Grammarly's rewrite suggestions are canonical examples. The interaction is frictionless when done well — a user finishing a sentence accepts a completion with a single Tab key press and never loses flow.

The accept-modify-dismiss triad

Inline suggestion design lives or dies by how gracefully it handles three user responses:

  1. Accept — the suggestion is right or close enough. The interaction should take one gesture: Tab, right arrow, or a click. Any more friction and users stop accepting even when the suggestion is correct.
  2. Modify — the suggestion is partly right. The user needs to tweak a word or phrase without the AI fighting them. A good design accepts partial adoption: the user can accept the first three words and delete the rest.
  3. Dismiss — the suggestion is wrong or distracting. The dismiss interaction must be invisible-when-not-needed and instantly available when it is. A ghost-text suggestion that cannot be dismissed without an extra keystroke will train users to ignore the feature entirely.
> Example: a well-designed inline suggestion in a document editor
>
> User types: "The primary risk of this approach is"
> AI suggests (shown as greyed ghost text): "increased implementation
> complexity during the migration phase, which may affect timeline."
> User presses Tab to accept, or Escape to dismiss, or ignores and
> keeps typing — in which case the ghost text silently vanishes.

Calibration: suggestion frequency and quality

The biggest tuning problem with inline suggestions is the precision-recall tradeoff. A system that suggests constantly (high recall) creates noise and trains users to tune out even when suggestions are good. A system that only suggests when confident (high precision) feels like it is barely there — users forget it exists.

The right calibration depends on context. Code completion tools like Copilot lean toward high frequency because the cost of a bad suggestion is low — the programmer evaluates and dismisses in under a second. A financial report editor might lean toward high precision — an unexpected suggestion mid-sentence is disruptive when the user is concentrating on numbers.

Calibration Works well when Fails when
High frequency Fast-moving, low-stakes tasks (chat, code) Tasks requiring concentration; suggestions become interruptions
High precision High-focus, high-stakes tasks Suggestions feel absent; users stop noticing the feature
Adaptive Works for most contexts — but harder to build Poorly tuned adaptation creates unpredictable experience

Interaction cost and flow

The Google PAIR guideline on inline suggestions is worth internalising: the user should never feel that the AI is competing with them for control of their document. Suggestions must feel subservient — available when wanted, invisible when not. A suggestion that demands attention, hovers too long, or requires an action to dismiss is no longer assisting; it is interrupting.

NOTE
Ghost Text vs. Tooltip vs. Side Panel
The visual treatment of an inline suggestion carries strong expectations. Ghost text (greyed completion in-line) signals 'accept or ignore.' A tooltip signals 'hover to decide.' A side panel signals 'this is an alternative, not a continuation.' Mixing these signals in the same interface creates confusion about what accepting a suggestion will do.
TIP
Measure Acceptance Rate as a Health Metric — Carefully
A high suggestion acceptance rate can mean the AI is excellent — or that users have given up rejecting suggestions and are accepting them uncritically. The metric you want alongside acceptance rate is post-acceptance edit rate: how often do users immediately modify what they just accepted? If that number is high, the suggestions are nearly-but-not-quite right, which is more frustrating than wrong-and-obvious. Lesson 8 covers measurement in depth.

4 Pattern 3: Autonomous Agents

Definition: The AI acts independently to complete a multi-step task on the user's behalf — browsing, writing, scheduling, executing — and checks in periodically or when it encounters a decision requiring human input.

Autonomous agents are the most powerful and the most dangerous pattern in this library. A chat interface generates text the user then acts on. An autonomous agent takes the actions itself. This distinction has profound implications for trust, error recovery, and the design of checkpoints.

The delegation contract

When a user delegates a task to an autonomous agent, they enter into an implicit contract: I trust you to act on my behalf, within bounds I understand, and I expect you to stop when you hit something that requires my judgement. Every design decision in an agentic interface is about making that contract legible.

The bounds question is critical. The user saying "book me a flight to Edinburgh on Thursday" is implicitly authorising the agent to search, filter, and select — but are they authorising it to enter their credit card details? To book the hotel too? To email their colleagues with the travel plan? Each step that was not explicitly authorised is a potential trust violation, even if the agent got it right.

> Example: how a well-designed agent communicates its intended plan before acting
>
> User: "Book me a flight to Edinburgh on Thursday under £200."
>
> Agent (before taking action): "I found three options. My plan:
> 1. Select the 07:45 EasyJet flight (£149, returns Friday evening).
> 2. Add it to your calendar.
> I will NOT enter payment details — I'll hand that step to you.
> Proceed, or should I show you the alternatives first?"

This "plan before action" pattern — borrowed from the software engineering concept of a dry run — gives the user a chance to review scope before the agent commits to anything.

Interruption design: when to pause, when to proceed

The hardest design question in agentic interfaces is where to put the checkpoints. Too few checkpoints and the agent does things the user would not have sanctioned. Too many and the agent is just a slow UI with extra steps — the user would have been faster doing the task themselves.

A useful heuristic comes from the Microsoft Human-AI Interaction guidelines: the cost and reversibility of an action should govern how much autonomy the agent exercises. Reading and summarising is low cost and fully reversible — no checkpoint needed. Sending an email is moderate cost and barely reversible — a confirmation makes sense. Charging a credit card or deleting a file is high cost and irreversible — a checkpoint is mandatory.

Action type Examples Checkpoint approach
Read-only, reversible Search, summarise, compare Proceed without pause
Write, reversible Draft an email, create a document Show output, await approval
Write, not immediately reversible Send an email, post to a channel Explicit confirmation required
Irreversible or high-stakes Delete, charge, publish publicly Hard pause; no ambiguity allowed

When autonomous agents fail

The failure modes unique to agents are more severe than those of the other patterns because actions compound. A chat interface that misunderstands your question produces a wrong answer — you rephrase and try again. An agent that misunderstands your goal two steps in may have sent emails, booked calendar slots, and modified files before you notice. The error surface grows with each autonomous action taken.

This is why observability — the user's ability to see what the agent has done and is about to do — is the most important design property of an agentic system. A clear, readable action log ("Here's what I've done so far"), a pause-and-review mechanism, and a meaningful undo where possible are not nice-to-haves. They are the design foundation on which trust in agentic AI is built.

WARNING
Do Not Ship Agents Without Guardrails
The pattern is technically exciting and can feel like magic in demos. In production, agents operating beyond their intended scope — even in small ways — are among the most trust-damaging AI failures a product can have. Define explicit scope boundaries before building, not after the first incident.
NOTE
The 'Human-in-the-Loop' Spectrum
Fully autonomous (agent acts, reports afterwards) and fully manual (agent suggests, human confirms every step) are the poles. Most well-designed agents sit somewhere in the middle: autonomous for low-stakes actions, pausing for high-stakes ones. Where you position the product on that spectrum is a product decision, not just a technical one — and it should be made explicitly, not by default.

5 Pattern 4: Ambient Intelligence

Definition: The AI works in the background without the user actively directing it — filtering, ranking, detecting anomalies, personalising surfaces, flagging risks — and surfaces only what requires attention. The user experiences the results without necessarily seeing the AI's work.

Ambient intelligence is the oldest AI UX pattern and the most invisible. Spam filters have been ambient AI for two decades. Content recommendation systems, fraud detection, predictive search ranking, and automatic photo tagging are ambient. The user did not ask for any specific output; the AI is continuously running and occasionally surfacing a result.

The three ambient modes

Ambient intelligence manifests in three distinct interaction modes:

Silent filtering — the AI removes or suppresses content without the user explicitly requesting it. Spam filtering is the canonical example. The user experiences the absence of unwanted content, not the presence of a feature. When working correctly, the user barely knows it exists. When failing (false positives — legitimate emails marked as spam), the user is damaged without knowing why.

Ranked and personalised surfaces — the AI determines the order or composition of what the user sees: a news feed, a search results page, a "recommended for you" list. The AI is making editorial decisions without the user's moment-to-moment instruction. The user experiences the ordering as natural; the influence of the AI is invisible unless they look for it.

Anomaly and risk alerts — the AI monitors a stream of data and surfaces anything that matches a pattern of concern: unusual login activity, an unexpected financial transaction, a metric going outside normal range. Here the AI breaks its silence specifically to request the user's attention.

Visibility and the ambient trust problem

Because ambient AI is invisible by design, its failure modes are also invisible — which makes them uniquely dangerous for trust. When a spam filter incorrectly suppresses a business-critical email, the user may not discover the failure for days, if ever. When a recommendation algorithm creates a filter bubble, users may not realise they are seeing a curated slice of reality rather than a full view.

The visibility question — how much should you show users about what the ambient AI is doing — has no single right answer. The Google PAIR guideline is instructive: tell users when AI is involved in decisions that affect them, even if you do not show them the full workings. A small "personalised for you" label on a recommendations surface, or a "why is this in my spam folder?" link, costs almost nothing and substantially raises user awareness and trust calibration.

> Example: a well-designed ambient AI disclosure in a product
>
> Instead of: a recommendations section with no context
>
> Better design: "Recommended based on your recent searches and the
> topics your team reads most often. [Adjust preferences]"
>
> The link does not need to expose the algorithm. It signals:
> "The AI made this choice, and you have some influence over it."

Ambient intelligence and the loss-of-agency risk

The most subtle failure of ambient AI is the slow erosion of user agency. When users never see the filtering, ranking, or suppression happening, they lose the ability to notice what they are not seeing. A product manager who relies on an AI-curated analytics dashboard may stop checking the full dataset because the AI has always highlighted the right things — until the day it does not.

Ambient mode Benefit to user Trust risk if it fails silently
Silent filtering Clears noise; nothing to manage False positives cause invisible losses
Ranked surfaces Surfaces relevance without manual effort Filter bubbles; user sees a skewed slice
Anomaly alerts Catches problems before the user notices False negatives; user over-relies and stops watching
WARNING
Invisible AI, Invisible Failure
The same invisibility that makes ambient AI feel effortless makes its failures hard to catch and easy to misattribute. Design for how users will discover that something went wrong — not just how the AI will behave when things go right.
TIP
Progressive Disclosure for Ambient Decisions
You do not have to expose the full algorithm. A 'why did I see this?' explanation link, shown on hover, satisfies most users' curiosity and dramatically improves trust without requiring you to reveal anything proprietary about the model. Surface the intent, not the mechanics.

6 Choosing the Right Pattern: A Decision Framework

With all four patterns named, the practical question is: given a feature proposal, how do you select the right one? The answer depends on four variables you can evaluate for any feature in a few minutes.

The four evaluation variables

1. Task structure. Is the task open-ended and exploratory, or bounded with known inputs and outputs? Open-ended tasks favour chat. Bounded tasks with a primary action favour inline suggestions or automation.

2. User control appetite. How much does this user need to feel in control of the output? Users who are accountable for the result (a manager approving a report, a doctor reviewing a diagnosis) need high control — chat or high-confirmation agents. Users doing a low-stakes, high-frequency task (filtering search results) accept low control — ambient or frictionless inline.

3. Interruption tolerance. Is this a deep-focus task or a lightweight task? Code completions are designed for flow states; a suggestion that forces the programmer to stop and read is a failure, not a feature. A complex project plan is a context where an agent asking "before I send these six emails, shall I confirm the dates with you?" is valued, not annoying.

4. Error cost. If the AI is wrong, how bad is the consequence? An incorrect autocomplete adds one extra keystroke to fix. An agent that sent the wrong version of a proposal to fifty clients is a crisis. Error cost drives how much human confirmation you build into the design.

> Pattern selection exercise prompt (use this with your team):
>
> "For the feature we're discussing:
> 1. Is the user's goal clear before they start, or does it emerge
>    through interaction? (bounded → inline/ambient; exploratory → chat)
> 2. How often does this task repeat for a given user?
>    (daily/hourly → ambient or inline; occasional → chat or agent)
> 3. What is the worst realistic consequence of an AI error here?
>    (trivial → low confirmation; serious → high confirmation/agent)
> 4. Does the user need to see how the AI reached its output?
>    (yes → chat or agent with log; no → ambient is fine)"

Pattern decision table

Feature characteristic Best-fit pattern Second choice
Open-ended, exploratory, user-led Chat Agent with high confirmation
Fast, within-flow, low stakes Inline suggestions Ambient (if fully silent)
Multi-step, complex, delegatable Autonomous agent Chat with action buttons
Continuous, background, low interruption Ambient intelligence Inline (if surfaced reactively)
Mix of exploration + action Chat + agent hybrid Agent with preview step

Combination patterns in practice

Real products often layer patterns on the same surface. A writing tool might use inline suggestions while the user types (Pattern 2), offer a "Rewrite this section" button that opens a chat pane (Pattern 1), and silently flag potential compliance issues in the background (Pattern 4). The combinations work when each pattern retains its natural interaction vocabulary — users understand each mode because it behaves consistently with their expectations of that pattern.

The failure mode in combination patterns is collision: two patterns fighting for the user's attention at the same moment. A chat response appearing at the same time as a ghost-text inline suggestion appearing at the same time as an ambient alert is not rich AI assistance — it is noise. Design the hierarchy: which pattern has priority on this surface at this moment?

NOTE
When Your Pattern Choice Is a Competitive Decision
The pattern you choose is partly a product strategy decision. Building a chat interface where competitors have inline suggestions signals a different product philosophy — you are betting that users want more control and context, not faster friction-free help. Neither is inherently right; the question is whether your choice matches the task and the user.
TIP
Name Your Pattern in the Brief
When scoping an AI feature, include the pattern name in the product brief — 'this feature uses the inline suggestion pattern.' It instantly aligns engineers, designers, and stakeholders on the interaction model before any mockups are created, and prevents the implicit assumption that every AI feature is a chat box.

7 Applying the Pattern Library: Three Feature Proposals

The pattern library is only useful if you can apply it under real product pressure. This step walks through three hypothetical feature proposals and applies the decision framework from Step 6. Use these as a template for running the same exercise with your own team.

Proposal A: AI-assisted performance review drafts

Context: A mid-size company's HR platform wants to help managers write annual performance reviews. Managers rate employees on five dimensions and then write a two-paragraph narrative summary. This narrative takes most managers thirty to forty minutes and is frequently identified as a pain point.

Evaluation:

  • Task structure: bounded inputs (five ratings, a name, a role) but open-ended prose output
  • User control appetite: high — managers own the review; HR and legal may audit it
  • Interruption tolerance: moderate — this is a low-frequency, high-stakes task
  • Error cost: significant — a poorly worded review can affect someone's career

Pattern verdict: This is a chat + document hybrid case. The bounded inputs suggest a form-to-draft flow (not pure chat), but the manager will want to refine the output in a conversational way ("make the tone warmer," "add something about the project last quarter"). The right design is: structured form pre-fills context for the AI, AI generates a draft, manager refines via a lightweight chat pane alongside the draft. Inline suggestions within the text editor could layer on for smaller edits.

Not: Ambient (requires user direction), or fully autonomous (manager must own the output).


Proposal B: Smart scheduling for a project management tool

Context: A project management tool wants to automatically rebalance task assignments when a team member unexpectedly takes leave, shifting tasks to available team members based on skills and current load.

Evaluation:

  • Task structure: well-defined problem with enumerable constraints (skills, availability, priority)
  • User control appetite: moderate to high — project managers are accountable for deadlines
  • Interruption tolerance: low — this is an urgent, time-sensitive situation
  • Error cost: high — wrong assignments can cause missed deliverables

Pattern verdict: This is an autonomous agent with a mandatory confirmation checkpoint. The task has clear parameters the AI can evaluate faster than a human. But the stakes are high enough that the agent should present a rebalancing plan ("Here's what I propose: 3 tasks moved, here's why") before applying it. The confirmation step is not optional — it is the design feature that makes the speed of the agent trustworthy.

Not: Chat (user wants a solution, not a conversation), or ambient (the user needs to see and approve the plan).


Proposal C: Content moderation for a community platform

Context: A consumer community platform wants to reduce harmful content. They moderate about 50,000 posts per day; human review of everything is impossible. They want to use AI to flag or remove content that violates their community guidelines.

Evaluation:

  • Task structure: well-defined classification problem (violates guidelines / does not violate guidelines)
  • User control appetite: low from the user posting; high from trust and safety teams monitoring
  • Interruption tolerance: none for legitimate users; the moderation must be invisible
  • Error cost: high in both directions — false positives silence legitimate voices; false negatives let harm propagate

Pattern verdict: This is ambient intelligence with a human review tier. The AI works silently for the vast majority of content. It should remove only the clearest violations automatically; flag borderline cases for human review; and never be the final word on account-level actions. The appeal mechanism — "I believe my content was incorrectly removed" — is not an edge case feature, it is the trust foundation of the entire system.

Not: Chat (users should not have to justify their posts in a conversation), or inline suggestions (there is no user-facing workflow to augment).

WARNING
These Are Product Decisions, Not Just Design Decisions
Notice that in Proposal B and C, the pattern choice carries direct product liability implications — the checkpoint in B and the appeal mechanism in C are not design flourishes. They are the features that determine whether the product is responsible when things go wrong. Involve legal, compliance, and trust and safety stakeholders in pattern selection for high-stakes features.
TIP
Run This Exercise in a Product Jam
The four evaluation questions (task structure, user control, interruption tolerance, error cost) take under ten minutes to answer as a team. Running this exercise before any wireframing aligns the room on interaction model and surfaces disagreements early, when they are cheap to resolve.

Questions & Answers

Q: Our stakeholders are asking for a chat interface because ChatGPT is popular. How do I make the case for a different pattern?
Show them the decision table from Step 6. Ask them to walk through the four evaluation variables for the specific feature. If the task is bounded, high-frequency, and requires staying in the user's flow, an inline suggestion will perform better on every measurable dimension — faster time-to-value, higher adoption, lower error cost. Use the case studies in the pattern library: Gmail's Smart Compose outperforms a chat interface for completing emails because users should not have to leave their email to get writing help. Chat is the right choice when it is the right choice — not because it is visible and impressive in a demo.
Q: We built an autonomous agent and users aren't trusting it enough to let it act. They confirm every single step even the trivial ones. How do we fix that?
This is a calibration problem — users have not yet built a mental model of which actions are safe to delegate. Three design responses help: first, make the agent's track record visible ("this agent has completed 47 tasks for you with no issues") so users can update their trust on evidence, not faith. Second, differentiate the visual weight of checkpoints — trivial confirmations should be lighter-touch (a small banner, not a modal) than high-stakes ones, which signals that not all checkpoints are equally urgent. Third, offer an "always approve this type of action" option so users can progressively extend trust as they gain confidence, rather than making a single all-or-nothing autonomy decision on first use.
Q: How do I tell if our ambient AI is quietly failing users in ways we can't see in the aggregate metrics?
Aggregate metrics are the problem — they average over the users who are not affected by failures and dilute the signal from those who are. Three practices help: first, actively sample the suppressed or filtered content on a regular basis (what would users have seen if the AI had not intervened?) and review it with a human moderator. Second, instrument a "why did the AI do this?" path even if you do not surface it to all users, so your team can audit individual decisions. Third, run periodic user interviews specifically recruiting users whose behaviour suggests under-engagement — they may have stopped noticing the feature because it stopped serving them. Ambient failures are discovered by looking, not by watching dashboards.
Q: Our product uses three of the four patterns. Users seem confused about what the AI can and can't do. What's the root cause?
Confusion in multi-pattern products almost always comes from inconsistent interaction vocabulary. Each pattern carries its own implied contract with the user. When ghost-text inline suggestions look similar to agent-generated content, or when an ambient alert uses the same visual treatment as a chat response, users cannot build a reliable mental model of which AI is doing what, with what level of autonomy, and with what consequences for accepting. Audit your visual language: each pattern should have a distinct and consistent visual and interaction treatment so users always know which mode they are in. Document these distinctions in your design system so they do not drift as the product grows.
Q: We're at the earliest stage — just deciding whether to use AI at all. Is it worth learning these patterns before we've validated the concept?
Yes, and the order actually matters. Knowing which pattern you are building changes the prototype you make and the questions you test. A Wizard of Oz test for an inline suggestion is a different exercise from one for an autonomous agent — the tasks you give users, the failure scenarios you probe, and the trust signals you observe are all different. Picking the pattern first means your validation work is specific enough to give you actionable signal. If you test a vague "AI-assisted" prototype without committing to a pattern, you learn that users like the idea of AI help — which you probably already knew. Lesson 7 covers prototyping approaches for each pattern in detail.

Key Takeaways

  1. Four patterns cover most AI UX — chat, inline suggestions, autonomous agents, and ambient intelligence map to the vast majority of AI features in production. Start every feature design by naming the pattern.
  2. Patterns carry user expectations — each pattern implies a different user contract: chat implies dialogue and control, inline implies non-intrusive assistance, agents imply delegation within known bounds, ambient implies invisible work that surfaces only what matters.
  3. Four variables drive pattern selection — task structure, user control appetite, interruption tolerance, and error cost will point you to the right pattern before you open a design tool.
  4. Autonomous agents require explicit guardrails — checkpoints are not optional safety theatre; they are the design feature that makes speed trustworthy when the stakes are real.
  5. Ambient AI failures are invisible by design — build audit and discovery mechanisms into ambient features from day one, not after the first incident.
  6. Name the pattern in the product brief — a single sentence in a brief aligns designers, engineers, and stakeholders on the interaction model before any work begins, and prevents the default assumption that every AI feature is a chat box.

Next Steps: Lesson 3: User Trust & Transparency