Designing for Errors

55 min intermediate Lesson 4

Learning Outcomes

  • Distinguish between the three failure modes — hallucination, degradation, and silent error — and describe the design response for each
  • Apply hedging language and visual uncertainty cues to AI-generated content so users calibrate trust appropriately
  • Design graceful degradation flows that keep users productive when the AI is unavailable or produces unusable output
  • Build correction-friendly interfaces that surface edit, regenerate, and alternative-suggestion affordances at the right moment
  • Map complete error-recovery flows — from the happy path through the subtle error to the catastrophic failure — for a real feature

Lesson Plan

Segment Duration Topic
Intro 3 min Why errors are a design problem, not just an engineering problem
Three failure modes 7 min Hallucination, graceful degradation, silent error
Hallucination UX 10 min Hedging language, visual cues, source attribution
Graceful degradation 8 min Fallback experiences and partial-output patterns
Correction flows 10 min Edit, regenerate, alternatives — surfaces and timing
Error recovery 10 min Getting back on track without starting over
Flow mapping exercise 5 min Happy path, subtle error, catastrophic failure
Wrap-up 2 min Key takeaways and next lesson

Before You Begin

Pre-work:

  • Complete Lesson 3: User Trust & Transparency — the confidence indicators and attribution patterns introduced there are the foundation for this lesson's hallucination UX section
  • Re-read the Lesson 2 pattern library (AI UX Patterns) and note which patterns carry the highest error risk — agents and ambient intelligence will come up repeatedly here
  • Think of one AI feature you have used recently where the output was wrong or unhelpful; bring that experience as a working example throughout the lesson

Shopping List:

  • A whiteboard or design tool (Figma, FigJam, Miro, or even a sheet of paper) for sketching the flow maps in the final exercise
  • The UX Patterns reference open alongside for quick callouts to the pattern library
  • No logins or accounts required beyond what you already use

1 Why Errors Are a Design Problem First

Ask an engineer how to handle AI errors and the conversation quickly turns to model confidence scores, retrieval accuracy, and fallback logic. All of that matters. But the moment a user sees wrong output, reads a confident fabrication, or watches a feature silently produce nothing, the experience is entirely a design problem. The model may have produced a reasonable token sequence; the product failed to communicate its uncertainty, offer a path forward, or catch the user before they acted on bad information.

The Google PAIR People+AI Guidebook frames this well: AI mistakes are inevitable, so design for them as a first-class use case — not a footnote handled with a generic error banner. The Microsoft Human-AI Interaction guidelines go further, describing the need to help users understand the AI's limitations throughout their experience, not just when something breaks.

That framing changes the design brief. Instead of "what error message do we show?", the question becomes:

  • How do we set expectations before the AI responds, so the error is less surprising?
  • How do we signal during the response that certainty is not guaranteed?
  • How do we make it easy to correct when the output is wrong?
  • How do we recover when the user has already acted on a bad output?

These four questions map to the four main topics of this lesson. But first, we need a shared vocabulary for the different ways AI goes wrong, because the design response is different for each.

NOTE
Key Insight
Every AI feature will produce wrong output eventually. The design question is not 'how do we prevent errors?' but 'how does the product behave so well around errors that users remain confident and productive despite them?'

The stakes are not uniform. An AI writing assistant producing a slightly awkward sentence is a minor inconvenience. An AI medical summary omitting a contraindication is a safety issue. An AI financial analysis inventing a figure that ends up in a board presentation is a reputational crisis. Error design must be proportionate to consequence — and your first task as a product designer is to decide, honestly, which category your feature sits in.

Error consequence Example feature Design priority
Cosmetic — easy to see, trivial to fix Style suggestions in a writing tool Low: show correction affordance, move on
Functional — changes what the user does Scheduling assistant booking the wrong time Medium: confirm before acting, easy undo
Consequential — output leaves the product Summary shared in a report or email High: confidence indicators, review prompts
Safety-critical — outcome affects health or safety Clinical decision support, legal advice Highest: hard human-review gates, disclaimers
WARNING
Stakes calibration
If your AI feature produces output that users will forward, publish, sign, or act on without re-reading — it is in the Consequential or Safety-critical row. Design accordingly. Most teams underestimate this until something goes wrong publicly.

2 The Three Failure Modes (and Their Design Signatures)

AI errors are not random noise. They fall into recognisable categories, each with a different design signature — the kind of thing that makes it worse or better from a user experience perspective.

Failure mode 1: Hallucination. The AI generates plausible-sounding content that is factually wrong, invented, or unsupported by the source material. The output looks correct; it reads fluently; it may be entirely fabricated. This is the highest-trust failure because users have no natural alarm signal — nothing looks broken. A research assistant confidently citing a journal article that does not exist; a chatbot giving a wrong product price with full confidence; a summarisation tool inventing a statistic that was not in the document.

Failure mode 2: Graceful degradation. The AI is unavailable, times out, or produces output so poor it cannot be used. Unlike hallucination, the product can detect this — the API returned nothing, the confidence score fell below threshold, the output failed a validation check. The design challenge is not signalling uncertainty in the output; it is deciding what the product does instead, and making sure the user can still accomplish their goal.

Failure mode 3: Silent error. The AI produces output that is wrong or unhelpful in a way that is difficult to detect — a subtly incorrect tone in a generated email, a recommendation that happens to be outdated, a categorisation that is systematically biased for a particular user group. These errors do not announce themselves. They accumulate. The design challenge is building correction and feedback surfaces that catch them over time, rather than in the moment.

NOTE
A useful test
Ask your team: 'If this feature produced wrong output, would the user know immediately, eventually, or never?' The answer determines your primary design investment: hallucination UX (immediately-visible errors), feedback loops (eventually-caught errors), or ethics and audit (never-noticed errors). Lesson 5 covers feedback in depth.

These three modes rarely appear in isolation. A scheduling agent might hallucinate the conference room booking policy (mode 1), then fall back to a manual booking flow (mode 2), while silently building an incorrect model of the user's preferred meeting times (mode 3). Good error design addresses all three.

Failure mode User experience of the failure Primary design response
Hallucination Output looks right but is wrong Hedging, attribution, confidence cues
Graceful degradation Output is absent or unusable Fallback path, partial output, status messaging
Silent error Output is wrong in ways hard to detect Correction surfaces, feedback loops, audit

The rest of this lesson works through each mode in depth, but the table above is worth committing to memory. When a new AI feature is on the table, run through all three rows and ask: does our design handle this?

TIP
Using this in reviews
In your next design review for an AI feature, call out which failure mode each error state addresses. If two of the three rows have no design response, the feature is not ready to ship — you are relying on the model never failing in those ways, which is not a product strategy.

3 Hallucination UX: Communicating Uncertainty Without Destroying Trust

Hallucination is the most dangerous failure mode because users cannot see it coming. The design goal is not to eliminate hallucination — you cannot control the model — but to design the interface so that users appropriately calibrate their trust in AI output, check what matters, and are not surprised when something is wrong.

There are four main tools in the hallucination UX toolkit: hedging language, visual uncertainty cues, source attribution, and human-review prompts. Used well, they reinforce each other without making the product feel anxious or unusable.

Hedging language is the most accessible tool. It means writing the interface copy — labels, response prefixes, empty states, and onboarding text — so that users understand the AI is generating, not confirming. The difference between "The answer is..." and "Based on the information provided, this appears to be..." is small in word count and significant in trust calibration.

A before-and-after comparison makes this concrete. Imagine an AI assistant in an HR product that answers policy questions.

Before (overconfident):

"Your notice period is 3 months."

After (appropriately hedged):

"Based on the standard employment contract template, the notice period is typically 3 months. Verify this against your specific contract before acting on it."

The second version is still useful — it answers the question — but it prompts the verification behaviour that matters for a consequential output.

> Examples of hedging language for an AI product interface:
> 
> "Based on [source], this appears to be..."
> "I wasn't able to confirm this — you may want to check..."
> "Here's a suggested answer. Review before sending."
> "This is a draft — some details may need verifying."
> "I'm not certain about this. Here's what I found..."
TIP
Calibrate hedging to stakes
Not every output needs a hedge. A spell-check suggestion does not need a disclaimer. A financial projection does. The mental test: if a user acted on this output without checking it, how bad could the outcome be? Match the hedge intensity to that answer.

Visual uncertainty cues communicate confidence without adding words to the output. The approaches that work in practice include:

Visual cue How it works Good for
Confidence band A percentage or colour-coded range on the output Data extraction, classification, scoring
Source chip A clickable attribution badge linking to the source material Research, summarisation, Q&A
"Verify this" badge An inline icon inviting review on specific claims Factual assertions, figures, dates
Italic or muted styling Visually distinguishes AI output from user-confirmed content Mixed-content documents and forms
Strikethrough + suggestion Shows original alongside proposed change Editing and rewriting tools

Source attribution is the single most trust-building pattern available when the AI is working from a document corpus or database. When a user can click through to the exact passage the AI cited, two things happen: they can verify the claim, and they viscerally understand the AI is drawing on real sources rather than confabulating. Perplexity and Bing Chat popularised this pattern; enterprise search and document intelligence products have followed.

Human-review prompts are appropriate when the output is consequential. These are interface elements — not error states — that appear as part of the normal flow: "Review before sending," a confirmation modal before an AI-drafted message is submitted, or a sidebar checklist of things to verify before publishing an AI-assisted report. The key is that they must feel like a natural part of the workflow, not a speed bump. If users find ways around them, the prompt is poorly designed, not the behaviour it is trying to encourage.

WARNING
The trust paradox
Over-hedging is as harmful as under-hedging. If every output comes with three disclaimers, users learn to ignore them — and then miss the one that mattered. Hedging language should be reserved for outputs where acting without checking carries real risk. Apply it selectively, or it stops working.

4 Graceful Degradation: What Happens When the AI Cannot Help

Graceful degradation is what happens when the AI component of your product is unavailable, returns unusable output, or falls below a confidence threshold that your system can detect. Unlike hallucination, this is a known failure state — and that means the product can respond purposefully rather than silently.

The worst response to degradation is a broken-looking interface: a spinner that never resolves, an empty results panel with no explanation, or a generic "something went wrong" banner that gives the user nowhere to go. The best response keeps the user productive — either through a fallback experience that uses non-AI means to accomplish the same goal, or through partial output that is clearly labelled as incomplete.

The degradation design matrix. Before writing fallback logic, map each AI-powered surface in your product to this matrix:

AI component fails User goal is still achievable without AI? Design response
Smart compose Yes — user can type manually Hide the suggestion, no message needed
Intelligent search ranking Yes — fall back to keyword ranking Silently degrade, optionally surface "Showing basic results"
AI summarisation Partial — user can read the original Show original with "Summary unavailable" message and read link
Autonomous action (e.g., booking) No — this was the whole feature Clear error state, manual path, status ETA if known

The key design insight is that degradation is a spectrum, not a binary. At one end, the AI component is a nice-to-have enhancement and its absence is invisible. At the other end, the AI component is the product, and its absence means the user has no path forward. Every point on that spectrum needs a different design response.

NOTE
The enhancement vs. core distinction
Products where AI is an enhancement (Grammarly's tone suggestions, Gmail's Smart Reply) can degrade silently — the product still works. Products where AI is the core experience (an AI scheduling assistant, a code review bot) need explicit fallback paths because there is no non-AI fallback. Know which one you are building.

Partial output patterns. Sometimes the AI starts a response but cannot complete it — a generation that times out mid-way, a summarisation that covers the first half of a document but not the second, an extraction that identifies three of five required fields. Partial output is often better than no output, but only if the product is honest about what is missing.

Design principles for partial output:

  • Make it visually obvious where the output ends and the gap begins
  • Label the gap explicitly: "Could not process the remaining sections"
  • Give the user a clear action: manual completion, retry, or contact support
  • Never allow partial output to be submitted or shared as if it were complete — build a confirmation or gate

Status messaging during degradation. If the AI is processing asynchronously and the user needs to wait, the interface should communicate progress honestly. "Analysing your document (usually under 30 seconds)" is better than a generic spinner. "This is taking longer than expected — we'll notify you when it's ready" is better than letting the user wonder if they should refresh. The Microsoft Human-AI Interaction guidelines call this "making it easy for users to understand the system's state" — a principle that applies especially when that state is 'not working as expected'.

TIP
Design the fallback path first
When scoping a new AI feature, ask the team: 'What does a user do if this produces nothing?' If the answer is 'they're stuck', build the fallback before the AI. You will ship it faster, and the AI enhancement layer will feel like a bonus rather than a crutch.

5 Correction Flows: Making It Easy to Fix What the AI Got Wrong

The moment a user spots an AI error, you have a design decision to make: how easy do you make it to fix it, and what do you offer them beyond a blank edit field? Correction flow design is one of the highest-leverage areas in AI product UX because it directly determines whether users keep trying or give up.

There are three correction mechanisms to design for, and most AI features need all three:

1. Inline editing. The simplest and most universal correction surface. The AI produces output and the user edits it directly, as they would any piece of text. The design questions are subtle: Is the output visually editable (does it look like a text field or a static label)? Does edit history work so users can undo? If the user edits, does the system treat the output as confirmed, or does it try to regenerate over the user's changes?

Common mistake: AI outputs that regenerate automatically when a parameter changes, overwriting the user's manual edits. This destroys the sense of control. Once a user has edited output, their version should be treated as authoritative unless they explicitly ask for a new generation.

2. Regenerate with guidance. The user wants a fresh attempt but wants to steer it differently, without starting from scratch. This is the "try again but..." pattern. The interface offers a way to provide corrective instruction — a prompt field, a set of adjustment sliders, or a list of refinement options — and regenerates with that guidance applied.

> Examples of regeneration guidance surfaces:
> 
> A text field: "Tell us what to change..." (open-ended)
> Preset buttons: [Shorter] [More formal] [Different approach]
> A slider: Tone: Casual ←→ Formal
> A checklist: "What was wrong?" [Too long] [Wrong tone] [Missing detail]

The checklist approach is especially valuable because it gives the product signal about why the output was wrong — implicit feedback that can inform model improvement over time. Lesson 5 covers this feedback loop in depth.

3. Alternative suggestions. Rather than regenerating one new output, the product offers multiple alternatives and lets the user choose. This pattern works well for short outputs where variance is valuable — subject line suggestions, reply options, image captions, product name ideas. GitHub Copilot uses this pattern for code completions; many writing tools use it for sentence rewrites.

The design principle is to show enough alternatives to surface genuine variety (typically three) without overwhelming the user with a wall of options. Each alternative should be meaningfully different — not three paraphrases of the same idea. If the model cannot generate genuinely distinct options, this pattern should not be used; three near-identical suggestions teaches users that the feature is not very powerful.

WARNING
The edit-vs-regenerate tension
Every time you add a regeneration affordance, you're implicitly telling users 'try asking again rather than fixing it yourself.' This creates a learned helplessness risk — users who never learn what makes a good output because they keep hitting regenerate. For learning-oriented or skill-building products, make editing the primary correction path and regeneration a secondary option.

Surface and timing. Correction affordances must appear at the right moment. Too early (before the user has read the output) and they are invisible. Too late (buried in a settings panel) and the user has already given up. The patterns that work:

Output type When to show correction Where to show it
Short text (subject line, title) Immediately, inline Below or alongside the output
Long-form draft After the user interacts (scrolls, selects) Floating toolbar or sidebar
Structured output (table, form) Per-cell or per-row on hover Inline in the cell/field
Autonomous action result In the action confirmation modal Before the action is confirmed
TIP
The 'ghost edit' pattern
Some products show editable AI output with a slightly different visual treatment — lighter background, italic text, a pencil icon — to signal 'this came from AI, you can change it.' This single visual cue can dramatically increase correction rates because users understand immediately that the output is their starting point, not their final answer.

6 Error Recovery: Getting Back on Track After a Failure

Correction flows handle the case where the user spots an error before acting on it. Error recovery handles the harder case: the user acted on bad AI output, the downstream consequences are now real, and the product needs to help them get back to a good state.

This is the domain most AI product teams under-design. The assumption is that because the error happened outside the product (the user sent the wrong email, booked the wrong meeting, submitted the wrong data), recovery is the user's problem. But products that earn long-term trust are those that take some responsibility for the consequences of their failures — even when those failures happened at the model level, not the application level.

Recovery has three phases: detection, triage, and repair.

Detection. The product needs a way to learn that something went wrong. For explicit failures (the API returned an error, the validation rule failed), detection is automatic. For semantic failures (the AI answered confidently but incorrectly), detection depends on user signals: they came back to edit, they clicked "this was unhelpful," they filed a support ticket, or they undid an action. Build these signals into your product as first-class events, not afterthoughts.

Triage. Not all errors need the same recovery path. A usability framework borrowed from incident response works well here:

Error severity Characteristics Recovery design
P3 — cosmetic Output quality low but goal still achieved Log, iterate, no user-facing recovery needed
P2 — functional Goal not achieved; user can retry Clear error message, retry affordance, support link
P1 — consequential User acted on wrong output outside the product Undo path if available; support escalation; proactive outreach
P0 — safety-critical Output caused or risked harm Immediate human review; suspension if systemic; regulatory notification if required

Repair. The repair path varies by error type, but the design principles are consistent:

  • Make the path to recovery as short as the path to the mistake. If a user can publish an AI-generated article in three clicks, they should be able to retract it in three clicks.
  • Where undo is not possible, offer the next-best thing: a template to communicate the correction, a log the user can export, or a human contact they can escalate to.
  • Acknowledge the failure without burying the user in apology. "We generated incorrect information — here's how to correct it" is better than three paragraphs of corporate hedging.
NOTE
Proactive recovery
The most trust-building recovery pattern is proactive: the product detects a likely error and surfaces it to the user before they discover it themselves. An AI scheduling tool that realises it double-booked a room should notify the user immediately, not wait for the conflict to cause an embarrassing moment in the meeting. If your product can detect confidence degradation after the fact — through feedback signals or quality checks — build proactive recovery flows even if they are rare.

Designing the full error arc. A complete error-recovery design covers three scenarios:

The happy path: The AI produces correct, useful output. The user acts on it with confidence. No friction, no doubt, efficient task completion. Design this experience first — it should feel effortless.

The subtle error: The AI produces output that is mostly right but has a factual mistake, a wrong tone, or a missing detail. The user catches it before acting on it (or just after). Correction flows come into play. The goal is that fixing the error takes less effort than the error caused — so the user's net experience is still positive.

The catastrophic failure: The AI produced output that was badly wrong, the user acted on it, and the consequences are real. Recovery flows come into play. The goal is to restore the user's confidence in the product — not just fix this instance, but demonstrate that the product learns from its failures and takes them seriously.

> A framework for auditing your error arc:
>
> Happy path: Is the output clearly useful and appropriately confident?
>             Does it invite correction without demanding it?
>
> Subtle error: Does the user have an obvious path to correct?
>              Is correction faster than deletion and retyping?
>              Does the correction provide signal back to the system?
>
> Catastrophic failure: Is there an undo path or next-best alternative?
>                       Is there a clear escalation route?
>                       Does the product acknowledge the failure honestly?
WARNING
The silent catastrophe
The failure mode most teams miss is the error that compounds silently — wrong output is acted on, the consequence is not noticed immediately, and the product never learns about it. Build monitoring for post-action user signals: did they come back to edit something they submitted? Did they contact support within 24 hours of an AI interaction? These are lagging indicators of silent catastrophes, and they are worth tracking even if you cannot act on every instance.

7 Putting It Together: The Three-Scenario Flow Map

The capstone exercise in this lesson is drawing the three-scenario flow map for an AI feature you are building or evaluating. This is a design artefact that makes error handling explicit, reviewable, and testable — rather than an implicit assumption that "the model will usually be right."

The map has three columns, one per scenario, and five rows:

Row Question
Trigger What causes the AI to respond? (User action, background event, scheduled task)
AI output What does the AI produce in this scenario? (Correct / subtly wrong / catastrophically wrong)
User signal How does the user interact with the output? (Accepts / corrects / reports)
Product response What does the interface do next? (Confirm / offer correction / recover)
System learning What signal does this interaction provide? (Positive, corrective, or failure signal)

A completed map for an AI email reply assistant might look like this:

Row Happy path Subtle error Catastrophic failure
Trigger User clicks "Suggest a reply" User clicks "Suggest a reply" User clicks "Send" on a draft they didn't review
AI output Correct, well-toned draft reply Grammatically fine but wrong recipient name Reply sent with confidential pricing in a public thread
User signal User edits lightly and sends User spots error before sending, corrects name User realises after sending, contacts support
Product response Log acceptance; surface "Accepted" signal Correction recorded; update user's name context Support notified; user offered template to recall the message; retraction prompt
System learning Positive signal; this template works Implicit correction signal; name context updated Failure event logged; incident review triggered
TIP
Run this in a design review
Print or share this map in your next design review for an AI feature and fill it in together with the team. The exercise reliably surfaces undesigned states — rows where the product response column is blank because 'the model won't do that.' Assume it will. Design the response.

When to build a lighter-weight version. If a full three-scenario map is too heavyweight for your current stage, start with a single question for each scenario:

For the happy path: does the interface help users act on correct output with confidence, without friction?

For the subtle error: if the user spots a mistake, what is the single next action the product offers?

For the catastrophic failure: if something goes badly wrong, is there a human path — a support route, an undo, a way to escalate — that is easy to find?

These three questions, answered honestly, will surface the majority of undesigned error states in most AI features.

NOTE
Connecting to Lesson 5
The 'system learning' row in your flow map is where this lesson's work connects directly to feedback loops. Every correction, regeneration request, and error report is a feedback signal. Lesson 5 covers how to design the collection, storage, and loop-back mechanisms that turn these signals into product improvements.

Questions & Answers

Q: Our AI vendor tells us the model's accuracy is above 95%. Doesn't that make hallucination UX unnecessary?
No, and this framing is one of the most common mistakes in AI product planning. A 95% accuracy rate means 1 in 20 outputs is wrong — and in a product used by thousands of people daily, that is thousands of errors per day. More importantly, accuracy rates are measured on benchmark datasets that may not match your users' actual queries. In production, accuracy degrades on edge cases, unusual requests, and inputs that differ from the training distribution. Hallucination UX is not a hedge against a bad model; it is recognition that every model fails, at some rate, on some inputs. Design for it regardless of the accuracy number.
Q: If we add too many disclaimers and hedges, won't users lose faith in the product entirely?
Yes — and that is exactly why this lesson distinguishes between calibrated hedging and over-hedging. The goal is not maximum disclaimers but appropriate trust calibration. Users should feel confident about outputs that are reliable and appropriately cautious about outputs that require verification. A spell-checker does not need a confidence disclaimer; an AI legal summary does. The test is: if a user acts on this output without checking it, what is the worst realistic outcome? Match the hedge intensity to that answer. Reserve strong hedges for consequential outputs; use minimal or no hedges for low-stakes features. If you find yourself hedging everything, the problem is probably that the AI is not ready for production on that surface, not that the hedge copy needs more work.
Q: We have an autonomous agent that takes actions on the user's behalf. How do you design error recovery when the action has already happened outside your product?
This is the hardest error design problem in AI products, and the answer has two parts. The first is to design for prevention: autonomous agents should confirm before taking irreversible actions, use progressive permission (start with actions the user can easily undo, earn the right to take harder-to-reverse actions over time), and maintain a detailed action log the user can audit. The second is to design the recovery path for the actions that do go wrong: what can the product do to help the user communicate the error to downstream systems, reach affected parties, or re-do the action correctly? The agent's action log becomes the recovery artefact. For high-stakes automations, always build a way for a human to step in — and make sure that path is visible and easy to reach before the first action is taken, not after the first failure.
Q: Our product team argues that error flows are engineering work, not design work. Who owns this?
Both, but the design team must lead the requirements. Engineering can implement a fallback state; only design can decide what the fallback state communicates to the user, what affordances it offers, and whether the experience maintains trust through the failure. In practice, error states are often under-designed because they are not on the happy-path user journey that gets most of the design attention. The solution is to include error scenarios explicitly in every design review and user research session — not as edge cases to handle later, but as first-class product experiences. The three-scenario flow map from Step 7 is a practical tool for making this a shared expectation between design and engineering from the start of a feature, not a retrofit at the end.
Q: How do we prioritise which error flows to design first when we have limited time?
Use the stakes table from Step 1 as your triage guide. Safety-critical and consequential features — any AI output that leaves your product or informs a real-world decision — must have complete error flows before launch. Functional errors (the AI fails to complete the task) need a fallback path at minimum. Cosmetic errors (output quality varies but the task still gets done) can ship with a correction affordance and iterate from there. The common mistake is prioritising based on likelihood rather than consequence. A one-in-a-thousand failure that causes a user to send wrong medical information matters more than a one-in-ten failure that produces a slightly awkward sentence. Design for consequence, then likelihood.

Key Takeaways

  1. Errors are a design problem first — The model will be wrong at some rate. The product decides whether users lose trust, stay productive, and can recover. Design all three outcomes explicitly, not as edge cases.
  2. Three failure modes, three design responses — Hallucination needs hedging and attribution; graceful degradation needs fallback paths; silent errors need correction surfaces and feedback loops. A complete AI feature design addresses all three.
  3. Calibrate hedging to consequences — Disclaimers and confidence cues should be proportionate to the stakes of acting on wrong output. Over-hedging trains users to ignore warnings; under-hedging lets them act on errors without realising it.
  4. Correction must be faster than the mistake — Edit, regenerate, and alternatives are the three correction mechanisms; each surface has a right moment and a right placement. Once a user has edited AI output, treat their version as authoritative.
  5. Recovery requires detection, triage, and repair — Design the undo or next-best path for consequential failures before they happen. Proactive detection — surfacing a likely error before the user finds it — is the highest-trust recovery pattern available.
  6. The three-scenario flow map makes error design explicit — Happy path, subtle error, catastrophic failure: map all three for every AI feature, with a product response and a system learning signal for each. Blank cells in the product response column are undesigned states waiting to fail in production.

Next Steps: Lesson 5: Feedback Loops