Troubleshooting
Six recurring problems in production AI features. Each entry lists causes and fixes — no code required. See also: UX Patterns, Rules, Cheat Sheet, Glossary.
Problem 1: Users Do Not Trust the Feature
Symptom: Low adoption despite good usability scores. Users call the feature "unreliable," try it once, see a surprising output, and revert permanently.
| Cause | Fix |
|---|---|
| Every output looks equally certain — no confidence signal | Use calibrated language: "Here is one interpretation" signals latitude; reserve authoritative phrasing for high-confidence outputs. |
| No provenance — users cannot verify where output came from | Show the source trail: "Summarised from the document you uploaded." Visible input anchors trust. |
| Unpredictable scope | Add a one-sentence capability claim before first use: "This summarises themes, not verbatim responses." |
| Feature launched before quality gates were met | Treat high first-session abandon as a quality problem. Gate rollout behind a threshold on representative inputs. |
Problem 2: Users Over-Trust and Miss Errors
Symptom: Users accept outputs without review. Errors reach downstream work and are discovered only after consequences.
| Cause | Fix |
|---|---|
| Polished presentation signals the output is finished | Use an "AI draft" label or lighter background to signal the output is a starting point. |
| High-stakes actions require only one click | Add friction proportional to stakes. A contract summary should require opening the output before forwarding. |
| No guidance on where errors cluster | Mark likely error zones: dates, numbers, named entities, real-time claims. Focus hasty reviewers where it matters. |
| Automation bias in high-load contexts | Require explicit human approval before consequential actions execute. Frame the UI as "review and send," not "approve." |
Problem 3: Low Feedback Participation
Symptom: Thumbs-up/thumbs-down buttons exist but fewer than 2% of users engage. High usage, almost no improvement signal.
| Cause | Fix |
|---|---|
| Feedback feels disconnected from outcomes | Close the loop visibly: "Thanks — we'll use this to improve suggestions." Better: show aggregate impact ("Your corrections reduced errors here by 30%"). |
| Too much friction | Use micro-interactions: inline rating that fades if ignored, no confirmation required. |
| Prompt appears after the task ends | Embed feedback in the edit flow. When a user edits an AI output, surface a prompt at that moment — peak motivation. |
| Users do not believe feedback changes anything | Make improvement visible: a "what's improved" note referencing user signals. |
Implicit signal: Instrument acceptance, edits, and deletions. Edit distance reveals what was wrong — implicit data supplements sparse explicit ratings.
Problem 4: Personalisation Feels Creepy
Symptom: Users describe the feature as "watching me" or "knowing too much." Some disable it. The personalisation is accurate but emotionally wrong.
| Cause | Fix |
|---|---|
| Accurate inference reveals deep data collection | Prefer explicit preference collection ("What topics do you want more of?") over pure inference. Asked-for personalisation feels helpful; inferred-only feels like surveillance. |
| No "why am I seeing this?" affordance | Add a one-line explanation: "Based on your recent activity in the marketing workspace." One sentence removes the "how did it know?" feeling. |
| Personalisation reveals sensitive inferences | Never surface outputs revealing inferences about health, finances, or political views — even if accurate. |
| No user control over the inferred profile | Provide a profile control panel: show inferences, allow corrections and deletion. Reduces creepiness and improves quality. |
Progressive approach: Start generic; personalise gradually. Highly tailored outputs on the first visit feel like a boundary was crossed. Time-paced personalisation is less uncanny. (PAIR Guidebook: "respecting mental models of data collection.")
Problem 5: The Feature Is Used But Not Valued
Symptom: Usage is healthy but NPS is flat. Users call the feature "fine" and would not notice if it were removed.
| Cause | Fix |
|---|---|
| Feature solves a task users did not find painful | Run task-completion research: compare AI-assisted vs. unassisted. If similar, no UX polish fixes a missing value proposition. |
| Outputs are acceptable but never impressive | Identify the "wow moment" — output that exceeds expectation — and optimise to reach it faster. Time-to-first-wow predicts retention. |
| Feature is buried; not discoverable at the moment of need | Surface features in context, at the problem moment. An inline assistant beats one requiring separate navigation. |
| Users do not know the full capability | Expose depth progressively: "Did you know you can ask this to compare options?" Most users use only a fraction of a feature's capability. |
Problem 6: Metrics Look Good But Users Churn
Symptom: Weekly active usage and task-completion are on target, but 90-day retention is poor. Heavy early engagement drops off by week six.
| Cause | Fix |
|---|---|
| Novelty-driven adoption — curiosity, not utility | Segment by acquisition cohort. Viral-moment users have a different curve than workflow-problem users; declining overall numbers can hide a healthy cohort. |
| No learning loop — quality is static | Make improvement visible: "Based on team feedback, we improved how this handles ambiguous requests." Users who see progress return. |
| Capability ceiling — no depth for expert users | Audit for power-user pathways: context-tuning, output chaining, advanced modes. Features with no depth churn fluent users. |
| Measuring activity, not value | Replace "outputs generated per week" with outcomes: downstream task completed, output shared. See Lesson 8 for the full measurement framework. |
Retention timeline: Commit to 90-day cohort analysis. AI features show a novelty spike, then a decay, then a stabilised habit curve (if genuinely useful). Shorter review cycles cannot tell novelty from habit.
Quick-Reference Diagnosis Table
| Symptom | Most Likely Problem | First Thing to Check |
|---|---|---|
| Low adoption despite good usability | Trust deficit | Onboarding explanation and first-run output quality |
| Errors reaching downstream work | Over-trust | Review friction and uncertainty signal design |
| Feedback rate below 2% | Low participation | Feedback timing and loop-closure visibility |
| "Watching me" or "knows too much" | Creepiness | Inference transparency and sensitive-attribute exposure |
| Usage healthy, NPS flat | Not valued | Task-completion comparison: AI-assisted vs. unassisted |
| 30-day retention good, 90-day poor | Novelty churn | Outcome metrics and expert-depth audit |
Key References
| Resource | What it covers |
|---|---|
| Google PAIR People+AI Guidebook — pair.withgoogle.com | Trust calibration, mental models, feedback design, UX patterns |
| Microsoft Human-AI Interaction Guidelines (Amershi et al.) | 18 empirically-validated design principles for AI-powered features |
| Nielsen Norman Group — UX for AI series — nngroup.com | Usability research on chat, inline AI, and autonomous agents |
| "The Alignment Problem" — Brian Christian | Why AI systems behave unexpectedly; accessible to non-technical readers |