Prototyping AI Features
Learning Outcomes
- Apply Wizard-of-Oz testing to validate AI feature assumptions before writing a line of code
- Design a prompt prototype that reveals user reactions to AI-generated output in under a week
- Choose the right validation stage — concept, behaviour, reliability, or scale — for where your feature is today
- Distinguish AI-in-the-loop from AI-in-the-lead operation and decide which to start with
- Create a full prototype plan that moves a proposed AI feature from napkin sketch to confident rollout decision
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why AI features break traditional prototyping methods |
| Stage 1 | 8 min | Wizard-of-Oz testing: faking the AI before building it |
| Stage 2 | 8 min | Prompt prototyping: quick LLM demos to surface reactions |
| Stage 3 | 7 min | AI-in-the-loop vs AI-in-the-lead: the autonomy ladder |
| Stage 4 | 8 min | Progressive rollout: feature flags, A/B tests, and guardrails |
| Evaluating non-determinism | 7 min | What to test when you can't write a traditional spec |
| Prototype plan exercise | 6 min | Pulling all four stages into one coherent plan |
| Wrap-up | 3 min | Key takeaways and next steps |
Before You Begin
Pre-work:
- Complete Lesson 1: When to Use AI (and When Not To) so you have a clear feature idea worth prototyping — the frameworks here assume you have already decided AI is appropriate
- Review Lesson 2: AI UX Patterns to understand which UX pattern your feature uses (chat, inline suggestions, autonomous agent, or ambient intelligence), because the right prototype technique depends on the pattern
- Skim Lesson 3: User Trust & Transparency and Lesson 4: Designing for Errors — prototyping is also where you stress-test the error and trust designs you planned there
Shopping List:
- A proposed AI feature to prototype — it can be real (something your team is debating) or hypothetical
- A whiteboard, a slide deck tool, or sticky notes for mapping the stages
- Access to any general-purpose AI chat tool (free tier is sufficient) for the prompt prototyping exercises in Steps 2 and 3
- A small set of real or representative users you can recruit for at least one round of testing — five participants is enough to start
Standard prototyping wisdom says: build the simplest thing that tests your riskiest assumption. For most features, that riskiest assumption is "will users want this?" — answered with a clickable Figma mockup and five user interviews.
AI features have a different structure. The riskiest assumption is usually not whether users want better recommendations or a smarter assistant — they almost always do. The riskiest assumptions are:
- Output quality: Will the AI produce outputs that are actually useful, not just impressive in a demo?
- User trust calibration: Will users trust the AI the right amount — neither ignoring it nor over-relying on it?
- Workflow fit: Will the AI suggestion arrive at the right moment in the user's process, or will it interrupt?
- Edge case tolerance: How will users react when the AI is confidently wrong?
These assumptions cannot be tested with a static mockup, because a static mockup cannot produce non-deterministic output. You can draw what a suggestion looks like, but you cannot draw the full distribution of what the AI will actually say to your actual users on their actual tasks.
This means AI prototyping requires a staged approach. Each stage answers a different question before you commit to the next level of investment.
| Stage | Question answered | Investment |
|---|---|---|
| Wizard of Oz | Do users find this interaction model valuable at all? | Very low — human labour |
| Prompt prototype | Does real AI output meet user expectations? | Low — API credits |
| AI-in-the-loop | Is the AI accurate enough to show users without embarrassing yourself? | Medium — partial build |
| Progressive rollout | Does it hold up across your real user base? | High — production-ready |
Wizard-of-Oz testing (named after the scene behind the curtain) is an old usability technique that works extraordinarily well for AI features. The idea: a human operator pretends to be the AI while real users interact with what looks like a working product. The user sees a finished-looking interface; behind the scenes, a team member reads the user's input and types a response.
This sounds low-tech because it is. That is the point. You can run a Wizard-of-Oz study for an AI feature in a week with no engineering, no model integration, and no infrastructure.
What you learn that you cannot learn any other way:
- Whether users phrase their requests in ways the AI will actually understand
- Whether users feel the interaction pattern is natural or awkward
- Whether the latency of a response matters (users often notice even a two-second wait)
- Whether users trust the output enough to act on it — or whether they immediately question it
- What users do when they disagree with a response
How to set one up:
The minimum version requires two people: one playing the user (your participant), one playing the AI (a team member in a back room or on a shared screen). The interface can be as simple as a shared document, a mockup in a prototyping tool where you type into text fields, or a basic chat interface. The participant does not know a human is responding.
> [Session brief for your Wizard-of-Oz operator]
Your job is to play the AI assistant in our email drafting feature.
When the user pastes a rough note and asks for a draft, you have
30 seconds to write a plausible email draft in the style we described.
Do NOT be perfect — produce output at roughly the quality level we
expect from the real AI. Do NOT reveal you are human.
Focus: draft emails only. If asked anything else, say "I can only
help with drafting emails right now."
What Wizard-of-Oz tests best:
| Good questions for WoZ | Poor questions for WoZ |
|---|---|
| Is the interaction model natural? | Is the AI fast enough? |
| Do users trust outputs enough to use them? | Will the model handle edge cases? |
| Are user inputs readable/usable? | Does the feature scale? |
| Does the UI surface the AI at the right moment? | What will the real model produce? |
A Wizard-of-Oz study does not tell you whether your AI model will actually work. It tells you whether the concept is worth building. That is the only question Stage One needs to answer.
Once your Wizard-of-Oz study has confirmed the concept is worth pursuing, you face a new question: will the real AI produce output good enough for users? This is where prompt prototyping comes in.
Prompt prototyping means using a general-purpose AI chat tool — Claude, Gemini, or similar — as a stand-in for your eventual production model, testing it against real user inputs, and observing user reactions, before you build any integration. You are not writing production code. You are running a structured experiment to understand the gap between what the AI produces and what your users need.
The three-part prompt prototype:
A useful prompt prototype has three components: the system context (what role the AI is playing), the user input (what a real user would type), and an evaluation rubric (what counts as "good enough").
> [System context for your prototype session]
You are an inline writing assistant in a project management tool.
When a user pastes rough notes from a meeting, you produce a
structured action-items list with owner, task, and due date columns.
You only output the table — no preamble, no sign-off.
If information is missing, leave the cell blank rather than guessing.
[User input to test]
"had a good mtg w/ Sarah and Devraj. sarah will sort out the vendor
contract stuff prob by end of month. devraj is on the eng timeline
should be done like 2-3 weeks. we need to do comms but no one took
it. also budget approval still outstanding."
[Evaluate: does this output meet the bar for showing users?]
You then run this prompt — or variations of it — against 20-30 real user inputs that you have collected from research, from your Wizard-of-Oz sessions, or from team members' own notes. The goal is not to achieve perfection. The goal is to understand the distribution of output quality: what proportion of inputs produce good output, what proportion produce mediocre output, and what proportion produce something that would damage user trust.
Documenting your prompt prototype findings:
Capture results in a simple evaluation table:
| Input type | Output quality | User reaction (observed or predicted) | Action |
|---|---|---|---|
| Well-structured note | High | Accepts and uses directly | Proceed |
| Ambiguous input | Medium | Likely to question one item | Needs uncertainty signalling |
| Very short / vague | Low | Frustration — unhelpful output | Needs fallback UX |
| Edge case (unusual format) | Very low | Trust damage | Block or add detection |
This table becomes the input to your error design (which you covered in Lesson 4: Designing for Errors) and your trust design (from Lesson 3: User Trust & Transparency). Prompt prototyping surfaces exactly the scenarios where your error and trust designs will be tested.
You have confirmed users want the concept (Wizard of Oz) and that the AI produces acceptable output often enough (prompt prototyping). Now you need to decide how much autonomy to give the AI when real users see it. This is the AI-in-the-loop versus AI-in-the-lead distinction, and it is one of the most consequential architectural decisions in AI product design.
AI-in-the-loop means the AI proposes; a human approves before anything happens. The AI's output is always a suggestion, never an action. Users see every output and have an explicit approve-or-reject step. Nothing irreversible happens without human sign-off.
AI-in-the-lead means the AI acts autonomously within defined boundaries, and humans review after the fact or only when the AI signals uncertainty. The AI does the work; humans can override but are not required to pre-approve.
These are not binary states. Think of them as positions on an autonomy ladder:
| Rung | Description | Example |
|---|---|---|
| 1. Show only | AI surfaces information; no action possible | "This email is likely spam" label |
| 2. Suggest | AI proposes; user initiates action explicitly | "Insert this draft" button |
| 3. Recommend with one click | AI proposes with a prominent default accept | Autocomplete with Tab to accept |
| 4. Act with easy undo | AI acts immediately; undo is prominent | Auto-categorised inbox (Primary/Promotions); user can move any message |
| 5. Act with notification | AI acts silently; user is notified after | Spam auto-moved to folder, shown in count |
| 6. Act autonomously | AI acts; user only sees outcomes | Fraud detection blocks transaction |
Why start lower on the ladder:
The Microsoft Human-AI Interaction guidelines and the Google PAIR Guidebook both recommend starting at a lower rung than you think you need, then moving up based on evidence. The reasons are practical:
- Users calibrate trust gradually. A feature that starts at Rung 2 and earns trust can move to Rung 4 based on user behaviour data — but a feature that starts at Rung 5 and loses trust is difficult to recover.
- Errors at higher rungs are more expensive. An incorrect suggestion at Rung 2 costs one extra click. An autonomous action at Rung 5 may require human effort, an apology, or a support ticket to reverse.
- Starting lower gives you data. When users approve or reject AI suggestions at Rung 2, you see exactly which suggestions are accepted (signal that the AI is correct and trusted) and which are rejected or edited (signal of error or trust gap).
Choosing the right starting rung:
| If the AI is wrong, the consequence is... | And user trust is... | Start at rung... |
|---|---|---|
| Minor (easy to undo, low stakes) | Established by prior research | 4 or 5 |
| Minor (easy to undo, low stakes) | Unestablished or unknown | 2 or 3 |
| Moderate (reversible, some effort) | Established | 3 or 4 |
| Moderate | Unestablished | 2 |
| Significant (hard to reverse, high stakes) | Any level | 1 or 2 — never higher until proven |
Progressive rollout is not specific to AI — most mature product teams use feature flags and A/B tests routinely. But AI features have characteristics that make progressive rollout more important, not less, even after your earlier prototype stages have passed.
Why AI features need extra caution in rollout:
- Variance at scale: A prompt prototype tested on 30 inputs may look fine. At a million inputs, the tail of unusual edge cases becomes a significant absolute number. What was a rare failure in testing becomes a recurring user complaint in production.
- Emergent user behaviour: Users invent uses for AI features that your team did not anticipate. Some of these are positive (unexpected value discovered). Some reveal failure modes you did not test for.
- Non-determinism makes regressions subtle: Traditional features either work or don't. AI features can subtly degrade — producing slightly worse outputs after a model update, a prompt change, or a shift in the underlying data distribution — in ways that do not trigger an error alert.
The progressive rollout sequence:
A robust AI feature rollout moves through at least four exposure stages:
| Stage | Who sees it | What you measure | Exit criteria |
|---|---|---|---|
| Internal (dogfood) | Your team and willing employees | Does it break? Does it embarrass? | No blocking bugs; team would recommend it |
| Private beta | Hand-picked power users or researchers | Do real users find it useful? What surprises them? | Satisfaction threshold met; key edge cases documented |
| Limited production (1-5%) | A small slice of real traffic, usually opt-in | Task completion rate, error correction rate, trust signals | Metrics meet baseline targets; no anomalies |
| Broad rollout (graduated %) | Increasing real traffic to 100% | Retention, error rates, NPS, guardrail metrics | Stable metrics across cohorts; team confident to proceed |
Feature flags and A/B configuration:
Feature flags let you show the AI feature to a defined subset of users without a code deployment. For AI features, configure your flags to capture:
- Which model version a user is experiencing (important when you update the model)
- Which prompt variant they are seeing (important for prompt experiments)
- Which autonomy rung they are on (important if you are testing two rungs simultaneously)
A simple A/B test structure for an AI suggestion feature:
| Group | Treatment | Primary metric |
|---|---|---|
| Control | No AI suggestion | Task completion time, baseline |
| Variant A | AI suggestion shown (Rung 2) | Suggestion acceptance rate, time delta |
| Variant B | AI suggestion with confidence indicator | Acceptance rate vs Variant A |
Traditional feature testing asks a binary question: does the feature behave as specified? AI features cannot be fully specified in the same way. Given the same input twice, an AI feature may produce two different outputs — both potentially valid, or one valid and one subtly wrong. This is non-determinism, and it fundamentally changes how you evaluate quality.
You cannot write a complete test suite for an AI feature. You can write partial test suites (sometimes called "evals"), define quality rubrics, sample outputs over time, and use human raters. But you cannot enumerate every correct output for every input and check against it the way you would with deterministic logic.
This is not a reason to give up on evaluation. It is a reason to use different tools.
The three evaluation tools for AI features:
1. Rubric-based evaluation. Instead of specifying the correct output, specify the properties of a good output. A rubric for an email drafting assistant might be:
| Property | Description | How to check |
|---|---|---|
| Accurate | Reflects the facts the user provided | Spot-check: can you trace every claim to the input? |
| Appropriately toned | Matches the requested register | Human rater: 1-5 scale |
| Concise | Not substantially longer than needed | Word count relative to input |
| Complete | Does not drop a key fact from the input | Human rater: did anything important get omitted? |
| Structurally correct | Follows the expected format | Automated: template matching |
You can apply a rubric to a sample of outputs — 50 or 100 — using human raters, and track rubric scores over time as a quality signal.
2. Regression sampling. When you change a prompt, swap a model, or modify a feature, run your new version against a fixed set of reference inputs and compare rubric scores against the previous version. You are not looking for identical outputs; you are looking for quality that is at least as good. A decline in rubric scores is a regression, even if no error was thrown.
3. User-signal proxies. When your feature is live, user behaviour tells you a great deal about output quality. High edit rates (users substantially rewriting AI output) signal poor quality. High acceptance rates (users using AI output with minimal changes) signal good quality. The Lesson 5: Feedback Loops framework covers how to collect and interpret these signals systematically.
Setting your quality bar before launch:
Define minimum acceptable quality before you ship, not after. A useful framing from the PAIR Guidebook: set your bar relative to the alternatives available to users, not relative to perfection. If your AI email drafter produces a genuinely useful first draft in eight out of ten cases, and the alternative is a blank page, that is a meaningful improvement worth shipping. If the alternative is a well-tested template that already works reliably, an eight-out-of-ten AI is a harder sell.
> [Quality bar decision template]
Feature: [Name the AI feature]
User's alternative without AI: [Blank page / template / manual process / nothing]
Minimum useful output quality: [Describe what "useful" looks like for this feature]
Acceptable failure rate: [What % of outputs falling below the minimum bar is tolerable?]
What happens when the AI fails: [Fallback experience — revisit Lesson 4]
Quality bar source: [Prompt prototype results / user research / competitive benchmark]
A prototype plan pulls the four stages into one coherent document that your team can align on before anyone writes a line of code. It is not a technical spec — it is a validation roadmap that answers, for each stage: what assumption are we testing, how will we test it, who will test it, and what result will make us proceed to the next stage?
The prototype plan template:
A one-page prototype plan covers the following:
Feature summary. One sentence: what does this feature do, for whom, in what context?
The riskiest assumptions (in priority order). List the top three things that would cause this feature to fail in the market — not engineering risks, but product/user risks. Common ones: "users will phrase requests in ways the AI understands," "users will trust the output enough to act on it without verifying everything," "the AI output quality is high enough on the range of inputs our users actually produce."
Stage 1 — Wizard of Oz plan:
| Element | Your plan |
|---|---|
| Assumption being tested | [e.g. users find the interaction model natural] |
| Interface (mockup, doc, tool) | |
| Operator instructions | |
| Participant profile | |
| Session length | |
| Pass/fail criterion |
Stage 2 — Prompt prototype plan:
| Element | Your plan |
|---|---|
| System context prompt | |
| Input sample (how many, how collected) | |
| Evaluation rubric | |
| User reaction method (observation, interview, survey) | |
| Pass/fail criterion (e.g. 80% rubric score) |
Stage 3 — Autonomy rung decision:
| Element | Your plan |
|---|---|
| Starting rung (1-6 from Step 4) | |
| Rationale (consequence of error × trust level) | |
| Criteria for moving to a higher rung |
Stage 4 — Progressive rollout plan:
| Stage | Audience | Duration | Metrics | Exit criteria |
|---|---|---|---|---|
| Internal | ||||
| Private beta | ||||
| Limited production | ||||
| Broad rollout |
Team checkpoints. At what point does the team review the data from each stage and make a Go/No-Go decision? Name the decision-maker for each checkpoint. A prototype plan without a named decision-maker is a document that will be ignored under shipping pressure.
Questions & Answers
Key Takeaways
- AI features require staged prototyping — four stages answer four different questions in order: should we build it (Wizard of Oz), does the AI output work (prompt prototype), can users trust it (autonomy rung), does it scale (progressive rollout).
- Wizard of Oz tests the concept, not the model — a human operator simulating AI in a basic interface reveals whether the interaction model is natural and whether users find the concept valuable, before any engineering investment.
- Prompt prototyping surfaces the quality distribution — running real user inputs through a general-purpose AI tool and evaluating outputs against a rubric shows you where the AI fails, so you can design errors before users encounter them.
- Start lower on the autonomy ladder than you think you need to — AI-in-the-loop (user approves before action) builds trust, generates quality signal, and is much easier to upgrade than to downgrade after a trust incident.
- Non-deterministic features need rubrics, not specs — define the properties of good output before testing, track rubric scores over time, and use user behaviour signals like edit rate and acceptance rate as ongoing quality proxies.
- Progressive rollout is not optional for AI — variance at scale, emergent user behaviour, and subtle quality regressions make gradual exposure and guardrail metrics essential even after earlier prototype stages have passed.
Next Steps: Lesson 8: Measuring AI Product Success