Prototyping AI Features

50 min advanced Lesson 7

Learning Outcomes

  • Apply Wizard-of-Oz testing to validate AI feature assumptions before writing a line of code
  • Design a prompt prototype that reveals user reactions to AI-generated output in under a week
  • Choose the right validation stage — concept, behaviour, reliability, or scale — for where your feature is today
  • Distinguish AI-in-the-loop from AI-in-the-lead operation and decide which to start with
  • Create a full prototype plan that moves a proposed AI feature from napkin sketch to confident rollout decision

Lesson Plan

Segment Duration Topic
Intro 3 min Why AI features break traditional prototyping methods
Stage 1 8 min Wizard-of-Oz testing: faking the AI before building it
Stage 2 8 min Prompt prototyping: quick LLM demos to surface reactions
Stage 3 7 min AI-in-the-loop vs AI-in-the-lead: the autonomy ladder
Stage 4 8 min Progressive rollout: feature flags, A/B tests, and guardrails
Evaluating non-determinism 7 min What to test when you can't write a traditional spec
Prototype plan exercise 6 min Pulling all four stages into one coherent plan
Wrap-up 3 min Key takeaways and next steps

Before You Begin

Pre-work:

Shopping List:

  • A proposed AI feature to prototype — it can be real (something your team is debating) or hypothetical
  • A whiteboard, a slide deck tool, or sticky notes for mapping the stages
  • Access to any general-purpose AI chat tool (free tier is sufficient) for the prompt prototyping exercises in Steps 2 and 3
  • A small set of real or representative users you can recruit for at least one round of testing — five participants is enough to start

1 Why AI Features Break Traditional Prototyping

Standard prototyping wisdom says: build the simplest thing that tests your riskiest assumption. For most features, that riskiest assumption is "will users want this?" — answered with a clickable Figma mockup and five user interviews.

AI features have a different structure. The riskiest assumption is usually not whether users want better recommendations or a smarter assistant — they almost always do. The riskiest assumptions are:

  • Output quality: Will the AI produce outputs that are actually useful, not just impressive in a demo?
  • User trust calibration: Will users trust the AI the right amount — neither ignoring it nor over-relying on it?
  • Workflow fit: Will the AI suggestion arrive at the right moment in the user's process, or will it interrupt?
  • Edge case tolerance: How will users react when the AI is confidently wrong?

These assumptions cannot be tested with a static mockup, because a static mockup cannot produce non-deterministic output. You can draw what a suggestion looks like, but you cannot draw the full distribution of what the AI will actually say to your actual users on their actual tasks.

This means AI prototyping requires a staged approach. Each stage answers a different question before you commit to the next level of investment.

Stage Question answered Investment
Wizard of Oz Do users find this interaction model valuable at all? Very low — human labour
Prompt prototype Does real AI output meet user expectations? Low — API credits
AI-in-the-loop Is the AI accurate enough to show users without embarrassing yourself? Medium — partial build
Progressive rollout Does it hold up across your real user base? High — production-ready
NOTE
A Mental Model Worth Keeping
Think of these four stages as answering four different questions in order: should we build this at all? → does the AI output actually work? → can we trust it enough to show users? → does it scale? Each stage only makes sense after the previous one has passed.
WARNING
The Demo Trap
AI is unusually good at impressing in a demo. A hand-curated response in a slide presentation will always look better than real production output on a random user's real input. Invest early in seeing what the AI actually produces, not what you can make it produce under ideal conditions.

2 Stage One: Wizard-of-Oz Testing

Wizard-of-Oz testing (named after the scene behind the curtain) is an old usability technique that works extraordinarily well for AI features. The idea: a human operator pretends to be the AI while real users interact with what looks like a working product. The user sees a finished-looking interface; behind the scenes, a team member reads the user's input and types a response.

This sounds low-tech because it is. That is the point. You can run a Wizard-of-Oz study for an AI feature in a week with no engineering, no model integration, and no infrastructure.

What you learn that you cannot learn any other way:

  • Whether users phrase their requests in ways the AI will actually understand
  • Whether users feel the interaction pattern is natural or awkward
  • Whether the latency of a response matters (users often notice even a two-second wait)
  • Whether users trust the output enough to act on it — or whether they immediately question it
  • What users do when they disagree with a response

How to set one up:

The minimum version requires two people: one playing the user (your participant), one playing the AI (a team member in a back room or on a shared screen). The interface can be as simple as a shared document, a mockup in a prototyping tool where you type into text fields, or a basic chat interface. The participant does not know a human is responding.

> [Session brief for your Wizard-of-Oz operator]
Your job is to play the AI assistant in our email drafting feature.
When the user pastes a rough note and asks for a draft, you have
30 seconds to write a plausible email draft in the style we described.
Do NOT be perfect — produce output at roughly the quality level we
expect from the real AI. Do NOT reveal you are human.
Focus: draft emails only. If asked anything else, say "I can only
help with drafting emails right now."
TIP
Calibrating the Operator's Output Quality
A common mistake is having your operator write better responses than the real AI will produce. Brief your operator on the expected quality level: 'roughly the quality of a competent but slightly generic assistant — useful, but not brilliant.' Over-performing operators give you false confidence.

What Wizard-of-Oz tests best:

Good questions for WoZ Poor questions for WoZ
Is the interaction model natural? Is the AI fast enough?
Do users trust outputs enough to use them? Will the model handle edge cases?
Are user inputs readable/usable? Does the feature scale?
Does the UI surface the AI at the right moment? What will the real model produce?

A Wizard-of-Oz study does not tell you whether your AI model will actually work. It tells you whether the concept is worth building. That is the only question Stage One needs to answer.

WARNING
Disclosure and Consent
Always debrief participants after a Wizard-of-Oz session and explain a human was responding. In a research context, participants consent to the study before it begins with language like 'the system may use simulated responses.' Deceiving users about AI involvement without consent crosses an ethical line the Google PAIR Guidebook explicitly flags.

3 Stage Two: Prompt Prototyping

Once your Wizard-of-Oz study has confirmed the concept is worth pursuing, you face a new question: will the real AI produce output good enough for users? This is where prompt prototyping comes in.

Prompt prototyping means using a general-purpose AI chat tool — Claude, Gemini, or similar — as a stand-in for your eventual production model, testing it against real user inputs, and observing user reactions, before you build any integration. You are not writing production code. You are running a structured experiment to understand the gap between what the AI produces and what your users need.

The three-part prompt prototype:

A useful prompt prototype has three components: the system context (what role the AI is playing), the user input (what a real user would type), and an evaluation rubric (what counts as "good enough").

> [System context for your prototype session]
You are an inline writing assistant in a project management tool.
When a user pastes rough notes from a meeting, you produce a
structured action-items list with owner, task, and due date columns.
You only output the table — no preamble, no sign-off.
If information is missing, leave the cell blank rather than guessing.

[User input to test]
"had a good mtg w/ Sarah and Devraj. sarah will sort out the vendor
contract stuff prob by end of month. devraj is on the eng timeline
should be done like 2-3 weeks. we need to do comms but no one took
it. also budget approval still outstanding."

[Evaluate: does this output meet the bar for showing users?]

You then run this prompt — or variations of it — against 20-30 real user inputs that you have collected from research, from your Wizard-of-Oz sessions, or from team members' own notes. The goal is not to achieve perfection. The goal is to understand the distribution of output quality: what proportion of inputs produce good output, what proportion produce mediocre output, and what proportion produce something that would damage user trust.

NOTE
The 80/20 Threshold
A rough rule from practitioners who have shipped AI features: if the AI produces useful output on about 80% of real user inputs, you have enough quality to proceed to the next stage with good error design. If it is below 60%, the model or the prompt needs more work before users see it. Between 60-80%, you need especially strong error and correction UX — revisit Lesson 4.

Documenting your prompt prototype findings:

Capture results in a simple evaluation table:

Input type Output quality User reaction (observed or predicted) Action
Well-structured note High Accepts and uses directly Proceed
Ambiguous input Medium Likely to question one item Needs uncertainty signalling
Very short / vague Low Frustration — unhelpful output Needs fallback UX
Edge case (unusual format) Very low Trust damage Block or add detection

This table becomes the input to your error design (which you covered in Lesson 4: Designing for Errors) and your trust design (from Lesson 3: User Trust & Transparency). Prompt prototyping surfaces exactly the scenarios where your error and trust designs will be tested.

TIP
Involve Real Users in Prompt Prototyping
Show the actual AI outputs to five real users without explaining they came from a prototype. Ask: 'How useful is this? What would you do with it? What would you change?' User reactions to real outputs are far more predictive than team opinions about whether the outputs look right.

4 Stage Three: AI-in-the-Loop vs AI-in-the-Lead

You have confirmed users want the concept (Wizard of Oz) and that the AI produces acceptable output often enough (prompt prototyping). Now you need to decide how much autonomy to give the AI when real users see it. This is the AI-in-the-loop versus AI-in-the-lead distinction, and it is one of the most consequential architectural decisions in AI product design.

AI-in-the-loop means the AI proposes; a human approves before anything happens. The AI's output is always a suggestion, never an action. Users see every output and have an explicit approve-or-reject step. Nothing irreversible happens without human sign-off.

AI-in-the-lead means the AI acts autonomously within defined boundaries, and humans review after the fact or only when the AI signals uncertainty. The AI does the work; humans can override but are not required to pre-approve.

These are not binary states. Think of them as positions on an autonomy ladder:

Rung Description Example
1. Show only AI surfaces information; no action possible "This email is likely spam" label
2. Suggest AI proposes; user initiates action explicitly "Insert this draft" button
3. Recommend with one click AI proposes with a prominent default accept Autocomplete with Tab to accept
4. Act with easy undo AI acts immediately; undo is prominent Auto-categorised inbox (Primary/Promotions); user can move any message
5. Act with notification AI acts silently; user is notified after Spam auto-moved to folder, shown in count
6. Act autonomously AI acts; user only sees outcomes Fraud detection blocks transaction

Why start lower on the ladder:

The Microsoft Human-AI Interaction guidelines and the Google PAIR Guidebook both recommend starting at a lower rung than you think you need, then moving up based on evidence. The reasons are practical:

  • Users calibrate trust gradually. A feature that starts at Rung 2 and earns trust can move to Rung 4 based on user behaviour data — but a feature that starts at Rung 5 and loses trust is difficult to recover.
  • Errors at higher rungs are more expensive. An incorrect suggestion at Rung 2 costs one extra click. An autonomous action at Rung 5 may require human effort, an apology, or a support ticket to reverse.
  • Starting lower gives you data. When users approve or reject AI suggestions at Rung 2, you see exactly which suggestions are accepted (signal that the AI is correct and trusted) and which are rejected or edited (signal of error or trust gap).
NOTE
The Autonomy Decision Is Reversible Upward, Rarely Downward
Moving a feature from Rung 2 to Rung 4 after users have come to trust it is a smooth upgrade. Moving it from Rung 5 back to Rung 2 after a trust incident feels like a demotion — users notice. Start conservative and promote the feature; don't start ambitious and demote it.

Choosing the right starting rung:

If the AI is wrong, the consequence is... And user trust is... Start at rung...
Minor (easy to undo, low stakes) Established by prior research 4 or 5
Minor (easy to undo, low stakes) Unestablished or unknown 2 or 3
Moderate (reversible, some effort) Established 3 or 4
Moderate Unestablished 2
Significant (hard to reverse, high stakes) Any level 1 or 2 — never higher until proven
WARNING
Autonomous Action in High-Stakes Domains
In domains where errors are expensive — financial transactions, healthcare data, legal documents, communications sent on behalf of users — stay at Rung 1 or 2 regardless of how accurate your internal tests say the model is. The asymmetry between the cost of one bad autonomous action and the benefit of saved clicks is rarely in favour of autonomy. This connects directly to the ethics framework in Lesson 9.

5 Stage Four: Progressive Rollout

Progressive rollout is not specific to AI — most mature product teams use feature flags and A/B tests routinely. But AI features have characteristics that make progressive rollout more important, not less, even after your earlier prototype stages have passed.

Why AI features need extra caution in rollout:

  • Variance at scale: A prompt prototype tested on 30 inputs may look fine. At a million inputs, the tail of unusual edge cases becomes a significant absolute number. What was a rare failure in testing becomes a recurring user complaint in production.
  • Emergent user behaviour: Users invent uses for AI features that your team did not anticipate. Some of these are positive (unexpected value discovered). Some reveal failure modes you did not test for.
  • Non-determinism makes regressions subtle: Traditional features either work or don't. AI features can subtly degrade — producing slightly worse outputs after a model update, a prompt change, or a shift in the underlying data distribution — in ways that do not trigger an error alert.

The progressive rollout sequence:

A robust AI feature rollout moves through at least four exposure stages:

Stage Who sees it What you measure Exit criteria
Internal (dogfood) Your team and willing employees Does it break? Does it embarrass? No blocking bugs; team would recommend it
Private beta Hand-picked power users or researchers Do real users find it useful? What surprises them? Satisfaction threshold met; key edge cases documented
Limited production (1-5%) A small slice of real traffic, usually opt-in Task completion rate, error correction rate, trust signals Metrics meet baseline targets; no anomalies
Broad rollout (graduated %) Increasing real traffic to 100% Retention, error rates, NPS, guardrail metrics Stable metrics across cohorts; team confident to proceed

Feature flags and A/B configuration:

Feature flags let you show the AI feature to a defined subset of users without a code deployment. For AI features, configure your flags to capture:

  • Which model version a user is experiencing (important when you update the model)
  • Which prompt variant they are seeing (important for prompt experiments)
  • Which autonomy rung they are on (important if you are testing two rungs simultaneously)

A simple A/B test structure for an AI suggestion feature:

Group Treatment Primary metric
Control No AI suggestion Task completion time, baseline
Variant A AI suggestion shown (Rung 2) Suggestion acceptance rate, time delta
Variant B AI suggestion with confidence indicator Acceptance rate vs Variant A
TIP
Define Guardrail Metrics Before You Launch
A guardrail metric is a measurement you will not allow to worsen as you optimise for your primary metric. Common guardrail metrics for AI features include: error correction rate (if users are correcting AI more than X% of the time, something is wrong), support ticket volume related to AI outputs, and opt-out rate (if users are turning off the AI feature, they are telling you something). Define these before launch; discovering them after is much harder.
WARNING
Don't Skip the Gradual Part
Teams under shipping pressure often jump from internal testing directly to full production rollout. For AI features, this is especially risky because errors at scale can move faster than your ability to respond. A single day at 1% traffic with monitoring gives you data you cannot get any other way, and costs almost nothing.

6 Evaluating Non-Deterministic Features

Traditional feature testing asks a binary question: does the feature behave as specified? AI features cannot be fully specified in the same way. Given the same input twice, an AI feature may produce two different outputs — both potentially valid, or one valid and one subtly wrong. This is non-determinism, and it fundamentally changes how you evaluate quality.

You cannot write a complete test suite for an AI feature. You can write partial test suites (sometimes called "evals"), define quality rubrics, sample outputs over time, and use human raters. But you cannot enumerate every correct output for every input and check against it the way you would with deterministic logic.

This is not a reason to give up on evaluation. It is a reason to use different tools.

The three evaluation tools for AI features:

1. Rubric-based evaluation. Instead of specifying the correct output, specify the properties of a good output. A rubric for an email drafting assistant might be:

Property Description How to check
Accurate Reflects the facts the user provided Spot-check: can you trace every claim to the input?
Appropriately toned Matches the requested register Human rater: 1-5 scale
Concise Not substantially longer than needed Word count relative to input
Complete Does not drop a key fact from the input Human rater: did anything important get omitted?
Structurally correct Follows the expected format Automated: template matching

You can apply a rubric to a sample of outputs — 50 or 100 — using human raters, and track rubric scores over time as a quality signal.

2. Regression sampling. When you change a prompt, swap a model, or modify a feature, run your new version against a fixed set of reference inputs and compare rubric scores against the previous version. You are not looking for identical outputs; you are looking for quality that is at least as good. A decline in rubric scores is a regression, even if no error was thrown.

3. User-signal proxies. When your feature is live, user behaviour tells you a great deal about output quality. High edit rates (users substantially rewriting AI output) signal poor quality. High acceptance rates (users using AI output with minimal changes) signal good quality. The Lesson 5: Feedback Loops framework covers how to collect and interpret these signals systematically.

NOTE
A Practical Evaluation Cadence
For a live AI feature, a sustainable evaluation cadence looks like this: weekly automated sampling (100 random inputs run through a rubric-scoring prompt), monthly human rater review (20-30 outputs rated on your full rubric by a team member or external rater), and triggered review after any model or prompt change. This is not exhaustive — it is a smoke detector, not a fire prevention system.

Setting your quality bar before launch:

Define minimum acceptable quality before you ship, not after. A useful framing from the PAIR Guidebook: set your bar relative to the alternatives available to users, not relative to perfection. If your AI email drafter produces a genuinely useful first draft in eight out of ten cases, and the alternative is a blank page, that is a meaningful improvement worth shipping. If the alternative is a well-tested template that already works reliably, an eight-out-of-ten AI is a harder sell.

> [Quality bar decision template]
Feature: [Name the AI feature]
User's alternative without AI: [Blank page / template / manual process / nothing]
Minimum useful output quality: [Describe what "useful" looks like for this feature]
Acceptable failure rate: [What % of outputs falling below the minimum bar is tolerable?]
What happens when the AI fails: [Fallback experience — revisit Lesson 4]
Quality bar source: [Prompt prototype results / user research / competitive benchmark]
TIP
Get the Failure Case Right First
Before finalising your quality bar, walk through what the worst plausible AI output looks like — not the average output, the worst tail case that real users will encounter. If your error design and fallback experience (from Lesson 4) handles that failure gracefully, your quality bar can afford to be pragmatic. If a bad output creates an embarrassing or trust-destroying user experience, your quality bar needs to be higher, or your error design needs more work.

7 Building Your Prototype Plan

A prototype plan pulls the four stages into one coherent document that your team can align on before anyone writes a line of code. It is not a technical spec — it is a validation roadmap that answers, for each stage: what assumption are we testing, how will we test it, who will test it, and what result will make us proceed to the next stage?

The prototype plan template:

A one-page prototype plan covers the following:

Feature summary. One sentence: what does this feature do, for whom, in what context?

The riskiest assumptions (in priority order). List the top three things that would cause this feature to fail in the market — not engineering risks, but product/user risks. Common ones: "users will phrase requests in ways the AI understands," "users will trust the output enough to act on it without verifying everything," "the AI output quality is high enough on the range of inputs our users actually produce."

Stage 1 — Wizard of Oz plan:

Element Your plan
Assumption being tested [e.g. users find the interaction model natural]
Interface (mockup, doc, tool)
Operator instructions
Participant profile
Session length
Pass/fail criterion

Stage 2 — Prompt prototype plan:

Element Your plan
System context prompt
Input sample (how many, how collected)
Evaluation rubric
User reaction method (observation, interview, survey)
Pass/fail criterion (e.g. 80% rubric score)

Stage 3 — Autonomy rung decision:

Element Your plan
Starting rung (1-6 from Step 4)
Rationale (consequence of error × trust level)
Criteria for moving to a higher rung

Stage 4 — Progressive rollout plan:

Stage Audience Duration Metrics Exit criteria
Internal
Private beta
Limited production
Broad rollout

Team checkpoints. At what point does the team review the data from each stage and make a Go/No-Go decision? Name the decision-maker for each checkpoint. A prototype plan without a named decision-maker is a document that will be ignored under shipping pressure.

NOTE
The Plan Is a Living Document
Prototype plans are not contracts — they are hypotheses. When Stage 1 reveals a surprise (users interact with the feature in a way you did not expect), update the plan rather than force the remaining stages to test the wrong thing. The value of the plan is the shared understanding it creates before testing begins, not the rigidity it imposes after.
TIP
One Page Is a Discipline, Not a Limit
Keep the prototype plan to one page. This forces you to be specific about what actually matters. A ten-page prototype plan will not be read by the engineer, designer, and PM in the room — a one-pager will be. The detail lives in your individual test scripts and evaluation rubrics; the plan is the shared map.

Questions & Answers

Q: We don't have time for four stages of prototyping before shipping. Can we skip some?
Yes — the stages are not equally important for every feature. If your feature is low-autonomy (a suggestion users can ignore), low-stakes (errors are easy to correct), and in a domain where you have existing user research, you might compress Stages 1 and 2 into a combined two-day study. What you should almost never skip is progressive rollout at Stage 4 — shipping an untested AI feature to 100% of users simultaneously is where most publicly embarrassing AI failures come from. A week at 1% traffic is a small investment relative to the risk of a bad launch.
Q: Our Wizard-of-Oz test showed users loved the concept, but real AI output in prompt prototyping disappointed them. What do we do?
This is one of the most important signals you can get, and it is much better to get it at Stage 2 than at launch. It means the concept is right but the AI is not yet good enough to deliver it. Your options, in order of preference: improve the system prompt (often dramatically improves output quality); constrain the input space (narrow the range of tasks or inputs the feature accepts, focusing on the ones where AI does well); revisit the UX to set more conservative user expectations and stronger error design; or delay the feature until the underlying model capability improves. Do not paper over a Stage 2 failure with hope — if 40% of outputs disappointed real users in a controlled test, they will disappoint users in production too.
Q: How do I evaluate AI output quality when my team disagrees about what "good" looks like?
Team disagreement about output quality is a sign that you have not yet written an explicit rubric. Before running any evaluation, convene the team and define, in writing, the properties of a good output — and then define what each level (excellent, acceptable, not acceptable) looks like on each property. Run three or four sample outputs through the rubric together until ratings converge. Calibration takes thirty minutes and eliminates most disagreements. The remaining disagreements are usually about which user need the feature is optimising for — which is a product strategy conversation, not an evaluation conversation.
Q: Our engineers say we can't do Wizard-of-Oz without building an interface first, and building an interface isn't low-cost. How do we handle this?
Wizard-of-Oz does not require a custom interface. For a chat or assistant feature, an existing AI chat tool with a custom system prompt simulates the feature convincingly. For an inline suggestion feature, a shared document where the "AI" is a team member making edits in real time works surprisingly well. For an email or scheduling feature, a simple script that forwards user inputs to the operator's phone works for a one-day study. The goal is not fidelity to the final UI — it is enough fidelity to observe whether the interaction model is natural. You can usually achieve that with tools you already have.
Q: What do I do if my A/B test shows no statistically significant difference between AI and no-AI? Does that mean the feature has failed?
Not necessarily — but it is a serious signal worth investigating before proceeding. Common causes: the metric you chose is not sensitive to what the AI actually changes (a task completion rate metric will not detect that AI made a task feel easier); the exposure period was too short (AI features often show benefits over repeated use, not in a first session); or the feature genuinely does not move the needle for this user population. Before concluding failure, check qualitative data — do users who accepted AI suggestions report a different experience than those who did not? If there is qualitative signal but no quantitative signal, you may need to refine your metric before making a shipping decision. If both are flat, the feature may not be delivering enough value to ship.

Key Takeaways

  1. AI features require staged prototyping — four stages answer four different questions in order: should we build it (Wizard of Oz), does the AI output work (prompt prototype), can users trust it (autonomy rung), does it scale (progressive rollout).
  2. Wizard of Oz tests the concept, not the model — a human operator simulating AI in a basic interface reveals whether the interaction model is natural and whether users find the concept valuable, before any engineering investment.
  3. Prompt prototyping surfaces the quality distribution — running real user inputs through a general-purpose AI tool and evaluating outputs against a rubric shows you where the AI fails, so you can design errors before users encounter them.
  4. Start lower on the autonomy ladder than you think you need to — AI-in-the-loop (user approves before action) builds trust, generates quality signal, and is much easier to upgrade than to downgrade after a trust incident.
  5. Non-deterministic features need rubrics, not specs — define the properties of good output before testing, track rubric scores over time, and use user behaviour signals like edit rate and acceptance rate as ongoing quality proxies.
  6. Progressive rollout is not optional for AI — variance at scale, emergent user behaviour, and subtle quality regressions make gradual exposure and guardrail metrics essential even after earlier prototype stages have passed.

Next Steps: Lesson 8: Measuring AI Product Success