Build an AI-Assisted Experiment Brief: Hypothesis to RICE Score
Objective: Given real funnel data, use AI at two stages of the workflow, hypothesis generation from session data and RICE prioritisation scoring, to produce a complete, ready-to-review experiment brief for one growth idea.
You're the growth marketer at MVMT, the DTC watch brand that sold to Movado Group for a reported $100M in 2018. Checkout abandonment is up 6 points this month and you have 12 candidate ideas competing for one dev sprint.
Use an LLM to turn raw session notes into ranked hypotheses, then use RICE scoring (also AI-assisted) to defend which one idea gets the sprint, and write the one-page brief a developer can build from.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free tier handles a 40-session text block and a 12-row scoring table in one conversation
Free, shareable with the dev team without export friction
Paid upgrades (optional, faster/deeper)
Built-in sample size calculator and Stats Engine for early significance detection
No access? Google Sheets with a manual sample-size formula
The process
2 steps
Step 01 of 02
The lesson contrasts traditional hypothesis generation (a PM eyeballs a heatmap) with AI-assisted generation: paste session recordings and support tickets into an LLM and ask for the top friction points, compressing 10 hours of manual review into seconds.
You have 40 session-recording notes and 15 support tickets mentioning checkout. What are the top 3 friction points worth turning into hypotheses?
Procedure
- Compile session notes and support-ticket excerpts into one text block
- Prompt: 'What are the top 5 friction points in this checkout flow, ranked by how many sessions show the pattern?'
- Review Claude's output against the raw notes for at least the top 2 patterns
- Convert the top pattern into a testable hypothesis with a clear metric
TOP FRICTION PATTERNS (from 40 sessions + 15 tickets) 1. Re-entering shipping address after a failed promo code (14 sessions) -- HYPOTHESIS: preserving form state across a failed promo-code submit will reduce checkout abandonment. 2. Watch band size chart opens in a new tab and loses cart context (9 sessions) 3. Order confirmation email delayed 20+ minutes, driving repeat-purchase attempts (6 tickets)
Healthy
One hypothesis is grounded in 14 of 40 sessions showing the same pattern.
Unhealthy
A hypothesis invented from a single anecdotal ticket with no session-data pattern behind it.
What this means
A hypothesis is only as strong as the number of independent sessions that show the same friction; AI's job is surfacing the pattern fast, not deciding it matters.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Team debates which of 12 ideas to run without evidence | Require every hypothesis to cite a session or ticket count before it enters the prioritisation step | 5 min |
Step 02 of 02
The lesson notes RICE (Reach x Impact x Confidence / Effort) is where AI excels because it multiplies across dimensions fast and surfaces ideas that look average on one metric but exceptional on others.
Your form-state hypothesis competes against 11 other ideas. Reach is high (checkout touches 100% of buyers), but is it the highest-RICE idea, or does a lower-reach idea with near-zero effort win?
Procedure
- List all 12 ideas with your best-guess reach, impact, confidence, and effort estimates
- Prompt Claude to compute RICE score per idea and rank the list
- Sanity-check the top-ranked idea against what you know the dev team can ship this sprint
- Write the one-page brief: hypothesis, RICE score and reasoning, success metric, sample size
RICE RANKING (top 3 of 12) 1. Preserve form state on failed promo code -- Reach 100% x Impact 2 x Confidence 80% / Effort 1 = 160. Ships in 2 days, touches every checkout session. 2. Fix band-size-chart tab context loss -- Reach 22% x Impact 1 x Confidence 60% / Effort 2 = 6.6 3. Speed up confirmation email -- Reach 15% x Impact 1 x Confidence 50% / Effort 3 = 2.5 BRIEF: Ship the form-state fix this sprint; RICE score is 24x the next idea and effort is a 2-day dev ticket.
Healthy
The top RICE idea also matches the dev team's actual sprint capacity.
Unhealthy
Picking the highest-reach idea without checking whether effort makes it undeliverable this sprint.
What this means
RICE surfaces the idea, but a human still checks that the winning score is buildable in the time available.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Team ships the loudest idea instead of the highest-RICE idea | Require a RICE table with all 12 ideas visible before any single idea gets greenlit | 30 min |
Final deliverable
A one-page experiment brief: the top hypothesis, its RICE score and reasoning against the other 11 ideas, the success metric, and the required sample size.
See a reference example
Chewy checkout brief (excerpt) HYPOTHESIS: Preserving cart contents when a shipping-zip lookup fails will reduce cart abandonment. EVIDENCE: 11 of 30 reviewed sessions show a zip-lookup failure immediately preceding exit. RICE: Reach 100% x Impact 2 x Confidence 75% / Effort 1 = 150 (highest of 9 candidate ideas). METRIC: checkout completion rate. SAMPLE SIZE: ~18,000 sessions per arm at 95% confidence for a 2pp lift.
Success criteria
You're done when you can:
- Hypothesis cites a specific session or ticket count as evidence
- RICE table shows all competing ideas, not just the winner
- Brief includes a stated success metric and sample size