Skip to content
Academy
Marketing Academy · Field Work●Growth Marketing
CoreBuild the Asset· 50 minutes

Build an AI-Assisted Experiment Brief: Hypothesis to RICE Score

MVMT

Objective: Given real funnel data, use AI at two stages of the workflow, hypothesis generation from session data and RICE prioritisation scoring, to produce a complete, ready-to-review experiment brief for one growth idea.

You're the growth marketer at MVMT, the DTC watch brand that sold to Movado Group for a reported $100M in 2018. Checkout abandonment is up 6 points this month and you have 12 candidate ideas competing for one dev sprint.

Use an LLM to turn raw session notes into ranked hypotheses, then use RICE scoring (also AI-assisted) to defend which one idea gets the sprint, and write the one-page brief a developer can build from.

Before you start

What you'll need

Free path (everything below is enough to finish)

Claude
FreemiumGenerate hypotheses from session data and compute RICE scores

Free tier handles a 40-session text block and a 12-row scoring table in one conversation

FreeHold the 12-idea RICE table and the final brief draft

Free, shareable with the dev team without export friction

Paid upgrades (optional, faster/deeper)

Optimizely(optional)
PaidCalculate the required sample size and run the experiment once it ships

Built-in sample size calculator and Stats Engine for early significance detection

No access? Google Sheets with a manual sample-size formula

The process

2 steps

Step 01 of 02

Hypothesis generation from qualitative session data

The lesson contrasts traditional hypothesis generation (a PM eyeballs a heatmap) with AI-assisted generation: paste session recordings and support tickets into an LLM and ask for the top friction points, compressing 10 hours of manual review into seconds.

You have 40 session-recording notes and 15 support tickets mentioning checkout. What are the top 3 friction points worth turning into hypotheses?

Claude— Paste the session notes and ticket excerpts into Claude in one message.

Procedure

  1. Compile session notes and support-ticket excerpts into one text block
  2. Prompt: 'What are the top 5 friction points in this checkout flow, ranked by how many sessions show the pattern?'
  3. Review Claude's output against the raw notes for at least the top 2 patterns
  4. Convert the top pattern into a testable hypothesis with a clear metric
Sample output
TOP FRICTION PATTERNS (from 40 sessions + 15 tickets)
1. Re-entering shipping address after a failed promo code (14 sessions) -- HYPOTHESIS: preserving form state across a failed promo-code submit will reduce checkout abandonment.
2. Watch band size chart opens in a new tab and loses cart context (9 sessions)
3. Order confirmation email delayed 20+ minutes, driving repeat-purchase attempts (6 tickets)

Healthy

One hypothesis is grounded in 14 of 40 sessions showing the same pattern.

Unhealthy

A hypothesis invented from a single anecdotal ticket with no session-data pattern behind it.

What this means

A hypothesis is only as strong as the number of independent sessions that show the same friction; AI's job is surfacing the pattern fast, not deciding it matters.

So what do I do about it?

SymptomActionEffort
Team debates which of 12 ideas to run without evidenceRequire every hypothesis to cite a session or ticket count before it enters the prioritisation step5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Automated experiment prioritisation with RICE

The lesson notes RICE (Reach x Impact x Confidence / Effort) is where AI excels because it multiplies across dimensions fast and surfaces ideas that look average on one metric but exceptional on others.

Your form-state hypothesis competes against 11 other ideas. Reach is high (checkout touches 100% of buyers), but is it the highest-RICE idea, or does a lower-reach idea with near-zero effort win?

Claude— Feed Claude your 12-idea list with reach/impact/confidence/effort estimates for each.

Procedure

  1. List all 12 ideas with your best-guess reach, impact, confidence, and effort estimates
  2. Prompt Claude to compute RICE score per idea and rank the list
  3. Sanity-check the top-ranked idea against what you know the dev team can ship this sprint
  4. Write the one-page brief: hypothesis, RICE score and reasoning, success metric, sample size
Sample output
RICE RANKING (top 3 of 12)
1. Preserve form state on failed promo code -- Reach 100% x Impact 2 x Confidence 80% / Effort 1 = 160. Ships in 2 days, touches every checkout session.
2. Fix band-size-chart tab context loss -- Reach 22% x Impact 1 x Confidence 60% / Effort 2 = 6.6
3. Speed up confirmation email -- Reach 15% x Impact 1 x Confidence 50% / Effort 3 = 2.5

BRIEF: Ship the form-state fix this sprint; RICE score is 24x the next idea and effort is a 2-day dev ticket.

Healthy

The top RICE idea also matches the dev team's actual sprint capacity.

Unhealthy

Picking the highest-reach idea without checking whether effort makes it undeliverable this sprint.

What this means

RICE surfaces the idea, but a human still checks that the winning score is buildable in the time available.

So what do I do about it?

SymptomActionEffort
Team ships the loudest idea instead of the highest-RICE ideaRequire a RICE table with all 12 ideas visible before any single idea gets greenlit30 min
YouYou can do this yourself, no engineering access required.

Final deliverable

A one-page experiment brief: the top hypothesis, its RICE score and reasoning against the other 11 ideas, the success metric, and the required sample size.

See a reference example
Sample output
Chewy checkout brief (excerpt)

HYPOTHESIS: Preserving cart contents when a shipping-zip lookup fails will reduce cart abandonment.
EVIDENCE: 11 of 30 reviewed sessions show a zip-lookup failure immediately preceding exit.
RICE: Reach 100% x Impact 2 x Confidence 75% / Effort 1 = 150 (highest of 9 candidate ideas).
METRIC: checkout completion rate.
SAMPLE SIZE: ~18,000 sessions per arm at 95% confidence for a 2pp lift.

Success criteria

You're done when you can:

  • Hypothesis cites a specific session or ticket count as evidence
  • RICE table shows all competing ideas, not just the winner
  • Brief includes a stated success metric and sample size