AI Hypothesis Audit: Separating Real Signal from Plausible Guesses
Objective: Given 8 AI-generated growth hypotheses and 3 known business constraints, apply a domain-context sanity check to flag which hypotheses address a real, evidenced user problem versus which are plausible-sounding guesses an LLM produced without knowing your business.
You're the growth analyst at RXBAR, the Chicago-founded protein bar brand ATT (acquired by Kellanova/Kellogg for $600M in 2017). Claude read your checkout session recordings and produced 8 hypotheses. Sprint planning is in an hour.
Score each hypothesis against 3 constraints you know and Claude doesn't: subscription customers already get free shipping, the mobile checkout was rebuilt last quarter, and first-time buyers are price-anchored by a 12-bar variety pack, not single flavors.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free tier handles an 8-item hypothesis list and a short back-and-forth without hitting usage limits
The process
1 step
Step 01 of 01
The lesson's pitfall section warns that LLMs generate plausible hypotheses, not ground-truth ones, and that a PM who runs all 10 AI-generated ideas in parallel without a domain filter watches most of them fail.
Claude's 8 hypotheses include 'add free shipping threshold banner', 'redesign mobile checkout button', and 'bundle flavors at a discount for first-time buyers'. Which of these survive contact with what you already know?
Procedure
- List the 8 AI-generated hypotheses in one column
- List your 3 known constraints in a second reference block
- For each hypothesis, ask: does this conflict with a constraint I already know is true?
- Kill any hypothesis that targets an already-solved problem (free shipping banner, mobile checkout)
- Keep hypotheses that target an unaddressed, evidenced friction point (flavor bundling for anchoring)
HYPOTHESIS AUDIT 1. Add free shipping threshold banner -- KILL. Subscription customers already get free shipping; banner targets a solved problem. 2. Redesign mobile checkout button -- KILL. Checkout was rebuilt last quarter; re-testing the same surface wastes a cycle. 3. Bundle flavors at a discount for first-time buyers -- KEEP. First-time buyers are anchored to the 12-bar variety pack; a flavor bundle addresses real anchoring friction Claude correctly inferred from session data. ...5 more rows
Healthy
3 of 8 hypotheses survive the constraint check and go into the sprint.
Unhealthy
All 8 hypotheses get greenlit because the AI confidence scores looked high.
What this means
An LLM's hypothesis list is a starting menu, not a ranked verdict; only a human who knows what's already shipped can tell which items are real.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Sprint burns a cycle re-testing an already-solved checkout problem | Keep a running list of 'already shipped or already known' constraints and paste it alongside every AI hypothesis prompt | 5 min |
Final deliverable
A 3-hypothesis shortlist for this sprint, plus the 5 killed hypotheses each paired with the specific constraint that killed it.
See a reference example
Halo Top hypothesis audit (excerpt) KEEP: 'Show calorie count above the fold on the pint PDP' -- addresses a real, unaddressed comparison-shopping friction seen in session recordings. KILL: 'Add a countdown timer to the flavor drop' -- Halo Top already ran and killed this exact test last quarter; re-testing wastes the cycle. KILL: 'Simplify the newsletter signup form' -- newsletter conversion isn't a business goal this quarter; hypothesis is plausible but off-target.
Success criteria
You're done when you can:
- Correctly kills every hypothesis that conflicts with a stated constraint
- Keeps only hypotheses that target a real, unaddressed friction point
- States the specific conflicting constraint for each killed hypothesis, not just 'not a priority'