From Raw Signals to a Prioritized Hypothesis Backlog
Objective: Given raw heatmap notes, 20 exit-survey responses, and 10 post-purchase-survey responses for one page, build a prioritized, ICE-scored hypothesis backlog using the lesson's 'because' clause format.
You're the growth researcher at Nubank. The credit-card signup page converts below benchmark, and you've just finished a round of qualitative research. Now you have to turn it into something the team can actually test.
Synthesize the raw research notes into 3-5 hypotheses, each with a because clause and an ICE score, ranked by priority.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, sufficient to synthesize and sort a backlog by ICE score
Paid upgrades (optional, faster/deeper)
Faster ongoing qualitative-data collection than free tools at scale
No access? Microsoft Clarity (free) covers the same heatmap and recording data collection
The process
3 steps
Step 01 of 03
Stage 3 synthesizes findings from heatmaps, recordings, and surveys into a prioritized list of test ideas, not a report.
Given raw notes (heatmap: 70% of sessions never scroll past the fee-schedule table; exit survey: 12 of 20 responses mention 'not sure about fees'; post-purchase: 6 of 10 mention 'almost gave up on the fee page'), what pattern do all three sources agree on?
Procedure
- List each finding from the heatmap notes, exit survey, and post-purchase survey in separate rows
- Group findings that point at the same page element or moment
- Discard single-source findings that no other tool corroborates for this round
Pattern found in all 3 sources: fee-schedule table on the signup page Heatmap: 70% never scroll to it Exit survey: 12/20 mention fee confusion Post-purchase: 6/10 nearly abandoned over fees -> Strongest candidate for a hypothesis
Healthy
The strongest hypothesis is the one corroborated across multiple research sources, not the loudest single comment.
Unhealthy
A hypothesis is written off one exit-survey comment with no corroboration from heatmap or post-purchase data.
What this means
Cross-source agreement is what separates a strong hypothesis from a guess dressed up as research.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A candidate hypothesis is based on only one research source | Check the other two sources for corroboration before writing the hypothesis, or mark it lower confidence | 5 min |
Step 02 of 03
'Because [specific finding], we believe [change] will [outcome] for [segment].' The because clause has to trace back to real evidence.
Turn the fee-schedule pattern from step 1 into a valid hypothesis using the because-clause format.
Procedure
- Write the finding into the 'because' clause with the specific numbers from step 1
- State the proposed change
- State the expected outcome and the user segment it applies to
Because heatmap data shows 70% of sessions never scroll to the fee-schedule table, and 12/20 exit-survey and 6/10 post-purchase responses cite fee confusion, we believe moving a summarized fee callout above the fold will increase signup completion for first-time applicants.
Healthy
Every backlog row can be traced back to specific research numbers from step 1.
Unhealthy
The row states the change and outcome but drops the because clause under time pressure.
What this means
The because clause is what makes this a hypothesis instead of a design opinion, it has to survive being written down.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A backlog row is missing its because clause | Go back to the raw notes and cite the specific finding before adding the row | 5 min |
Step 03 of 03
Score each hypothesis 1-10 on Impact, Confidence, and Ease, then sort the backlog by total score.
Given 2 other draft hypotheses in the backlog with rough impact/confidence/ease estimates, where does the fee-callout hypothesis from step 2 rank?
Procedure
- Score the fee-callout hypothesis on Impact, Confidence, Ease (1-10 each), using the corroboration strength from step 1 as the Confidence input
- Score the other backlog rows the same way
- Sum each row and sort descending
Fee-callout: Impact 8, Confidence 9 (3-source corroboration), Ease 7 = 24 Trust badges: Impact 6, Confidence 5, Ease 8 = 19 Button color: Impact 3, Confidence 4, Ease 9 = 16 Ranked: Fee-callout (24) > Trust badges (19) > Button color (16)
Healthy
The highest-corroboration hypothesis scores highest on Confidence and rises to the top of the backlog.
Unhealthy
A low-confidence, single-source idea outranks a well-corroborated one because Ease was overweighted.
What this means
ICE scoring only works if Confidence reflects real evidence strength, not a gut feeling separate from the research.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A low-evidence hypothesis is ranked above a well-corroborated one | Re-score Confidence using the source count from step 1, then re-sort the backlog | 5 min |
Final deliverable
A 3-5 row prioritized hypothesis backlog, each row with a because-clause hypothesis and an ICE score.
See a reference example
Wise signup-page hypothesis backlog (excerpt) 1. (Score 24) Because heatmap data shows 65% never scroll to the fee table, and exit/post-purchase surveys corroborate fee confusion, we believe a fee callout above the fold will increase completions for first-time users. 2. (Score 19) Because 8/20 exit-survey responses cite trust concerns, we believe adding a security badge near the submit button will increase completions. 3. (Score 16) Because heatmap clicks cluster on the current button color with no drop-off pattern, we believe a color change alone will not move completions much, low priority.
Success criteria
You're done when you can:
- Every backlog row is corroborated by at least one specific data point from the raw notes
- Every row includes a complete because-clause hypothesis
- ICE scores are consistent with the corroboration strength found in step 1, and the backlog is sorted correctly