Score the Backlog: An ICE Calibration Sprint
Objective: Score a real 8-idea growth backlog with ICE, then audit each Confidence score against the evidence actually cited for it before finalizing the rank order.
You're the growth marketer at Snowflake heading into a 45-minute backlog session with a designer and an engineer. Eight experiment ideas need an ICE score before anyone commits a week of work.
Apply the lesson's ICE scales, then flag and downgrade any Confidence score that isn't backed by a prior test or a cited benchmark.
Before you start
What you'll need
Free path (everything below is enough to finish)
Fast enough to fill live in a 45-minute backlog session with the whole team watching
The process
2 steps
Step 01 of 02
ICE scores each idea 1-10 on Impact, Confidence, and Ease, then multiplies the three for a single rank-order number; it's fast because it skips Reach, which makes it best for ideas you can ship in under a week.
Idea: 'Add social proof (customer logos) to the free-trial signup page.' The team scores Impact 6 (a plausible but unproven 5-10% lift), Confidence 7 (a similar test at a comparable SaaS company showed a real lift), Ease 9 (a one-day design change, no engineering). What is the ICE score, and where does it likely rank against a 'redesign the entire trial flow' idea scored Impact 9, Confidence 4, Ease 2?
Procedure
- Social proof idea: 6 x 7 x 9 = 378
- Trial flow redesign: 9 x 4 x 2 = 72
- Sort the table descending by score; the social proof idea ranks well above the redesign despite the redesign's higher Impact score
ICE backlog (excerpt, sorted) 1. Add social proof to trial signup 6 x 7 x 9 = 378 2. Shorten trial signup form 7 x 6 x 8 = 336 ... 8. Redesign entire trial flow 9 x 4 x 2 = 72
Healthy
Low-Ease, low-Confidence big bets naturally sink to the bottom of a fast-lane ICE list, which is correct: they belong in a separate roadmap conversation, not this week's sprint.
Unhealthy
The team overrides the score to run the trial flow redesign anyway because it 'feels bigger,' defeating the purpose of scoring at all.
What this means
ICE is deliberately blunt: a high-Impact idea with low Ease and low Confidence should lose to a smaller, well-evidenced, cheap-to-ship idea in a fast-lane backlog, that's the tool working as designed.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A big, exciting idea scores low and someone wants to override the ranking | Let it rank low for the ICE fast-lane sprint, and route it to a RICE-based roadmap review instead | 5 min |
Step 02 of 02
Inflating Confidence to win the argument corrupts the whole backlog; if you can't cite a prior test, an industry benchmark, or qualitative evidence, Confidence belongs below 60%, or below 6 on the 1-10 ICE scale.
Two backlog rows both have Confidence 8. Row A cites 'a similar social proof test that lifted conversion 12% at a comparable SaaS company, per a public case study.' Row B cites 'the whole team feels good about this one.' Which Confidence score should be downgraded, and to roughly what value?
Procedure
- Row A: keep Confidence at 8, cited external benchmark supports it
- Row B: downgrade Confidence to 3-4, no cited test or benchmark exists
- Recalculate Row B's ICE score with the corrected Confidence and re-sort the backlog
Row A: Impact 6, Confidence 8 (cited case study), Ease 8 -> 384 Row B, before: Impact 7, Confidence 8 (no citation), Ease 7 -> 392 Row B, after correction: Impact 7, Confidence 3, Ease 7 -> 147, drops from rank 1 to rank 6
Healthy
The corrected backlog moves the uncited idea down several ranks, and it gets routed to a quick research step instead of the top of the sprint.
Unhealthy
The uncited idea stays at Confidence 8 because downgrading it would mean losing the argument in the room, which is exactly the score inflation the lesson warns about.
What this means
A one-sentence citation requirement next to every Confidence score is a cheap, mechanical way to catch inflation before it corrupts the whole backlog's rank order.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A Confidence score of 8+ has no citation attached | Downgrade it to 3-4 and route the idea to a research step (interview, teardown, or survey) before it competes for a sprint slot | 5 min |
Final deliverable
A ranked ICE backlog of 8 ideas with a Confidence-evidence column, and a corrected rank order after the calibration pass.
See a reference example
Robinhood growth backlog, ICE pass (excerpt) 1. Simplify the deposit confirmation screen 8 x 8 x 8 = 512 (cited: prior A/B test) 2. Add referral status to the home feed 7 x 6 x 8 = 336 (cited: comparable fintech benchmark) ... 6. Gamify the watchlist (downgraded) 6 x 3 x 6 = 108, was 288 before Confidence correction
Success criteria
You're done when you can:
- Correctly computes ICE scores for all 8 backlog rows
- Identifies at least one uncited Confidence score and downgrades it with a stated new value
- Re-sorts the backlog after the correction and shows the resulting rank change