Grading a Growth Team's Experimentation Maturity
Objective: Given a raw backlog export from a real growth team, score their experimentation process against the five-gear framework and identify the single weakest gear.
You're a growth analyst embedded with Nubank's app growth team for a quarter. They've handed you a backlog export and asked you to diagnose why velocity is stuck at 2-3 tests a quarter.
Score hypothesis compliance and ICE-scoring coverage from the raw backlog, then name the weakest gear with a concrete fix.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, no account friction, sortable columns for ICE ranking
The process
2 steps
Step 01 of 02
Gear 1 requires every backlog submission to state 'If we change X for users Y, metric Z will move because rationale R.' Without that structure, ideas are wishes, not experiments.
Nubank's backlog has 18 entries. Only 4 use the full hypothesis format; the other 14 are one-line feature requests like 'Test a new home screen banner.' What's the actual state of Gear 1?
Procedure
- Import backlog-export.csv into Sheets
- Add a column 'Has Hypothesis' and mark TRUE only for rows naming a specific change, a specific user segment, and a specific metric
- Count TRUE vs FALSE to get the compliance rate
18 backlog rows 4 rows: full hypothesis (change + audience + metric + rationale) 14 rows: title only, no metric named Compliance rate: 22%
Healthy
Compliance rate above 80% — most ideas already name a specific metric and audience before they reach the backlog.
Unhealthy
Compliance rate under 30% — the backlog is really a wishlist, and prioritization scoring later will be guessing, not measuring.
What this means
A low compliance rate means Gear 1 is broken even if Gear 2's scoring template looks polished; you can't score a hypothesis that was never written.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Backlog entries are one-line feature requests with no named metric | Add a required 'hypothesis' field to the intake form and reject submissions missing it | 30 min |
Step 02 of 02
Gear 2 requires every hypothesis to be scored (ICE or PXL) before anyone commits engineering time. 58% of teams skip this entirely, and that single gap kills velocity more than anything else.
None of Nubank's 18 backlog rows carry an Impact, Confidence, or Ease score. Engineering picks the next test from whichever PM asked most recently. What does this tell you about Gear 2, and what's the fix?
Procedure
- Add three columns: Impact, Confidence, Ease (1-10 scale)
- Score only the 4 rows with a real hypothesis, since the other 14 can't be scored honestly yet
- Sort descending by ICE score to produce a ranked shortlist
ICE-scored shortlist (of 4 valid hypotheses) 1. Simplify KYC upload flow I:8 C:7 E:6 = 336 2. Reorder home screen cards I:5 C:4 E:8 = 160 Remaining 14 backlog rows: unscored, unranked
Healthy
A ranked shortlist exists and engineering pulls the next test from the top of that list, not from whoever asked last.
Unhealthy
Zero rows carry a score, and the next test is chosen by seniority or urgency of the request, not evidence.
What this means
Missing ICE scores plus a HiPPO-driven pick order is the exact failure mode the lesson calls out: 58% of teams skip prioritization, and it is the single most common reason velocity collapses.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Engineering picks tests by whoever asked most recently | Require an ICE score on every backlog row before it can be assigned to a sprint | 30 min |
Final deliverable
A one-page maturity scorecard for Nubank's experimentation program: hypothesis compliance rate, ICE-scoring coverage, and the single weakest gear with a named fix.
See a reference example
Grab experimentation maturity scorecard (reference) Gear 1 (Intake): 91% hypothesis compliance — healthy Gear 2 (Prioritization): 100% ICE-scored before sprint assignment — healthy Gear 4 (Review): no weekly council, SRM checks ad hoc — WEAKEST GEAR Recommended fix: stand up a 30-minute weekly experimentation council.
Success criteria
You're done when you can:
- Correctly computes hypothesis compliance rate from the raw backlog
- Identifies missing ICE scores as the primary velocity blocker, not a tooling problem
- Names one concrete fix tied to the weakest gear