Skip to content
Academy
Marketing Academy · Field Work●Growth Marketing
CoreBuild the Asset· 70 minutes

Standing Up an Experimentation Program From Zero

Grab Holdings

Objective: Build the founding artifacts of a working experimentation program for a team that currently runs ad-hoc tests: a structured hypothesis backlog, an ICE-scored prioritization sheet, a power/sample-size check, and a weekly review council agenda with an SRM check built in.

Grab's food-delivery growth team runs 2-3 tests a quarter with no shared backlog and no review ritual. You've been asked to stand up the process before the next planning cycle.

Produce a backlog template, a scored shortlist, a power check, and a review ritual — the gears most teams skip first.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeBacklog, ICE scoring, and review agenda in one workbook

Free, shareable, sortable, no account friction for a cross-functional council

Google Analytics 4(optional)
FreePull weekly session counts per surface for the power calculation

Free tier covers the session-volume data needed for sample-size math

No access? Estimate weekly traffic from any existing analytics export

The process

4 steps

Step 01 of 04

Writing hypotheses in the if/because format

Gear 1: every submission must state 'If we change [X] for users [Y], metric [Z] will move because [rationale R].' Without that structure, ideas are wishes, not experiments.

Grab's team has one raw idea note: 'riders drop off during payment, maybe simplify it.' Turn this into a real hypothesis using the if/because format.

Google Sheets— Create a Backlog tab with columns: Hypothesis, Audience, Metric, Rationale, Submitted By.

Procedure

  1. Create a Backlog tab with columns Hypothesis / Audience / Metric / Rationale / Submitted By
  2. Rewrite the raw note as: 'If we reduce payment steps from 4 to 2 for first-time riders, checkout completion rate will increase because fewer steps reduce drop-off from decision fatigue'
  3. Add 4 more raw ideas from the team and rewrite each the same way
Sample output
Backlog tab (row 1)
Hypothesis: If we reduce payment steps from 4 to 2 for first-time riders, checkout completion will increase
Audience: First-time riders, mobile app
Metric: Checkout completion rate
Rationale: Fewer steps reduce drop-off from decision fatigue

Healthy

Every row names a specific change, a specific audience segment, and one measurable metric.

Unhealthy

Rows that just restate the raw idea ('simplify checkout') with no named metric or audience.

What this means

A hypothesis without a named metric can't be scored in Gear 2 or measured in Gear 4 — the format isn't bureaucracy, it's what makes the next two gears possible.

So what do I do about it?

SymptomActionEffort
Backlog rows are one-line feature ideasReject any backlog submission missing an audience or metric field30 min
YouYou can do this yourself, no engineering access required.

Step 02 of 04

Scoring backlog items with ICE before committing engineering time

Score every hypothesis on Impact x Confidence x Ease before committing engineering time. 58% of teams skip this, and it's the single biggest reason velocity collapses.

You now have 5 hypotheses. Score each 1-10 on Impact, Confidence, and Ease, then rank them. Which one ships first?

Google Sheets— Add Impact / Confidence / Ease columns to the Backlog tab and a formula column for ICE score.

Procedure

  1. Add columns Impact, Confidence, Ease (1-10), and an ICE formula column (=Impact*Confidence*Ease)
  2. Score all 5 hypotheses as a team, not solo, to avoid one person's bias driving the ranking
  3. Sort descending by ICE score
Sample output
Ranked shortlist
1. Reduce payment steps 4->2   I:8 C:7 E:6 = 336
2. Add saved-card autofill    I:6 C:8 E:7 = 336
3. Reorder promo carousel     I:4 C:5 E:9 = 180

Healthy

The top-ranked test by ICE score is the one that gets engineering time next sprint, no exceptions.

Unhealthy

A lower-scored test jumps the queue because a senior stakeholder asked for it — that's HiPPO-driven guessing, not a program.

What this means

The score only has teeth if it actually determines the sprint order; a scoring sheet nobody follows is theater, not Gear 2.

So what do I do about it?

SymptomActionEffort
A stakeholder request jumps the ICE-ranked queueRequire a documented ICE override reason before any queue-jump5 min
YouYou can do this yourself, no engineering access required.

Step 03 of 04

Calculating whether a surface has enough traffic to reach significance in four weeks

Before launch, calculate minimum detectable effect and required sample size. If you can't reach 95% confidence within four weeks at current traffic, don't run the test on that surface.

The checkout page gets 40,000 weekly sessions. The team wants to detect a 3% lift in completion rate. Should this test run on checkout, or does it need a higher-traffic surface?

Google Analytics 4— Pull weekly session counts for the checkout page from GA4, then run the numbers through a sample-size calculator.

Procedure

  1. Pull the checkout page's weekly session count from GA4 (assume ~22% baseline completion rate)
  2. Plug baseline rate, 3% relative MDE, and 95% confidence into a sample-size calculator
  3. Compare the required sample per variant to 4 weeks of available traffic (40,000/week x 4 = 160,000 sessions, split two ways = 80,000/variant)
Sample output
Required sample per variant: ~65,000
Available per variant in 4 weeks: ~80,000
Verdict: sufficient power, test can launch on checkout as-is

Healthy

Available traffic per variant exceeds the required sample size before the test ever launches.

Unhealthy

Available traffic falls short, and the test runs anyway, then gets extended indefinitely chasing significance that will never arrive.

What this means

Power belongs before launch, not as an excuse invented after a test stalls at week 6 with an inconclusive read.

So what do I do about it?

SymptomActionEffort
A test has been running 6+ weeks with no significant resultCheck power before extending; if underpowered, either raise the MDE or move to a higher-traffic surface30 min
YouYou can do this yourself, no engineering access required.

Step 04 of 04

Catching sample-ratio mismatch in a weekly review before shipping

A weekly experimentation council — PM, engineer, analyst, designer, 30 minutes — reviews queued tests and catches sample-ratio mismatches (SRM) before they happen. This ritual is the structural difference between 5 tests a year and 5 tests a week.

Write the standing agenda for Grab's first weekly council. What gets checked on every live test before anyone reads the results?

Google Sheets— Add a Review tab with a fixed agenda template and an SRM check row for each live test.

Procedure

  1. Create a Review tab with 4 fixed agenda items: new tests to approve, live tests' SRM check, tests ready to read out, backlog re-prioritization
  2. For each live test, log actual split (e.g. 51.8/48.2) against the expected 50/50
  3. Flag any split beyond roughly 1-2% deviation for investigation before reading results
Sample output
Weekly council agenda, Aug 18
1. Approve: payment-steps test (ICE 336)
2. SRM check: saved-card-autofill test — split 52.4/47.6, FLAGGED, investigate before reading
3. Ready to read out: none this week
4. Re-prioritize: promo-carousel test dropped after new data

Healthy

Every live test gets an SRM check logged every week, before anyone looks at the primary metric.

Unhealthy

Results get read and shared the moment they look significant, with no check on whether the split itself is broken.

What this means

SRM is not a nice-to-have footnote — a broken 52/48 split invalidates every downstream number regardless of what the dashboard's p-value says.

So what do I do about it?

SymptomActionEffort
A test's actual split has drifted more than ~2% from the intended splitFreeze the readout and route to engineering to find the allocation bug before trusting any result30 min
DeveloperNeeds a developer/engineer to ship the fix.

Final deliverable

A working experimentation program starter kit for Grab: a Backlog tab (hypotheses in if/because format), an ICE-scored ranked shortlist, a power check for the target surface, and a Review tab with a standing weekly council agenda that includes an SRM check.

See a reference example
Sample output
Nubank experimentation starter kit (reference)

Backlog: 6 hypotheses, all if/because format
ICE shortlist: KYC upload simplification ranked #1 (I:8 C:7 E:6=336)
Power check: checkout surface cleared at 3% MDE, 95% confidence
Review agenda: 4 fixed items, SRM logged weekly, one test flagged at 53/47 split in week 2

Success criteria

You're done when you can:

  • All backlog rows follow the if/because hypothesis format with a named metric
  • ICE scores are computed and used to rank the shortlist, not just recorded
  • A power/sample-size check is done before the top-ranked test launches
  • The review agenda includes an explicit SRM check step for every live test