Skip to content
Academy
Marketing Academy · Field Work●Conversion Rate Optimization
CoreAudit· 40 minutes

The Stop/Continue Call: Auditing a Live Test Dashboard Before You Decide

Wise (formerly TransferWise)

Objective: Given a live-test dashboard export (day-by-day p-values, metric list, and traffic-allocation log), apply the lesson's three checks to decide whether the result can be trusted.

You're the CRO lead at Wise. A pricing-page test has been running for 9 of its planned 14 days, and the PM is asking whether you can call it today.

Walk the dashboard export through the peeking check, the multiple comparisons check, and the Simpson's Paradox check before answering the PM.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeReview the dashboard export, metrics tab, and allocation log

Free, sufficient to filter and compare columns from an exported test dashboard

The process

3 steps

Step 01 of 03

Peeking / optional stopping

A p-value calculated for a fixed sample size assumes one look at the end; checking daily and stopping at first significance can push a nominal 5% false positive rate to 20%+.

The dashboard shows day 9 of a planned 14-day/30,000-visitor test crossing p=0.048 today after being non-significant on days 3-8. Should you call it now?

Google Sheets— Open dashboard-export.csv, day-by-day p-value column.

Procedure

  1. Open dashboard-export.csv and check the 'planned end date/sample size' field logged at launch
  2. Compare today's date and sample against that plan
  3. If the plan isn't reached, do not stop, regardless of today's p-value
Sample output
Planned: 14 days / 30,000 visitors (logged at launch)
Current: Day 9 / 19,400 visitors
p-value history: Day3 0.61, Day5 0.34, Day7 0.19, Day9 0.048
Status: NOT yet at plan -> do not stop

Healthy

The team holds the line and waits for day 14 / 30,000 visitors regardless of today's p-value.

Unhealthy

The team calls the test today because p crossed 0.05, five days and roughly 10,600 visitors early.

What this means

A p-value crossing 0.05 before the pre-registered endpoint is exactly the pattern peeking produces, it is not evidence.

So what do I do about it?

SymptomActionEffort
PM is asking to call the test early because today's number looks goodShow the pre-registered end date/sample size and hold until it's reached5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 03

Multiple comparisons problem

Testing many metrics or variants without correction means the odds of at least one false 'win' climb fast, roughly 64% across 20 tests at alpha 0.05.

The dashboard logs 7 tracked metrics for this test. Which one was declared the primary metric before launch, and does the 'win' show up there?

Google Sheets— Metrics tab of dashboard-export.csv

Procedure

  1. Find the 'primary metric' field logged at launch
  2. Check whether that specific metric, not any of the other 6, is the one showing significance
  3. If a secondary metric is significant but the primary isn't, treat it as a diagnostic lead, not a result
Sample output
Primary metric (logged at launch): Checkout conversion rate -> p=0.41 (not significant)
Secondary metrics: AOV p=0.03, email opt-in p=0.09, scroll depth p=0.22, ...
Status: Primary metric not significant -> no win to call

Healthy

The team reports 'primary metric not significant, test continues or fails' regardless of which secondary metric turned green.

Unhealthy

The team reports a win on AOV because it's the only metric under 0.05, ignoring that 7 metrics were tracked.

What this means

A secondary metric turning significant among 7 tracked metrics is the multiple comparisons problem showing up in real data.

So what do I do about it?

SymptomActionEffort
A secondary metric is significant but the primary metric is flatReport the primary metric result as the decision, and log the secondary finding as a new hypothesis to test on its own5 min
YouYou can do this yourself, no engineering access required.

Step 03 of 03

Simpson's Paradox

Simpson's Paradox happens when a trend inside every subgroup reverses once you combine the subgroups, usually from uneven traffic mix between variants.

The traffic-allocation log shows the treatment split moved from 10% to 50% on day 6 to speed up data collection. Does the combined day 1-9 conversion number mean anything on its own?

Google Sheets— Allocation-log tab of dashboard-export.csv

Procedure

  1. Check the allocation-log tab for any change in the treatment/control traffic split during the test
  2. If the split changed, break the conversion numbers out by day/allocation period instead of trusting the combined total
  3. Confirm treatment's per-period result agrees with the combined result before reporting either
Sample output
Days 1-5 (10% split): Control 2.1% CVR, Treatment 2.4% CVR -> Treatment wins
Days 6-9 (50% split): Control 1.3% CVR, Treatment 1.5% CVR -> Treatment wins
Combined days 1-9: Control 1.68%, Treatment 1.52% -> looks like Control wins

Healthy

The team reports the per-period breakdown and flags the allocation change as the reason the combined number is misleading.

Unhealthy

The team reports 'Control wins' off the combined total without checking whether the allocation changed mid-test.

What this means

The allocation change on day 6 is exactly the trigger the lesson names for Simpson's Paradox; the combined total is not trustworthy on its own.

So what do I do about it?

SymptomActionEffort
Combined result disagrees with every per-period resultHold traffic allocation constant for future tests, and re-run this one cleanly instead of trusting the combined number30 min
EitherYou or a developer can handle this, depending on your access.

Final deliverable

A written stop/continue recommendation to the PM covering all 3 checks, with the reasoning shown, not just the verdict.

See a reference example
Sample output
Freshworks pricing-test stop/continue memo (excerpt)

1. Peeking check: Day 8/14, 21,000/28,000 visitors -> not at plan, hold.
2. Multiple comparisons check: Primary metric (trial starts) p=0.22, not significant -> no win yet.
3. Simpson's check: Allocation stable at 50/50 throughout -> combined number is trustworthy.
Recommendation: Continue to day 14.

Success criteria

You're done when you can:

  • Correctly applies all 3 checks using the dashboard's own logged plan/allocation data, not just today's p-value
  • Recommendation matches what the 3 checks actually show
  • Explains reasoning for the stop/continue call, not just a verdict