The Stop/Continue Call: Auditing a Live Test Dashboard Before You Decide
Objective: Given a live-test dashboard export (day-by-day p-values, metric list, and traffic-allocation log), apply the lesson's three checks to decide whether the result can be trusted.
You're the CRO lead at Wise. A pricing-page test has been running for 9 of its planned 14 days, and the PM is asking whether you can call it today.
Walk the dashboard export through the peeking check, the multiple comparisons check, and the Simpson's Paradox check before answering the PM.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, sufficient to filter and compare columns from an exported test dashboard
The process
3 steps
Step 01 of 03
A p-value calculated for a fixed sample size assumes one look at the end; checking daily and stopping at first significance can push a nominal 5% false positive rate to 20%+.
The dashboard shows day 9 of a planned 14-day/30,000-visitor test crossing p=0.048 today after being non-significant on days 3-8. Should you call it now?
Procedure
- Open dashboard-export.csv and check the 'planned end date/sample size' field logged at launch
- Compare today's date and sample against that plan
- If the plan isn't reached, do not stop, regardless of today's p-value
Planned: 14 days / 30,000 visitors (logged at launch) Current: Day 9 / 19,400 visitors p-value history: Day3 0.61, Day5 0.34, Day7 0.19, Day9 0.048 Status: NOT yet at plan -> do not stop
Healthy
The team holds the line and waits for day 14 / 30,000 visitors regardless of today's p-value.
Unhealthy
The team calls the test today because p crossed 0.05, five days and roughly 10,600 visitors early.
What this means
A p-value crossing 0.05 before the pre-registered endpoint is exactly the pattern peeking produces, it is not evidence.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| PM is asking to call the test early because today's number looks good | Show the pre-registered end date/sample size and hold until it's reached | 5 min |
Step 02 of 03
Testing many metrics or variants without correction means the odds of at least one false 'win' climb fast, roughly 64% across 20 tests at alpha 0.05.
The dashboard logs 7 tracked metrics for this test. Which one was declared the primary metric before launch, and does the 'win' show up there?
Procedure
- Find the 'primary metric' field logged at launch
- Check whether that specific metric, not any of the other 6, is the one showing significance
- If a secondary metric is significant but the primary isn't, treat it as a diagnostic lead, not a result
Primary metric (logged at launch): Checkout conversion rate -> p=0.41 (not significant) Secondary metrics: AOV p=0.03, email opt-in p=0.09, scroll depth p=0.22, ... Status: Primary metric not significant -> no win to call
Healthy
The team reports 'primary metric not significant, test continues or fails' regardless of which secondary metric turned green.
Unhealthy
The team reports a win on AOV because it's the only metric under 0.05, ignoring that 7 metrics were tracked.
What this means
A secondary metric turning significant among 7 tracked metrics is the multiple comparisons problem showing up in real data.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A secondary metric is significant but the primary metric is flat | Report the primary metric result as the decision, and log the secondary finding as a new hypothesis to test on its own | 5 min |
Step 03 of 03
Simpson's Paradox happens when a trend inside every subgroup reverses once you combine the subgroups, usually from uneven traffic mix between variants.
The traffic-allocation log shows the treatment split moved from 10% to 50% on day 6 to speed up data collection. Does the combined day 1-9 conversion number mean anything on its own?
Procedure
- Check the allocation-log tab for any change in the treatment/control traffic split during the test
- If the split changed, break the conversion numbers out by day/allocation period instead of trusting the combined total
- Confirm treatment's per-period result agrees with the combined result before reporting either
Days 1-5 (10% split): Control 2.1% CVR, Treatment 2.4% CVR -> Treatment wins Days 6-9 (50% split): Control 1.3% CVR, Treatment 1.5% CVR -> Treatment wins Combined days 1-9: Control 1.68%, Treatment 1.52% -> looks like Control wins
Healthy
The team reports the per-period breakdown and flags the allocation change as the reason the combined number is misleading.
Unhealthy
The team reports 'Control wins' off the combined total without checking whether the allocation changed mid-test.
What this means
The allocation change on day 6 is exactly the trigger the lesson names for Simpson's Paradox; the combined total is not trustworthy on its own.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Combined result disagrees with every per-period result | Hold traffic allocation constant for future tests, and re-run this one cleanly instead of trusting the combined number | 30 min |
Final deliverable
A written stop/continue recommendation to the PM covering all 3 checks, with the reasoning shown, not just the verdict.
See a reference example
Freshworks pricing-test stop/continue memo (excerpt) 1. Peeking check: Day 8/14, 21,000/28,000 visitors -> not at plan, hold. 2. Multiple comparisons check: Primary metric (trial starts) p=0.22, not significant -> no win yet. 3. Simpson's check: Allocation stable at 50/50 throughout -> combined number is trustworthy. Recommendation: Continue to day 14.
Success criteria
You're done when you can:
- Correctly applies all 3 checks using the dashboard's own logged plan/allocation data, not just today's p-value
- Recommendation matches what the 3 checks actually show
- Explains reasoning for the stop/continue call, not just a verdict