Two Test Plans, One Approval: Auditing for Peeking Risk
Objective: Given two competing test plans for the same experiment, identify which one is statistically sound and quantify why the other one is not, using the lesson's false-positive-inflation evidence.
Two PMs at Zendesk each propose a different way to run the same support-ticket upsell banner test, and only one plan can be approved this sprint.
Score both plans against the lesson's peeking and validation rules, then write the approval memo.
Before you start
What you'll need
Free path (everything below is enough to finish)
Fast to build, easy to attach as evidence in the approval memo
Paid upgrades (optional, faster/deeper)
If daily monitoring is a hard business requirement, this is the correct way to get it without inflating false positives
No access? Pre-register a fixed sample and single end-of-test read in Google Sheets instead
The process
2 steps
Step 01 of 02
The lesson cites a 2025 simulation study finding that checking results 10+ times and stopping at the first significant read inflates false positives above 40%, versus the nominal 5% alpha.
Plan A checks the p-value every morning starting Day 3 and stops the moment it crosses p<0.05. Plan B pre-registers a fixed sample of 24,000 total and a 14-day end date with no interim stops. Which plan's approach to 'significance' can actually be trusted?
Procedure
- Enter the two plans' rules side by side: check frequency, stopping condition, sample target
- Log the given 12-day p-value sequence for Plan A: it dips below 0.05 on Day 6, back above on Day 9, below again on Day 13
- Flag every day Plan A's rule would have triggered a stop, and count how many separate 'wins' it would have declared
Plan A daily p-value log Day 6: p=0.041 <- would stop and declare a winner here Day 9: p=0.11 (already 'shipped' by Day 6 under Plan A's rule) Day 13: p=0.038 Plan B: no interim checks, single read on Day 14, p=0.06 (not significant)
Healthy
The memo flags that Plan A's Day 6 stop is exactly the peeking pattern the lesson's cited study measured at 40%+ false positive rates.
Unhealthy
Approving Plan A because it 'found a winner faster' without checking whether that reflects the peeking risk.
What this means
Plan A's early stop on Day 6 is not evidence of a real effect, it is the predictable artifact of checking daily and stopping at first significance; Plan B's single Day 14 read (p=0.06) is the trustworthy number.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A test plan checks significance more than once and stops at the first pass | Reject the plan or require it switch to a pre-registered fixed sample and single end-of-test read | 5 min |
Step 02 of 02
The lesson's Step 5 checklist requires confirming the planned sample was hit, the test ran a full business cycle, and the result clears both the significance threshold and the MDE before calling a winner.
Plan B's Day 14 read shows p=0.06 and only 19,000 of the planned 24,000 sample. Should this be approved as a completed test?
Procedure
- Check sample hit: 19,000 of 24,000 planned, not yet met
- Check business cycle: 14 days run, meets the 14-day preferred minimum
- Check significance and MDE: p=0.06 is above alpha 0.05, does not clear the bar
Validation checklist, Plan B Day 14 Sample hit? NO (19,000 / 24,000) Business cycle? YES (14 days) Significant + MDE? NO (p=0.06) Verdict: extend the test, do not call a winner yet
Healthy
The memo recommends extending Plan B to hit its planned sample rather than approving an incomplete read.
Unhealthy
Calling Plan B 'basically significant' and shipping the change because p=0.06 is close to 0.05.
What this means
Neither plan currently has a trustworthy result: Plan A's is inflated by peeking, Plan B's is real but incomplete.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A test hasn't hit its planned sample size at the scheduled end date | Extend the runtime rather than approve a partial-sample result | 5 min |
Final deliverable
A one-page approval memo recommending which plan to run, with the peeking-risk math and validation checklist shown as evidence.
See a reference example
Snowflake trial-signup banner test, plan audit memo (excerpt) Plan A (daily checks, stop at first p<0.05): REJECTED, peeking pattern matches the 40%+ false-positive study cited in CRO training. Plan B (fixed 22,000 sample, 14-day single read): APPROVED, add 3 more days to reach planned sample before reading results.
Success criteria
You're done when you can:
- Correctly identifies which plan's stopping rule inflates false positives
- Applies the Step 5 validation checklist before recommending approval of either plan
- Recommends extending an incomplete test rather than approving a borderline p-value