Skip to content
Academy
Marketing Academy · Field Work●Conversion Rate Optimization
MiniHead-to-Head· 25 minutes

Two Test Plans, One Approval: Auditing for Peeking Risk

Zendesk

Objective: Given two competing test plans for the same experiment, identify which one is statistically sound and quantify why the other one is not, using the lesson's false-positive-inflation evidence.

Two PMs at Zendesk each propose a different way to run the same support-ticket upsell banner test, and only one plan can be approved this sprint.

Score both plans against the lesson's peeking and validation rules, then write the approval memo.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeLog daily p-values and score each plan against the validation checklist

Fast to build, easy to attach as evidence in the approval memo

Paid upgrades (optional, faster/deeper)

Optimizely(optional)
PaidSequential-testing and alpha-spending safeguards that make daily peeking statistically safe

If daily monitoring is a hard business requirement, this is the correct way to get it without inflating false positives

No access? Pre-register a fixed sample and single end-of-test read in Google Sheets instead

The process

2 steps

Step 01 of 02

False positive inflation from peeking without a pre-set stopping rule

The lesson cites a 2025 simulation study finding that checking results 10+ times and stopping at the first significant read inflates false positives above 40%, versus the nominal 5% alpha.

Plan A checks the p-value every morning starting Day 3 and stops the moment it crosses p<0.05. Plan B pre-registers a fixed sample of 24,000 total and a 14-day end date with no interim stops. Which plan's approach to 'significance' can actually be trusted?

Google Sheets— Log each plan's daily p-value sequence in a sheet to see how often 'significant' appears before Day 14.

Procedure

  1. Enter the two plans' rules side by side: check frequency, stopping condition, sample target
  2. Log the given 12-day p-value sequence for Plan A: it dips below 0.05 on Day 6, back above on Day 9, below again on Day 13
  3. Flag every day Plan A's rule would have triggered a stop, and count how many separate 'wins' it would have declared
Sample output
Plan A daily p-value log
  Day 6:  p=0.041  <- would stop and declare a winner here
  Day 9:  p=0.11   (already 'shipped' by Day 6 under Plan A's rule)
  Day 13: p=0.038

Plan B: no interim checks, single read on Day 14, p=0.06 (not significant)

Healthy

The memo flags that Plan A's Day 6 stop is exactly the peeking pattern the lesson's cited study measured at 40%+ false positive rates.

Unhealthy

Approving Plan A because it 'found a winner faster' without checking whether that reflects the peeking risk.

What this means

Plan A's early stop on Day 6 is not evidence of a real effect, it is the predictable artifact of checking daily and stopping at first significance; Plan B's single Day 14 read (p=0.06) is the trustworthy number.

So what do I do about it?

SymptomActionEffort
A test plan checks significance more than once and stops at the first passReject the plan or require it switch to a pre-registered fixed sample and single end-of-test read5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Validating results only after reaching planned sample size and a full business cycle

The lesson's Step 5 checklist requires confirming the planned sample was hit, the test ran a full business cycle, and the result clears both the significance threshold and the MDE before calling a winner.

Plan B's Day 14 read shows p=0.06 and only 19,000 of the planned 24,000 sample. Should this be approved as a completed test?

Google Sheets— Same sheet, add a validation checklist row per the lesson's Step 5 criteria.

Procedure

  1. Check sample hit: 19,000 of 24,000 planned, not yet met
  2. Check business cycle: 14 days run, meets the 14-day preferred minimum
  3. Check significance and MDE: p=0.06 is above alpha 0.05, does not clear the bar
Sample output
Validation checklist, Plan B Day 14
  Sample hit?         NO  (19,000 / 24,000)
  Business cycle?     YES (14 days)
  Significant + MDE?  NO  (p=0.06)
  Verdict: extend the test, do not call a winner yet

Healthy

The memo recommends extending Plan B to hit its planned sample rather than approving an incomplete read.

Unhealthy

Calling Plan B 'basically significant' and shipping the change because p=0.06 is close to 0.05.

What this means

Neither plan currently has a trustworthy result: Plan A's is inflated by peeking, Plan B's is real but incomplete.

So what do I do about it?

SymptomActionEffort
A test hasn't hit its planned sample size at the scheduled end dateExtend the runtime rather than approve a partial-sample result5 min
YouYou can do this yourself, no engineering access required.

Final deliverable

A one-page approval memo recommending which plan to run, with the peeking-risk math and validation checklist shown as evidence.

See a reference example
Sample output
Snowflake trial-signup banner test, plan audit memo (excerpt)

Plan A (daily checks, stop at first p<0.05): REJECTED, peeking pattern matches the 40%+ false-positive study cited in CRO training.
Plan B (fixed 22,000 sample, 14-day single read): APPROVED, add 3 more days to reach planned sample before reading results.

Success criteria

You're done when you can:

  • Correctly identifies which plan's stopping rule inflates false positives
  • Applies the Step 5 validation checklist before recommending approval of either plan
  • Recommends extending an incomplete test rather than approving a borderline p-value