Ship It or Kill It: Auditing a Suspiciously Good Test Result
Objective: Given a supplied A/B test results table (sample sizes, conversion rates, p-value, dates), decide whether the test is actually conclusive or whether peeking, underpowering, or the novelty effect makes the declared winner unreliable.
You're the growth analyst at Casper. A PM is pushing to ship a new checkout variant after 4 days because 'it's already significant.' You've been handed the raw results table and asked to sign off before it ships.
Check the sample size against the pre-registered calculation, check test duration against a full business cycle, and check whether the p-value was read mid-test or at the planned end date.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, no account friction, sufficient for a sample-size and peeking-pattern audit
The process
2 steps
Step 01 of 02
The lesson's Step 2 requires calculating sample size from baseline rate, MDE, 95% confidence, and 80% power before launch, not checking it after the fact.
The pre-launch plan called for 30,000 visitors per variant. The results table shows 6,200 in control and 6,050 in variant, collected over 4 days. Is this test adequately powered?
Procedure
- Import results-table.csv into Google Sheets
- Sum visitors per variant and compare against the pre-registered target of 30,000
- Flag the test as underpowered if actual visitors are under 25% of the target
Pre-registered target: 30,000 visitors/variant Actual after 4 days: Control: 6,200 visitors, 248 conversions (4.0%) Variant: 6,050 visitors, 278 conversions (4.6%) Actual sample = 20% of target
Healthy
The test is flagged as underpowered and not shipped until it reaches the pre-registered sample size.
Unhealthy
The PM ships the variant because the dashboard shows a green checkmark, ignoring that the sample is one-fifth of the planned size.
What this means
A p-value computed on 20% of the required sample is not evidence, it is noise that happened to cross a threshold.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A stakeholder wants to ship based on an early significant reading | Pull the pre-registered sample size and show the gap in writing before any ship decision | 5 min |
Step 02 of 02
The lesson's Common Mistakes section states early stopping inflates the false positive rate from 5% to over 30%, and that the end date must be locked before launch and not moved.
The results table has a 'checked_at' column showing the team looked at p-values on day 1, day 2, day 3, and day 4, and shipped the moment day 4 crossed p < 0.05. What does this pattern tell you about the reported significance?
Procedure
- Filter the checked_at column to count how many times the team looked at results before shipping
- Cross-reference the day the decision was made against the planned 2-week end date
- Recompute the effective false positive rate implied by 4 looks (roughly 20%+, not 5%)
checked_at log: Day 1 (p=0.31), Day 2 (p=0.14), Day 3 (p=0.09), Day 4 (p=0.048, shipped) Planned end date: Day 14 Looks before ship: 4
Healthy
The test is declared inconclusive and restarted with a locked end date and a single significance check at the finish line.
Unhealthy
The team treats the day-4 p=0.048 as the real result and ships, unaware that 4 sequential looks pushed the true false-positive rate well above 5%.
What this means
Every additional peek at an unfinished test is another roll of the dice; the reported p-value at the moment of shipping is not the test's true error rate.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A dashboard shows daily p-value snapshots being screenshotted and shared | Disable public significance dashboards during the test, or restrict access until the locked end date | 30 min |
Final deliverable
A one-page audit verdict (ship / do not ship / extend test) with the specific sample-size gap and peeking pattern cited as evidence.
See a reference example
Allbirds, checkout CTA test audit (excerpt) VERDICT: DO NOT SHIP Sample size: 8,400 / 28,000 required visitors per variant (30% of target) Peeking: 5 significance checks logged before the day-5 ship decision Recommendation: restart with a locked 2-week end date, no interim dashboard access
Success criteria
You're done when you can:
- Correctly identifies the sample-size shortfall against the pre-registered target
- Correctly identifies the peeking pattern from the checked_at log
- Recommends not shipping, with both the sample-size and peeking evidence cited