Skip to content
Academy
Marketing Academy · Field Work●Growth Marketing
MiniAudit· 25 minutes

Ship It or Kill It: Auditing a Suspiciously Good Test Result

Casper Sleep

Objective: Given a supplied A/B test results table (sample sizes, conversion rates, p-value, dates), decide whether the test is actually conclusive or whether peeking, underpowering, or the novelty effect makes the declared winner unreliable.

You're the growth analyst at Casper. A PM is pushing to ship a new checkout variant after 4 days because 'it's already significant.' You've been handed the raw results table and asked to sign off before it ships.

Check the sample size against the pre-registered calculation, check test duration against a full business cycle, and check whether the p-value was read mid-test or at the planned end date.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeImport and audit the raw results table

Free, no account friction, sufficient for a sample-size and peeking-pattern audit

The process

2 steps

Step 01 of 02

Calculating required sample size before launch

The lesson's Step 2 requires calculating sample size from baseline rate, MDE, 95% confidence, and 80% power before launch, not checking it after the fact.

The pre-launch plan called for 30,000 visitors per variant. The results table shows 6,200 in control and 6,050 in variant, collected over 4 days. Is this test adequately powered?

Google Sheets— Open results-table.csv, compare the 'visitors' column against the pre-registered sample size note in row 1.

Procedure

  1. Import results-table.csv into Google Sheets
  2. Sum visitors per variant and compare against the pre-registered target of 30,000
  3. Flag the test as underpowered if actual visitors are under 25% of the target
Sample output
Pre-registered target: 30,000 visitors/variant
Actual after 4 days:
  Control: 6,200 visitors, 248 conversions (4.0%)
  Variant: 6,050 visitors, 278 conversions (4.6%)
Actual sample = 20% of target

Healthy

The test is flagged as underpowered and not shipped until it reaches the pre-registered sample size.

Unhealthy

The PM ships the variant because the dashboard shows a green checkmark, ignoring that the sample is one-fifth of the planned size.

What this means

A p-value computed on 20% of the required sample is not evidence, it is noise that happened to cross a threshold.

So what do I do about it?

SymptomActionEffort
A stakeholder wants to ship based on an early significant readingPull the pre-registered sample size and show the gap in writing before any ship decision5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Peeking and early stopping inflates false positives

The lesson's Common Mistakes section states early stopping inflates the false positive rate from 5% to over 30%, and that the end date must be locked before launch and not moved.

The results table has a 'checked_at' column showing the team looked at p-values on day 1, day 2, day 3, and day 4, and shipped the moment day 4 crossed p < 0.05. What does this pattern tell you about the reported significance?

Google Sheets— Filter the 'checked_at' log column in results-table.csv.

Procedure

  1. Filter the checked_at column to count how many times the team looked at results before shipping
  2. Cross-reference the day the decision was made against the planned 2-week end date
  3. Recompute the effective false positive rate implied by 4 looks (roughly 20%+, not 5%)
Sample output
checked_at log: Day 1 (p=0.31), Day 2 (p=0.14), Day 3 (p=0.09), Day 4 (p=0.048, shipped)
Planned end date: Day 14
Looks before ship: 4

Healthy

The test is declared inconclusive and restarted with a locked end date and a single significance check at the finish line.

Unhealthy

The team treats the day-4 p=0.048 as the real result and ships, unaware that 4 sequential looks pushed the true false-positive rate well above 5%.

What this means

Every additional peek at an unfinished test is another roll of the dice; the reported p-value at the moment of shipping is not the test's true error rate.

So what do I do about it?

SymptomActionEffort
A dashboard shows daily p-value snapshots being screenshotted and sharedDisable public significance dashboards during the test, or restrict access until the locked end date30 min
YouYou can do this yourself, no engineering access required.

Final deliverable

A one-page audit verdict (ship / do not ship / extend test) with the specific sample-size gap and peeking pattern cited as evidence.

See a reference example
Sample output
Allbirds, checkout CTA test audit (excerpt)

VERDICT: DO NOT SHIP

Sample size: 8,400 / 28,000 required visitors per variant (30% of target)
Peeking: 5 significance checks logged before the day-5 ship decision
Recommendation: restart with a locked 2-week end date, no interim dashboard access

Success criteria

You're done when you can:

  • Correctly identifies the sample-size shortfall against the pre-registered target
  • Correctly identifies the peeking pattern from the checked_at log
  • Recommends not shipping, with both the sample-size and peeking evidence cited