The Launch Forecast: Sample Size and Runtime for a Real Test
Objective: Given a real baseline conversion rate, MDE, and daily traffic figure, calculate the required sample size per variant and forecast the exact runtime before a test launches.
You're a growth marketer at Slack proposing a new pricing-page CTA test. Your manager will not approve the launch without a sample size and end date attached to the request.
Apply the lesson's sample size formula to real numbers, then convert the required sample into a calendar runtime using daily traffic.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, no account friction, easy to attach to an approval doc
Paid upgrades (optional, faster/deeper)
Automates the calculation and locks the end date at launch, matching Booking.com's automated-power-analysis approach
No access? Google Sheets calculator covers the same math manually
The process
2 steps
Step 01 of 02
The lesson's Step 3 uses baseline conversion rate, MDE, alpha, and power to compute the visitors needed per variant, roughly 14,000 per variant for a 3% baseline and a 10% relative MDE.
Slack's pricing page converts at 4.2% baseline. You want to detect a 12% relative improvement (to 4.7%) at 95% confidence and 80% power. Using the same calculator logic as the lesson's worked example, is 8,000 visitors per variant enough?
Procedure
- Enter baseline 4.2%, MDE 12% relative (target 4.7%), alpha 0.05, power 80% as labeled input cells
- Cross-check the output against Evan Miller's standard sample size calculator methodology referenced in the lesson
- Record the required sample per variant and total sample across both variants
Sample size calculator, pricing CTA test Baseline: 4.2% Target: 4.7% (12% relative MDE) Alpha / Power: 0.05 / 80% Required n/variant: ~11,400 Total sample: ~22,800
Healthy
The calculated requirement (~11,400 per variant) is compared honestly against the proposed 8,000 and the gap is flagged before launch.
Unhealthy
Rounding down the requirement or assuming 8,000 is 'close enough' because the team is eager to launch this week.
What this means
8,000 per variant is short of the ~11,400 needed; launching anyway means the test is underpowered and a real 12% lift could go undetected.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Proposed sample (8,000/variant) is below the calculated requirement (~11,400/variant) | Either widen the MDE to a level 8,000 can detect, or extend runtime to reach the real requirement | 30 min |
Step 02 of 02
The lesson's Step 4 divides total required sample by daily traffic to the test page to get a runtime in days, and requires at least one full business cycle regardless of when sample is hit.
The pricing page gets 1,050 daily visitors. Given the ~22,800 total sample from Step 1, what launch date and end date should go in the test-plan doc?
Procedure
- Divide total required sample (22,800) by daily traffic (1,050) to get raw days needed
- Round up to the nearest full week and apply the 14-day business-cycle minimum from the lesson
- Write the launch date and locked end date directly into the test-plan doc before requesting approval
Runtime forecast Total sample needed: 22,800 Daily traffic: 1,050 Raw days: 21.7 -> round to 22 days Business-cycle floor: 14 days (already exceeded) Locked end date: Day 22, no early stopping
Healthy
A specific end date is written into the plan before the test launches, and the plan states no peeking-based early stop.
Unhealthy
Leaving the end date as 'until significant' so the team can stop whenever the metric looks good.
What this means
22 days is the honest runtime; anything shorter risks the 40%+ false-positive inflation the lesson documents for early stopping.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Test plan has no fixed end date | Add a locked end date (Day 22) to the approval doc before requesting sign-off | 5 min |
Final deliverable
A one-page test plan showing the required sample size per variant, total sample, and a locked launch/end date.
See a reference example
Adyen checkout-page CTA test plan (excerpt) Baseline: 2.8% | MDE: 15% relative (target 3.2%) Required sample: 9,600/variant, 19,200 total Daily traffic: 1,400 | Runtime: 14 days (business-cycle floor applied) Locked end date: Day 14. No early stopping regardless of interim significance.
Success criteria
You're done when you can:
- Correctly calculates required sample size per variant from the given baseline, MDE, alpha, and power
- Correctly converts total sample into a runtime using daily traffic, applying the 14-day business-cycle floor
- States a locked end date rather than an 'until significant' condition