Pre-Register the Test: Building a Properly Powered A/B Test Plan
Objective: Given a real baseline conversion rate and traffic volume, build a complete pre-registered test plan: hypothesis, sample size calculation, primary and guardrail metrics, and a locked end date, before any variant is built.
You're the growth lead at Nykaa. The team wants to test a simplified 2-step checkout against the current 4-step checkout. Before anyone touches design, you own the test plan.
Turn a vague 'let's test the checkout' idea into a fully pre-registered plan with a real sample size number, a locked duration, and guardrail metrics that would catch a hidden regression.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, shareable, sufficient for hypothesis, sample size math, and metric definitions
Paid upgrades (optional, faster/deeper)
Handles randomization, sample size tracking, and guardrail dashboards without manual tracking
No access? Google Sheets can log manual daily counts if no testing platform is available yet
The process
3 steps
Step 01 of 03
The lesson's Step 1 requires the template 'Because we observed [data], we believe [change] will cause [metric] to improve by [amount] for [audience].' If you cannot fill every blank, it is a guess, not a hypothesis.
Analytics shows 38% of users abandon at the 3rd of 4 checkout steps, mostly on mobile. Write a hypothesis that fills every blank in the template using this data.
Procedure
- Pull the step-by-step checkout funnel drop-off from the analytics export
- Identify the specific step and audience segment with the highest abandonment
- Fill in every blank of the hypothesis template with a specific, falsifiable claim
Because we observed 38% of mobile users abandon at the shipping-address step (step 3 of 4), we believe collapsing steps 2-4 into a single scrollable screen will cause checkout completion rate to improve by 8% for mobile users.
Healthy
Every blank in the template is filled with a specific number and audience, making the hypothesis falsifiable.
Unhealthy
The hypothesis reads 'we believe a simpler checkout will convert better,' which cannot fail and is not testable.
What this means
A hypothesis that cannot be wrong is not a hypothesis; it is a preference restated as a claim.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A test brief has no specific numbers in the hypothesis line | Reject the brief and require a filled-in template before sample size work begins | 5 min |
Step 02 of 03
The lesson's Step 2 requires four inputs before launch: baseline conversion rate, minimum detectable effect, 95% confidence, and 80% power.
Baseline checkout completion is 62%. You want to detect the 8% relative lift from your hypothesis (62% to 66.96%) at 95% confidence and 80% power. Using a sample size calculator, roughly how many visitors per variant do you need, and does current traffic support it?
Procedure
- Enter baseline rate 62% and MDE 8% relative into a sample size calculator (e.g. Evan Miller's)
- Record the required sample size per variant, typically several thousand for an 8% relative lift on a 62% baseline
- Divide by current daily checkout traffic to estimate how many days the test needs to run
Baseline: 62% MDE: 8% relative Confidence: 95% Power: 80% Required: ~3,900 visitors per variant Current daily checkout traffic: ~650/day total, ~325/variant Estimated runtime: ~12 days minimum
Healthy
The plan locks a 14-day minimum runtime, covering both the calculated sample size and a full two-week business cycle.
Unhealthy
The team launches with no runtime estimate and starts checking results after 2 days because 'it felt long enough.'
What this means
Sample size and business-cycle duration are two separate checks; a plan needs both satisfied, not whichever comes first.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A test plan has no calculated sample size or runtime estimate | Block the launch until the calculator output and a locked end date are both in the plan doc | 30 min |
Step 03 of 03
The lesson's Step 3 requires committing to one primary decision metric before launch, plus guardrail metrics tracked to catch regressions without deciding the winner.
Checkout completion rate is the obvious primary metric. What guardrail metrics would catch a hidden regression, like the 2-step checkout lifting completion but quietly increasing return/refund requests?
Procedure
- Set checkout completion rate as the single primary metric
- Add revenue per session as a guardrail against a low-value completion spike
- Add 30-day return rate and support ticket volume as guardrails against a rushed, error-prone flow
PRIMARY: Checkout completion rate GUARDRAILS: Revenue per session (catches low-AOV completions) 30-day return rate (catches rushed/mistaken orders) Checkout-related support tickets (catches confusing new flow)
Healthy
Guardrails are defined and monitored before launch so a completion-rate win that breaks returns or support volume gets caught, not shipped.
Unhealthy
Only completion rate is tracked; the team ships an 8% lift and discovers a return-rate spike two months later.
What this means
A win on the primary metric that breaks a guardrail is not a win, per the lesson's own framing.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A test plan lists only one metric with no guardrails | Add at least one revenue and one downstream quality guardrail before the plan is approved | 30 min |
Final deliverable
A complete pre-registered test plan document: filled hypothesis template, sample size calculation with inputs and output, primary metric, 3 guardrail metrics, and a locked minimum runtime.
See a reference example
Lenskart, checkout redesign test plan (excerpt) HYPOTHESIS: Because we observed 41% of users abandon at the prescription-upload step, we believe making upload optional-until-payment will cause checkout completion to improve by 6% for first-time buyers. SAMPLE SIZE: baseline 58%, MDE 6% relative, 95%/80% power = ~4,800/variant RUNTIME: 15 days minimum (traffic-bound) PRIMARY: Checkout completion rate GUARDRAILS: Revenue per session, prescription-verification failure rate, support tickets
Success criteria
You're done when you can:
- Hypothesis fills every blank of the template with specific numbers
- Sample size is calculated from real baseline/MDE inputs, not guessed
- Exactly one primary metric is named, with at least 2 guardrail metrics
- Runtime accounts for both sample size and the full-business-cycle minimum