Teardown: The Test Writeup That Broke Every Rule
Objective: Given a real-style CRO test writeup, identify which statistical pitfalls (peeking, multiple comparisons, Simpson's Paradox) it committed and which claims are actually sound.
You're a growth analyst at Freshworks. A teammate is about to email leadership a test writeup recommending a full rollout. Your job is to catch anything wrong before it ships.
Read the underlying specimen writeup snippets and flag every statistical defect, citing the lesson concept each one violates.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, no account friction, sufficient for a written peer review
The process
Specimens to review
What's wrong with this recommendation, and what should the team have done instead?
Day 3 update: Treatment is already beating control by +4% (p=0.04). We've been checking the dashboard every morning since launch. Given the strong early signal, recommend we call this test now and ship to 100% today, no need to wait for the original 14-day/25,000-visitor plan.
Specimen: synthetic, realistic
Is declaring a win on average order value here justified? What's the actual issue?
We tracked 6 metrics on this test: primary conversion rate, add-to-cart rate, average order value, email signup rate, page scroll depth, and time on page. Average order value came back significant at p=0.03, so we're calling this a win on AOV even though conversion rate itself was flat.
Specimen: synthetic, realistic
Should the team kill treatment based on this combined number? What else needs checking first?
Combined across the full test, treatment converted at 1.20% versus control's 1.68%, so we're planning to kill treatment. Note: we ramped treatment traffic from 1% on day 1 to 50% by day 2 to speed up data collection.
Specimen: synthetic, realistic
Final deliverable
A written peer-review comment thread flagging each statistical defect, its severity, and what should happen instead before the writeup ships.
See a reference example
Wise CRO peer review (excerpt) Defect 1 (critical): Stopped test on day 3 after daily peeking hit p=0.04. Fix: hold to the pre-registered 14-day/25,000-visitor plan. Defect 2 (critical): Declared win on AOV, a secondary metric, while the primary metric (conversion rate) was flat across 6 tracked metrics. Fix: report against the primary metric only. Defect 3 (critical): Traffic ramped 1% -> 50% mid-test; combined number reverses the per-day result. Fix: break out by allocation period before trusting the total.
Success criteria
You're done when you can:
- Identifies all 3 statistical defects and correctly labels which lesson concept each one violates
- Does not flag either distractor per item as a real defect
- Recommends the correct fix for each (pre-registered plan, primary metric, stable allocation)