Forecasting Which Q3 Test Wins Will Hold at 10x Traffic
Objective: Given a portfolio of 6 completed A/B tests with sample sizes, p-values, and effect sizes, forecast which results will replicate at national scale and which are noise, using proportional Bayesian updating instead of ranking by p-value alone.
You're a growth analyst at Zomato reviewing a quarter's worth of completed onboarding-flow experiments before recommending which changes get rolled out from a 3-city pilot to the entire country.
Score each test's prior plausibility and evidence strength, then forecast which changes will hold up when traffic scales roughly 10x.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, and a portfolio table is exactly what a spreadsheet is built for
Free tier covers this volume of event data
The process
3 steps
Step 01 of 03
The lesson warns that some priors deserve high confidence while others just feel familiar. Confusing 'true' with 'familiar' is a common marketing failure.
For each of the 6 tests below, rate your prior confidence (before seeing results) that the change would help, and note whether that confidence is based on real prior data or just 'it seems obviously better.'
Procedure
- List all 6 tests: OTP-vs-social login, single-page vs. multi-step signup, restaurant-photo-first vs. menu-first browse, location-permission timing, referral-code prompt placement, welcome-offer banner copy
- For each, write a prior confidence (%) and mark the basis as 'prior data' or 'assumption only'
- Flag any test where confidence is high but the basis is 'assumption only', these are the highest-risk overreaction candidates
TEST PRIOR% BASIS OTP-vs-social login 55% assumption only Single-page signup 70% prior data (2 past tests) Photo-first browse 50% assumption only Location-permission timing 60% prior data (1 past test) Referral-code placement 45% assumption only Welcome-offer banner copy 50% assumption only
Healthy
4 of 6 tests are flagged 'assumption only,' so the team plans to weight their results more cautiously regardless of how large the effect looks.
Unhealthy
All 6 tests get equal trust because 'we ran a proper A/B test,' ignoring that most priors were pure guesses.
What this means
A test built on an assumption-only prior needs stronger evidence to earn the same updated confidence as one built on real prior data.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Every winning test gets rolled out at the same speed regardless of how it was set up | Tag each test's prior basis before results are read, and slow-roll assumption-only wins | 30 min |
Step 02 of 03
Step 2 of the lesson: a test with 500 visitors is much weaker evidence than one with 50,000, regardless of the p-value it produced.
Given the results below, rate each test's evidence strength as weak, moderate, or strong, factoring in BOTH sample size and p-value, not p-value alone.
Procedure
- Record each test's sample size per arm and p-value from the results tab
- Rate evidence strength: weak (n<2,000 or p>0.05), moderate (n=2,000-10,000 and p<0.05), strong (n>10,000 and p<0.01)
- Cross-reference against Step 1's prior basis to find tests where weak evidence met an assumption-only prior
TEST N/ARM P-VALUE STRENGTH OTP-vs-social login 1,400 0.04 weak (small n) Single-page signup 22,000 0.001 strong Photo-first browse 980 0.03 weak (small n) Location-permission timing 18,500 0.02 strong Referral-code placement 3,200 0.06 weak (not sig.) Welcome-offer banner copy 15,200 0.15 weak (not sig.)
Healthy
Photo-first browse (weak evidence, assumption-only prior) is flagged as the riskiest 'win' to scale, even though its p-value looks significant.
Unhealthy
Photo-first browse gets rolled out nationally first because its p=0.03 looked cleaner than the 'strong' tests' more complex readouts.
What this means
A statistically significant result from 980 users on an assumption-only prior deserves a small update and a re-test, not a national rollout.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Rollout order is decided by whichever test 'felt' most convincing in the readout meeting | Rank rollout order by combined prior-basis + evidence-strength score, not by p-value alone | 30 min |
Step 03 of 03
The lesson's Mistake 2: treating p<0.05 as truth ignores that running 20 tests should produce roughly one false positive by chance alone, exactly what a 6-test portfolio risks.
With 6 tests run this quarter and 4 showing p<0.05, roughly how many of those 4 'wins' would you expect to be false positives by chance alone, and which specific test is the most likely candidate to not replicate at 10x traffic?
Procedure
- Apply the 1-in-20 false-positive base rate to the 4 significant results to estimate expected false positives
- Cross-reference the false-positive risk against Step 1's assumption-only priors and Step 2's weak-evidence flags
- Name the single test most likely to fail to replicate, and propose a confirming re-test before national rollout
4 of 6 tests hit p<0.05. At a 1-in-20 base rate, roughly 0.2-0.3 of these could be chance, but 'photo-first browse' carries the compounded risk: assumption-only prior AND small sample (n=980) AND is the only 'weak' test that still cleared p<0.05. FORECAST: Single-page signup and location-permission timing (both strong evidence, prior-data-backed) are safe to roll out nationally now. Photo-first browse needs a confirming re-test at 10,000+ sessions per arm before national rollout, it is the most likely false positive in this batch.
Healthy
The team ships the 2 strong-evidence tests immediately and holds the weak-evidence 'win' for a confirming test before it touches the whole country.
Unhealthy
All 4 significant tests roll out to 100% of national traffic simultaneously because each individually cleared p<0.05.
What this means
A portfolio view catches false-positive risk that reading each test in isolation misses entirely.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Multiple 'winning' tests from the same quarter later get quietly reverted | Review each quarter's wins as a portfolio, not one at a time, and flag the false-positive-risk math before rollout | 30 min |
Final deliverable
A portfolio forecast table ranking all 6 tests by prior basis + evidence strength, with an explicit national-rollout recommendation per test.
See a reference example
Squarespace, Q2 onboarding test portfolio, forecast excerpt Test: 'Free trial length: 14 vs. 30 days' Prior: 65% (prior data from 2023 pricing test) Evidence: n=19,000/arm, p=0.008, strong Forecast: High confidence this replicates at scale. Roll out nationally. Test: 'Template gallery sort order' Prior: 50% (assumption only) Evidence: n=1,100/arm, p=0.04, weak (small n) Forecast: Likely a false positive. Re-test at 10x sample before touching the default.
Success criteria
You're done when you can:
- Every test gets a prior-basis tag (prior data vs. assumption only)
- Evidence strength rating uses both sample size and p-value
- Applies the 1-in-20 false-positive base rate to the set of significant results
- Names a specific highest-risk test with a reasoned justification, not just the lowest p-value