Skip to content
Academy
Marketing Academy · Field Work●Mental Models
CoreForecast· 45 minutes

Forecasting Which Q3 Test Wins Will Hold at 10x Traffic

Zomato

Objective: Given a portfolio of 6 completed A/B tests with sample sizes, p-values, and effect sizes, forecast which results will replicate at national scale and which are noise, using proportional Bayesian updating instead of ranking by p-value alone.

You're a growth analyst at Zomato reviewing a quarter's worth of completed onboarding-flow experiments before recommending which changes get rolled out from a 3-city pilot to the entire country.

Score each test's prior plausibility and evidence strength, then forecast which changes will hold up when traffic scales roughly 10x.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeTrack priors, evidence ratings, and rollout forecasts across the whole test portfolio

Free, and a portfolio table is exactly what a spreadsheet is built for

Google Analytics 4(optional)
FreePull actual session counts per test arm to confirm the sample sizes used in the forecast

Free tier covers this volume of event data

The process

3 steps

Step 01 of 03

Distinguishing a plausible prior from a familiar-but-untested assumption

The lesson warns that some priors deserve high confidence while others just feel familiar. Confusing 'true' with 'familiar' is a common marketing failure.

For each of the 6 tests below, rate your prior confidence (before seeing results) that the change would help, and note whether that confidence is based on real prior data or just 'it seems obviously better.'

Google Sheets— Import onboarding-tests-q3.csv into a new sheet, columns A-F.

Procedure

  1. List all 6 tests: OTP-vs-social login, single-page vs. multi-step signup, restaurant-photo-first vs. menu-first browse, location-permission timing, referral-code prompt placement, welcome-offer banner copy
  2. For each, write a prior confidence (%) and mark the basis as 'prior data' or 'assumption only'
  3. Flag any test where confidence is high but the basis is 'assumption only', these are the highest-risk overreaction candidates
Sample output
TEST                          PRIOR%  BASIS
OTP-vs-social login            55%    assumption only
Single-page signup              70%    prior data (2 past tests)
Photo-first browse               50%    assumption only
Location-permission timing      60%    prior data (1 past test)
Referral-code placement          45%    assumption only
Welcome-offer banner copy        50%    assumption only

Healthy

4 of 6 tests are flagged 'assumption only,' so the team plans to weight their results more cautiously regardless of how large the effect looks.

Unhealthy

All 6 tests get equal trust because 'we ran a proper A/B test,' ignoring that most priors were pure guesses.

What this means

A test built on an assumption-only prior needs stronger evidence to earn the same updated confidence as one built on real prior data.

So what do I do about it?

SymptomActionEffort
Every winning test gets rolled out at the same speed regardless of how it was set upTag each test's prior basis before results are read, and slow-roll assumption-only wins30 min
YouYou can do this yourself, no engineering access required.

Step 02 of 03

Rating evidence strength by sample size and p-value together

Step 2 of the lesson: a test with 500 visitors is much weaker evidence than one with 50,000, regardless of the p-value it produced.

Given the results below, rate each test's evidence strength as weak, moderate, or strong, factoring in BOTH sample size and p-value, not p-value alone.

Google Sheets— Same sheet, columns G-I.

Procedure

  1. Record each test's sample size per arm and p-value from the results tab
  2. Rate evidence strength: weak (n<2,000 or p>0.05), moderate (n=2,000-10,000 and p<0.05), strong (n>10,000 and p<0.01)
  3. Cross-reference against Step 1's prior basis to find tests where weak evidence met an assumption-only prior
Sample output
TEST                          N/ARM   P-VALUE   STRENGTH
OTP-vs-social login            1,400   0.04      weak (small n)
Single-page signup             22,000  0.001     strong
Photo-first browse               980   0.03      weak (small n)
Location-permission timing     18,500  0.02      strong
Referral-code placement        3,200   0.06      weak (not sig.)
Welcome-offer banner copy      15,200  0.15      weak (not sig.)

Healthy

Photo-first browse (weak evidence, assumption-only prior) is flagged as the riskiest 'win' to scale, even though its p-value looks significant.

Unhealthy

Photo-first browse gets rolled out nationally first because its p=0.03 looked cleaner than the 'strong' tests' more complex readouts.

What this means

A statistically significant result from 980 users on an assumption-only prior deserves a small update and a re-test, not a national rollout.

So what do I do about it?

SymptomActionEffort
Rollout order is decided by whichever test 'felt' most convincing in the readout meetingRank rollout order by combined prior-basis + evidence-strength score, not by p-value alone30 min
YouYou can do this yourself, no engineering access required.

Step 03 of 03

Forecasting replication before a national rollout

The lesson's Mistake 2: treating p<0.05 as truth ignores that running 20 tests should produce roughly one false positive by chance alone, exactly what a 6-test portfolio risks.

With 6 tests run this quarter and 4 showing p<0.05, roughly how many of those 4 'wins' would you expect to be false positives by chance alone, and which specific test is the most likely candidate to not replicate at 10x traffic?

Google Sheets— Same sheet, summary row at the bottom.

Procedure

  1. Apply the 1-in-20 false-positive base rate to the 4 significant results to estimate expected false positives
  2. Cross-reference the false-positive risk against Step 1's assumption-only priors and Step 2's weak-evidence flags
  3. Name the single test most likely to fail to replicate, and propose a confirming re-test before national rollout
Sample output
4 of 6 tests hit p<0.05. At a 1-in-20 base rate, roughly 0.2-0.3 of these could be chance, but 'photo-first browse' carries the compounded risk: assumption-only prior AND small sample (n=980) AND is the only 'weak' test that still cleared p<0.05.

FORECAST: Single-page signup and location-permission timing (both strong evidence, prior-data-backed) are safe to roll out nationally now.
Photo-first browse needs a confirming re-test at 10,000+ sessions per arm before national rollout, it is the most likely false positive in this batch.

Healthy

The team ships the 2 strong-evidence tests immediately and holds the weak-evidence 'win' for a confirming test before it touches the whole country.

Unhealthy

All 4 significant tests roll out to 100% of national traffic simultaneously because each individually cleared p<0.05.

What this means

A portfolio view catches false-positive risk that reading each test in isolation misses entirely.

So what do I do about it?

SymptomActionEffort
Multiple 'winning' tests from the same quarter later get quietly revertedReview each quarter's wins as a portfolio, not one at a time, and flag the false-positive-risk math before rollout30 min
YouYou can do this yourself, no engineering access required.

Final deliverable

A portfolio forecast table ranking all 6 tests by prior basis + evidence strength, with an explicit national-rollout recommendation per test.

See a reference example
Sample output
Squarespace, Q2 onboarding test portfolio, forecast excerpt

Test: 'Free trial length: 14 vs. 30 days'
Prior: 65% (prior data from 2023 pricing test)
Evidence: n=19,000/arm, p=0.008, strong
Forecast: High confidence this replicates at scale. Roll out nationally.

Test: 'Template gallery sort order'
Prior: 50% (assumption only)
Evidence: n=1,100/arm, p=0.04, weak (small n)
Forecast: Likely a false positive. Re-test at 10x sample before touching the default.

Success criteria

You're done when you can:

  • Every test gets a prior-basis tag (prior data vs. assumption only)
  • Evidence strength rating uses both sample size and p-value
  • Applies the 1-in-20 false-positive base rate to the set of significant results
  • Names a specific highest-risk test with a reasoned justification, not just the lowest p-value