The Calibration Call: Reading Glossybox's Test Results
Objective: Given real-shaped result data from a Glossybox creative test, correctly read statistical significance, calculate the budget needed for the next test, and pick the next highest-priority variable to test.
You're the paid social analyst on Glossybox's growth team. This month's creative test just wrapped and the team is deciding what to do next based on the numbers, not gut feel.
Work through the actual result numbers, calculate what the next test needs, and rank what to test next using the lesson's priority list.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, no account friction, handles all the calculations in this project
Paid upgrades (optional, faster/deeper)
Useful once you're running multiple simultaneous tests and need a shared view for the team
No access? Google Sheets with manual daily entry covers a single test cleanly
The process
4 steps
Step 01 of 04
After 7-14 days with enough conversions, check for statistical significance. Below 95%, the difference could be random noise, extend the test or accept neither variant is clearly better.
After 9 days, Variant A (problem-first hook) has 68 conversions at $19.40 CPA; Variant B (product-first hook) has 71 conversions at $19.90 CPA. Platform-reported significance is 88%. What's the correct action?
Procedure
- Log daily conversions and CPA for both variants into the sheet
- Check the platform's reported significance against the 95% threshold
- If below 95%, calculate how many more days at the current daily conversion rate are needed to reach the 50-100 conversions per variant range with more confidence
Day 9 results: Variant A: 68 conversions, $19.40 CPA Variant B: 71 conversions, $19.90 CPA Significance: 88% Decision: Extend test 4-5 more days; 88% is below the 95% threshold and both variants are still short of a clearly separated result.
Healthy
The team extends the test because 88% is below 95%, even though B looks slightly ahead on raw conversions.
Unhealthy
The team declares B the winner on day 9 because it has more raw conversions, ignoring that significance is only 88%.
What this means
A variant with more conversions isn't automatically the real winner; below 95% significance, the gap could still be noise.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Team wants to call a winner below 95% significance | Extend the test window rather than declaring a winner on raw conversion count alone | 5 min |
Step 02 of 04
Aim for 50-100 conversions per variant within your test window. Rough formula: divide the target conversion count by your target cost per purchase to get budget per variant.
Glossybox's target CPA is $20. Next month's test has 3 variants. What total budget is needed to reach at least 50 conversions per variant?
Procedure
- Multiply target CPA ($20) by minimum conversions per variant (50) to get the minimum per-variant budget
- Multiply that per-variant budget by the number of variants (3)
- Add a buffer for the top end of the range (100 conversions) to see the full budget window
Minimum per variant: $20 x 50 = $1,000 Minimum total (3 variants): $1,000 x 3 = $3,000 Upper end per variant: $20 x 100 = $2,000 Budget window: $3,000-$6,000 total for the 3-variant test
Healthy
The team allocates at least $3,000 across the 3 variants before launching.
Unhealthy
The team launches a 3-variant test with a $900 total budget, which can't reach 50 conversions per variant at a $20 target CPA.
What this means
An underfunded test window guarantees an underpowered read, no matter how clean the variable isolation is.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Test budget is below the calculated minimum | Either raise the budget to the minimum or cut the variant count to fit the available budget | 5 min |
Step 03 of 04
Not all creative elements move the needle equally. Test in priority order: format, then hook, then core message angle, then visual style, then CTA text, then headline copy.
Glossybox's last 6 tests were all headline copy variations (priority 6). Video format has never been tested against static images. What should next month's test target instead?
Procedure
- List every test run in the last quarter and the priority tier each one belongs to
- Identify the highest-priority element (lowest number) that has never been tested
- Recommend that element for the next test instead of another headline variant
Tests run this quarter: 6 headline copy tests (priority 6), 0 format tests, 0 hook tests. Recommendation: Test video vs. static image (priority 1) next month — it has the biggest potential variance and has never been tested.
Healthy
The team pivots to a format test even though headline testing feels 'safer' and more familiar.
Unhealthy
The team runs a 7th headline test because the team already has a workflow for it.
What this means
Testing what's easy instead of what's highest-priority wastes months chasing small gains while a bigger lever goes untested.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Recent tests cluster at the bottom of the priority list | Schedule the next test against the highest untested priority tier, not the most familiar one | 30 min |
Step 04 of 04
Change exactly one element between control and challenger. Common elements: hook, visual format, problem vs. product framing, social proof, CTA.
The proposed format test brief says: 'Test UGC video with a customer testimonial vs. static image without a testimonial.' Is this a clean one-variable test?
Procedure
- Compare the two proposed variants line by line
- Identify every element that differs between them, not just the one the brief names
- Flag any additional variable and rewrite the brief to isolate the intended one
Brief as written changes 2 variables: video format AND presence of testimonial. Fix: Test UGC video with testimonial vs. UGC video without testimonial (holds format constant, isolates testimonial); run format vs. static as a separate test.
Healthy
The reviewer catches the hidden second variable before launch and splits it into two clean tests.
Unhealthy
The test launches as written, and a win can't be attributed to format or to the testimonial.
What this means
A brief can look like a one-variable test on the surface while quietly bundling two changes into 'format.'
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A test brief bundles two changes under one label | Split into two sequential single-variable tests before launch | 30 min |
Final deliverable
A completed test-results memo: significance verdict for the current test, required budget for the next test, the top-priority variable to test next, and a corrected one-variable brief.
See a reference example
Allbirds test-results memo (excerpt) CURRENT TEST: Wool Runner hook test, day 8 Significance: 97% — Variant B (problem-first hook) is a confirmed winner Action: Pause Variant A, scale Variant B, document 'problem-first outperforms product-first for cold audiences' NEXT TEST BUDGET: $20 target CPA x 60 conversions x 2 variants = $2,400 minimum NEXT PRIORITY: Format (video vs. static) — never tested, highest potential variance
Success criteria
You're done when you can:
- Correctly recommends extending the test at 88% significance rather than declaring a winner
- Budget calculation matches target CPA x minimum conversions x variant count
- Recommends the highest-priority untested element, not another headline test
- Catches the hidden second variable in the format test brief