Read the Test Results: Picking Allbirds' Next Subject Line Hypothesis
Objective: Given real-shaped A/B test results from two completed subject line tests, apply the lesson's one-variable-at-a-time rule to declare a winner and design the next hypothesis correctly.
You're the email marketer at Allbirds reviewing last month's subject line tests before planning next month's test calendar.
Two tests ran last month. Decide which one is trustworthy, declare the winner where valid, and write the next hypothesis to test.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free tier includes built-in A/B testing for subject lines
Free, matches the lesson's recommended tracking spreadsheet
The process
2 steps
Step 01 of 02
A valid A/B test changes exactly one variable between version A and B. Changing wording and length together makes the result impossible to attribute.
Test 1: A = 'Your shoes are back in stock' (35% open) vs B = 'Restocked: the shoe everyone's been asking about, get yours today' (38% open). Test 2: A = 'New arrivals' (29% open) vs B = 'What's new?' (41% open). Which test result can you actually trust, and why?
Procedure
- Check whether each test changed exactly one variable
- Note that Test 1 changed both wording and length at once, its result cannot be attributed to either change alone
- Confirm Test 2 changed only the format (statement vs. question) with matching length, making its winner trustworthy
Test 1: INVALID (confounded: wording + length both changed) Test 2: VALID (single variable: statement vs. question) -> Question wins, 41% vs 29%
Healthy
A test where only one variable changed between A and B.
Unhealthy
Declaring a winner from a test that changed wording and length simultaneously.
What this means
Test 1's 3-point lift could be from the extra detail, the length, or both, there is no way to know which.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A test changed two variables at once | Discard the result and rerun with only one variable different | 5 min |
Step 02 of 02
A simple testing rhythm: pick one hypothesis, write A and B identical except for that one variable, split 45/45/10, wait 4 hours, send the winner to the held-back 10%.
Using Test 2's confirmed result (questions beat statements), write the next single-variable hypothesis and test pair to run.
Procedure
- Write a hypothesis that builds on the confirmed finding without introducing a second variable
- Draft version A and B identical except for that one new variable
- Set the split to 45% A / 45% B / 10% held back for the winner send
Hypothesis: "Personalized questions outperform generic questions." A: "What's new?" B: "Surya, what's new for you?" Split: 45/45/10, check at 4 hours
Healthy
A hypothesis that changes exactly one new variable from the last confirmed win.
Unhealthy
A hypothesis that also changes length or adds an emoji at the same time.
What this means
Building tests sequentially on confirmed single-variable wins is how a real swipe file of audience-specific patterns gets built over 10+ tests.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| The next test changes personalization and adds an emoji together | Split into two separate sequential tests | 5 min |
Final deliverable
A written verdict on both tests (trustworthy or confounded) plus a properly single-variable hypothesis and A/B pair for the next test.
See a reference example
ThredUp test log (excerpt) Test 4: A 'Free shipping today' vs B 'Free shipping, today only' -> CONFOUNDED (added urgency + changed length) Test 5: A 'New drops' vs B 'New drops just for you' -> VALID (personalization only), B wins 44% vs 33%
Success criteria
You're done when you can:
- Correctly identifies Test 1 as confounded and Test 2 as valid
- Next hypothesis changes exactly one variable from the confirmed win
- Split proposed matches the lesson's 45/45/10 rhythm