Trust It or Toss It: Calibrating Three Past Test Calls
Objective: Given three past decisions pulled from a real test log, predict whether each one was actually statistically sound before checking the numbers, then compare your call against the lesson's rules.
Klaviyo's client, a mid-size ecommerce brand, wants its test log audited before its next quarterly planning session. Three decisions in the log look confident. Your job is to predict which ones actually held up before you check the underlying numbers.
Predict first (sound / not sound), then check your prediction against the numbers and the lesson's rule that applies.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, matches how the brand's real test log is already kept
Paid upgrades (optional, faster/deeper)
Its multivariate reporting breaks out CTR by variant directly, useful once you need to re-check a mismeasured test
No access? Export the raw per-variant click data from whatever platform ran the original test
The process
2 steps
Step 01 of 02
The lesson's minimum viable test needs 200-300 opens per variant; below that, a 20% difference between variants could easily be random chance.
Log row 1: "Sender name test, Version A: 180 opens, Version B: 224 opens. Winner: B, applied to all future sends." Predict: was this decision sound, before you check the math?
Procedure
- State your prediction (sound / not sound) before doing the check
- Compare both variants' opens against the lesson's 200-300 minimum
- Reveal: Version A's 180 opens falls under the minimum, so the 224 vs. 180 gap is not reliable evidence
Version A: 180 opens. Version B: 224 opens. Both under or barely at the 200-300 floor.
Healthy
A test log entry with both variants clearing 200-300 opens before a winner is applied everywhere
Unhealthy
A winner locked in from a variant that never reached the opens minimum
What this means
180 opens is below the lesson's own floor, this decision was not statistically sound no matter which variant technically had more opens.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A past winner was applied list-wide from an under-powered test | Re-run the sender-name test with both variants held open until each clears 200-300 opens | 30 min |
Step 02 of 02
The lesson's rule: match the metric to the element tested. A CTA test should be measured on click-through rate, not open rate, optimizing for the wrong metric can make a test 'win' while revenue doesn't move.
Log row 2: "CTA button text test, measured by open rate. Version A ('Start my trial') had a 2% higher open rate, declared winner." Predict: was this decision sound?
Procedure
- State your prediction before checking
- Identify which element was actually tested (the CTA button)
- Identify which metric was used to judge it (open rate) and whether that's the right pairing
Element tested: CTA button text. Metric used to declare a winner: open rate.
Healthy
A CTA test judged on click-through rate, the metric the CTA can actually move
Unhealthy
A CTA test judged on open rate, a metric the CTA button can't influence since it isn't seen until after the email is already open
What this means
Open rate is set by the subject line and sender name, not the CTA button; this decision measured the wrong thing and the '2% higher open rate' is very likely unrelated to which CTA text won.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A CTA test was declared using open-rate data | Re-pull the same test's click-through-rate numbers per variant before trusting the winner | 5 min |
Final deliverable
A two-row calibration note: your prediction, the revealed answer, and the one lesson rule each log entry actually broke.
See a reference example
Mailchimp test-log audit note (reference) Row 1 (sender name): predicted sound, actual: NOT sound, 180 opens under the 200-300 floor. Row 2 (CTA test): predicted sound, actual: NOT sound, wrong metric (open rate) used to judge a CTA change.
Success criteria
You're done when you can:
- States a prediction before checking either row
- Correctly identifies the sample-size floor violation in row 1
- Correctly identifies the metric-mismatch violation in row 2