Skip to content
Academy
Marketing Academy · Field Work●Email & Lifecycle
MiniForecast· 20 minutes

Trust It or Toss It: Calibrating Three Past Test Calls

Klaviyo

Objective: Given three past decisions pulled from a real test log, predict whether each one was actually statistically sound before checking the numbers, then compare your call against the lesson's rules.

Klaviyo's client, a mid-size ecommerce brand, wants its test log audited before its next quarterly planning session. Three decisions in the log look confident. Your job is to predict which ones actually held up before you check the underlying numbers.

Predict first (sound / not sound), then check your prediction against the numbers and the lesson's rule that applies.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeHold the test log and record each prediction next to the revealed answer

Free, matches how the brand's real test log is already kept

Paid upgrades (optional, faster/deeper)

ActiveCampaign(optional)
PaidPull the underlying per-variant click-through-rate numbers for the misjudged CTA test

Its multivariate reporting breaks out CTR by variant directly, useful once you need to re-check a mismeasured test

No access? Export the raw per-variant click data from whatever platform ran the original test

The process

2 steps

Step 01 of 02

Sample Size Minimums

The lesson's minimum viable test needs 200-300 opens per variant; below that, a 20% difference between variants could easily be random chance.

Log row 1: "Sender name test, Version A: 180 opens, Version B: 224 opens. Winner: B, applied to all future sends." Predict: was this decision sound, before you check the math?

Google Sheets— The brand's shared test log spreadsheet, one row per past test

Procedure

  1. State your prediction (sound / not sound) before doing the check
  2. Compare both variants' opens against the lesson's 200-300 minimum
  3. Reveal: Version A's 180 opens falls under the minimum, so the 224 vs. 180 gap is not reliable evidence
Sample output
Version A: 180 opens. Version B: 224 opens. Both under or barely at the 200-300 floor.

Healthy

A test log entry with both variants clearing 200-300 opens before a winner is applied everywhere

Unhealthy

A winner locked in from a variant that never reached the opens minimum

What this means

180 opens is below the lesson's own floor, this decision was not statistically sound no matter which variant technically had more opens.

So what do I do about it?

SymptomActionEffort
A past winner was applied list-wide from an under-powered testRe-run the sender-name test with both variants held open until each clears 200-300 opens30 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Matching Metric to Element

The lesson's rule: match the metric to the element tested. A CTA test should be measured on click-through rate, not open rate, optimizing for the wrong metric can make a test 'win' while revenue doesn't move.

Log row 2: "CTA button text test, measured by open rate. Version A ('Start my trial') had a 2% higher open rate, declared winner." Predict: was this decision sound?

Google Sheets— The same shared test log spreadsheet

Procedure

  1. State your prediction before checking
  2. Identify which element was actually tested (the CTA button)
  3. Identify which metric was used to judge it (open rate) and whether that's the right pairing
Sample output
Element tested: CTA button text. Metric used to declare a winner: open rate.

Healthy

A CTA test judged on click-through rate, the metric the CTA can actually move

Unhealthy

A CTA test judged on open rate, a metric the CTA button can't influence since it isn't seen until after the email is already open

What this means

Open rate is set by the subject line and sender name, not the CTA button; this decision measured the wrong thing and the '2% higher open rate' is very likely unrelated to which CTA text won.

So what do I do about it?

SymptomActionEffort
A CTA test was declared using open-rate dataRe-pull the same test's click-through-rate numbers per variant before trusting the winner5 min
YouYou can do this yourself, no engineering access required.

Final deliverable

A two-row calibration note: your prediction, the revealed answer, and the one lesson rule each log entry actually broke.

See a reference example
Sample output
Mailchimp test-log audit note (reference)

Row 1 (sender name): predicted sound, actual: NOT sound, 180 opens under the 200-300 floor.
Row 2 (CTA test): predicted sound, actual: NOT sound, wrong metric (open rate) used to judge a CTA change.

Success criteria

You're done when you can:

  • States a prediction before checking either row
  • Correctly identifies the sample-size floor violation in row 1
  • Correctly identifies the metric-mismatch violation in row 2