Skip to content
Academy
Marketing Academy · Field Work●Mental Models
MiniReverse-Engineer· 25 minutes

Name Your Prior: A Calibration Drill on a Real Test Result

Klaviyo

Objective: Given a real A/B test result with sample size and p-value, state a numeric prior before seeing the outcome framing, then apply the lesson's three-step update process to land on a calibrated new belief.

You're a lifecycle marketing analyst at Klaviyo evaluating whether a new abandoned-cart email subject line should become the default template offered to every client account.

State your prior with a number, weigh the evidence honestly, and update proportionally instead of shipping on the raw win.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeLog the prior, the result, and the updated probability side by side

Free, no account friction, and forces the prior to be written down before the result is seen

The process

2 steps

Step 01 of 02

Naming your prior with a probability

The lesson's Step 1 says: before evidence lands, put a percentage on what you already believed and how strongly.

Before you look at the test data below, what is your prior? Current subject line: 'You left something behind.' Proposed: 'Still thinking it over? Your cart's waiting.' Internal copywriters are confident the new one wins because it feels more conversational. State a prior probability that the new line beats the control on open rate, before reading further.

Google Sheets— Open the shared 'Q3 Email Tests' sheet, tab 'abandoned-cart-subject-line'.

Procedure

  1. Write your prior probability (e.g. 55%) in cell B2 before opening the results tab
  2. Open the results tab: control 21.4% open rate (n=8,200), variant 23.1% open rate (n=8,150), p=0.09
  3. Note that 'conversational' subject line rewrites have a mixed track record in the company's own test archive: 6 wins, 9 losses over two years
Sample output
PRIOR (before results): 55% confident new line wins

RESULT
  Control:  21.4% open rate, n=8,200
  Variant:  23.1% open rate, n=8,150
  p-value:  0.09 (not significant at p<0.05)

HOUSE TRACK RECORD
  Conversational rewrites: 6 wins / 15 tested (40%) over 2 years

Healthy

Recognizing that a p=0.09 result on a subject line category with a 40% historical win rate is weak evidence against a roughly 50/50 prior, and holding off on a full rollout.

Unhealthy

Treating the 1.7-point open-rate lift as a confirmed win because it 'felt right' to the copywriters, and rolling it out to every client account.

What this means

A non-significant result on a category with a below-50% historical win rate should nudge your prior only slightly, not flip it.

So what do I do about it?

SymptomActionEffort
A test result with p>0.05 gets shipped as 'the new default' anywayRequire p<0.05 or a second confirming test before any subject-line default changes5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Weighing evidence strength before updating

Step 2 and 3 of the lesson: rate how much the evidence actually tells you, then move your belief proportionally, not to 0% or 100%.

Given the prior you stated, the p=0.09 result, and the 40% historical win rate for this style of rewrite, what is your updated probability that the new subject line is genuinely better? Propose a next action.

Google Sheets— Same sheet, tab 'abandoned-cart-subject-line', cell B3.

Procedure

  1. Rate the evidence strength as weak, moderate, or strong given n≈8,200 per arm and p=0.09
  2. Write your updated probability in cell B3, showing the move from your Step 1 prior
  3. Recommend either: ship now, re-run at 3x sample size, or discard, with a one-sentence reason
Sample output
UPDATE
  Prior: 55%
  Evidence strength: weak-to-moderate (p=0.09, in-line with historical base rate)
  Updated belief: 58%

RECOMMENDATION: Re-run at ~25,000 sessions per arm before touching the default template.
Reason: the move from 55% to 58% is too small to justify rolling out to every client account.

Healthy

An 8,200-session test with p=0.09 moves the prior by single digits, and the team re-tests at scale before touching a shared default.

Unhealthy

The same result gets summarized in a Slack message as 'new subject line wins' and shipped as the default that day.

What this means

Small proportional updates are the correct output of weak evidence, not a failure of the test.

So what do I do about it?

SymptomActionEffort
Team debates feel like 'my opinion vs. your opinion' with no shared numberRequire every test readout to state prior%, evidence strength, and updated% in writing30 min
YouYou can do this yourself, no engineering access required.

Final deliverable

A one-page calibration memo: stated prior, evidence-strength rating, updated probability, and a ship / re-test / discard decision.

See a reference example
Sample output
Warby Parker, homepage hero-image test, calibration memo (excerpt)

Prior: 40% confident the lifestyle photo beats the product-on-white shot
Evidence: n=14,000/arm, +2.1pp add-to-cart rate, p=0.03
Evidence strength: moderate (significant, but a single test)
Updated belief: 62%
Decision: Ship to 50% of traffic for one more cycle before full rollout, given only one confirming test exists.

Success criteria

You're done when you can:

  • States a numeric prior before citing the result
  • Rates evidence strength using sample size and p-value, not p-value alone
  • Updated probability moves proportionally, not to 0% or 100%
  • Recommendation matches the size of the update (small update -> no full rollout)