Name Your Prior: A Calibration Drill on a Real Test Result
Objective: Given a real A/B test result with sample size and p-value, state a numeric prior before seeing the outcome framing, then apply the lesson's three-step update process to land on a calibrated new belief.
You're a lifecycle marketing analyst at Klaviyo evaluating whether a new abandoned-cart email subject line should become the default template offered to every client account.
State your prior with a number, weigh the evidence honestly, and update proportionally instead of shipping on the raw win.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, no account friction, and forces the prior to be written down before the result is seen
The process
2 steps
Step 01 of 02
The lesson's Step 1 says: before evidence lands, put a percentage on what you already believed and how strongly.
Before you look at the test data below, what is your prior? Current subject line: 'You left something behind.' Proposed: 'Still thinking it over? Your cart's waiting.' Internal copywriters are confident the new one wins because it feels more conversational. State a prior probability that the new line beats the control on open rate, before reading further.
Procedure
- Write your prior probability (e.g. 55%) in cell B2 before opening the results tab
- Open the results tab: control 21.4% open rate (n=8,200), variant 23.1% open rate (n=8,150), p=0.09
- Note that 'conversational' subject line rewrites have a mixed track record in the company's own test archive: 6 wins, 9 losses over two years
PRIOR (before results): 55% confident new line wins RESULT Control: 21.4% open rate, n=8,200 Variant: 23.1% open rate, n=8,150 p-value: 0.09 (not significant at p<0.05) HOUSE TRACK RECORD Conversational rewrites: 6 wins / 15 tested (40%) over 2 years
Healthy
Recognizing that a p=0.09 result on a subject line category with a 40% historical win rate is weak evidence against a roughly 50/50 prior, and holding off on a full rollout.
Unhealthy
Treating the 1.7-point open-rate lift as a confirmed win because it 'felt right' to the copywriters, and rolling it out to every client account.
What this means
A non-significant result on a category with a below-50% historical win rate should nudge your prior only slightly, not flip it.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A test result with p>0.05 gets shipped as 'the new default' anyway | Require p<0.05 or a second confirming test before any subject-line default changes | 5 min |
Step 02 of 02
Step 2 and 3 of the lesson: rate how much the evidence actually tells you, then move your belief proportionally, not to 0% or 100%.
Given the prior you stated, the p=0.09 result, and the 40% historical win rate for this style of rewrite, what is your updated probability that the new subject line is genuinely better? Propose a next action.
Procedure
- Rate the evidence strength as weak, moderate, or strong given n≈8,200 per arm and p=0.09
- Write your updated probability in cell B3, showing the move from your Step 1 prior
- Recommend either: ship now, re-run at 3x sample size, or discard, with a one-sentence reason
UPDATE Prior: 55% Evidence strength: weak-to-moderate (p=0.09, in-line with historical base rate) Updated belief: 58% RECOMMENDATION: Re-run at ~25,000 sessions per arm before touching the default template. Reason: the move from 55% to 58% is too small to justify rolling out to every client account.
Healthy
An 8,200-session test with p=0.09 moves the prior by single digits, and the team re-tests at scale before touching a shared default.
Unhealthy
The same result gets summarized in a Slack message as 'new subject line wins' and shipped as the default that day.
What this means
Small proportional updates are the correct output of weak evidence, not a failure of the test.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Team debates feel like 'my opinion vs. your opinion' with no shared number | Require every test readout to state prior%, evidence strength, and updated% in writing | 30 min |
Final deliverable
A one-page calibration memo: stated prior, evidence-strength rating, updated probability, and a ship / re-test / discard decision.
See a reference example
Warby Parker, homepage hero-image test, calibration memo (excerpt) Prior: 40% confident the lifestyle photo beats the product-on-white shot Evidence: n=14,000/arm, +2.1pp add-to-cart rate, p=0.03 Evidence strength: moderate (significant, but a single test) Updated belief: 62% Decision: Ship to 50% of traffic for one more cycle before full rollout, given only one confirming test exists.
Success criteria
You're done when you can:
- States a numeric prior before citing the result
- Rates evidence strength using sample size and p-value, not p-value alone
- Updated probability moves proportionally, not to 0% or 100%
- Recommendation matches the size of the update (small update -> no full rollout)