The Calibration Check: Scoring Your Own Odds
Objective: Given 5 real marketing scenarios with known eventual outcomes, assign a probability estimate to each before the outcome is revealed, then score how well your stated confidence actually matches your hit rate.
You're a lifecycle marketer at Zendesk building the habit of writing down odds before every campaign bet, instead of only ever remembering the ones that worked.
Score five campaign scenarios with an explicit percentage before seeing the outcome, then check whether your 70% calls actually land around 70% of the time.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, no account friction, and hiding/locking columns is enough structure for this drill
The process
2 steps
Step 01 of 02
The lesson's calibration question is: would you bet a month's budget on this at these odds? A probability you can't say out loud with a straight face is theater, not an estimate.
Given only the campaign brief (channel, audience, and the account's past benchmark performance), what single percentage chance would you give this push-notification campaign of beating its 12% open-rate benchmark?
Procedure
- Read each of the 5 campaign briefs without unhiding column C
- Write one number, not a word like 'likely' or 'probably', in column B for every row
- Lock or protect column B before unhiding outcomes, no revising after the fact
Scenario Your odds 1. Re-engagement push, dormant 90d users 45% 2. New-feature announce, active users 75% 3. Price-change notice, all users 30% 4. Referral-program push, power users 60% 5. Holiday sale push, full list 55%
Healthy
Every row has a specific number written down before column C is ever unhidden.
Unhealthy
Rows get skipped, hedged with a range ('50-70%'), or filled in after peeking at the outcomes.
What this means
A probability that's still vague after you've forced yourself to write one number down was never really an estimate, it was a mood.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| You keep writing ranges instead of a single number | Force yourself to pick the midpoint and write only that | 5 min |
Step 02 of 02
Reviewing a decision journal quarterly tells you whether your 70% calls actually hit around 70% of the time, that's calibration, not just confidence.
Column C is now unhidden with the real outcomes (hit or miss against benchmark). Grouped by your stated probability, does your hit rate roughly match the number you wrote?
Procedure
- Unhide column C and mark each row hit or miss in column D
- Bucket the 5 rows by stated probability (below 50%, 50-69%, 70%+) and compute hit rate per bucket
- Flag any bucket where your hit rate is more than 20 points off your stated number
Bucket Rows Hits Hit rate Stated 30-45% 2 0 0% ~38% 55-60% 2 1 50% ~58% 75% 1 1 100% 75% Flag: 30-45% bucket hit 0/2, consistent with stated odds, not miscalibrated. 55-60% bucket is too small a sample to judge yet.
Healthy
Hit rates land within roughly 20 points of stated odds, or the sample is explicitly flagged as too small to judge.
Unhealthy
Every 70%+ call turns out to be a coin flip in practice, a sign of systematic overconfidence nobody had measured before.
What this means
One 5-row sheet won't prove calibration, the value is in doing this every quarter until the pattern is undeniable.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 70%+ calls hit closer to 50% across several quarters | Discount your own high-confidence estimates by a fixed margin until the gap closes | 30 min |
Final deliverable
A 5-row calibration scorecard showing your stated odds, the real outcome, and a bucketed hit-rate comparison.
See a reference example
Robinhood, Q2 calibration scorecard (excerpt) Bucket Rows Hits Hit rate Stated 40-50% 3 1 33% ~45% 70-80% 4 3 75% ~74% Note: the 70-80% bucket is well-calibrated; the 40-50% bucket ran slightly optimistic and gets a 10-point discount next quarter.
Success criteria
You're done when you can:
- All 5 rows have a single stated percentage recorded before outcomes are revealed
- Outcomes are bucketed by stated probability, not just listed individually
- At least one bucket is flagged as over- or under-confident, or explicitly too small to judge