Skip to content
Academy
Marketing Academy · Field Work●Analytics & Attribution
CoreForecast· 40 minutes

Trust But Verify: Validating a Churn Model Against Real Outcomes

Klaviyo

Objective: Given a 25-customer export with model-predicted churn-risk scores and their actual 30-day outcomes, calculate precision and recall to decide whether the model is trustworthy enough to trigger retention campaigns.

You support a Klaviyo customer, a mid-size DTC brand, that just deployed a churn-risk model. Before it's allowed to trigger automated retention offers, you have to confirm it's actually right, not just confident.

Don't just eyeball the scores, calculate precision and recall the way the lesson defines them, and make a go/no-go call on live deployment.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeBuild the confusion matrix and test threshold scenarios

Free, transparent formulas that a non-technical stakeholder can audit

Paid upgrades (optional, faster/deeper)

Amplitude(optional)
FreemiumTrack cohort-level retention outcomes automatically once the model is live

Removes the need to manually export and match outcomes every cycle

No access? Manual 30-day outcome export in Google Sheets, as done in this project

The process

2 steps

Step 01 of 02

Validating predictions against real outcomes (precision and recall)

The lesson defines precision as: of everyone flagged high-risk, how many actually churned. Recall is: of everyone who actually churned, how many did the model catch in advance.

The model flagged 12 of 25 customers as 'high churn risk.' Of those 12, 9 actually churned within 30 days. Of the 13 customers NOT flagged, 2 also churned. What are precision and recall?

Google Sheets— Import churn-model-export.csv and build a 2x2 confusion matrix.

Procedure

  1. Import churn-model-export.csv with columns: customer_id, predicted_risk, actual_outcome
  2. Build a 2x2 table: flagged/not-flagged vs. churned/retained
  3. Calculate precision = true positives / (true positives + false positives)
  4. Calculate recall = true positives / (true positives + false negatives)
Sample output
                Actually Churned   Actually Retained
Flagged High-Risk       9                    3
Not Flagged             2                   11

Precision = 9/12 = 75%
Recall = 9/11 = 82%

Healthy

Both precision and recall land above 70%, meaning most flagged customers really are at risk, and the model catches most real churners before they leave.

Unhealthy

Precision is high but recall is low (e.g. 90% precision, 30% recall), meaning the model is overly cautious and misses most customers who actually churn, wasting the campaign's real opportunity.

What this means

A model can look accurate on paper while still failing the business goal; precision and recall catch that in a way a single 'accuracy' number hides.

So what do I do about it?

SymptomActionEffort
Recall is below 50% even though precision looks strongLower the risk-score threshold that triggers a 'high-risk' flag so more real churners get caught, then re-check precision30 min
Both precision and recall exceed 70%Approve the model for live retention-campaign triggers5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Interpreting a propensity score threshold

The lesson describes propensity models as scoring each customer 0-100 on likelihood of an action. The threshold you pick to call something 'high risk' is a business decision, not a fixed rule.

The model outputs a 0-100 churn score for every customer, not just a flag. At what score cutoff should 'high risk' start, and what changes if you move it from 70 to 50?

Google Sheets— Sort the export by raw churn score (0-100) rather than the pre-set flag.

Procedure

  1. Sort all 25 customers by raw churn score, descending
  2. Test a threshold of 70: count how many actual churners fall above it
  3. Test a threshold of 50: count how many actual churners fall above it
  4. Compare how many total customers get flagged (and get a retention offer) at each threshold
Sample output
Threshold 70: 8 customers flagged, 7 actual churners caught (recall 64%)
Threshold 50: 15 customers flagged, 10 actual churners caught (recall 91%)
Retention offer cost is fixed per customer flagged.

Healthy

The team picks the threshold deliberately, weighing the cost of more retention offers (lower threshold) against the cost of missed churners (higher threshold).

Unhealthy

The team keeps whatever default threshold the tool shipped with, without checking whether it fits the actual cost of a retention offer versus the cost of a lost customer.

What this means

There is no universally 'correct' threshold, only the one that matches what a false negative costs you versus what a retention offer costs you.

So what do I do about it?

SymptomActionEffort
Retention budget is limited but recall at the default threshold is lowRaise the threshold to flag fewer, higher-confidence customers rather than spreading budget thin30 min
YouYou can do this yourself, no engineering access required.

Final deliverable

A validation memo stating the model's precision and recall at its current threshold, a recommended threshold change if warranted, and a go/no-go call on live deployment.

See a reference example
Sample output
Robinhood, Churn Model Validation Memo (excerpt)

Current threshold (70): Precision 78%, Recall 61%
Tested threshold (55): Precision 68%, Recall 85%
RECOMMENDATION: Lower threshold to 55. Retention budget covers the extra flagged volume, and the 24-point recall gain catches significantly more real churners.

Success criteria

You're done when you can:

  • Correctly calculates precision and recall from the confusion matrix
  • Tests at least one alternate threshold and compares the tradeoff
  • States a clear go/no-go recommendation grounded in the numbers, not intuition