Pick and Defend PolicyBazaar's North Star Metric
Objective: Given four candidate North Star Metrics for an insurance-comparison marketplace, choose one and defend it in writing against the lesson's three-criteria test and its four-category input-metric decomposition, the way a growth lead would defend it to a skeptical leadership team.
You're the incoming growth lead at an insurance-comparison marketplace built on the PolicyBazaar model, SEO- and comparison-content-driven acquisition feeding an insurance and lending marketplace. Leadership has four different numbers they each personally favor as 'the metric that matters' and cannot agree. You have one week to bring back a single recommendation, decomposed into inputs teams can actually act on.
Four candidates, one slide's worth of criteria. Run each through the lesson's value-statement method and three-criteria test, eliminate the ones that fail, then decompose your winner into the four Amplitude input categories so it's an operating metric, not a decoration.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free for a single user, enough to produce every artifact this project asks for.
Free tier's Retention report produces the same shape of table used in Step 3, this exercise is complete without ever opening it.
Paid upgrades (optional, faster/deeper)
The free path (Notion plus GA4's free-tier retention report) is complete on its own. Amplitude is only worth paying for once the team wants the decomposition running as a live dashboard.
An upgrade for a team that wants this decomposition live and automated, never required to write the recommendation doc.
The process
4 steps
Step 01 of 04
Step 1 of the lesson's playbook: finish the sentence 'Our product is valuable when a customer ___.' The verb in your answer usually points directly to your NSM.
Given the four candidates, which value statement does each one actually describe, and does it match what a comparison-and-purchase insurance marketplace is actually for?
Procedure
- List all four candidates: (A) total registered users (families and agents), (B) quotes generated per month, (C) policies purchased within 30 days of a quote, (D) monthly premium revenue processed
- For each candidate, write the one-sentence value statement it implies, e.g. 'valuable when a customer registers an account' for A
- Flag any candidate whose implied value statement stops short of the real outcome (a customer being covered, not just browsing)
Candidate Implied value statement Stops short of real outcome? A '...creates an account.' Yes, no product use implied at all B '...receives a comparison quote.' Yes, still browsing, not covered C '...buys the policy they compared.' No, this is the actual outcome D '...generates revenue for the business.' Yes, this is the company's result, not the customer's
Healthy
The candidate list narrows to one or two once you write out what each one actually implies about customer behavior.
Unhealthy
All four still look equally plausible after this step, meaning the value statements weren't written specifically enough to differentiate them.
What this means
Candidate C is the only one whose value statement describes the customer actually receiving what they came for (coverage), not just interacting with the product (A, B) or the company's own result (D).
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A value statement describes an account action, not a customer outcome | Rewrite it one step further downstream: what does the customer get, not what do they click | 30 min |
| A value statement is actually describing company revenue, not customer value | Separate it out, revenue is a lagging result per Step 3's criteria, not a candidate NSM itself | 5 min |
Step 02 of 04
The lesson's three questions: does it reflect real customer value, does it predict revenue (leading, not lagging), and can multiple teams move it with their daily work?
Score all four candidates against the three criteria and see which one is the only clean pass.
Procedure
- For each candidate, answer yes/no: reflects real customer value? leading indicator of revenue? multi-team actionable?
- Eliminate any candidate with even one 'no'
- For the survivor, write one sentence on why each 'yes' is actually true, not just asserted
Candidate Real value? Leading indicator? Multi-team actionable? A, registered users No No (vanity) Yes B, quotes generated Partial Weak (browsing) Yes C, policies purchased in 30 days Yes Yes Yes D, monthly premium revenue Yes (result) No (lagging) Partially (mostly sales)
Healthy
Exactly one candidate clears all three criteria with a genuine 'yes,' not a stretched one.
Unhealthy
Two or more candidates clear all three, or the survivor only clears them on a technicality, e.g. 'multi-team actionable' because everyone can theoretically influence revenue.
What this means
Candidate C is the only clean pass. B fails the 'real value' test because a quote alone proves interest, not coverage. D fails the leading-indicator test explicitly, it's the lesson's named Mistake 1: picking revenue itself as the NSM.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Revenue (D) is the candidate leadership keeps defaulting back to | Point at the lesson's Mistake 1 directly: revenue is a lagging result you can inflate short-term with discounts while damaging retention | 5 min |
| Quotes generated (B) looks tempting because it's already the team's biggest existing dashboard number | Name explicitly what it's missing: it measures interest, not the moment of actual coverage | 5 min |
Step 03 of 04
The lesson's caveat: before obsessing over an NSM, confirm retention curves are flattening at a healthy level, that's the real proof of product-market fit an NSM choice should build on.
Before locking in Candidate C, does the real cohort-retention data support treating this marketplace as past the product-market-fit stage, or is the NSM choice premature?
Procedure
- Read the month-0 through month-5 retention percentages for each monthly cohort
- Check whether the curves are flattening (retention loss slowing by month 4-5) or still falling steeply with no floor
- Compare the newest complete cohorts (Feb-Apr 2026) against the oldest (Nov-Dec 2025) to see whether retention is improving, flat, or worsening over time
Cohort M0 M1 M2 M3 M4 M5 2025-11 100.0 47.0 34.3 26.1 20.3 16.8 2025-12 100.0 50.0 40.9 34.0 28.2 24.9 2026-01 100.0 50.4 40.8 33.9 30.2 27.1 2026-02 100.0 48.4 39.3 33.4 29.6 25.8 2026-03 100.0 50.5 40.7 32.5 29.1 25.9 2026-04 100.0 46.0 37.0 28.5 24.9 22.5
Healthy
Later cohorts retain at or above the level of earlier cohorts by month 4-5, and each cohort's month-to-month drop-off visibly slows down rather than falling in a straight line toward zero.
Unhealthy
Later cohorts retain worse than earlier ones, or the curve keeps falling at close to the same rate all the way to month 5 with no sign of leveling off.
What this means
The Nov 2025 cohort is the clear outlier at 16.8% by month 5, every cohort from Dec 2025 onward retains meaningfully higher (22.5%-27.1% by month 5) and the month-4-to-5 drop is a few points, not a cliff. That's a flattening curve on the cohorts that matter, product-market fit looks confirmed enough to build an NSM on.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| An NSM proposal with no retention-curve check behind it | Always attach this table before locking a candidate, a plausible-sounding NSM built on pre-PMF data will mislead every team that acts on it | 30 min |
| Nov 2025 cohort's weak 16.8% used to argue against the whole marketplace | Treat it as the pre-fix baseline, not the current state, compare against Dec 2025 onward instead | 5 min |
Step 04 of 04
Step 4 of the lesson's playbook: break the NSM into 3-5 input metrics across Amplitude's four categories, breadth, depth, frequency, efficiency, so individual teams have something to move each sprint.
What are this marketplace's breadth, depth, frequency, and efficiency inputs underneath Candidate C, and which team owns each one?
Procedure
- Breadth: what share of quote-generators go on to purchase within 30 days, who owns it (conversion/CRO team)
- Depth: how many insurers/policy types does a buyer compare before purchasing, who owns it (content/comparison-engine team)
- Frequency: what share of buyers return to renew or add a second policy, who owns it (lifecycle/retention team)
- Efficiency: how many days from first quote to purchase, who owns it (product/UX team)
Input Metric Owning team Breadth Quote-to-purchase rate within 30 days Conversion/CRO Depth Insurers compared per buyer before purchase Content/comparison engine Frequency Buyers who renew or add a 2nd policy within 12 months Lifecycle/retention Efficiency Days from first quote to completed purchase Product/UX
Healthy
Each of the four inputs has exactly one clear owning team and a concrete, trackable number, not a vague description.
Unhealthy
An input with no clear owner, or two inputs that are really the same metric restated, e.g. listing both 'quote-to-purchase rate' and 'purchases' as separate inputs.
What this means
This decomposition gives four different teams a weekly number to move, none of which is 'increase total registered users,' exactly what makes Candidate C an operating metric instead of a dashboard vanity number.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| An input metric with no owning team named | Assign it before shipping the recommendation, an NSM without owned inputs is decoration per the lesson's Mistake 5 | 30 min |
| Leadership asks why 'total registered users' isn't one of the four inputs | Point back to Steps 1-2: it failed the value-statement and criteria test, it doesn't re-enter through the back door as an input | 5 min |
Final deliverable
A one-page NSM recommendation: chosen candidate, the value-statement and three-criteria reasoning that eliminated the other three, the retention-curve evidence confirming product-market fit, and the four-input decomposition table with owning teams.
See a reference example
Recommendation: Candidate C, 'policies purchased within 30 days of a quote.' It's the only candidate whose value statement describes the customer actually receiving coverage, not just browsing (B) or registering (A), and the only one that passes all three criteria cleanly, unlike revenue (D), which fails the leading-indicator test the same way Airbnb rejected signups in favor of nights booked, a metric that captures both sides of the marketplace getting value at once. Retention curves from Dec 2025 onward flatten in the 22-27% range by month 5, confirming enough product-market fit to build an NSM on. Four inputs assigned: quote-to-purchase rate (CRO), insurers compared per buyer (content), 12-month renewal/second-policy rate (lifecycle), days-to-purchase (product).
Success criteria
You're done when you can:
- Eliminates registered users and revenue as NSM candidates using the lesson's own reasons (vanity; lagging indicator)
- Correctly identifies policies-purchased-within-30-days as the only candidate that passes all three criteria
- Uses the real cohort-retention numbers to confirm, not assume, product-market fit before finalizing the NSM
- Decomposes the chosen NSM into four input metrics, each with a named owning team