Skip to content
Academy
Marketing Academy · Field Work●AI in Marketing
CoreAudit· 45 minutes

The Calibration Check: Auditing Synthetic Feedback Against a Real Benchmark

Concord Biotech

Objective: Given a table comparing eight synthetic persona predictions to the real post-campaign survey answers they were meant to forecast, calculate the agreement rate, decide whether it clears a usable trust threshold, and recommend whether to scale the synthetic-only workflow or run another real validation round.

You're the marketing analyst at Concord Biotech, the Ahmedabad-founded fermentation-based API manufacturer, three weeks after testing a new value proposition against a synthetic 'R&D Director evaluating a new API supplier' persona. A real 8-account customer survey has just come back, and you need to decide whether the synthetic panel is trustworthy enough to run solo on the next three campaigns.

Score the synthetic-vs-real agreement rate against the industry benchmark cited in the lesson, and recommend scale-up, re-grounding, or another real round.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeTabulate synthetic vs. real answers and calculate the agreement rate

Free, built-in COUNTIF formulas are enough for an 8-row comparison

Claude(optional)
FreemiumDraft the one-page calibration memo summarizing the recommendation

Free tier is sufficient for a short internal memo

The process

1 step

Step 01 of 01

Validating synthetic feedback against a real benchmark before scaling trust in it

The lesson's hybrid workflow is explicit: start with synthetic testing to iterate fast, then run one small real validation round and compare, if synthetic and real feedback align you've validated the setup; if they diverge wildly, the personas need more grounding data before you scale.

6 of 8 synthetic predictions matched what the real survey respondents actually said (75% agreement). One platform in this lesson claims a 78% correlation benchmark from its own validation study. Does 75% clear the bar to trust this persona for the next three campaigns solo?

Google Sheets— Build a two-column comparison table (synthetic prediction vs. real answer) and a match/mismatch flag column.

Procedure

  1. List the 8 questions asked, with the synthetic persona's predicted top objection in column B and the real survey respondent's actual top objection in column C
  2. Flag each row MATCH or MISMATCH in column D
  3. Calculate the percentage of MATCH rows (=COUNTIF(D:D,"MATCH")/8)
  4. Compare the calculated rate against the 78% benchmark the lesson cites for a mature synthetic-testing platform
  5. Read the 2 mismatched rows for a pattern, not just a number
Sample output
Synthetic vs. real, 8-question comparison (Concord Biotech, R&D Director persona)

1. Top switching objection: synthetic="regulatory filing continuity" real="regulatory filing continuity" MATCH
2. Price sensitivity: synthetic="moderate, tied to volume" real="moderate, tied to volume" MATCH
3. Preferred contact channel: synthetic="technical webinar" real="in-person plant audit" MISMATCH
4. Sample-review timeline: synthetic="2 weeks acceptable" real="2 weeks acceptable" MATCH
5. Compliance documentation priority: synthetic="DMF status first" real="DMF status first" MATCH
6. Emotional response to new-vendor risk: synthetic="neutral, data-driven" real="cautious, wants a plant visit" MISMATCH
7. Contract length preference: synthetic="annual" real="annual" MATCH
8. Deal-breaker: synthetic="missed regulatory deadline" real="missed regulatory deadline" MATCH

Agreement: 6/8 = 75%. Both mismatches involve in-person, sensory trust signals (plant visits, audits), not documented in the persona's training data.

Healthy

75% agreement, below the 78% benchmark, with mismatches concentrated in exactly the emotional/sensory category the lesson already flags as a synthetic-testing blind spot.

Unhealthy

Treating 75% as 'close enough' and running the next three campaigns on synthetic feedback alone with no real check.

What this means

A below-benchmark agreement rate with mismatches clustered in one known blind spot is a specific, fixable gap, not a reason to abandon synthetic testing entirely.

So what do I do about it?

SymptomActionEffort
Team wants to skip the next real validation round to save two weeksKeep synthetic testing for message and pricing iteration; add one real 15-account round before any campaign that hinges on in-person trust signalshalf day
YouYou can do this yourself, no engineering access required.

Final deliverable

A calibration memo stating the measured agreement rate against the cited benchmark, the pattern in the mismatches, and a scale-up-or-revalidate recommendation.

See a reference example
Sample output
Five-Star Business Finance, calibration memo (excerpt)

Agreement rate: 7/8 = 87.5% vs. 78% platform benchmark. Persona cleared the threshold.

Recommendation: Scale this persona to the next two campaign tests without a real validation round; re-check calibration again after the third campaign.

Success criteria

You're done when you can:

  • Correctly calculates the agreement rate from the comparison table
  • Compares the calculated rate against the lesson's cited benchmark rather than judging it in isolation
  • Recommendation follows from where the mismatches cluster, not just the raw percentage