Three Engines, One Brand: Comparing AI Citation Rates Head-to-Head
Objective: Given citation results from the same prompt set run across three AI engines, correctly interpret the differences without falling into the raw-comparison trap the lesson warns against.
You're a digital marketing analyst at Go Digit General Insurance, and leadership wants a monthly readout on whether the brand shows up when people ask AI assistants about buying car or health insurance in India.
Build a small prompt panel, log citation results across three engines, then normalize before concluding which engine 'likes' the brand least.
Which AI engine actually needs attention once citation rate is compared against each engine's own baseline, not the raw numbers?
Before you start
What you'll need
- —Understanding of what a citation rate measures
- —Basic comfort building and reading a comparison table in a spreadsheet
- Citation rate
- the percent of a defined prompt set where an AI engine links to or names your content.
- Baseline mention rate
- how often an engine names any brand at all for a category, used to normalize citation rate before comparing engines.
Free path (everything below is enough to finish)
Free tier covers a 15-30 prompt weekly panel
Free tier available, different citation behavior than ChatGPT
Free, no account friction
Paid upgrades (optional, faster/deeper)
Removes the manual weekly re-run once the panel grows past 30 prompts
No access? Manual weekly Google Sheets panel
The process
2 steps
Step 01 of 02
The lesson defines citation rate as the percent of a defined prompt set where an engine links to your content, and recommends 20-30 real buyer-style prompts run weekly to build a trend line rather than a single snapshot.
You ran the same 15 buyer-style prompts ('best car insurance for a first-time driver in India', etc.) across ChatGPT, Perplexity, and Google AI Overviews. Go Digit was cited 6/15 times by ChatGPT, 3/15 by Perplexity, and 2/15 by Google AI Overviews. What's the one thing you should NOT conclude yet?
Procedure
- Write 15-30 full-sentence buyer questions, not keyword fragments
- Run each prompt against ChatGPT, Perplexity, and Google AI Overviews and log whether Go Digit is cited, named-only, or absent
- Calculate citation rate per engine (cited count / total prompts)
- Hold off on any 'engine X likes us least' conclusion until baseline rates are checked (next step)
CITATION RATE BY ENGINE (15 prompts) Engine ChatGPT Perplexity Google AI Overviews Cited 6 3 2 Citation rate 40% 20% 13%
Healthy
A rising citation rate trend across repeated weekly runs of the same prompt set.
Unhealthy
Treating one week's raw citation count as a final verdict on an engine's attitude toward the brand.
What this means
40% vs 13% looks like Google is ignoring Go Digit, but that comparison is meaningless until you know each engine's own baseline citation rate for the insurance category.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| One engine's raw citation rate looks much lower than another's | Hold the comparison and check each engine's baseline mention rate for the category before reporting a conclusion | 5 min |
Step 02 of 02
The lesson warns that different models mention brands at very different baseline rates, and that raw citation counts must be normalized against each model's own baseline for the category before concluding anything.
A quick baseline check across 15 generic insurance-category prompts (no brand name mentioned) shows ChatGPT names some insurer 90% of the time, Perplexity 45% of the time, and Google AI Overviews names any insurer only 20% of the time. Given Go Digit's 40% / 20% / 13% citation rates, which engine is actually underperforming its own baseline the most?
Procedure
- Run 15 generic category prompts with no brand name and log whether ANY insurer is named per engine
- Divide Go Digit's citation rate by each engine's baseline rate to get a normalized share
- Rank engines by normalized share, not raw citation rate
- Flag the engine with the biggest gap between its baseline and Go Digit's actual share as the real priority
NORMALIZED COMPARISON Engine Go Digit rate Engine baseline Normalized share ChatGPT 40% 90% 44% Perplexity 20% 45% 44% Google AI Overviews 13% 20% 65%
Healthy
A normalized share roughly in line with, or above, the brand's fair share of the category conversation.
Unhealthy
Reading raw citation rate as the final answer and deprioritizing the engine with the lowest baseline, which may actually be the strongest relative performer.
What this means
Once normalized, Google AI Overviews is actually Go Digit's best-performing engine (65% of its own baseline), while ChatGPT and Perplexity are tied at a weaker 44% share, the opposite of what the raw numbers suggested.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Raw citation numbers suggest one engine is hostile to the brand | Recalculate as a share of that engine's own baseline mention rate before reporting to leadership | 30 min |
Analyze your findings
What to look for
- Raw vs normalized
- Is the comparison based on raw citation counts, or each engine's own baseline mention rate for the category?
- Sample size
- Is the prompt panel large enough (15-30 prompts) to be a reasonable trend snapshot, not a single query?
- Trend, not snapshot
- Is this framed as one week's reading, with a plan to re-run weekly, or treated as a final verdict?
- Direction of the gap
- Does normalizing the numbers actually flip which engine looks like the real priority?
Make the call
Raw citation rates are ChatGPT 40%, Perplexity 20%, Google AI Overviews 13%. Category baselines are ChatGPT 90%, Perplexity 45%, Google AI Overviews 20%. Which engine should leadership actually prioritize?
Recommendation · Priority: Medium
“Report both raw and normalized numbers to leadership, but base the actual prioritization on normalized share: focus improvement work on ChatGPT and Perplexity, not Google AI Overviews, and re-run this same prompt panel weekly to confirm the pattern holds before committing further resources.”
Common mistakes
What trips people up
Comparing raw citation rates across engines without normalizing — different models mention any brand at very different baseline rates, so a low raw number can still represent strong relative performance.
Drawing a conclusion from a single week's run — AI answers are non-deterministic; one week's numbers are a data point, not a trend, until repeated.
Skipping the baseline check because it takes extra time — the baseline run is the step that prevents a completely backwards conclusion about which engine is underperforming.
Final deliverable
A citation-rate comparison sheet with both raw and baseline-normalized numbers, and a one-line recommendation on which engine actually needs attention.
See a reference example
TBO Tek, AI citation head-to-head (excerpt) Raw citation rate: ChatGPT 33%, Perplexity 27%, Google AI Overviews 7% Category baseline: ChatGPT 80%, Perplexity 50%, Google AI Overviews 10% Normalized share: ChatGPT 41%, Perplexity 54%, Google AI Overviews 70% Recommendation: ChatGPT is the real priority despite having the highest raw number, its normalized share is the weakest
Success criteria
You're done when you can:
- Calculates both raw citation rate and baseline-normalized share correctly for all three engines
- Does not recommend deprioritizing the engine with the lowest raw citation rate without checking its baseline first
Key takeaway
Raw citation counts across AI engines are not directly comparable, each engine has its own baseline rate for how often it names any brand in a category. Normalizing against that baseline before ranking engines can completely reverse which one actually needs attention.