Skip to content
Academy
Marketing Academy · Field Work●SEO
MiniHead-to-Head· 35 minutes

Three Engines, One Brand: Comparing AI Citation Rates Head-to-Head

Go Digit General Insurance

Objective: Given citation results from the same prompt set run across three AI engines, correctly interpret the differences without falling into the raw-comparison trap the lesson warns against.

You're a digital marketing analyst at Go Digit General Insurance, and leadership wants a monthly readout on whether the brand shows up when people ask AI assistants about buying car or health insurance in India.

Build a small prompt panel, log citation results across three engines, then normalize before concluding which engine 'likes' the brand least.

Which AI engine actually needs attention once citation rate is compared against each engine's own baseline, not the raw numbers?

AI Search Measurement/Data Normalization/Cross-Engine Comparison

Before you start

What you'll need

  • —Understanding of what a citation rate measures
  • —Basic comfort building and reading a comparison table in a spreadsheet
Citation rate
the percent of a defined prompt set where an AI engine links to or names your content.
Baseline mention rate
how often an engine names any brand at all for a category, used to normalize citation rate before comparing engines.

Free path (everything below is enough to finish)

FreemiumRun the buyer prompt panel and log citations

Free tier covers a 15-30 prompt weekly panel

FreemiumRun the same prompt panel for cross-engine comparison

Free tier available, different citation behavior than ChatGPT

FreeLog citations, calculate citation rate and normalized share

Free, no account friction

Paid upgrades (optional, faster/deeper)

Brandwatch(optional)
PaidAutomate prompt-panel runs and baseline tracking at scale

Removes the manual weekly re-run once the panel grows past 30 prompts

No access? Manual weekly Google Sheets panel

The process

2 steps

Step 01 of 02

Building a repeatable prompt panel to measure citation rate and share of voice

The lesson defines citation rate as the percent of a defined prompt set where an engine links to your content, and recommends 20-30 real buyer-style prompts run weekly to build a trend line rather than a single snapshot.

You ran the same 15 buyer-style prompts ('best car insurance for a first-time driver in India', etc.) across ChatGPT, Perplexity, and Google AI Overviews. Go Digit was cited 6/15 times by ChatGPT, 3/15 by Perplexity, and 2/15 by Google AI Overviews. What's the one thing you should NOT conclude yet?

Google Sheets— A shared prompt-tracking sheet, one row per prompt per engine

Procedure

  1. Write 15-30 full-sentence buyer questions, not keyword fragments
  2. Run each prompt against ChatGPT, Perplexity, and Google AI Overviews and log whether Go Digit is cited, named-only, or absent
  3. Calculate citation rate per engine (cited count / total prompts)
  4. Hold off on any 'engine X likes us least' conclusion until baseline rates are checked (next step)
Sample output
CITATION RATE BY ENGINE (15 prompts)
Engine                ChatGPT   Perplexity   Google AI Overviews
Cited                 6         3            2
Citation rate         40%       20%          13%

Healthy

A rising citation rate trend across repeated weekly runs of the same prompt set.

Unhealthy

Treating one week's raw citation count as a final verdict on an engine's attitude toward the brand.

What this means

40% vs 13% looks like Google is ignoring Go Digit, but that comparison is meaningless until you know each engine's own baseline citation rate for the insurance category.

So what do I do about it?

SymptomActionEffort
One engine's raw citation rate looks much lower than another'sHold the comparison and check each engine's baseline mention rate for the category before reporting a conclusion5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Normalizing citation rate against each model's own baseline mention rate before comparing

The lesson warns that different models mention brands at very different baseline rates, and that raw citation counts must be normalized against each model's own baseline for the category before concluding anything.

A quick baseline check across 15 generic insurance-category prompts (no brand name mentioned) shows ChatGPT names some insurer 90% of the time, Perplexity 45% of the time, and Google AI Overviews names any insurer only 20% of the time. Given Go Digit's 40% / 20% / 13% citation rates, which engine is actually underperforming its own baseline the most?

Google Sheets— Same tracking sheet, add a baseline column per engine

Procedure

  1. Run 15 generic category prompts with no brand name and log whether ANY insurer is named per engine
  2. Divide Go Digit's citation rate by each engine's baseline rate to get a normalized share
  3. Rank engines by normalized share, not raw citation rate
  4. Flag the engine with the biggest gap between its baseline and Go Digit's actual share as the real priority
Sample output
NORMALIZED COMPARISON
Engine                Go Digit rate   Engine baseline   Normalized share
ChatGPT               40%             90%               44%
Perplexity             20%             45%               44%
Google AI Overviews    13%             20%               65%

Healthy

A normalized share roughly in line with, or above, the brand's fair share of the category conversation.

Unhealthy

Reading raw citation rate as the final answer and deprioritizing the engine with the lowest baseline, which may actually be the strongest relative performer.

What this means

Once normalized, Google AI Overviews is actually Go Digit's best-performing engine (65% of its own baseline), while ChatGPT and Perplexity are tied at a weaker 44% share, the opposite of what the raw numbers suggested.

So what do I do about it?

SymptomActionEffort
Raw citation numbers suggest one engine is hostile to the brandRecalculate as a share of that engine's own baseline mention rate before reporting to leadership30 min
YouYou can do this yourself, no engineering access required.

Analyze your findings

What to look for

Raw vs normalized
Is the comparison based on raw citation counts, or each engine's own baseline mention rate for the category?
Sample size
Is the prompt panel large enough (15-30 prompts) to be a reasonable trend snapshot, not a single query?
Trend, not snapshot
Is this framed as one week's reading, with a plan to re-run weekly, or treated as a final verdict?
Direction of the gap
Does normalizing the numbers actually flip which engine looks like the real priority?

Make the call

Raw citation rates are ChatGPT 40%, Perplexity 20%, Google AI Overviews 13%. Category baselines are ChatGPT 90%, Perplexity 45%, Google AI Overviews 20%. Which engine should leadership actually prioritize?

Recommendation · Priority: Medium

“Report both raw and normalized numbers to leadership, but base the actual prioritization on normalized share: focus improvement work on ChatGPT and Perplexity, not Google AI Overviews, and re-run this same prompt panel weekly to confirm the pattern holds before committing further resources.”

Common mistakes

What trips people up

  • Comparing raw citation rates across engines without normalizing — different models mention any brand at very different baseline rates, so a low raw number can still represent strong relative performance.

  • Drawing a conclusion from a single week's run — AI answers are non-deterministic; one week's numbers are a data point, not a trend, until repeated.

  • Skipping the baseline check because it takes extra time — the baseline run is the step that prevents a completely backwards conclusion about which engine is underperforming.

Final deliverable

A citation-rate comparison sheet with both raw and baseline-normalized numbers, and a one-line recommendation on which engine actually needs attention.

See a reference example
Sample output
TBO Tek, AI citation head-to-head (excerpt)

Raw citation rate: ChatGPT 33%, Perplexity 27%, Google AI Overviews 7%
Category baseline: ChatGPT 80%, Perplexity 50%, Google AI Overviews 10%
Normalized share: ChatGPT 41%, Perplexity 54%, Google AI Overviews 70%
Recommendation: ChatGPT is the real priority despite having the highest raw number, its normalized share is the weakest

Success criteria

You're done when you can:

  • Calculates both raw citation rate and baseline-normalized share correctly for all three engines
  • Does not recommend deprioritizing the engine with the lowest raw citation rate without checking its baseline first

Key takeaway

Raw citation counts across AI engines are not directly comparable, each engine has its own baseline rate for how often it names any brand in a category. Normalizing against that baseline before ranking engines can completely reverse which one actually needs attention.