Skip to content
Academy
Marketing Academy · Field Work●SEO
CoreAudit· 45 minutes

Flat Traffic, Full Index? Diagnosing a Real Indexing Export

Yatra Online

Objective: Given a real 5,280-page Search Console indexing export, rank the five not-indexed reasons by share of the problem, decide which one risks silently blocking AI crawlers too, and separate normal churn from something that needs a monthly-cadence fix.

Yatra Online, one of India's earliest online travel agencies, built its early growth on SEO dominance for flight and hotel search queries. The content team says 'AI Overviews traffic looks fine but organic clicks feel flat' and hands you a fresh Search Console indexing export (5,280 pages) with no further context. You need to turn that export into a prioritized list of what to check in raw server logs next.

Four passes over one export: rank the five not-indexed reasons by size, check whether the robots.txt block doubles as an accidental AI-crawler block, separate architecture problems from content problems in the largest bucket, and decide whether the smaller buckets are normal churn or need this month's attention.

Across five different not-indexed reasons on a 5,280-page export, which ones actually deserve this week's attention, and which are normal churn?

Search Console Indexing Reports/Log-Based Prioritization/Robots.txt Scope Auditing/Orphan Page Diagnosis

Before you start

What you'll need

  • —Comfort working with a CSV export in a spreadsheet (summing and grouping a column)
  • —Familiarity with Search Console's indexing-reason categories
Crawled, Currently Not Indexed
a Search Console status meaning Google fetched the page but chose not to add it to the index, usually a content-quality or discoverability signal rather than a technical block.

Free path (everything below is enough to finish)

FreeSource of the indexing export and the robots.txt Tester used to confirm rule scope

Free, and the only tool that shows Google's own indexing-reason classification per URL.

FreemiumCrawl the site to confirm internal-link counts for the orphaned-page check in Step 3

Free for up to 500 URLs, sufficient to sample-check a bucket this size before deciding it needs a bigger fix.

Paid upgrades (optional, faster/deeper)

This diagnostic is fully completable on the free path for a site of this size.

Ahrefs(optional)
PaidSite-wide internal-link and orphaned-page auditing at a scale beyond the free Screaming Frog crawl limit

The free path is complete for a 5,280-page export sampled by hand; Ahrefs earns its cost once the site is large enough that sampling isn't reliable anymore.

Download project dataset

The process

4 steps

Step 01 of 04

Which pages get crawled most and least often, revealing a mismatch between what you consider important and what bots actually prioritize

The lesson lists this as the first thing raw logs can tell you that no other tool can: which pages bots actually prioritize versus which pages you consider important. A Search Console export is the fastest first pass at that mismatch before opening a single raw log line.

Of Yatra's 5,280 exported pages, 1,389 aren't indexed, split across five reasons. Ranked by share of that 1,389, which single reason deserves the first raw-log check, and which two reasons combined already explain nearly 80% of the problem?

Google Search Console— Indexing > Pages report, or the exported gsc-indexing-export.csv opened in a spreadsheet.

Procedure

  1. Open gsc-indexing-export.csv and sum the 'pages' column for every non-'Indexed' row (should total 1,389)
  2. Compute each reason's share of that 1,389 total
  3. Rank the five reasons from largest to smallest share
  4. Note which two reasons combined account for the largest chunk of the problem
Sample output
gsc-indexing-export.csv, not-indexed breakdown (1,389 pages total)
  Crawled, currently not indexed        690 pages   49.7%
  Blocked by robots.txt                  412 pages   29.7%
  Duplicate, no user-selected canonical  182 pages   13.1%
  Not found (404)                         94 pages    6.8%
  Server error (5xx)                      11 pages    0.8%

Top two reasons combined: 79.4% of all not-indexed pages.

Healthy

The largest reason (Crawled, currently not indexed) gets investigated first, since it's nearly half the problem on its own.

Unhealthy

Treating all five reasons as equally urgent and spreading a single week's dev time evenly across them, instead of spending most of it on the 690-page bucket.

What this means

'Crawled, currently not indexed' and 'Blocked by robots.txt' together explain 79.4% of the gap. Everything downstream in this diagnostic should focus on those two first; the remaining three reasons (287 pages, 20.6%) are a monthly-habit check, not this week's fire drill.

So what do I do about it?

SymptomActionEffort
The dev team wants to fix all five reasons in one sprintScope the sprint to the two largest buckets only (690 + 412 = 1,102 pages, 79.4% of the problem)5 min
Nobody has quantified how big '412 blocked pages' actually is relative to the whole problemAlways report reasons as a percentage of not-indexed, not raw counts, so priority is obvious at a glance5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 04

Exactly which AI crawlers visit, how often, and which pages they touch, data that exists nowhere else

The lesson's 2026 callout: non-Google AI crawler traffic now averages roughly 4% of all HTML requests network-wide, and Search Console cannot show you anything about GPTBot, ClaudeBot, or PerplexityBot activity. Only raw logs can confirm whether a robots.txt rule meant for one bot accidentally blocks another.

412 pages (29.7% of the not-indexed total) are 'Blocked by robots.txt'. Before assuming this is fine because it was written for Googlebot, what does a raw-log check need to confirm about GPTBot, ClaudeBot, and PerplexityBot traffic to the same 412 URLs?

Google Search Console— robots.txt Tester (or the live robots.txt file) cross-referenced against raw access logs for the same URL list.

Procedure

  1. Pull the current robots.txt and list every Disallow rule with its scoped user-agent(s)
  2. Export the 412 blocked URLs from the indexing report
  3. Check raw logs for GPTBot, ClaudeBot, and PerplexityBot hits against that same URL list over the past 30 days
  4. Note whether the rule is scoped narrowly (e.g. a specific bot, a specific path) or written broadly enough to catch AI bots by accident
Sample output
robots.txt (excerpt)
  User-agent: *
  Disallow: /flights/fare-calendar/

Raw log check, 30 days, /flights/fare-calendar/* paths
  Googlebot hits:      0 (expected, rule is working as intended for Google)
  GPTBot hits:         0
  ClaudeBot hits:      0
  PerplexityBot hits:  0

Finding: the rule uses 'User-agent: *', a wildcard that blocks every bot, not just Googlebot.
412 pages are invisible to every AI retrieval bot too, not only Google.

Healthy

A robots.txt rule scoped to the specific bot(s) it's meant to affect, verified by logs showing exactly those bots (and no others) staying away.

Unhealthy

A wildcard 'User-agent: *' rule that was written to solve one crawl-budget problem for Googlebot years ago, now silently blocking every AI retrieval bot from 412 pages without anyone deciding that on purpose.

What this means

A wildcard rule is a blunt instrument: it can't distinguish 'stop wasting Googlebot's crawl budget on fare-calendar pages' from 'never let any AI system cite these pages.' If those fare-calendar pages have any AI-search value, this rule is quietly costing Yatra visibility nobody chose to give up.

So what do I do about it?

SymptomActionEffort
412 pages are blocked by a rule nobody has reviewed since it was writtenRewrite the wildcard rule to name specific bots, splitting Googlebot's crawl-budget rule from any AI-bot decision30 min
No one currently checks whether robots.txt rules affect AI crawlers differently than GooglebotAdd 'check AI-bot log hits against every Disallow rule' to the monthly log-review habit5 min
DeveloperNeeds a developer/engineer to ship the fix.

Step 03 of 04

Orphaned pages that bots reach only via the sitemap, never through an internal link, a sign of weak site architecture

The lesson names this as a distinct failure mode from a content problem: a page bots can technically reach (via the sitemap) but never through the site's own internal links, which is an architecture signal a content review alone won't catch.

690 pages (49.7% of not-indexed, the single largest bucket) are 'Crawled, currently not indexed.' A raw-log referrer check on a sample of these separates two different root causes. What does each sample URL's referrer evidence actually tell you?

Screaming Frog SEO Spider— Crawl the site, then check Internal Links and Referring Pages for a sample of the 690 flagged URLs.

Procedure

  1. Pull a sample of 10-15 URLs from the 690-page 'Crawled, currently not indexed' bucket
  2. Crawl the site with Screaming Frog SEO Spider and check each sample URL's internal-link count
  3. Separate the sample into 'zero internal links, sitemap-only' versus 'many internal links, still not indexed'
  4. Treat each group as a different fix, architecture for the first, content quality for the second
Sample output
Sample of 3 URLs from the 690-page bucket
  /hotels/pune-budget-stays/         internal links: 0   found only in sitemap.xml   -> orphaned
  /flights/delhi-goa-cheap-fares/    internal links: 0   found only in sitemap.xml   -> orphaned
  /guides/goa-monsoon-travel-tips/   internal links: 14  linked from 3 category pages -> NOT orphaned, likely a thin-content issue instead

Healthy

Orphaned pages get an internal-linking fix; well-linked-but-unindexed pages get a content-depth review instead of the same generic fix.

Unhealthy

Treating all 690 pages as one problem and running a single blanket 'request indexing' pass in Search Console, which does nothing for the structural cause behind most of them.

What this means

Roughly two-thirds of a sample this size showing zero internal links suggests the majority of the 690-page bucket is an architecture problem (pages nobody links to internally), not a content-quality problem. That changes the fix from 'rewrite these pages' to 'link to these pages from somewhere real.'

So what do I do about it?

SymptomActionEffort
690 unindexed pages and no clear fix prioritySample-check internal links first; route the orphaned subset to an internal-linking sprint, not a rewrite sprinthalf day
The 'well-linked-but-unindexed' pages (like the Goa travel guide) don't respond to internal-linking fixesRoute those specifically to a content-depth review instead, they're a different problem wearing the same symptomhalf day
EitherYou or a developer can handle this, depending on your access.

Step 04 of 04

Crawl behavior shifts after every major release, a monthly cadence tells you whether things are trending in the right direction or quietly breaking

The lesson's closing point: a single analysis tells you today's picture, but a monthly cadence tells you the trend. The three smallest buckets here are the ones a monthly habit is specifically built to catch before they grow into next quarter's largest bucket.

The remaining 287 pages (182 duplicate-canonical + 94 not-found + 11 server-error, 20.6% of not-indexed) are individually small. What would make this month's check different from just noting the numbers and moving on?

Google Search Console— Save this month's five-reason breakdown, compare it against next month's export side by side.

Procedure

  1. Record this month's exact counts for all five reasons in a running tracker
  2. Flag any reason that grows month-over-month, even a small one, as a trend worth investigating early
  3. Specifically watch whether 'Not found (404)' or 'Server error (5xx)' grows, both usually point to a recent site change
  4. Re-run this exact breakdown next month against the same tracker
Sample output
Monthly tracker (illustrative, first two entries)
  Aug 2026: Duplicate 182 | 404s 94 | 5xx 11
  Sep 2026: Duplicate 190 | 404s 141 | 5xx 12

404 count jumped from 94 to 141 (+50%) in one month, worth investigating before it becomes next quarter's largest bucket.

Healthy

A month-over-month tracker catches a 50% jump in 404s within 30 days, while the count is still small enough to fix cheaply.

Unhealthy

Running this analysis once, filing it, and re-opening Search Console only when someone in a meeting asks why traffic dropped.

What this means

Small buckets are cheap to fix and expensive to ignore. A 404 count that quietly doubles every month for two quarters becomes a 690-page problem exactly like the one already dominating this export; the monthly habit is what keeps it from getting there.

So what do I do about it?

SymptomActionEffort
There's no standing process for re-checking this breakdown after this one-time diagnosticPut 'export and compare gsc-indexing-export' on a recurring monthly calendar entry, not a one-off task5 min
A small month-over-month increase in one reason gets dismissed as noiseSet a simple trigger: any reason growing more than 25% month-over-month gets a same-week root-cause check5 min
YouYou can do this yourself, no engineering access required.

Analyze your findings

What to look for

Share of problem, not raw count
Does a reason's percentage of the total, not just its headline number, actually justify the priority given to it?
Rule scope vs. rule intent
Does a robots.txt rule written for one bot (Googlebot) accidentally cover bots it was never meant to affect?
Root cause behind a shared symptom
Within one large bucket, are all the pages failing for the same reason, or do they split into an architecture problem and a content problem?
Trend, not just snapshot
Is a small bucket actually stable, or growing month over month toward becoming next quarter's biggest problem?

Make the call

The 690-page 'Crawled, currently not indexed' bucket is the largest of five reasons. A sample shows some URLs have zero internal links (sitemap-only) while others have 14+ internal links but are still unindexed. What's the correct next step?

Recommendation · Priority: High

“Prioritize the two largest not-indexed reasons first, 'Crawled, currently not indexed' (49.7%) and 'Blocked by robots.txt' (29.7%), which together explain 79.4% of the gap. Within the first, split the fix by root cause using the internal-link sample evidence: an internal-linking sprint for the orphaned subset, a content-depth review for the well-linked subset. For the robots.txt bucket, rewrite the wildcard 'User-agent: *' rule to name Googlebot specifically once log evidence confirms AI bots are also being blocked by it. Treat the remaining 287 pages as a monthly-tracker item, not this week's fire drill.”

Common mistakes

What trips people up

  • Spreading effort evenly across all five not-indexed reasons — two reasons account for 79.4% of the problem; equal effort across all five wastes most of a sprint on the smallest 20.6%.

  • Assuming a robots.txt rule written for Googlebot only affects Googlebot — a wildcard 'User-agent: *' rule blocks every bot, including AI retrieval crawlers nobody intended to exclude, unless the log evidence is actually checked.

  • Treating 'Crawled, currently not indexed' as always a content-quality problem — the sample evidence here shows a real architecture-vs-content split; assuming one cause for the whole bucket misdirects the fix for half of it.

  • Noting a small bucket's number without tracking it over time — a small reason like 404s can grow 50% in a single month and become next quarter's dominant problem if nobody is watching the trend.

Final deliverable

A prioritized indexing/log-check memo: the two reasons that explain 79.4% of the problem, whether the robots.txt block also silently blocks AI crawlers, which subset of the largest bucket is architecture vs. content, and which smaller reason needs this month's watch-list entry.

See a reference example
Sample output
Applying the same lens to a different travel platform's own export (illustrative): a 'Blocked by robots.txt' bucket turned out to be scoped narrowly to a single staging-leftover path rather than a wildcard rule, closing that line of inquiry in five minutes so the team could spend the rest of the sprint on their much larger 'Crawled, currently not indexed' bucket instead.

Success criteria

You're done when you can:

  • Correctly ranks 'Crawled, currently not indexed' (49.7%) and 'Blocked by robots.txt' (29.7%) as the two priority reasons, not treated as five equal problems
  • Identifies that a wildcard 'User-agent: *' robots.txt rule blocks AI bots too, not only Googlebot, as a distinct finding from the raw indexing count
  • Separates the 690-page bucket into an orphaned-page (architecture) subset and a well-linked (content) subset using real internal-link evidence, not assumption
  • Proposes a specific month-over-month trigger (e.g. 25%+ growth in any reason) rather than a vague 'keep monitoring' recommendation

Key takeaway

A single indexing export contains at least three different diagnostic questions in one table: which reasons matter most by share, whether a rule meant for one bot is silently affecting others, and whether a large bucket is actually one problem or two disguised as one. Answering all three, not just reading the top-line numbers, is what turns an export into a prioritized fix list.