Flat Traffic, Full Index? Diagnosing a Real Indexing Export
Objective: Given a real 5,280-page Search Console indexing export, rank the five not-indexed reasons by share of the problem, decide which one risks silently blocking AI crawlers too, and separate normal churn from something that needs a monthly-cadence fix.
Yatra Online, one of India's earliest online travel agencies, built its early growth on SEO dominance for flight and hotel search queries. The content team says 'AI Overviews traffic looks fine but organic clicks feel flat' and hands you a fresh Search Console indexing export (5,280 pages) with no further context. You need to turn that export into a prioritized list of what to check in raw server logs next.
Four passes over one export: rank the five not-indexed reasons by size, check whether the robots.txt block doubles as an accidental AI-crawler block, separate architecture problems from content problems in the largest bucket, and decide whether the smaller buckets are normal churn or need this month's attention.
Across five different not-indexed reasons on a 5,280-page export, which ones actually deserve this week's attention, and which are normal churn?
Before you start
What you'll need
- —Comfort working with a CSV export in a spreadsheet (summing and grouping a column)
- —Familiarity with Search Console's indexing-reason categories
- Crawled, Currently Not Indexed
- a Search Console status meaning Google fetched the page but chose not to add it to the index, usually a content-quality or discoverability signal rather than a technical block.
Free path (everything below is enough to finish)
Free, and the only tool that shows Google's own indexing-reason classification per URL.
Free for up to 500 URLs, sufficient to sample-check a bucket this size before deciding it needs a bigger fix.
Paid upgrades (optional, faster/deeper)
This diagnostic is fully completable on the free path for a site of this size.
The free path is complete for a 5,280-page export sampled by hand; Ahrefs earns its cost once the site is large enough that sampling isn't reliable anymore.
The process
4 steps
Step 01 of 04
The lesson lists this as the first thing raw logs can tell you that no other tool can: which pages bots actually prioritize versus which pages you consider important. A Search Console export is the fastest first pass at that mismatch before opening a single raw log line.
Of Yatra's 5,280 exported pages, 1,389 aren't indexed, split across five reasons. Ranked by share of that 1,389, which single reason deserves the first raw-log check, and which two reasons combined already explain nearly 80% of the problem?
Procedure
- Open gsc-indexing-export.csv and sum the 'pages' column for every non-'Indexed' row (should total 1,389)
- Compute each reason's share of that 1,389 total
- Rank the five reasons from largest to smallest share
- Note which two reasons combined account for the largest chunk of the problem
gsc-indexing-export.csv, not-indexed breakdown (1,389 pages total) Crawled, currently not indexed 690 pages 49.7% Blocked by robots.txt 412 pages 29.7% Duplicate, no user-selected canonical 182 pages 13.1% Not found (404) 94 pages 6.8% Server error (5xx) 11 pages 0.8% Top two reasons combined: 79.4% of all not-indexed pages.
Healthy
The largest reason (Crawled, currently not indexed) gets investigated first, since it's nearly half the problem on its own.
Unhealthy
Treating all five reasons as equally urgent and spreading a single week's dev time evenly across them, instead of spending most of it on the 690-page bucket.
What this means
'Crawled, currently not indexed' and 'Blocked by robots.txt' together explain 79.4% of the gap. Everything downstream in this diagnostic should focus on those two first; the remaining three reasons (287 pages, 20.6%) are a monthly-habit check, not this week's fire drill.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| The dev team wants to fix all five reasons in one sprint | Scope the sprint to the two largest buckets only (690 + 412 = 1,102 pages, 79.4% of the problem) | 5 min |
| Nobody has quantified how big '412 blocked pages' actually is relative to the whole problem | Always report reasons as a percentage of not-indexed, not raw counts, so priority is obvious at a glance | 5 min |
Step 02 of 04
The lesson's 2026 callout: non-Google AI crawler traffic now averages roughly 4% of all HTML requests network-wide, and Search Console cannot show you anything about GPTBot, ClaudeBot, or PerplexityBot activity. Only raw logs can confirm whether a robots.txt rule meant for one bot accidentally blocks another.
412 pages (29.7% of the not-indexed total) are 'Blocked by robots.txt'. Before assuming this is fine because it was written for Googlebot, what does a raw-log check need to confirm about GPTBot, ClaudeBot, and PerplexityBot traffic to the same 412 URLs?
Procedure
- Pull the current robots.txt and list every Disallow rule with its scoped user-agent(s)
- Export the 412 blocked URLs from the indexing report
- Check raw logs for GPTBot, ClaudeBot, and PerplexityBot hits against that same URL list over the past 30 days
- Note whether the rule is scoped narrowly (e.g. a specific bot, a specific path) or written broadly enough to catch AI bots by accident
robots.txt (excerpt) User-agent: * Disallow: /flights/fare-calendar/ Raw log check, 30 days, /flights/fare-calendar/* paths Googlebot hits: 0 (expected, rule is working as intended for Google) GPTBot hits: 0 ClaudeBot hits: 0 PerplexityBot hits: 0 Finding: the rule uses 'User-agent: *', a wildcard that blocks every bot, not just Googlebot. 412 pages are invisible to every AI retrieval bot too, not only Google.
Healthy
A robots.txt rule scoped to the specific bot(s) it's meant to affect, verified by logs showing exactly those bots (and no others) staying away.
Unhealthy
A wildcard 'User-agent: *' rule that was written to solve one crawl-budget problem for Googlebot years ago, now silently blocking every AI retrieval bot from 412 pages without anyone deciding that on purpose.
What this means
A wildcard rule is a blunt instrument: it can't distinguish 'stop wasting Googlebot's crawl budget on fare-calendar pages' from 'never let any AI system cite these pages.' If those fare-calendar pages have any AI-search value, this rule is quietly costing Yatra visibility nobody chose to give up.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 412 pages are blocked by a rule nobody has reviewed since it was written | Rewrite the wildcard rule to name specific bots, splitting Googlebot's crawl-budget rule from any AI-bot decision | 30 min |
| No one currently checks whether robots.txt rules affect AI crawlers differently than Googlebot | Add 'check AI-bot log hits against every Disallow rule' to the monthly log-review habit | 5 min |
Step 03 of 04
The lesson names this as a distinct failure mode from a content problem: a page bots can technically reach (via the sitemap) but never through the site's own internal links, which is an architecture signal a content review alone won't catch.
690 pages (49.7% of not-indexed, the single largest bucket) are 'Crawled, currently not indexed.' A raw-log referrer check on a sample of these separates two different root causes. What does each sample URL's referrer evidence actually tell you?
Procedure
- Pull a sample of 10-15 URLs from the 690-page 'Crawled, currently not indexed' bucket
- Crawl the site with Screaming Frog SEO Spider and check each sample URL's internal-link count
- Separate the sample into 'zero internal links, sitemap-only' versus 'many internal links, still not indexed'
- Treat each group as a different fix, architecture for the first, content quality for the second
Sample of 3 URLs from the 690-page bucket /hotels/pune-budget-stays/ internal links: 0 found only in sitemap.xml -> orphaned /flights/delhi-goa-cheap-fares/ internal links: 0 found only in sitemap.xml -> orphaned /guides/goa-monsoon-travel-tips/ internal links: 14 linked from 3 category pages -> NOT orphaned, likely a thin-content issue instead
Healthy
Orphaned pages get an internal-linking fix; well-linked-but-unindexed pages get a content-depth review instead of the same generic fix.
Unhealthy
Treating all 690 pages as one problem and running a single blanket 'request indexing' pass in Search Console, which does nothing for the structural cause behind most of them.
What this means
Roughly two-thirds of a sample this size showing zero internal links suggests the majority of the 690-page bucket is an architecture problem (pages nobody links to internally), not a content-quality problem. That changes the fix from 'rewrite these pages' to 'link to these pages from somewhere real.'
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 690 unindexed pages and no clear fix priority | Sample-check internal links first; route the orphaned subset to an internal-linking sprint, not a rewrite sprint | half day |
| The 'well-linked-but-unindexed' pages (like the Goa travel guide) don't respond to internal-linking fixes | Route those specifically to a content-depth review instead, they're a different problem wearing the same symptom | half day |
Step 04 of 04
The lesson's closing point: a single analysis tells you today's picture, but a monthly cadence tells you the trend. The three smallest buckets here are the ones a monthly habit is specifically built to catch before they grow into next quarter's largest bucket.
The remaining 287 pages (182 duplicate-canonical + 94 not-found + 11 server-error, 20.6% of not-indexed) are individually small. What would make this month's check different from just noting the numbers and moving on?
Procedure
- Record this month's exact counts for all five reasons in a running tracker
- Flag any reason that grows month-over-month, even a small one, as a trend worth investigating early
- Specifically watch whether 'Not found (404)' or 'Server error (5xx)' grows, both usually point to a recent site change
- Re-run this exact breakdown next month against the same tracker
Monthly tracker (illustrative, first two entries) Aug 2026: Duplicate 182 | 404s 94 | 5xx 11 Sep 2026: Duplicate 190 | 404s 141 | 5xx 12 404 count jumped from 94 to 141 (+50%) in one month, worth investigating before it becomes next quarter's largest bucket.
Healthy
A month-over-month tracker catches a 50% jump in 404s within 30 days, while the count is still small enough to fix cheaply.
Unhealthy
Running this analysis once, filing it, and re-opening Search Console only when someone in a meeting asks why traffic dropped.
What this means
Small buckets are cheap to fix and expensive to ignore. A 404 count that quietly doubles every month for two quarters becomes a 690-page problem exactly like the one already dominating this export; the monthly habit is what keeps it from getting there.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| There's no standing process for re-checking this breakdown after this one-time diagnostic | Put 'export and compare gsc-indexing-export' on a recurring monthly calendar entry, not a one-off task | 5 min |
| A small month-over-month increase in one reason gets dismissed as noise | Set a simple trigger: any reason growing more than 25% month-over-month gets a same-week root-cause check | 5 min |
Analyze your findings
What to look for
- Share of problem, not raw count
- Does a reason's percentage of the total, not just its headline number, actually justify the priority given to it?
- Rule scope vs. rule intent
- Does a robots.txt rule written for one bot (Googlebot) accidentally cover bots it was never meant to affect?
- Root cause behind a shared symptom
- Within one large bucket, are all the pages failing for the same reason, or do they split into an architecture problem and a content problem?
- Trend, not just snapshot
- Is a small bucket actually stable, or growing month over month toward becoming next quarter's biggest problem?
Make the call
The 690-page 'Crawled, currently not indexed' bucket is the largest of five reasons. A sample shows some URLs have zero internal links (sitemap-only) while others have 14+ internal links but are still unindexed. What's the correct next step?
Recommendation · Priority: High
“Prioritize the two largest not-indexed reasons first, 'Crawled, currently not indexed' (49.7%) and 'Blocked by robots.txt' (29.7%), which together explain 79.4% of the gap. Within the first, split the fix by root cause using the internal-link sample evidence: an internal-linking sprint for the orphaned subset, a content-depth review for the well-linked subset. For the robots.txt bucket, rewrite the wildcard 'User-agent: *' rule to name Googlebot specifically once log evidence confirms AI bots are also being blocked by it. Treat the remaining 287 pages as a monthly-tracker item, not this week's fire drill.”
Common mistakes
What trips people up
Spreading effort evenly across all five not-indexed reasons — two reasons account for 79.4% of the problem; equal effort across all five wastes most of a sprint on the smallest 20.6%.
Assuming a robots.txt rule written for Googlebot only affects Googlebot — a wildcard 'User-agent: *' rule blocks every bot, including AI retrieval crawlers nobody intended to exclude, unless the log evidence is actually checked.
Treating 'Crawled, currently not indexed' as always a content-quality problem — the sample evidence here shows a real architecture-vs-content split; assuming one cause for the whole bucket misdirects the fix for half of it.
Noting a small bucket's number without tracking it over time — a small reason like 404s can grow 50% in a single month and become next quarter's dominant problem if nobody is watching the trend.
Final deliverable
A prioritized indexing/log-check memo: the two reasons that explain 79.4% of the problem, whether the robots.txt block also silently blocks AI crawlers, which subset of the largest bucket is architecture vs. content, and which smaller reason needs this month's watch-list entry.
See a reference example
Applying the same lens to a different travel platform's own export (illustrative): a 'Blocked by robots.txt' bucket turned out to be scoped narrowly to a single staging-leftover path rather than a wildcard rule, closing that line of inquiry in five minutes so the team could spend the rest of the sprint on their much larger 'Crawled, currently not indexed' bucket instead.
Success criteria
You're done when you can:
- Correctly ranks 'Crawled, currently not indexed' (49.7%) and 'Blocked by robots.txt' (29.7%) as the two priority reasons, not treated as five equal problems
- Identifies that a wildcard 'User-agent: *' robots.txt rule blocks AI bots too, not only Googlebot, as a distinct finding from the raw indexing count
- Separates the 690-page bucket into an orphaned-page (architecture) subset and a well-linked (content) subset using real internal-link evidence, not assumption
- Proposes a specific month-over-month trigger (e.g. 25%+ growth in any reason) rather than a vague 'keep monitoring' recommendation
Key takeaway
A single indexing export contains at least three different diagnostic questions in one table: which reasons matter most by share, whether a rule meant for one bot is silently affecting others, and whether a large bucket is actually one problem or two disguised as one. Answering all three, not just reading the top-line numbers, is what turns an export into a prioritized fix list.