Skip to content
Academy
Marketing Academy · Field Work●SEO
CoreAudit· 40 minutes

The Crawl Budget Audit: Finding Where Snowflake's Docs Section Loses Google's Attention

Snowflake

Objective: Run a three-part crawl-budget audit, Crawl Stats, orphan-page detection, and the Indexing report, to find exactly where Google's limited daily crawl allowance is being wasted on a large documentation site, and reclaim it for pages that actually need it.

Snowflake's technical marketing team maintains thousands of docs pages. New docs for a just-shipped feature sometimes take three weeks to appear in Google, while old deprecated pages get recrawled daily. You're asked to find out why, using the lesson's Stage 1 and Stage 2 framework.

Three checks in sequence: read Google's own account of where it spends its daily crawl budget on the site, find pages no internal link points to, then confirm which of the resulting problems are crawl issues versus indexing issues.

Where is Google's limited daily crawl budget being wasted on a large site, and how does that connect to pages not getting indexed?

Crawl Budget Management/Orphan Page Detection/Google Search Console/Site Architecture

Before you start

What you'll need

  • —Familiarity with Google Search Console's Crawl Stats and Indexing reports
  • —Basic understanding of internal linking and sitemaps
Crawl Budget
the rough daily limit on how many pages Googlebot will crawl on a given site, which can be consumed by low-value pages at the expense of important ones.
Orphan Page
a published page that no other page on the site links to, making it much harder for crawlers to discover even if it's listed in a sitemap.

Free path (everything below is enough to finish)

FreeRead Crawl Stats and the Page Indexing report

Free, and the only tool showing Google's own account of what it crawled and why it didn't index something.

FreemiumCross-reference sitemap URLs against a real crawl to find orphan pages

Free tier crawls up to 500 URLs at once and supports List mode for exactly this sitemap-vs-crawl comparison.

Paid upgrades (optional, faster/deeper)

This audit is complete on Search Console and Screaming Frog's free tiers. Ahrefs is only useful afterward, to track whether the fix worked.

Ahrefs(optional)
PaidMonitor whether the crawl-budget fixes actually shift indexing speed for future launches over time

The free path fully diagnoses the problem; Ahrefs is only useful afterward, to track whether the fix worked over following months.

The process

3 steps

Step 01 of 03

Crawl budget as a finite daily allocation

The lesson explains Google allocates a crawl budget, a rough limit on how many pages Googlebot will crawl per day, and that sites with thousands of low-quality or duplicate pages can see important pages crawled infrequently or skipped.

Snowflake's docs site has thousands of URLs. Which page types is Googlebot actually spending its daily crawl budget on, and is it the pages that matter?

Google Search Console— Search Console > Settings > Crawl stats.

Procedure

  1. Open Settings > Crawl stats in Search Console
  2. Review the 'By response' and 'By file type' breakdowns
  3. Check the 'By purpose' split between Discovery (new URLs) and Refresh (recrawling known URLs)
  4. Note which page type is consuming the largest share of daily crawl requests
Sample output
Crawl Stats, past 90 days
------------------------------------------
Total crawl requests: 84,200/day average
By purpose: Refresh 71%, Discovery 29%
By response: 200 OK 82%, 301 redirect 11%, 404 not found 5%, 5xx 2%
Top crawled path pattern: /docs/deprecated/* (19% of all requests)
                          /docs/v1-legacy/* (14% of all requests)

New feature docs published this week: crawled within 48 hours for
only 3 of 11 new pages.

Healthy

The majority of crawl requests hit current, high-value docs paths, and new pages get crawled within a day or two of publishing.

Unhealthy

33% of daily crawl requests go to deprecated and legacy doc paths that no longer need frequent recrawling, while 8 of 11 new pages wait longer than 48 hours.

What this means

Google spends crawl budget on what it has historically found valuable to recrawl, deprecated paths that still get linked internally keep pulling budget away from new content. This is exactly the lesson's warning about low-quality or duplicate pages starving important pages of crawl attention.

So what do I do about it?

SymptomActionEffort
19% + 14% = 33% of daily crawl budget goes to deprecated/legacy doc pathsnoindex or 410 the deprecated paths and remove remaining internal links to themdev ticket
8 of 11 new feature docs took longer than 48 hours to be crawledAdd new docs to the XML sitemap immediately at publish time and request indexing via URL Inspection for launch-critical pages30 min
EitherYou or a developer can handle this, depending on your access.

Step 02 of 03

Orphan pages that crawlers rarely find

The lesson defines an orphan page as one no other page links to, and warns crawlers rarely find pages like this even if they're technically published.

Are any of Snowflake's new feature docs orphaned, published but not linked from anywhere else on the site?

Screaming Frog SEO Spider— Screaming Frog > List mode, crawl the sitemap URLs, then cross-reference against a standard crawl from the homepage.

Procedure

  1. Crawl the site normally from the homepage in Spider mode (up to 500 URLs free)
  2. Switch to List mode and upload the XML sitemap's URL list
  3. Compare the two lists: any sitemap URL that never appeared in the homepage-start crawl is an orphan candidate
  4. Manually check the top 5 orphan candidates for at least one internal link pointing to them
Sample output
Orphan Page Check
------------------------------------------
Sitemap URLs: 3,140
Found via homepage-start crawl: 3,047
Orphan candidates (in sitemap, not reached by crawl): 93

Sample of 5 orphan candidates:
  /docs/features/dynamic-tables-preview   0 internal links found
  /docs/features/cortex-search-beta       0 internal links found
  /docs/reference/api-v3-changelog        0 internal links found
  /docs/features/data-clean-rooms         1 internal link (from a
                                            deprecated page, itself
                                            unlinked)
  /docs/features/snowpark-container       0 internal links found

Healthy

Every new feature doc has at least one internal link from a currently-linked, non-deprecated page within a day of publishing.

Unhealthy

93 sitemap URLs, including several just-launched feature docs, have zero internal links pointing to them anywhere on the crawlable site.

What this means

A page in the sitemap is a hint to Google, not a guarantee, but an orphan page has no link equity flowing to it and depends entirely on the sitemap hint. That's a much weaker discovery signal than an actual internal link, which is exactly why these pages take three weeks instead of three days.

So what do I do about it?

SymptomActionEffort
93 sitemap URLs have zero internal links pointing to themAdd each new feature doc to its category's docs index page and at least one related-docs sidebar the day it publisheshalf day
DeveloperNeeds a developer/engineer to ship the fix.

Step 03 of 03

Not every crawled page gets indexed

The lesson lists four reasons Google skips indexing a crawled page: duplicate/near-duplicate content, thin pages, pages blocked by robots.txt or noindex, and pages it deems low-quality or spammy.

Of the pages that ARE being crawled, how many are failing to convert into indexed pages, and why?

Google Search Console— Search Console > Indexing > Pages report.

Procedure

  1. Open Indexing > Pages in Search Console
  2. Review the 'Why pages aren't indexed' table, sorted by page count
  3. Cross-reference the largest 'Duplicate' bucket against the deprecated/legacy paths found in Step 1
  4. Note whether any of the 93 orphan candidates from Step 2 also appear in the 'Discovered, currently not indexed' bucket
Sample output
Page Indexing Report
------------------------------------------
Indexed: 2,690
Not indexed: 450
  Duplicate, Google chose different canonical: 210 (mostly /docs/v1-legacy/*)
  Crawled, currently not indexed: 140 (mix of thin API-reference stubs)
  Discovered, currently not indexed: 65 (overlaps with 41 of the 93
                                          orphan candidates from Step 2)
  Blocked by robots.txt: 35 (an old /docs/internal/* disallow rule)

Healthy

Not-indexed pages are a small, explainable fraction, mostly intentional (internal-only paths correctly blocked).

Unhealthy

210 legacy pages are creating duplicate-content confusion, and 41 of the same orphan pages flagged in Step 2 are stuck in 'Discovered, currently not indexed' limbo, exactly the outcome an orphan page produces.

What this means

The three checks now tell one consistent story: crawl budget is going to legacy paths (Step 1), new docs are orphaned (Step 2), and those same orphaned docs are stuck as 'discovered but not indexed' because Google found the URL via the sitemap but never found a reason to prioritize crawling it (Step 3). Fixing internal linking solves all three at once.

So what do I do about it?

SymptomActionEffort
41 new docs are stuck in 'Discovered, currently not indexed'Once internal links are added (Step 2 fix), request indexing for these specific URLs via URL Inspection to speed up the first recrawl30 min
210 legacy pages are flagged as duplicate with a Google-chosen canonicalAdd explicit self-referential or cross-referential canonical tags to the legacy docs instead of leaving the choice to Googledev ticket
EitherYou or a developer can handle this, depending on your access.

Analyze your findings

What to look for

Budget allocation
Is crawl budget going disproportionately to low-value paths (deprecated, legacy) instead of current content?
Discovery signal strength
Does a new page have a real internal link, or only a sitemap listing?
Causal chain
Do the findings from separate reports (Crawl Stats, orphan check, Indexing report) point to the same root cause?
Intentional vs accidental exclusion
Is a page not indexed because of a deliberate robots.txt or noindex rule, or an accidental orphaning?

Make the call

Crawl Stats shows 33% of budget going to deprecated/legacy paths. The orphan check finds 93 sitemap URLs with zero internal links. The Indexing report shows 41 of those same 93 pages stuck in 'Discovered, currently not indexed'. What's the single highest-leverage fix?

Recommendation · Priority: High

“Reclaim Snowflake's crawl budget with three parallel fixes: noindex or 410 the deprecated and legacy doc paths currently consuming 33% of daily crawl requests, add each new feature doc to its category index page and a related-docs sidebar at publish time to eliminate orphaning, and add explicit canonical tags to the 210 legacy pages currently flagged as duplicate with a Google-chosen canonical. These three fixes address the full causal chain found in this audit, wasted budget, undiscoverable new pages, and unindexed duplicates, rather than treating each report finding as an isolated issue.”

Common mistakes

What trips people up

  • Treating each Search Console report in isolation — Crawl Stats, orphan detection, and the Indexing report describe the same underlying problem here; reading them separately misses the causal chain.

  • Requesting indexing for stuck pages without fixing internal linking — a manual index request can nudge one URL, but the same orphaning problem recurs for every future doc published the same way.

  • Assuming a sitemap listing is enough to guarantee discovery — a sitemap is a hint, not a guarantee; an actual internal link is a much stronger discovery signal.

  • Fixing only the highest-percentage finding — the 33% crawl-budget stat looks like the biggest number, but leaving orphaning and duplicate-canonical issues unaddressed means new content keeps landing in the same limbo.

Final deliverable

A crawl-budget remediation memo covering three fixes: deprecate/noindex the legacy paths draining budget, internal-link every new doc at publish time, and add explicit canonicals to duplicate-flagged legacy pages.

See a reference example
Sample output
The same three-part audit on Delhivery's logistics-API documentation found a similar pattern: 27% of crawl budget going to a retired v1 API reference, 60+ new endpoint docs with zero internal links, and a matching spike in 'Discovered, not indexed' pages. The fix list looked almost identical.

Success criteria

You're done when you can:

  • Identified the specific page-type pattern consuming disproportionate crawl budget (not just 'crawl budget is wasted somewhere')
  • Found and listed real orphan-page candidates by cross-referencing sitemap vs. crawl-discovered URLs
  • Connected the orphan pages to the matching 'Discovered, not indexed' Search Console reason, showing the causal chain
  • Proposed fixes for all three findings, not just the most obvious one

Key takeaway

A crawl budget problem rarely shows up as one clean finding, it surfaces as related symptoms across multiple Search Console reports: wasted budget on stale paths, undiscoverable new pages, and pages stuck unindexed. Connecting those symptoms into one causal chain, then fixing the root cause, internal linking and budget reclamation, rather than patching individual URLs, is what actually closes the gap.