The Crawl Budget Audit: Finding Where Snowflake's Docs Section Loses Google's Attention
Objective: Run a three-part crawl-budget audit, Crawl Stats, orphan-page detection, and the Indexing report, to find exactly where Google's limited daily crawl allowance is being wasted on a large documentation site, and reclaim it for pages that actually need it.
Snowflake's technical marketing team maintains thousands of docs pages. New docs for a just-shipped feature sometimes take three weeks to appear in Google, while old deprecated pages get recrawled daily. You're asked to find out why, using the lesson's Stage 1 and Stage 2 framework.
Three checks in sequence: read Google's own account of where it spends its daily crawl budget on the site, find pages no internal link points to, then confirm which of the resulting problems are crawl issues versus indexing issues.
Where is Google's limited daily crawl budget being wasted on a large site, and how does that connect to pages not getting indexed?
Before you start
What you'll need
- —Familiarity with Google Search Console's Crawl Stats and Indexing reports
- —Basic understanding of internal linking and sitemaps
- Crawl Budget
- the rough daily limit on how many pages Googlebot will crawl on a given site, which can be consumed by low-value pages at the expense of important ones.
- Orphan Page
- a published page that no other page on the site links to, making it much harder for crawlers to discover even if it's listed in a sitemap.
Free path (everything below is enough to finish)
Free, and the only tool showing Google's own account of what it crawled and why it didn't index something.
Free tier crawls up to 500 URLs at once and supports List mode for exactly this sitemap-vs-crawl comparison.
Paid upgrades (optional, faster/deeper)
This audit is complete on Search Console and Screaming Frog's free tiers. Ahrefs is only useful afterward, to track whether the fix worked.
The free path fully diagnoses the problem; Ahrefs is only useful afterward, to track whether the fix worked over following months.
The process
3 steps
Step 01 of 03
The lesson explains Google allocates a crawl budget, a rough limit on how many pages Googlebot will crawl per day, and that sites with thousands of low-quality or duplicate pages can see important pages crawled infrequently or skipped.
Snowflake's docs site has thousands of URLs. Which page types is Googlebot actually spending its daily crawl budget on, and is it the pages that matter?
Procedure
- Open Settings > Crawl stats in Search Console
- Review the 'By response' and 'By file type' breakdowns
- Check the 'By purpose' split between Discovery (new URLs) and Refresh (recrawling known URLs)
- Note which page type is consuming the largest share of daily crawl requests
Crawl Stats, past 90 days
------------------------------------------
Total crawl requests: 84,200/day average
By purpose: Refresh 71%, Discovery 29%
By response: 200 OK 82%, 301 redirect 11%, 404 not found 5%, 5xx 2%
Top crawled path pattern: /docs/deprecated/* (19% of all requests)
/docs/v1-legacy/* (14% of all requests)
New feature docs published this week: crawled within 48 hours for
only 3 of 11 new pages.Healthy
The majority of crawl requests hit current, high-value docs paths, and new pages get crawled within a day or two of publishing.
Unhealthy
33% of daily crawl requests go to deprecated and legacy doc paths that no longer need frequent recrawling, while 8 of 11 new pages wait longer than 48 hours.
What this means
Google spends crawl budget on what it has historically found valuable to recrawl, deprecated paths that still get linked internally keep pulling budget away from new content. This is exactly the lesson's warning about low-quality or duplicate pages starving important pages of crawl attention.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 19% + 14% = 33% of daily crawl budget goes to deprecated/legacy doc paths | noindex or 410 the deprecated paths and remove remaining internal links to them | dev ticket |
| 8 of 11 new feature docs took longer than 48 hours to be crawled | Add new docs to the XML sitemap immediately at publish time and request indexing via URL Inspection for launch-critical pages | 30 min |
Step 02 of 03
The lesson defines an orphan page as one no other page links to, and warns crawlers rarely find pages like this even if they're technically published.
Are any of Snowflake's new feature docs orphaned, published but not linked from anywhere else on the site?
Procedure
- Crawl the site normally from the homepage in Spider mode (up to 500 URLs free)
- Switch to List mode and upload the XML sitemap's URL list
- Compare the two lists: any sitemap URL that never appeared in the homepage-start crawl is an orphan candidate
- Manually check the top 5 orphan candidates for at least one internal link pointing to them
Orphan Page Check
------------------------------------------
Sitemap URLs: 3,140
Found via homepage-start crawl: 3,047
Orphan candidates (in sitemap, not reached by crawl): 93
Sample of 5 orphan candidates:
/docs/features/dynamic-tables-preview 0 internal links found
/docs/features/cortex-search-beta 0 internal links found
/docs/reference/api-v3-changelog 0 internal links found
/docs/features/data-clean-rooms 1 internal link (from a
deprecated page, itself
unlinked)
/docs/features/snowpark-container 0 internal links foundHealthy
Every new feature doc has at least one internal link from a currently-linked, non-deprecated page within a day of publishing.
Unhealthy
93 sitemap URLs, including several just-launched feature docs, have zero internal links pointing to them anywhere on the crawlable site.
What this means
A page in the sitemap is a hint to Google, not a guarantee, but an orphan page has no link equity flowing to it and depends entirely on the sitemap hint. That's a much weaker discovery signal than an actual internal link, which is exactly why these pages take three weeks instead of three days.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 93 sitemap URLs have zero internal links pointing to them | Add each new feature doc to its category's docs index page and at least one related-docs sidebar the day it publishes | half day |
Step 03 of 03
The lesson lists four reasons Google skips indexing a crawled page: duplicate/near-duplicate content, thin pages, pages blocked by robots.txt or noindex, and pages it deems low-quality or spammy.
Of the pages that ARE being crawled, how many are failing to convert into indexed pages, and why?
Procedure
- Open Indexing > Pages in Search Console
- Review the 'Why pages aren't indexed' table, sorted by page count
- Cross-reference the largest 'Duplicate' bucket against the deprecated/legacy paths found in Step 1
- Note whether any of the 93 orphan candidates from Step 2 also appear in the 'Discovered, currently not indexed' bucket
Page Indexing Report
------------------------------------------
Indexed: 2,690
Not indexed: 450
Duplicate, Google chose different canonical: 210 (mostly /docs/v1-legacy/*)
Crawled, currently not indexed: 140 (mix of thin API-reference stubs)
Discovered, currently not indexed: 65 (overlaps with 41 of the 93
orphan candidates from Step 2)
Blocked by robots.txt: 35 (an old /docs/internal/* disallow rule)Healthy
Not-indexed pages are a small, explainable fraction, mostly intentional (internal-only paths correctly blocked).
Unhealthy
210 legacy pages are creating duplicate-content confusion, and 41 of the same orphan pages flagged in Step 2 are stuck in 'Discovered, currently not indexed' limbo, exactly the outcome an orphan page produces.
What this means
The three checks now tell one consistent story: crawl budget is going to legacy paths (Step 1), new docs are orphaned (Step 2), and those same orphaned docs are stuck as 'discovered but not indexed' because Google found the URL via the sitemap but never found a reason to prioritize crawling it (Step 3). Fixing internal linking solves all three at once.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 41 new docs are stuck in 'Discovered, currently not indexed' | Once internal links are added (Step 2 fix), request indexing for these specific URLs via URL Inspection to speed up the first recrawl | 30 min |
| 210 legacy pages are flagged as duplicate with a Google-chosen canonical | Add explicit self-referential or cross-referential canonical tags to the legacy docs instead of leaving the choice to Google | dev ticket |
Analyze your findings
What to look for
- Budget allocation
- Is crawl budget going disproportionately to low-value paths (deprecated, legacy) instead of current content?
- Discovery signal strength
- Does a new page have a real internal link, or only a sitemap listing?
- Causal chain
- Do the findings from separate reports (Crawl Stats, orphan check, Indexing report) point to the same root cause?
- Intentional vs accidental exclusion
- Is a page not indexed because of a deliberate robots.txt or noindex rule, or an accidental orphaning?
Make the call
Crawl Stats shows 33% of budget going to deprecated/legacy paths. The orphan check finds 93 sitemap URLs with zero internal links. The Indexing report shows 41 of those same 93 pages stuck in 'Discovered, currently not indexed'. What's the single highest-leverage fix?
Recommendation · Priority: High
“Reclaim Snowflake's crawl budget with three parallel fixes: noindex or 410 the deprecated and legacy doc paths currently consuming 33% of daily crawl requests, add each new feature doc to its category index page and a related-docs sidebar at publish time to eliminate orphaning, and add explicit canonical tags to the 210 legacy pages currently flagged as duplicate with a Google-chosen canonical. These three fixes address the full causal chain found in this audit, wasted budget, undiscoverable new pages, and unindexed duplicates, rather than treating each report finding as an isolated issue.”
Common mistakes
What trips people up
Treating each Search Console report in isolation — Crawl Stats, orphan detection, and the Indexing report describe the same underlying problem here; reading them separately misses the causal chain.
Requesting indexing for stuck pages without fixing internal linking — a manual index request can nudge one URL, but the same orphaning problem recurs for every future doc published the same way.
Assuming a sitemap listing is enough to guarantee discovery — a sitemap is a hint, not a guarantee; an actual internal link is a much stronger discovery signal.
Fixing only the highest-percentage finding — the 33% crawl-budget stat looks like the biggest number, but leaving orphaning and duplicate-canonical issues unaddressed means new content keeps landing in the same limbo.
Final deliverable
A crawl-budget remediation memo covering three fixes: deprecate/noindex the legacy paths draining budget, internal-link every new doc at publish time, and add explicit canonicals to duplicate-flagged legacy pages.
See a reference example
The same three-part audit on Delhivery's logistics-API documentation found a similar pattern: 27% of crawl budget going to a retired v1 API reference, 60+ new endpoint docs with zero internal links, and a matching spike in 'Discovered, not indexed' pages. The fix list looked almost identical.
Success criteria
You're done when you can:
- Identified the specific page-type pattern consuming disproportionate crawl budget (not just 'crawl budget is wasted somewhere')
- Found and listed real orphan-page candidates by cross-referencing sitemap vs. crawl-discovered URLs
- Connected the orphan pages to the matching 'Discovered, not indexed' Search Console reason, showing the causal chain
- Proposed fixes for all three findings, not just the most obvious one
Key takeaway
A crawl budget problem rarely shows up as one clean finding, it surfaces as related symptoms across multiple Search Console reports: wasted budget on stale paths, undiscoverable new pages, and pages stuck unindexed. Connecting those symptoms into one causal chain, then fixing the root cause, internal linking and budget reclamation, rather than patching individual URLs, is what actually closes the gap.