The Indexability Triage: Reading a Real GSC Export Before Touching Code
Objective: Given a real 5,280-URL Google Search Console indexing export, separate the indexability pillar's two biggest failure reasons and decide which one is a five-minute robots.txt fix and which one needs a content-quality decision.
You're a technical SEO analyst supporting Squarespace's own Help Center documentation. A GSC export just landed showing 1,389 of 5,280 help-doc URLs are not indexed. Leadership wants to know: is this a quick technical fix, or a bigger content problem?
Two passes over the same export: quantify the robots.txt-blocked bucket and its likely one-line fix, then quantify the 'crawled but not indexed' bucket and explain why that one can't be fixed with a config change alone.
Given a real GSC indexing export, which non-indexed bucket is a same-day config fix and which needs a content decision?
Before you start
What you'll need
- —Basic spreadsheet filtering and sorting
- —Familiarity with reading a robots.txt file
- robots.txt
- a text file at a site's root that tells crawlers which paths they may or may not request; a Disallow rule blocks Google from crawling matching URLs entirely.
- Crawled, Currently Not Indexed
- a Search Console status meaning Google successfully crawled a page but chose not to store it in the index, usually due to thin or duplicate content.
Free path (everything below is enough to finish)
Free, and the source of both the export data and the robots.txt tester needed to confirm the fix.
The process
2 steps
Step 01 of 02
The lesson warns a single robots.txt typo, like 'Disallow: /' with nothing else, can block Google from an entire site overnight, and that this is one of the most common and damaging technical SEO errors.
412 of the 5,280 URLs in this export are 'Blocked by robots.txt'. Is that a content-quality problem or a one-line config fix?
Procedure
- Import gsc-indexing-export.csv and sort the 'reason' column
- Confirm 412 pages fall under 'Blocked by robots.txt'
- Open yourdomain.com/robots.txt directly in a browser and read every Disallow line
- Test one of the 412 blocked URLs against the live robots.txt using Search Console's tester
gsc-indexing-export.csv, reason column ------------------------------------------ Indexed: 3,891 Crawled, currently not indexed: 690 Blocked by robots.txt: 412 Duplicate, no user-selected canonical: 182 Not found (404): 94 Server error (5xx): 11 ------------------------------------------ Total: 5,280 robots.txt tester result on one blocked URL: Disallow: /help/archived/* <- matches the blocked URL's path This directive was added 14 months ago during a help-center reorganization and was never removed.
Healthy
The 412 blocked pages are intentionally excluded (truly archived, low-value content) and the robots.txt rule is doing its job.
Unhealthy
The 'archived' directory still contains help docs customers actively search for and link to, meaning a 14-month-old cleanup rule is now blocking live, useful content.
What this means
A robots.txt block is binary and total, Google will not index a blocked page no matter how good it is. When the blocked content turns out to still be valuable, this is the single highest-leverage fix in the whole export: one line change unblocks all 412 pages at once.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 412 pages blocked by a robots.txt rule nobody has revisited in 14 months | Audit the /help/archived/* directory, remove the Disallow line for any subpath still receiving organic traffic or backlinks | 30 min |
Step 02 of 02
The lesson explains a page can be crawled but still excluded from the index for reasons distinct from robots.txt: Google deeming it duplicate, thin, or low-quality even when nothing technically blocks it.
690 URLs are 'Crawled, currently not indexed', the single largest bucket. Why can't this one be fixed with a config change the way the robots.txt bucket can?
Procedure
- In Search Console, click into the 'Crawled, currently not indexed' reason group
- Open 5-10 sample URLs from the list
- For each, check word count and whether the page duplicates an existing help article's topic
- Categorize each sample as thin, duplicate, or genuinely unique but still excluded
Sample of 8 'Crawled, currently not indexed' URLs
------------------------------------------
/help/billing-faq-old 210 words, near-duplicate of
/help/billing-faq (thin + duplicate)
/help/domain-connect-v1 180 words, superseded by
/help/domain-connect (thin + duplicate)
/help/template-swap-legacy 240 words, covers a removed feature
/help/export-content-2019 95 words, outdated screenshots only
4 more samples: same pattern, short pages covering topics a newer,
longer article already covers
Pattern: 690 pages skew toward thin, superseded help articles left
live after their replacement article was published.Healthy
Every crawled page is unique enough and substantial enough that Google chooses to store it, no thin leftover duplicates competing with their own replacements.
Unhealthy
690 thin, superseded articles are sitting uncrawled-into-index because Google has already decided a newer, longer article on the same topic is the better version to store.
What this means
Unlike the robots.txt bucket, there's no config line to flip here. Google is making a quality judgment call, and the fix is a content decision: redirect or merge the old article into its replacement, not a technical toggle.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 690 thin, superseded help articles are being skipped at the indexing stage | 301 redirect each superseded article to its replacement instead of leaving both live | half day |
Analyze your findings
What to look for
- Fix type
- Is the cause a binary technical block (robots.txt) or a quality judgment Google is making (thin or duplicate content)?
- Intent behind the block
- Was the robots.txt rule intentional and still correct, or a stale leftover from an old reorganization?
- Sample before generalizing
- Does a real sample of URLs in the largest bucket actually confirm the suspected pattern (thin, superseded content)?
- Fix vs delete
- Should superseded content be redirected to its replacement, or simply removed?
Make the call
412 pages are blocked by robots.txt, tracing to a 14-month-old /help/archived/* rule. 690 pages are 'crawled, currently not indexed', sampled as thin, superseded articles. Leadership wants the fastest win logged first. Which is it?
Recommendation · Priority: High
“Ship the robots.txt fix first: audit the /help/archived/* directory and remove the Disallow line for any subpath still receiving organic traffic or backlinks, unblocking all 412 pages in a single same-day change. In parallel, begin a content consolidation plan for the 690 'crawled, currently not indexed' pages. Sampled evidence shows they skew toward thin, superseded articles; 301 redirect each to its replacement rather than leaving both versions live, since Google has already signaled it won't index the older version anyway.”
Common mistakes
What trips people up
Treating both buckets as the same type of problem — a robots.txt block and a content-quality exclusion have completely different fixes and timelines; bundling them into one 'fix indexing' ticket obscures which one is actually fast.
Assuming a robots.txt Disallow rule is still correct just because it exists — a 14-month-old rule from a past reorganization can silently block content that's since become valuable again.
Generalizing the crawled-not-indexed cause without sampling real URLs — the pattern of thin, superseded articles only becomes clear by opening several sample pages, not by assuming from the bucket name alone.
Deleting superseded pages instead of redirecting them — a 301 redirect preserves any existing backlinks and traffic pointing at the old URL; outright deletion throws that away.
Final deliverable
A two-line triage memo: which bucket is a same-day technical fix (robots.txt) and which bucket needs a content consolidation plan (crawled-not-indexed), with the exact page counts for each.
See a reference example
Applying the same triage to a Wise help-center export (illustrative) found a near-identical split: an old /support/eur-only/* robots.txt rule blocking 180 now-relevant multi-currency articles, and 310 thin FAQ stubs superseded by longer guides. The robots.txt fix shipped same-day; the content merge took three weeks.
Success criteria
You're done when you can:
- Correctly quantified both buckets from the real CSV (412 blocked, 690 crawled-not-indexed) rather than estimating
- Identified the robots.txt bucket as a config fix and the crawled-not-indexed bucket as a content decision, not the same kind of fix
- Sampled actual URLs in the crawled-not-indexed bucket rather than guessing at the cause
- Proposed 301 redirects (not just deletion) for the superseded articles
Key takeaway
Not every 'not indexed' page has the same cause or the same fix. Separating a binary technical block from a content-quality judgment, and sizing each bucket from the real export instead of guessing, is what turns a vague 1,389-page problem into two clearly prioritized workstreams.