Skip to content
Academy
Marketing Academy · Field Work●SEO
MiniAudit· 20 minutes

The Indexability Triage: Reading a Real GSC Export Before Touching Code

Squarespace

Objective: Given a real 5,280-URL Google Search Console indexing export, separate the indexability pillar's two biggest failure reasons and decide which one is a five-minute robots.txt fix and which one needs a content-quality decision.

You're a technical SEO analyst supporting Squarespace's own Help Center documentation. A GSC export just landed showing 1,389 of 5,280 help-doc URLs are not indexed. Leadership wants to know: is this a quick technical fix, or a bigger content problem?

Two passes over the same export: quantify the robots.txt-blocked bucket and its likely one-line fix, then quantify the 'crawled but not indexed' bucket and explain why that one can't be fixed with a config change alone.

Given a real GSC indexing export, which non-indexed bucket is a same-day config fix and which needs a content decision?

Technical SEO/robots.txt/Indexability Diagnosis/Content Consolidation

Before you start

What you'll need

  • —Basic spreadsheet filtering and sorting
  • —Familiarity with reading a robots.txt file
robots.txt
a text file at a site's root that tells crawlers which paths they may or may not request; a Disallow rule blocks Google from crawling matching URLs entirely.
Crawled, Currently Not Indexed
a Search Console status meaning Google successfully crawled a page but chose not to store it in the index, usually due to thin or duplicate content.

Free path (everything below is enough to finish)

FreeRead the indexing export and test robots.txt directives live

Free, and the source of both the export data and the robots.txt tester needed to confirm the fix.

Download project dataset

The process

2 steps

Step 01 of 02

robots.txt

The lesson warns a single robots.txt typo, like 'Disallow: /' with nothing else, can block Google from an entire site overnight, and that this is one of the most common and damaging technical SEO errors.

412 of the 5,280 URLs in this export are 'Blocked by robots.txt'. Is that a content-quality problem or a one-line config fix?

Google Search Console— Import gsc-indexing-export.csv into a spreadsheet, then verify the live robots.txt in Search Console's robots.txt tester.

Procedure

  1. Import gsc-indexing-export.csv and sort the 'reason' column
  2. Confirm 412 pages fall under 'Blocked by robots.txt'
  3. Open yourdomain.com/robots.txt directly in a browser and read every Disallow line
  4. Test one of the 412 blocked URLs against the live robots.txt using Search Console's tester
Sample output
gsc-indexing-export.csv, reason column
------------------------------------------
Indexed:                                     3,891
Crawled, currently not indexed:                690
Blocked by robots.txt:                         412
Duplicate, no user-selected canonical:         182
Not found (404):                                94
Server error (5xx):                             11
------------------------------------------
Total:                                        5,280

robots.txt tester result on one blocked URL:
  Disallow: /help/archived/*   <- matches the blocked URL's path
  This directive was added 14 months ago during a help-center
  reorganization and was never removed.

Healthy

The 412 blocked pages are intentionally excluded (truly archived, low-value content) and the robots.txt rule is doing its job.

Unhealthy

The 'archived' directory still contains help docs customers actively search for and link to, meaning a 14-month-old cleanup rule is now blocking live, useful content.

What this means

A robots.txt block is binary and total, Google will not index a blocked page no matter how good it is. When the blocked content turns out to still be valuable, this is the single highest-leverage fix in the whole export: one line change unblocks all 412 pages at once.

So what do I do about it?

SymptomActionEffort
412 pages blocked by a robots.txt rule nobody has revisited in 14 monthsAudit the /help/archived/* directory, remove the Disallow line for any subpath still receiving organic traffic or backlinks30 min
DeveloperNeeds a developer/engineer to ship the fix.

Step 02 of 02

Indexability

The lesson explains a page can be crawled but still excluded from the index for reasons distinct from robots.txt: Google deeming it duplicate, thin, or low-quality even when nothing technically blocks it.

690 URLs are 'Crawled, currently not indexed', the single largest bucket. Why can't this one be fixed with a config change the way the robots.txt bucket can?

Google Search Console— Search Console > Indexing > Pages, filter to 'Crawled, currently not indexed', sample 5-10 URLs.

Procedure

  1. In Search Console, click into the 'Crawled, currently not indexed' reason group
  2. Open 5-10 sample URLs from the list
  3. For each, check word count and whether the page duplicates an existing help article's topic
  4. Categorize each sample as thin, duplicate, or genuinely unique but still excluded
Sample output
Sample of 8 'Crawled, currently not indexed' URLs
------------------------------------------
  /help/billing-faq-old         210 words, near-duplicate of
                                 /help/billing-faq (thin + duplicate)
  /help/domain-connect-v1       180 words, superseded by
                                 /help/domain-connect (thin + duplicate)
  /help/template-swap-legacy    240 words, covers a removed feature
  /help/export-content-2019     95 words, outdated screenshots only
  4 more samples: same pattern, short pages covering topics a newer,
  longer article already covers

Pattern: 690 pages skew toward thin, superseded help articles left
live after their replacement article was published.

Healthy

Every crawled page is unique enough and substantial enough that Google chooses to store it, no thin leftover duplicates competing with their own replacements.

Unhealthy

690 thin, superseded articles are sitting uncrawled-into-index because Google has already decided a newer, longer article on the same topic is the better version to store.

What this means

Unlike the robots.txt bucket, there's no config line to flip here. Google is making a quality judgment call, and the fix is a content decision: redirect or merge the old article into its replacement, not a technical toggle.

So what do I do about it?

SymptomActionEffort
690 thin, superseded help articles are being skipped at the indexing stage301 redirect each superseded article to its replacement instead of leaving both livehalf day
YouYou can do this yourself, no engineering access required.

Analyze your findings

What to look for

Fix type
Is the cause a binary technical block (robots.txt) or a quality judgment Google is making (thin or duplicate content)?
Intent behind the block
Was the robots.txt rule intentional and still correct, or a stale leftover from an old reorganization?
Sample before generalizing
Does a real sample of URLs in the largest bucket actually confirm the suspected pattern (thin, superseded content)?
Fix vs delete
Should superseded content be redirected to its replacement, or simply removed?

Make the call

412 pages are blocked by robots.txt, tracing to a 14-month-old /help/archived/* rule. 690 pages are 'crawled, currently not indexed', sampled as thin, superseded articles. Leadership wants the fastest win logged first. Which is it?

Recommendation · Priority: High

“Ship the robots.txt fix first: audit the /help/archived/* directory and remove the Disallow line for any subpath still receiving organic traffic or backlinks, unblocking all 412 pages in a single same-day change. In parallel, begin a content consolidation plan for the 690 'crawled, currently not indexed' pages. Sampled evidence shows they skew toward thin, superseded articles; 301 redirect each to its replacement rather than leaving both versions live, since Google has already signaled it won't index the older version anyway.”

Common mistakes

What trips people up

  • Treating both buckets as the same type of problem — a robots.txt block and a content-quality exclusion have completely different fixes and timelines; bundling them into one 'fix indexing' ticket obscures which one is actually fast.

  • Assuming a robots.txt Disallow rule is still correct just because it exists — a 14-month-old rule from a past reorganization can silently block content that's since become valuable again.

  • Generalizing the crawled-not-indexed cause without sampling real URLs — the pattern of thin, superseded articles only becomes clear by opening several sample pages, not by assuming from the bucket name alone.

  • Deleting superseded pages instead of redirecting them — a 301 redirect preserves any existing backlinks and traffic pointing at the old URL; outright deletion throws that away.

Final deliverable

A two-line triage memo: which bucket is a same-day technical fix (robots.txt) and which bucket needs a content consolidation plan (crawled-not-indexed), with the exact page counts for each.

See a reference example
Sample output
Applying the same triage to a Wise help-center export (illustrative) found a near-identical split: an old /support/eur-only/* robots.txt rule blocking 180 now-relevant multi-currency articles, and 310 thin FAQ stubs superseded by longer guides. The robots.txt fix shipped same-day; the content merge took three weeks.

Success criteria

You're done when you can:

  • Correctly quantified both buckets from the real CSV (412 blocked, 690 crawled-not-indexed) rather than estimating
  • Identified the robots.txt bucket as a config fix and the crawled-not-indexed bucket as a content decision, not the same kind of fix
  • Sampled actual URLs in the crawled-not-indexed bucket rather than guessing at the cause
  • Proposed 301 redirects (not just deletion) for the superseded articles

Key takeaway

Not every 'not indexed' page has the same cause or the same fix. Separating a binary technical block from a content-quality judgment, and sizing each bucket from the real export instead of guessing, is what turns a vague 1,389-page problem into two clearly prioritized workstreams.