Skip to content
Academy
Marketing Academy · Field Work●SEO
MiniAudit· 20 minutes

Reading the Warning Signs: A Search Console Coverage Audit

HelloFresh

Objective: Given a real Search Console indexing export, separate the content-quality-driven exclusion reasons from routine technical ones, and map each content-quality reason to the correct fix.

HelloFresh's content team can't explain why hundreds of recipe and meal-plan pages sit in the Coverage report instead of the index. You've been handed the raw Search Console indexing export and asked to figure out what's actually going on before anyone starts rewriting pages at random.

Two passes over the same export: first isolate which exclusion reasons are actually about content quality versus routine crawl or server issues, then map each content-quality reason to the lesson's canonicalize/noindex/consolidate framework.

Of the hundreds of pages sitting outside the index, which are a real content-quality problem, and which fix, canonical, noindex, or consolidate, applies to each?

Search Console Auditing/Content Quality Diagnosis/Data Triage

Before you start

What you'll need

  • —Access to a Google Search Console Coverage/Indexing report or export
  • —Basic spreadsheet skills for sorting and sampling
Crawled, currently not indexed
Google's indexing status for a page it has read but chosen not to add to the index, often an early content-quality signal.
Duplicate, no user-selected canonical
an indexing status meaning Google found multiple similar URLs and no canonical tag told it which one to prefer.

Free path (everything below is enough to finish)

FreeSource the Coverage report data and confirm findings against live URLs

Free, and the only source of Google's own indexing decisions per URL, no export substitute replaces it for confirmation.

FreemiumCross-check the sampled duplicate and not-indexed URLs for near-duplicate content at scale

Free tier comfortably covers a 20-URL sample from a 5,280-page site export.

Paid upgrades (optional, faster/deeper)

The diagnosis and fix-mapping in this project are both complete on the free path; paid tools only save time once every flagged page (not just a sample) needs individual review.

Ahrefs(optional)
PaidRun near-duplicate content detection across the full 872 flagged pages instead of a 20-page manual sample

The free path (GSC plus a Screaming Frog sample) is complete for diagnosing the pattern; Ahrefs saves time once triaging the full page count instead of a representative sample.

Download project dataset

The process

2 steps

Step 01 of 02

Watch for 'Crawled, currently not indexed'

The lesson flags a rising 'Crawled, currently not indexed' count in Search Console's Coverage report as an increasingly reliable early warning sign of a thin or duplicate content problem, since Google's Helpful Content system is now fully built into core ranking and evaluates continuously rather than in periodic waves.

Of the six exclusion reasons in this export, which two are actually driven by content quality rather than a routine crawl or server issue, and how many pages does each represent?

Google Search Console— Search Console > Indexing > Pages report, or the exported gsc-indexing-export.csv opened in a spreadsheet.

Procedure

  1. Open gsc-indexing-export.csv in a spreadsheet
  2. Sort the reason column by page count, descending
  3. Isolate 'Crawled, currently not indexed' and 'Duplicate, no user-selected canonical' as the two content-quality-driven reasons
  4. Set aside 'Blocked by robots.txt', 'Not found (404)', and 'Server error (5xx)' as separate technical issues outside this audit's scope
Sample output
Google Search Console Indexing Export - HelloFresh
-----------------------------------------------
Indexed:                                       3,891 pages
Crawled, currently not indexed:                   690 pages  <- content-quality signal
Blocked by robots.txt:                            412 pages  (technical, out of scope)
Duplicate, no user-selected canonical:            182 pages  <- content-quality signal
Not found (404):                                   94 pages  (technical, out of scope)
Server error (5xx):                                11 pages  (technical, out of scope)

Healthy

'Crawled, currently not indexed' and 'Duplicate, no user-selected canonical' together represent a small, stable fraction of total pages, tracked quarter over quarter.

Unhealthy

690 pages sit in 'Crawled, currently not indexed', more than double the site's server errors and 404s combined, meaning Google is actively choosing not to index nearly 700 recipe and meal-plan pages it has already read.

What this means

The lesson is explicit that this specific reason is Google's own quality system quietly flagging pages it doesn't think are worth indexing, which makes 690 the number that actually needs a content decision, not just a technical fix.

So what do I do about it?

SymptomActionEffort
690 pages sit in 'Crawled, currently not indexed' with no plan to address themPull a sample of 20 of the 690 pages and manually check each for thin or duplicate content before deciding a fixhalf day
The team has been treating all six Coverage reasons as one undifferentiated 'indexing problem'Split the Coverage report into content-quality reasons and technical reasons before assigning any fix5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Fixing It: Canonicalize, Noindex, or Consolidate

The lesson's framework: use a canonical tag for legitimate technical duplicates, noindex for pages needed by users but not search, and consolidation when several thin or overlapping pages could genuinely become one comprehensive resource, the fix Google's Helpful Content system rewards most directly.

182 pages are flagged as duplicate with no canonical, and 690 are crawled but not indexed. Which fix applies to which group, and why isn't it the same fix for both?

Google Search Console— Same spreadsheet, cross-referenced against a manual sample check of URLs from each group.

Procedure

  1. For the 182 duplicate pages, sample 10 URLs and confirm they're legitimate technical duplicates (tracking parameters, sort orders on recipe filter pages)
  2. For the 690 not-indexed pages, sample 10 URLs and check whether each is a near-duplicate of another recipe page or genuinely thin on its own
  3. Assign canonical tags to the confirmed technical duplicates
  4. Assign consolidation to genuinely thin, overlapping recipe pages; reserve noindex only for pages that must exist for users but add nothing in search
Sample output
182 'Duplicate, no user-selected canonical' -> sampled 10, all 10 were the same recipe
page under different sort/filter URL parameters -> FIX: add self-referential and
parameter canonical tags.

690 'Crawled, currently not indexed' -> sampled 10: 6 were near-duplicate seasonal
variants of the same recipe ('Summer Chicken Salad' vs 'Chicken Salad, Summer
Edition') -> FIX: consolidate into one comprehensive recipe page. 4 were genuine
internal-use pages (old A/B test landing variants) -> FIX: noindex.

Healthy

Each group gets a fix chosen after actually sampling and reading real URLs, not a single blanket fix applied to all 872 flagged pages at once.

Unhealthy

Noindexing all 872 pages in one bulk operation because it's the fastest fix, including the near-duplicate recipe pages that would have made a genuinely stronger consolidated page.

What this means

The lesson's framework exists precisely because these two groups need different fixes: the 182 are a canonical-tag problem, cheap and mechanical, while a meaningful share of the 690 are a content-strategy problem that a canonical tag alone can't solve.

So what do I do about it?

SymptomActionEffort
182 pages lack a canonical tag on legitimate sort/filter duplicatesAdd canonical tags pointing to the default sort/filter view this sprintdev ticket
6 of 10 sampled 'not indexed' pages are near-duplicate seasonal recipe variantsConsolidate seasonal variants into one comprehensive recipe page with a note on seasonal availability, then 301 redirect the thin variantshalf day
EitherYou or a developer can handle this, depending on your access.

Analyze your findings

What to look for

Reason classification
Is each Coverage exclusion reason actually about content quality, or a routine technical/crawl issue?
Scale relative to the site
How large is the content-quality group relative to total indexed pages and to the technical-issue groups?
Sample before deciding
Has a real sample of URLs from each group been read, not just the reason label and count?
Fix differentiation
Does each sampled sub-group get the fix that matches its actual cause, rather than one blanket fix applied to everything?

Make the call

690 pages sit in 'Crawled, currently not indexed' and 182 in 'Duplicate, no user-selected canonical'. What's the correct approach?

Recommendation · Priority: High

“HelloFresh's 872 content-quality-flagged pages need two different remediation paths, not one. The 182 'Duplicate, no user-selected canonical' pages sampled as legitimate sort/filter parameter duplicates on the same recipe, add self-referential and parameter canonical tags this sprint. The 690 'Crawled, currently not indexed' pages are a mixed group: roughly 60% sampled as near-duplicate seasonal recipe variants that should be consolidated into single comprehensive recipe pages with 301 redirects from the thin variants, while the remainder are internal-use pages (old A/B test landing variants) that should be noindexed rather than deleted.”

Common mistakes

What trips people up

  • Treating all six Coverage exclusion reasons as one undifferentiated 'indexing problem' — robots.txt blocks, 404s, and server errors are technical issues, distinct from the content-quality signal 'Crawled, currently not indexed' actually represents.

  • Applying one fix to an entire flagged group without sampling — the 690 'not indexed' pages contained two genuinely different sub-groups, seasonal duplicates and internal test pages, that needed opposite fixes.

  • Defaulting to noindex as the fast universal fix — noindexing near-duplicate seasonal recipes throws away pages that could have been consolidated into one stronger, more valuable resource.

  • Ignoring 'Crawled, currently not indexed' because it isn't a hard error — the lesson treats this specific status as an early content-quality warning sign, not a benign technicality to deprioritize.

Final deliverable

A one-page memo mapping each Coverage report reason to the correct fix (canonical, noindex, or consolidate), citing the exact page counts and the sampled evidence behind each call.

See a reference example
Sample output
Running the same two-pass audit against an Instacart-style category-page export (illustrative): 'Duplicate, no user-selected canonical' turned out to be sort-order parameters on grocery category pages (?sort=price-asc, ?sort=rating), fixed with parameter canonical tags in a single dev ticket. 'Crawled, currently not indexed' was mostly seasonal recipe-collection pages ('Thanksgiving Sides 2024', 'Thanksgiving Sides 2025') that were consolidated into one evergreen 'Thanksgiving Sides' page updated annually instead of recreated.

Success criteria

You're done when you can:

  • Correctly separated the two content-quality Coverage reasons (690 and 182) from the three technical ones before proposing any fix
  • Sampled real URLs from each flagged group rather than assuming every page in a group needs the same fix
  • Assigned canonical tags specifically to the confirmed sort/filter-parameter duplicates, not the whole 182 blindly
  • Assigned consolidation, not noindex, to genuinely near-duplicate content that could become one stronger page

Key takeaway

A Search Console Coverage report mixes technical issues and content-quality signals under one list, and treating them the same wastes effort. Sampling real URLs from each flagged group, rather than acting on the reason label and count alone, is what reveals whether a canonical tag, a noindex, or a genuine consolidation is the right fix.