Reading the Warning Signs: A Search Console Coverage Audit
Objective: Given a real Search Console indexing export, separate the content-quality-driven exclusion reasons from routine technical ones, and map each content-quality reason to the correct fix.
HelloFresh's content team can't explain why hundreds of recipe and meal-plan pages sit in the Coverage report instead of the index. You've been handed the raw Search Console indexing export and asked to figure out what's actually going on before anyone starts rewriting pages at random.
Two passes over the same export: first isolate which exclusion reasons are actually about content quality versus routine crawl or server issues, then map each content-quality reason to the lesson's canonicalize/noindex/consolidate framework.
Of the hundreds of pages sitting outside the index, which are a real content-quality problem, and which fix, canonical, noindex, or consolidate, applies to each?
Before you start
What you'll need
- —Access to a Google Search Console Coverage/Indexing report or export
- —Basic spreadsheet skills for sorting and sampling
- Crawled, currently not indexed
- Google's indexing status for a page it has read but chosen not to add to the index, often an early content-quality signal.
- Duplicate, no user-selected canonical
- an indexing status meaning Google found multiple similar URLs and no canonical tag told it which one to prefer.
Free path (everything below is enough to finish)
Free, and the only source of Google's own indexing decisions per URL, no export substitute replaces it for confirmation.
Free tier comfortably covers a 20-URL sample from a 5,280-page site export.
Paid upgrades (optional, faster/deeper)
The diagnosis and fix-mapping in this project are both complete on the free path; paid tools only save time once every flagged page (not just a sample) needs individual review.
The free path (GSC plus a Screaming Frog sample) is complete for diagnosing the pattern; Ahrefs saves time once triaging the full page count instead of a representative sample.
The process
2 steps
Step 01 of 02
The lesson flags a rising 'Crawled, currently not indexed' count in Search Console's Coverage report as an increasingly reliable early warning sign of a thin or duplicate content problem, since Google's Helpful Content system is now fully built into core ranking and evaluates continuously rather than in periodic waves.
Of the six exclusion reasons in this export, which two are actually driven by content quality rather than a routine crawl or server issue, and how many pages does each represent?
Procedure
- Open gsc-indexing-export.csv in a spreadsheet
- Sort the reason column by page count, descending
- Isolate 'Crawled, currently not indexed' and 'Duplicate, no user-selected canonical' as the two content-quality-driven reasons
- Set aside 'Blocked by robots.txt', 'Not found (404)', and 'Server error (5xx)' as separate technical issues outside this audit's scope
Google Search Console Indexing Export - HelloFresh ----------------------------------------------- Indexed: 3,891 pages Crawled, currently not indexed: 690 pages <- content-quality signal Blocked by robots.txt: 412 pages (technical, out of scope) Duplicate, no user-selected canonical: 182 pages <- content-quality signal Not found (404): 94 pages (technical, out of scope) Server error (5xx): 11 pages (technical, out of scope)
Healthy
'Crawled, currently not indexed' and 'Duplicate, no user-selected canonical' together represent a small, stable fraction of total pages, tracked quarter over quarter.
Unhealthy
690 pages sit in 'Crawled, currently not indexed', more than double the site's server errors and 404s combined, meaning Google is actively choosing not to index nearly 700 recipe and meal-plan pages it has already read.
What this means
The lesson is explicit that this specific reason is Google's own quality system quietly flagging pages it doesn't think are worth indexing, which makes 690 the number that actually needs a content decision, not just a technical fix.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 690 pages sit in 'Crawled, currently not indexed' with no plan to address them | Pull a sample of 20 of the 690 pages and manually check each for thin or duplicate content before deciding a fix | half day |
| The team has been treating all six Coverage reasons as one undifferentiated 'indexing problem' | Split the Coverage report into content-quality reasons and technical reasons before assigning any fix | 5 min |
Step 02 of 02
The lesson's framework: use a canonical tag for legitimate technical duplicates, noindex for pages needed by users but not search, and consolidation when several thin or overlapping pages could genuinely become one comprehensive resource, the fix Google's Helpful Content system rewards most directly.
182 pages are flagged as duplicate with no canonical, and 690 are crawled but not indexed. Which fix applies to which group, and why isn't it the same fix for both?
Procedure
- For the 182 duplicate pages, sample 10 URLs and confirm they're legitimate technical duplicates (tracking parameters, sort orders on recipe filter pages)
- For the 690 not-indexed pages, sample 10 URLs and check whether each is a near-duplicate of another recipe page or genuinely thin on its own
- Assign canonical tags to the confirmed technical duplicates
- Assign consolidation to genuinely thin, overlapping recipe pages; reserve noindex only for pages that must exist for users but add nothing in search
182 'Duplicate, no user-selected canonical' -> sampled 10, all 10 were the same recipe
page under different sort/filter URL parameters -> FIX: add self-referential and
parameter canonical tags.
690 'Crawled, currently not indexed' -> sampled 10: 6 were near-duplicate seasonal
variants of the same recipe ('Summer Chicken Salad' vs 'Chicken Salad, Summer
Edition') -> FIX: consolidate into one comprehensive recipe page. 4 were genuine
internal-use pages (old A/B test landing variants) -> FIX: noindex.Healthy
Each group gets a fix chosen after actually sampling and reading real URLs, not a single blanket fix applied to all 872 flagged pages at once.
Unhealthy
Noindexing all 872 pages in one bulk operation because it's the fastest fix, including the near-duplicate recipe pages that would have made a genuinely stronger consolidated page.
What this means
The lesson's framework exists precisely because these two groups need different fixes: the 182 are a canonical-tag problem, cheap and mechanical, while a meaningful share of the 690 are a content-strategy problem that a canonical tag alone can't solve.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| 182 pages lack a canonical tag on legitimate sort/filter duplicates | Add canonical tags pointing to the default sort/filter view this sprint | dev ticket |
| 6 of 10 sampled 'not indexed' pages are near-duplicate seasonal recipe variants | Consolidate seasonal variants into one comprehensive recipe page with a note on seasonal availability, then 301 redirect the thin variants | half day |
Analyze your findings
What to look for
- Reason classification
- Is each Coverage exclusion reason actually about content quality, or a routine technical/crawl issue?
- Scale relative to the site
- How large is the content-quality group relative to total indexed pages and to the technical-issue groups?
- Sample before deciding
- Has a real sample of URLs from each group been read, not just the reason label and count?
- Fix differentiation
- Does each sampled sub-group get the fix that matches its actual cause, rather than one blanket fix applied to everything?
Make the call
690 pages sit in 'Crawled, currently not indexed' and 182 in 'Duplicate, no user-selected canonical'. What's the correct approach?
Recommendation · Priority: High
“HelloFresh's 872 content-quality-flagged pages need two different remediation paths, not one. The 182 'Duplicate, no user-selected canonical' pages sampled as legitimate sort/filter parameter duplicates on the same recipe, add self-referential and parameter canonical tags this sprint. The 690 'Crawled, currently not indexed' pages are a mixed group: roughly 60% sampled as near-duplicate seasonal recipe variants that should be consolidated into single comprehensive recipe pages with 301 redirects from the thin variants, while the remainder are internal-use pages (old A/B test landing variants) that should be noindexed rather than deleted.”
Common mistakes
What trips people up
Treating all six Coverage exclusion reasons as one undifferentiated 'indexing problem' — robots.txt blocks, 404s, and server errors are technical issues, distinct from the content-quality signal 'Crawled, currently not indexed' actually represents.
Applying one fix to an entire flagged group without sampling — the 690 'not indexed' pages contained two genuinely different sub-groups, seasonal duplicates and internal test pages, that needed opposite fixes.
Defaulting to noindex as the fast universal fix — noindexing near-duplicate seasonal recipes throws away pages that could have been consolidated into one stronger, more valuable resource.
Ignoring 'Crawled, currently not indexed' because it isn't a hard error — the lesson treats this specific status as an early content-quality warning sign, not a benign technicality to deprioritize.
Final deliverable
A one-page memo mapping each Coverage report reason to the correct fix (canonical, noindex, or consolidate), citing the exact page counts and the sampled evidence behind each call.
See a reference example
Running the same two-pass audit against an Instacart-style category-page export (illustrative): 'Duplicate, no user-selected canonical' turned out to be sort-order parameters on grocery category pages (?sort=price-asc, ?sort=rating), fixed with parameter canonical tags in a single dev ticket. 'Crawled, currently not indexed' was mostly seasonal recipe-collection pages ('Thanksgiving Sides 2024', 'Thanksgiving Sides 2025') that were consolidated into one evergreen 'Thanksgiving Sides' page updated annually instead of recreated.Success criteria
You're done when you can:
- Correctly separated the two content-quality Coverage reasons (690 and 182) from the three technical ones before proposing any fix
- Sampled real URLs from each flagged group rather than assuming every page in a group needs the same fix
- Assigned canonical tags specifically to the confirmed sort/filter-parameter duplicates, not the whole 182 blindly
- Assigned consolidation, not noindex, to genuinely near-duplicate content that could become one stronger page
Key takeaway
A Search Console Coverage report mixes technical issues and content-quality signals under one list, and treating them the same wastes effort. Sampling real URLs from each flagged group, rather than acting on the reason label and count alone, is what reveals whether a canonical tag, a noindex, or a genuine consolidation is the right fix.