The robots.txt & Canonical Teardown: Spot the Defect Before It Ships
Objective: Given two real-world-realistic specimens, a robots.txt file and a pair of canonical tags, find the defect in each before it reaches production.
MapmyIndia's dev team is about to deploy a redesigned maps-API documentation site. You're the last technical SEO review before launch, checking the staging robots.txt and the new canonical tag setup on two API-plan pages.
Two specimens, one robots.txt file and one pair of canonical tags. Find the defect in each and explain what it would do to the site if it shipped as-is.
Which pre-launch defects would silently damage indexing if they shipped as-is?
Before you start
What you'll need
- —Ability to read a robots.txt file's User-agent, Disallow, and Allow syntax
- —Understanding of what a canonical tag does
- Canonical Tag
- an HTML tag (<link rel="canonical">) declaring which URL is the authoritative version of a page's content, consolidating ranking signals to that one URL.
- Disallow Directive
- a robots.txt rule telling a crawler it may not request matching paths; a bare 'Disallow: /' blocks an entire site for that user-agent.
Free path (everything below is enough to finish)
Free for up to 500 URLs, enough to catch exactly this class of pre-launch defect on a documentation site.
The process
Specimens to review
Review this robots.txt pulled from the staging environment, scheduled to go live with the redesign tonight. Find the line that would be catastrophic if it survives the deploy.
=== robots.txt, staging environment (scheduled to deploy tonight) === User-agent: * Disallow: / User-agent: Googlebot Allow: /docs/ Allow: /pricing/ Sitemap: https://developer.mapmyindia.com/sitemap.xml
Specimen: synthetic, realistic
Review the canonical tags on these two distinct API pricing plan pages. Find the defect that would remove one of them from Google's index.
=== /pricing/starter-plan (HTML head) === <link rel="canonical" href="https://developer.mapmyindia.com/pricing/starter-plan" /> === /pricing/enterprise-plan (HTML head) === <link rel="canonical" href="https://developer.mapmyindia.com/pricing/starter-plan" />
Specimen: synthetic, realistic
Analyze your findings
What to look for
- Scope of the block
- Does an Allow rule for one crawler actually cover every path, or only a subset?
- Self-reference vs cross-reference
- Does each page's canonical tag point to itself, or accidentally to a different page?
- Consequence, not just presence
- What specifically happens (deindexed, dropped from results, blocked entirely) if the defect ships?
- Distractor discipline
- Is a stylistic difference (URL format, rule ordering) being mistaken for a functional defect?
Make the call
The robots.txt has 'Disallow: /' under 'User-agent: *' and then 'Allow: /docs/' and 'Allow: /pricing/' under 'User-agent: Googlebot'. What's the actual scope of what stays blocked?
Recommendation · Priority: High
“Block this deploy until both defects are fixed. The staging robots.txt's global 'Disallow: /' combined with narrow Googlebot-only Allow rules for /docs/ and /pricing/ would shut out every other crawler entirely and block Googlebot from every path outside those two, including the homepage, exactly the 'block Google from your entire site overnight' scenario. Separately, the Enterprise plan page's canonical tag incorrectly points at the Starter plan page instead of self-referencing, which would cause Google to treat it as a duplicate and drop it from the index, leaving no distinct page to rank for enterprise-tier pricing searches.”
Common mistakes
What trips people up
Assuming a crawler-specific Allow rule fully cancels a wildcard Disallow — the Allow only applies to the specific paths listed; every other path remains blocked for that crawler too.
Flagging the sitemap's URL format or rule ordering as the defect — these are cosmetic and don't affect crawling or indexing; the real defect is the scope of the Disallow rule.
Missing which direction a canonical points — the defect is the Enterprise page pointing at the Starter page, not the reverse; misreading the direction leads to fixing the wrong page.
Describing the canonical defect only as 'a ranking issue' — the actual consequence is more severe, the Enterprise page gets dropped from the index entirely as a perceived duplicate, not merely ranked lower.
Final deliverable
A pre-launch review note flagging both defects with severity and the one-line fix for each, blocking the deploy until both are resolved.
See a reference example
The same pre-launch review on RateGain's pricing-page redesign caught a near-identical defect: a leftover 'Disallow: /api-docs/' rule from a six-month-old staging environment, still present in the file scheduled to go live. Same fix: delete the line before merging.
Success criteria
You're done when you can:
- Identified the 'Disallow: /' rule as the critical blocker, not the sitemap or ordering distractors
- Explained specifically why the Googlebot Allow rules don't rescue every other path
- Identified the Enterprise page's canonical pointing at the Starter page as the defect, not the reverse
- Explained the specific index consequence (Enterprise page dropped, not just 'ranking loss')
Key takeaway
A pre-launch technical SEO review exists to catch defects that look like small syntax details but have site-wide or page-wide consequences if they ship. Reading exactly what a robots.txt rule's scope covers, and exactly which direction a canonical tag points, is the difference between catching a catastrophic block before launch and finding out about it in a traffic report weeks later.