Skip to content
Academy
Marketing Academy · Field Work●SEO
MiniTeardown· 20 minutes

The robots.txt & Canonical Teardown: Spot the Defect Before It Ships

MapmyIndia (CE Info Systems)

Objective: Given two real-world-realistic specimens, a robots.txt file and a pair of canonical tags, find the defect in each before it reaches production.

MapmyIndia's dev team is about to deploy a redesigned maps-API documentation site. You're the last technical SEO review before launch, checking the staging robots.txt and the new canonical tag setup on two API-plan pages.

Two specimens, one robots.txt file and one pair of canonical tags. Find the defect in each and explain what it would do to the site if it shipped as-is.

Which pre-launch defects would silently damage indexing if they shipped as-is?

robots.txt/Canonical Tags/Pre-Launch QA/Technical SEO Teardown

Before you start

What you'll need

  • —Ability to read a robots.txt file's User-agent, Disallow, and Allow syntax
  • —Understanding of what a canonical tag does
Canonical Tag
an HTML tag (<link rel="canonical">) declaring which URL is the authoritative version of a page's content, consolidating ranking signals to that one URL.
Disallow Directive
a robots.txt rule telling a crawler it may not request matching paths; a bare 'Disallow: /' blocks an entire site for that user-agent.

Free path (everything below is enough to finish)

FreemiumCrawl the staging environment to catch robots.txt and canonical defects before launch

Free for up to 500 URLs, enough to catch exactly this class of pre-launch defect on a documentation site.

The process

Specimens to review

Review this robots.txt pulled from the staging environment, scheduled to go live with the redesign tonight. Find the line that would be catastrophic if it survives the deploy.

Sample output
=== robots.txt, staging environment (scheduled to deploy tonight) ===
User-agent: *
Disallow: /

User-agent: Googlebot
Allow: /docs/
Allow: /pricing/

Sitemap: https://developer.mapmyindia.com/sitemap.xml

Specimen: synthetic, realistic

Review the canonical tags on these two distinct API pricing plan pages. Find the defect that would remove one of them from Google's index.

Sample output
=== /pricing/starter-plan (HTML head) ===
<link rel="canonical" href="https://developer.mapmyindia.com/pricing/starter-plan" />

=== /pricing/enterprise-plan (HTML head) ===
<link rel="canonical" href="https://developer.mapmyindia.com/pricing/starter-plan" />

Specimen: synthetic, realistic

Analyze your findings

What to look for

Scope of the block
Does an Allow rule for one crawler actually cover every path, or only a subset?
Self-reference vs cross-reference
Does each page's canonical tag point to itself, or accidentally to a different page?
Consequence, not just presence
What specifically happens (deindexed, dropped from results, blocked entirely) if the defect ships?
Distractor discipline
Is a stylistic difference (URL format, rule ordering) being mistaken for a functional defect?

Make the call

The robots.txt has 'Disallow: /' under 'User-agent: *' and then 'Allow: /docs/' and 'Allow: /pricing/' under 'User-agent: Googlebot'. What's the actual scope of what stays blocked?

Recommendation · Priority: High

“Block this deploy until both defects are fixed. The staging robots.txt's global 'Disallow: /' combined with narrow Googlebot-only Allow rules for /docs/ and /pricing/ would shut out every other crawler entirely and block Googlebot from every path outside those two, including the homepage, exactly the 'block Google from your entire site overnight' scenario. Separately, the Enterprise plan page's canonical tag incorrectly points at the Starter plan page instead of self-referencing, which would cause Google to treat it as a duplicate and drop it from the index, leaving no distinct page to rank for enterprise-tier pricing searches.”

Common mistakes

What trips people up

  • Assuming a crawler-specific Allow rule fully cancels a wildcard Disallow — the Allow only applies to the specific paths listed; every other path remains blocked for that crawler too.

  • Flagging the sitemap's URL format or rule ordering as the defect — these are cosmetic and don't affect crawling or indexing; the real defect is the scope of the Disallow rule.

  • Missing which direction a canonical points — the defect is the Enterprise page pointing at the Starter page, not the reverse; misreading the direction leads to fixing the wrong page.

  • Describing the canonical defect only as 'a ranking issue' — the actual consequence is more severe, the Enterprise page gets dropped from the index entirely as a perceived duplicate, not merely ranked lower.

Final deliverable

A pre-launch review note flagging both defects with severity and the one-line fix for each, blocking the deploy until both are resolved.

See a reference example
Sample output
The same pre-launch review on RateGain's pricing-page redesign caught a near-identical defect: a leftover 'Disallow: /api-docs/' rule from a six-month-old staging environment, still present in the file scheduled to go live. Same fix: delete the line before merging.

Success criteria

You're done when you can:

  • Identified the 'Disallow: /' rule as the critical blocker, not the sitemap or ordering distractors
  • Explained specifically why the Googlebot Allow rules don't rescue every other path
  • Identified the Enterprise page's canonical pointing at the Starter page as the defect, not the reverse
  • Explained the specific index consequence (Enterprise page dropped, not just 'ranking loss')

Key takeaway

A pre-launch technical SEO review exists to catch defects that look like small syntax details but have site-wide or page-wide consequences if they ship. Reading exactly what a robots.txt rule's scope covers, and exactly which direction a canonical tag points, is the difference between catching a catastrophic block before launch and finding out about it in a traffic report weeks later.