The Locked Door: Auditing Slack's robots.txt for AI Crawler Access
Objective: Given a real-style robots.txt snippet with a mix of AI-crawler directives, correctly identify which rules block LLM citation, which merely throttle it, and which are already fine, before recommending fixes.
You're a growth marketer at Slack. ChatGPT and Perplexity almost never mention Slack when users ask for the best team-chat tool, even though Slack ranks on page one of Google for the same query. Your first hypothesis: something in robots.txt is blocking the AI crawlers that would otherwise find and cite Slack's help center and comparison pages.
Read each robots.txt block, decide whether it's a critical defect, a moderate defect, or fine as-is, referencing the lesson's Step 1 checklist.
Which robots.txt rules actually block AI crawlers from citing the site, which merely slow them down, and which are unrelated and fine as-is?
Before you start
What you'll need
- —Basic familiarity with robots.txt syntax (User-agent, Disallow, Allow, Crawl-delay)
- —Understanding that named-crawler rules and wildcard rules can interact
- robots.txt
- a text file at a site's root that tells web crawlers, including AI crawlers like GPTBot and ClaudeBot, which paths they may or may not access.
- Crawl-delay
- a directive that limits how often a crawler may request pages, slowing discovery of new or updated content without fully blocking it.
Free path (everything below is enough to finish)
Its custom user-agent switcher and robots.txt checker let you simulate GPTBot, ClaudeBot, and PerplexityBot crawls for free
Free, shareable audit trail for the dev handoff
The process
Specimens to review
Slack's robots.txt has this block. What does it mean for LLMO, and is it a defect?
User-agent: GPTBot Disallow: /
Specimen: synthetic, realistic
ClaudeBot is technically allowed here. Is there still a problem?
User-agent: ClaudeBot Crawl-delay: 300 Allow: /
Specimen: synthetic, realistic
This wildcard rule applies to every crawler, including every AI bot not listed by name above. Is it a defect?
User-agent: * Disallow: /blog/ Disallow: /customer-stories/
Specimen: synthetic, realistic
Analyze your findings
What to look for
- Sitewide vs. scoped blocks
- Is an AI crawler disallowed entirely, or only from a specific path?
- Crawl-delay severity
- Does a crawl-delay value meaningfully slow discovery of new content, or is it negligible?
- Wildcard precedence
- Does a wildcard Disallow silently block AI crawlers not named elsewhere in the file, even if named-bot rules look fine?
- Path relevance
- Are the blocked paths ones that actually carry citation-worthy content (blog, comparison pages), or unrelated internal paths?
Make the call
The file has a named 'Allow: /' rule for GPTBot, but also a separate 'User-agent: *' block with 'Disallow: /blog/'. Does GPTBot get to crawl /blog/?
Recommendation · Priority: High
“Fix all three defects before the next deploy: change the sitewide GPTBot Disallow to Allow, lower or remove the 300-second ClaudeBot crawl-delay, and scope the wildcard Disallow so it no longer silently blocks /blog/ and /customer-stories/ from every AI crawler not named elsewhere in the file. The wildcard fix is the most urgent since it affects the widest set of citation-worthy content.”
Common mistakes
What trips people up
Assuming a named Allow rule for one bot cancels a wildcard Disallow elsewhere in the file — robots.txt directives don't merge across separate user-agent blocks this way; the wildcard rule still applies to any crawler without its own explicit exception.
Treating a crawl-delay as harmless since the bot is technically 'allowed' — a long crawl-delay can functionally block fresh content from being discovered in any useful timeframe, even without a literal Disallow.
Flagging correct, standard syntax as a defect — the distractors (a correctly spelled bot name, a valid Allow: / line) exist specifically to test this; only flag lines that actually change crawler behavior for the worse.
Fixing only the sitewide GPTBot block and missing the wildcard rule — the wildcard Disallow is easy to miss since it doesn't name any AI crawler directly, but it blocks the widest range of citation-worthy paths.
Final deliverable
A corrected robots.txt file with every AI-crawler block resolved or justified, plus a one-page severity-ranked defect log for the dev handoff.
See a reference example
Mailchimp robots.txt audit CRITICAL User-agent: GPTBot / Disallow: / -> fix: change to Allow: / User-agent: * / Disallow: /blog/ -> fix: scope wildcard disallow to /blog/drafts/ only, keep published posts open MODERATE User-agent: ClaudeBot / Crawl-delay: 300 -> fix: lower to Crawl-delay: 10 or remove NOT A DEFECT User-agent: * / Disallow: /admin/ -> leave as-is
Success criteria
You're done when you can:
- Correctly assigns severity to all 3 defects
- Does not flag either distractor line as a defect in any item
- Proposes a specific, actionable fix for each defect
Key takeaway
A robots.txt file can look mostly fine at a glance while still blocking AI crawlers through indirect means: a wildcard rule with no named exception, or a crawl-delay so long it functionally prevents fresh content discovery. Reading each rule for its actual effect on each specific crawler, not just scanning for an obvious sitewide block, is what a real audit requires.