Skip to content
Academy
Marketing Academy · Field Work●SEO
MiniTeardown· 25 minutes

The Locked Door: Auditing Slack's robots.txt for AI Crawler Access

Slack

Objective: Given a real-style robots.txt snippet with a mix of AI-crawler directives, correctly identify which rules block LLM citation, which merely throttle it, and which are already fine, before recommending fixes.

You're a growth marketer at Slack. ChatGPT and Perplexity almost never mention Slack when users ask for the best team-chat tool, even though Slack ranks on page one of Google for the same query. Your first hypothesis: something in robots.txt is blocking the AI crawlers that would otherwise find and cite Slack's help center and comparison pages.

Read each robots.txt block, decide whether it's a critical defect, a moderate defect, or fine as-is, referencing the lesson's Step 1 checklist.

Which robots.txt rules actually block AI crawlers from citing the site, which merely slow them down, and which are unrelated and fine as-is?

robots.txt auditing/AI crawler access diagnosis/Directive precedence reasoning

Before you start

What you'll need

  • —Basic familiarity with robots.txt syntax (User-agent, Disallow, Allow, Crawl-delay)
  • —Understanding that named-crawler rules and wildcard rules can interact
robots.txt
a text file at a site's root that tells web crawlers, including AI crawlers like GPTBot and ClaudeBot, which paths they may or may not access.
Crawl-delay
a directive that limits how often a crawler may request pages, slowing discovery of new or updated content without fully blocking it.

Free path (everything below is enough to finish)

FreemiumTest how each named user-agent is actually treated against real URLs before publishing the fix

Its custom user-agent switcher and robots.txt checker let you simulate GPTBot, ClaudeBot, and PerplexityBot crawls for free

FreeLog every rule with severity and proposed fix

Free, shareable audit trail for the dev handoff

The process

Specimens to review

Slack's robots.txt has this block. What does it mean for LLMO, and is it a defect?

Sample output
User-agent: GPTBot
Disallow: /

Specimen: synthetic, realistic

ClaudeBot is technically allowed here. Is there still a problem?

Sample output
User-agent: ClaudeBot
Crawl-delay: 300
Allow: /

Specimen: synthetic, realistic

This wildcard rule applies to every crawler, including every AI bot not listed by name above. Is it a defect?

Sample output
User-agent: *
Disallow: /blog/
Disallow: /customer-stories/

Specimen: synthetic, realistic

Analyze your findings

What to look for

Sitewide vs. scoped blocks
Is an AI crawler disallowed entirely, or only from a specific path?
Crawl-delay severity
Does a crawl-delay value meaningfully slow discovery of new content, or is it negligible?
Wildcard precedence
Does a wildcard Disallow silently block AI crawlers not named elsewhere in the file, even if named-bot rules look fine?
Path relevance
Are the blocked paths ones that actually carry citation-worthy content (blog, comparison pages), or unrelated internal paths?

Make the call

The file has a named 'Allow: /' rule for GPTBot, but also a separate 'User-agent: *' block with 'Disallow: /blog/'. Does GPTBot get to crawl /blog/?

Recommendation · Priority: High

“Fix all three defects before the next deploy: change the sitewide GPTBot Disallow to Allow, lower or remove the 300-second ClaudeBot crawl-delay, and scope the wildcard Disallow so it no longer silently blocks /blog/ and /customer-stories/ from every AI crawler not named elsewhere in the file. The wildcard fix is the most urgent since it affects the widest set of citation-worthy content.”

Common mistakes

What trips people up

  • Assuming a named Allow rule for one bot cancels a wildcard Disallow elsewhere in the file — robots.txt directives don't merge across separate user-agent blocks this way; the wildcard rule still applies to any crawler without its own explicit exception.

  • Treating a crawl-delay as harmless since the bot is technically 'allowed' — a long crawl-delay can functionally block fresh content from being discovered in any useful timeframe, even without a literal Disallow.

  • Flagging correct, standard syntax as a defect — the distractors (a correctly spelled bot name, a valid Allow: / line) exist specifically to test this; only flag lines that actually change crawler behavior for the worse.

  • Fixing only the sitewide GPTBot block and missing the wildcard rule — the wildcard Disallow is easy to miss since it doesn't name any AI crawler directly, but it blocks the widest range of citation-worthy paths.

Final deliverable

A corrected robots.txt file with every AI-crawler block resolved or justified, plus a one-page severity-ranked defect log for the dev handoff.

See a reference example
Sample output
Mailchimp robots.txt audit

CRITICAL
  User-agent: GPTBot / Disallow: / -> fix: change to Allow: /
  User-agent: * / Disallow: /blog/ -> fix: scope wildcard disallow to /blog/drafts/ only, keep published posts open

MODERATE
  User-agent: ClaudeBot / Crawl-delay: 300 -> fix: lower to Crawl-delay: 10 or remove

NOT A DEFECT
  User-agent: * / Disallow: /admin/ -> leave as-is

Success criteria

You're done when you can:

  • Correctly assigns severity to all 3 defects
  • Does not flag either distractor line as a defect in any item
  • Proposes a specific, actionable fix for each defect

Key takeaway

A robots.txt file can look mostly fine at a glance while still blocking AI crawlers through indirect means: a wildcard rule with no named exception, or a crawl-delay so long it functionally prevents fresh content discovery. Reading each rule for its actual effect on each specific crawler, not just scanning for an obvious sitewide block, is what a real audit requires.