Skip to content
Academy
Marketing Academy · Field Work●SEO
MiniTeardown· 20 minutes

Blocked Everything, Cited Nothing: A Broken robots.txt and llms.txt Pair

Coinbase

Objective: Given a company's actual robots.txt and llms.txt files, spot the defects that block the wrong bots, defeat the team's own stated AI-visibility goal, and turn llms.txt into stale, low-value markdown instead of a useful map.

Coinbase's content team wants its Learn education hub cited when someone asks an AI assistant how to set up two-factor authentication, but does not want its proprietary research pages used to train a competitor's model. Someone drafted a robots.txt and llms.txt pair last quarter and nobody has looked at either file since.

Two specimens: the live robots.txt rules for /learn/ and /research/, and the current llms.txt file. Find what actively works against the team's own stated goals.

Do this site's actual robots.txt and llms.txt files achieve the AI-visibility goal the team says it wants, or work against it?

AI Bot Access Policy/robots.txt Rule Scoping/llms.txt Auditing/Training vs Retrieval Bot Distinction

Before you start

What you'll need

  • —Understanding that different AI companies run separate training and retrieval crawlers under different user-agent names
  • —Basic familiarity with reading a robots.txt Disallow/Allow block
Training Bot vs Retrieval Bot
a training bot (e.g. GPTBot) collects content to train a model on; a retrieval bot (e.g. OAI-SearchBot) fetches a page live to answer a specific user question. The same AI company runs both under different user-agent names, and blocking one does not block the other.
llms.txt
a proposed (not officially enforced) markdown file at a site's root listing its most important pages with short descriptions, meant to give AI systems a clean map of the site.

Free path (everything below is enough to finish)

FreemiumCrawl the site to confirm the dead link and check which pages the current robots.txt rules actually scope

Free for up to 500 URLs, enough to verify both defects against the live site rather than the specimen alone.

Google Search Console(optional)
FreeConfirm which Learn pages actually drive organic traffic, to verify which page is genuinely 'highest-value'

Free, and the most direct source for the traffic ranking the llms.txt file should reflect.

The process

Specimens to review

Coinbase's stated goal: block AI training on /research/, stay citable everywhere else, especially /learn/. Compare that goal against the actual robots.txt rules below.

Sample output
=== robots.txt (excerpt, /learn/ and /research/) ===
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
Disallow: /research/

User-agent: ClaudeBot
User-agent: Claude-User
Disallow: /research/

User-agent: PerplexityBot
Disallow: /

User-agent: Google-Extended
Allow: /

Specimen: synthetic, realistic

Coinbase's /llms.txt file is below, alongside a note about what analytics actually shows as the top Learn page.

Sample output
=== /llms.txt ===
# Coinbase Learn
> Coinbase is the world's most trusted place to buy, sell, and manage crypto. We believe in a future powered by cryptocurrency and are here to guide you every step of the way, because your journey matters to us.

## Docs
- [Getting Started](https://coinbase.com/learn/getting-started-2024): everything you need to know
- [Security Tips](https://coinbase.com/learn/old-security-page): stay safe out there

=== ANALYTICS NOTE ===
/learn/two-factor-authentication-setup/ is the #1 most-visited Learn page by organic sessions. It is not listed above.
/learn/old-security-page/ returns a 404, it was renamed during last quarter's site refresh.

Specimen: synthetic, realistic

Analyze your findings

What to look for

Bot grouping
Are training and retrieval bots from the same company grouped under one shared rule, or scoped individually?
Rule scope vs. stated goal
Does a Disallow rule accidentally block a page the team explicitly wants cited?
Link freshness
Do the links inside llms.txt actually resolve, or point at renamed/removed pages?
Content selection
Does llms.txt list the pages that actually matter (by real traffic), or whatever someone thought to add?

Make the call

Coinbase's robots.txt groups GPTBot, OAI-SearchBot, and ChatGPT-User under one shared Disallow rule for /research/. The stated goal is: block AI training, stay citable when a user asks a live question. What's the correct fix?

Recommendation · Priority: High

“Split the robots.txt rule for /research/ so GPTBot remains disallowed for training while OAI-SearchBot and ChatGPT-User are explicitly allowed for retrieval, and remove the site-wide PerplexityBot block since it currently disallows the very /learn/ pages the team wants cited. Separately, fix the dead link in llms.txt and add the #1-traffic 2FA setup page, which is currently missing entirely. Both files currently work against the stated AI-visibility goal rather than for it.”

Common mistakes

What trips people up

  • Grouping all bots from one AI company under a single rule — training and retrieval crawlers serve different purposes and can be controlled independently; grouping them collapses a nuanced policy into an all-or-nothing block.

  • Missing a site-wide block buried among page-specific rules — the PerplexityBot 'Disallow: /' rule is easy to skim past next to the more detailed /research/-scoped rules, but it silently blocks the entire site including the pages the team most wants cited.

  • Treating llms.txt tone (marketing copy vs. plain description) as the main defect — a dead link and a missing top-traffic page are functionally worse than promotional language, since they actively mislead or omit rather than just read oddly.

Final deliverable

An annotated fix list: which bot rules need splitting, which link needs updating, and which page is missing from llms.txt.

See a reference example
Sample output
Running the same two-file check against a different fintech's setup (illustrative, The Trade Desk): the robots.txt correctly split GPTBot from OAI-SearchBot, but the llms.txt file listed a marketing landing page instead of the actual API documentation root, the page developers were most likely to want summarized for an AI coding assistant.

Success criteria

You're done when you can:

  • Identifies that grouping training and retrieval bots under one Disallow rule defeats the team's own stated citation goal
  • Flags the site-wide PerplexityBot block as directly contradicting the goal of keeping /learn/ citable
  • Catches the dead link in llms.txt as a credibility problem, not just a broken URL
  • Identifies the missing highest-traffic page as the more important defect than the marketing-copy tone

Key takeaway

A robots.txt and llms.txt pair can look deliberately configured while quietly working against the team's own stated goal. Checking each rule and each link against what the team actually wants, not just whether the files are syntactically valid, is what catches a policy that defeats itself.