Skip to content
Academy
Marketing Academy · Field Work●SEO
MiniAudit· 25 minutes

Write the Policy, Then Prove It Holds: An AI Bot Access Audit

Robinhood

Objective: Draft a training-vs-retrieval robots.txt split for a real site using the lesson's own checklist, then define exactly what a quarterly recheck of that policy needs to confirm.

Robinhood's marketing lead wants a clear, defensible AI-bot policy for the investing-education blog before the next content sprint: allow retrieval and citation, block training, and actually verify it every quarter instead of publishing it once and forgetting it exists.

Two passes: draft the specific bot-by-bot split using the lesson's checklist, then define the quarterly recheck that keeps the policy from going stale the way most robots.txt AI-bot rules do.

What does a defensible, bot-by-bot AI access policy look like from scratch, and what keeps it from going stale within a quarter?

AI Bot Policy Design/robots.txt Authoring/Quarterly Compliance Recheck/Log-Based Bot Discovery

Before you start

What you'll need

  • —Familiarity with editing a robots.txt file's User-agent/Disallow/Allow syntax
  • —Understanding of the training-vs-retrieval bot distinction (see llms-txt-ai-crawler-management-robots-teardown)

Free path (everything below is enough to finish)

FreemiumFetch and validate the live robots.txt file's syntax before and after edits

Free, built-in robots.txt checking under Configuration, catches syntax errors before the file goes live.

Google Search Console(optional)
FreeSanity-check that the new rules don't accidentally block Googlebot from indexing the blog itself

Free, and the fastest way to confirm the AI-bot rules didn't collide with normal Google indexing.

The process

2 steps

Step 01 of 02

Training bots and retrieval bots from the same AI company are different user-agents, you can allow one and block the other

The lesson's own robots.txt example splits GPTBot/ClaudeBot/PerplexityBot/Google-Extended (training) from OAI-SearchBot/Claude-User/ChatGPT-User (retrieval), naming each bot individually rather than using a wildcard rule.

Robinhood's live robots.txt currently has zero AI-bot-specific rules. Following the lesson's 'decide your training-data stance first' step, which bots get Disallowed for training and which get Allowed for retrieval on the investing-education blog?

Screaming Frog SEO Spider— The site's live robots.txt at robinhood.com/robots.txt, or Configuration > robots.txt > Custom in Screaming Frog.

Procedure

  1. Fetch the live robots.txt and confirm it currently has no AI-bot-specific rules
  2. List the 8 named bots from the lesson's checklist: GPTBot, ClaudeBot, PerplexityBot, Google-Extended (training) and OAI-SearchBot, Claude-User, ChatGPT-User, Perplexity-User (retrieval)
  3. Decide the training-data stance first, block all four training bots, matching the goal of not feeding proprietary investing-education content into model training
  4. Explicitly Allow the four retrieval bots so citation and live answers stay possible
Sample output
robots.txt, before (no AI-bot rules exist)
  User-agent: *
  Disallow:

robots.txt, after (this project's proposed split)
  User-agent: GPTBot
  Disallow: /

  User-agent: OAI-SearchBot
  Allow: /

  User-agent: ChatGPT-User
  Allow: /

  User-agent: ClaudeBot
  Disallow: /

  User-agent: Claude-User
  Allow: /

  User-agent: PerplexityBot
  Allow: /

  User-agent: Google-Extended
  Disallow: /

Healthy

A robots.txt file that names each bot individually and matches every rule to a stated training-vs-retrieval decision.

Unhealthy

Leaving the file with no AI-bot rules at all, which defaults to every bot, training and retrieval alike, being fully allowed by omission.

What this means

Silence in robots.txt is itself a decision, it just isn't an intentional one. Naming each bot explicitly turns an accidental default into a policy someone can actually defend to legal or leadership.

So what do I do about it?

SymptomActionEffort
robots.txt has no AI-bot rules and nobody remembers deciding that on purposePublish the 8-bot split above, scoped to the education blog section specifically30 min
Legal asks 'are we blocking AI training?' and there's no clear answerPoint to the explicit Disallow rules for GPTBot, ClaudeBot, and Google-Extended as the documented answer5 min
DeveloperNeeds a developer/engineer to ship the fix.

Step 02 of 02

Revisit the whole setup quarterly, this space moves fast and last quarter's setup may already be outdated

The lesson's checklist ends with a standing habit: new bots appear, existing ones split into sub-agents, and industry compliance norms keep shifting month to month, so a policy that was correct in Q1 can be stale by Q3.

Three months after publishing the split from Step 1, what exactly does the quarterly recheck need to confirm, beyond just re-reading the same file and assuming nothing changed?

Cloudflare— Server logs, or a hosting/CDN provider's bot-traffic reporting (Cloudflare, Fastly, Vercel) if raw logs aren't accessible.

Procedure

  1. Check server logs (or the CDN dashboard) for hits from each of the 8 named user-agents over the past 30 days
  2. Confirm no new AI-bot user-agent has appeared that isn't yet covered by an explicit rule
  3. Re-read current AI-crawler guidance for any bot that's split into new sub-agents since last quarter
  4. Update the robots.txt file and log the recheck date
Sample output
Quarterly recheck log (illustrative)
  Q1 2026: 8 bots covered, 0 unrecognized AI user-agents in logs. No changes needed.
  Q2 2026: 1 new user-agent detected in logs, 'Amazonbot' (not yet covered). Added Disallow rule.
  Q3 2026 (due): recheck scheduled.

Healthy

A dated log showing each quarter's recheck, including any new bot added, not just a policy that sits untouched.

Unhealthy

Publishing the robots.txt split once and never opening the file again, the exact failure mode the lesson warns about repeatedly.

What this means

'Robots.txt is a request, not a lock' cuts both ways: the rules only work if they cover the bots actually showing up in logs, and new bots appear often enough that a policy older than a quarter is already at risk of missing one.

So what do I do about it?

SymptomActionEffort
The robots.txt split was published once, six months ago, with no recheck sinceAdd a recurring quarterly calendar reminder: check logs, check for new bots, update the file5 min
A new AI bot user-agent shows up in logs with no matching robots.txt ruleAdd an explicit Allow/Disallow rule for it within the week, following the same training-vs-retrieval logic as Step 130 min
YouYou can do this yourself, no engineering access required.

Analyze your findings

What to look for

Individual bot naming
Does the policy name each of the 8 bots explicitly, or rely on a wildcard that can't distinguish training from retrieval?
Stance consistency
Does every training bot get the same Disallow decision, and every retrieval bot get the same Allow decision, based on one clear stance?
Silence as a default
Is the absence of AI-bot rules being treated as a real decision, or as an accident nobody noticed?
Recheck evidence
Does the quarterly process check actual log traffic for new bot user-agents, or just re-read the same static file?

Make the call

Robinhood's robots.txt currently has zero AI-bot-specific rules. What does that silence actually mean for AI training access to the investing-education blog?

Recommendation · Priority: Medium

“Publish the 8-bot training-vs-retrieval split (Disallow GPTBot, ClaudeBot, PerplexityBot, and Google-Extended; Allow OAI-SearchBot, Claude-User, Perplexity-User, and ChatGPT-User) scoped to the education blog, so the current silent-allow-by-default state becomes a documented, defensible policy. Pair it with a recurring quarterly calendar reminder to check server logs for new, uncovered AI-bot user-agents, since this space adds new bots and sub-agents faster than a policy written once can track.”

Common mistakes

What trips people up

  • Assuming no AI-bot rules means AI bots are already blocked — robots.txt is opt-out, not opt-in; an empty AI-bot section means full access by default, not protection.

  • Using a single wildcard rule instead of naming each bot — a wildcard can't express 'block training but allow retrieval' from the same company; only naming each user-agent individually can.

  • Publishing the policy once and treating it as permanent — new AI bots and sub-agents appear regularly; a policy that isn't rechecked quarterly against real log traffic goes stale within months.

Final deliverable

A dated robots.txt training/retrieval split for the education blog, plus a recurring quarterly-recheck log template.

See a reference example
Sample output
The same 8-bot split applied to a different fintech's blog (illustrative, TBO Tek's agent-training resources hub): the Q2 recheck caught 'Applebot-Extended' showing up in logs with no matching rule, closing the gap within the week instead of letting it sit uncovered for another quarter.

Success criteria

You're done when you can:

  • Names all 8 bots individually rather than using a wildcard rule
  • Correctly pairs each training bot with its own separate retrieval-bot counterpart from the same AI company
  • Defines a concrete quarterly recheck procedure, not just 'check it sometimes'
  • Includes a plan for handling a newly discovered AI bot user-agent found in logs

Key takeaway

An AI bot access policy isn't just a file to publish once, it's a stance (which bots get training access, which get retrieval access) that has to be explicit rather than left to robots.txt's silent opt-out default, and rechecked on a schedule since the bot landscape changes faster than most teams remember to look.