Write the Policy, Then Prove It Holds: An AI Bot Access Audit
Objective: Draft a training-vs-retrieval robots.txt split for a real site using the lesson's own checklist, then define exactly what a quarterly recheck of that policy needs to confirm.
Robinhood's marketing lead wants a clear, defensible AI-bot policy for the investing-education blog before the next content sprint: allow retrieval and citation, block training, and actually verify it every quarter instead of publishing it once and forgetting it exists.
Two passes: draft the specific bot-by-bot split using the lesson's checklist, then define the quarterly recheck that keeps the policy from going stale the way most robots.txt AI-bot rules do.
What does a defensible, bot-by-bot AI access policy look like from scratch, and what keeps it from going stale within a quarter?
Before you start
What you'll need
- —Familiarity with editing a robots.txt file's User-agent/Disallow/Allow syntax
- —Understanding of the training-vs-retrieval bot distinction (see llms-txt-ai-crawler-management-robots-teardown)
Free path (everything below is enough to finish)
Free, built-in robots.txt checking under Configuration, catches syntax errors before the file goes live.
Free, and the fastest way to confirm the AI-bot rules didn't collide with normal Google indexing.
The process
2 steps
Step 01 of 02
The lesson's own robots.txt example splits GPTBot/ClaudeBot/PerplexityBot/Google-Extended (training) from OAI-SearchBot/Claude-User/ChatGPT-User (retrieval), naming each bot individually rather than using a wildcard rule.
Robinhood's live robots.txt currently has zero AI-bot-specific rules. Following the lesson's 'decide your training-data stance first' step, which bots get Disallowed for training and which get Allowed for retrieval on the investing-education blog?
Procedure
- Fetch the live robots.txt and confirm it currently has no AI-bot-specific rules
- List the 8 named bots from the lesson's checklist: GPTBot, ClaudeBot, PerplexityBot, Google-Extended (training) and OAI-SearchBot, Claude-User, ChatGPT-User, Perplexity-User (retrieval)
- Decide the training-data stance first, block all four training bots, matching the goal of not feeding proprietary investing-education content into model training
- Explicitly Allow the four retrieval bots so citation and live answers stay possible
robots.txt, before (no AI-bot rules exist) User-agent: * Disallow: robots.txt, after (this project's proposed split) User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Google-Extended Disallow: /
Healthy
A robots.txt file that names each bot individually and matches every rule to a stated training-vs-retrieval decision.
Unhealthy
Leaving the file with no AI-bot rules at all, which defaults to every bot, training and retrieval alike, being fully allowed by omission.
What this means
Silence in robots.txt is itself a decision, it just isn't an intentional one. Naming each bot explicitly turns an accidental default into a policy someone can actually defend to legal or leadership.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| robots.txt has no AI-bot rules and nobody remembers deciding that on purpose | Publish the 8-bot split above, scoped to the education blog section specifically | 30 min |
| Legal asks 'are we blocking AI training?' and there's no clear answer | Point to the explicit Disallow rules for GPTBot, ClaudeBot, and Google-Extended as the documented answer | 5 min |
Step 02 of 02
The lesson's checklist ends with a standing habit: new bots appear, existing ones split into sub-agents, and industry compliance norms keep shifting month to month, so a policy that was correct in Q1 can be stale by Q3.
Three months after publishing the split from Step 1, what exactly does the quarterly recheck need to confirm, beyond just re-reading the same file and assuming nothing changed?
Procedure
- Check server logs (or the CDN dashboard) for hits from each of the 8 named user-agents over the past 30 days
- Confirm no new AI-bot user-agent has appeared that isn't yet covered by an explicit rule
- Re-read current AI-crawler guidance for any bot that's split into new sub-agents since last quarter
- Update the robots.txt file and log the recheck date
Quarterly recheck log (illustrative) Q1 2026: 8 bots covered, 0 unrecognized AI user-agents in logs. No changes needed. Q2 2026: 1 new user-agent detected in logs, 'Amazonbot' (not yet covered). Added Disallow rule. Q3 2026 (due): recheck scheduled.
Healthy
A dated log showing each quarter's recheck, including any new bot added, not just a policy that sits untouched.
Unhealthy
Publishing the robots.txt split once and never opening the file again, the exact failure mode the lesson warns about repeatedly.
What this means
'Robots.txt is a request, not a lock' cuts both ways: the rules only work if they cover the bots actually showing up in logs, and new bots appear often enough that a policy older than a quarter is already at risk of missing one.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| The robots.txt split was published once, six months ago, with no recheck since | Add a recurring quarterly calendar reminder: check logs, check for new bots, update the file | 5 min |
| A new AI bot user-agent shows up in logs with no matching robots.txt rule | Add an explicit Allow/Disallow rule for it within the week, following the same training-vs-retrieval logic as Step 1 | 30 min |
Analyze your findings
What to look for
- Individual bot naming
- Does the policy name each of the 8 bots explicitly, or rely on a wildcard that can't distinguish training from retrieval?
- Stance consistency
- Does every training bot get the same Disallow decision, and every retrieval bot get the same Allow decision, based on one clear stance?
- Silence as a default
- Is the absence of AI-bot rules being treated as a real decision, or as an accident nobody noticed?
- Recheck evidence
- Does the quarterly process check actual log traffic for new bot user-agents, or just re-read the same static file?
Make the call
Robinhood's robots.txt currently has zero AI-bot-specific rules. What does that silence actually mean for AI training access to the investing-education blog?
Recommendation · Priority: Medium
“Publish the 8-bot training-vs-retrieval split (Disallow GPTBot, ClaudeBot, PerplexityBot, and Google-Extended; Allow OAI-SearchBot, Claude-User, Perplexity-User, and ChatGPT-User) scoped to the education blog, so the current silent-allow-by-default state becomes a documented, defensible policy. Pair it with a recurring quarterly calendar reminder to check server logs for new, uncovered AI-bot user-agents, since this space adds new bots and sub-agents faster than a policy written once can track.”
Common mistakes
What trips people up
Assuming no AI-bot rules means AI bots are already blocked — robots.txt is opt-out, not opt-in; an empty AI-bot section means full access by default, not protection.
Using a single wildcard rule instead of naming each bot — a wildcard can't express 'block training but allow retrieval' from the same company; only naming each user-agent individually can.
Publishing the policy once and treating it as permanent — new AI bots and sub-agents appear regularly; a policy that isn't rechecked quarterly against real log traffic goes stale within months.
Final deliverable
A dated robots.txt training/retrieval split for the education blog, plus a recurring quarterly-recheck log template.
See a reference example
The same 8-bot split applied to a different fintech's blog (illustrative, TBO Tek's agent-training resources hub): the Q2 recheck caught 'Applebot-Extended' showing up in logs with no matching rule, closing the gap within the week instead of letting it sit uncovered for another quarter.
Success criteria
You're done when you can:
- Names all 8 bots individually rather than using a wildcard rule
- Correctly pairs each training bot with its own separate retrieval-bot counterpart from the same AI company
- Defines a concrete quarterly recheck procedure, not just 'check it sometimes'
- Includes a plan for handling a newly discovered AI bot user-agent found in logs
Key takeaway
An AI bot access policy isn't just a file to publish once, it's a stance (which bots get training access, which get retrieval access) that has to be explicit rather than left to robots.txt's silent opt-out default, and rechecked on a schedule since the bot landscape changes faster than most teams remember to look.