How do you separate AI search access from AI training access in robots.txt?
by•
I keep seeing teams treat GPTBot, OAI-SearchBot, ClaudeBot and Google-Extended as if they had one purpose. They do not. A site may want to allow AI search retrieval while restricting training, or make a different choice for each crawler. How are you documenting that policy for your team? Do you keep a tested robots.txt example, or just edit rules when a new crawler appears? The tricky part is that robots.txt only expresses a declared preference. It does not prove that a firewall allows the request, that a page is indexed, or that an AI system will cite it. I would be interested in practical examples of how others verify the separate layers.
1 view
Replies