Published by Qomvia, , 3 min read
Three jobs, three robots
OpenAI, Anthropic and Perplexity each document a split that looks the same from the outside. One robot builds a training corpus. One builds a search index that answers can draw on. One fetches a single page right now because a user asked. They carry different user-agent strings, publish different IP ranges, and respond differently to robots.txt.
Training crawlers
GPTBot is used, in OpenAI's words, to crawl content that may be used in training foundation models; disallowing it signals that the site's content should not be used that way. ClaudeBot plays the same role for Anthropic. Google-Extended is different in shape: it is a token in robots.txt rather than a separate crawler, controlling whether pages Googlebot already fetched may be used for Gemini training and grounding.
These are the bots most opt-out guides tell you to block. Blocking them has no effect on whether you appear in an AI answer today. It only affects what the next generation of models learned from.
Search indexers
OAI-SearchBot builds the index behind ChatGPT search. OpenAI is explicit: sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. Claude-SearchBot indexes content to improve Claude's search responses. PerplexityBot surfaces and links sites in Perplexity results and is not used for model training.
These bots obey robots.txt and take a day or so to notice a change. They are the bots to allow if you want to be quoted.
User-triggered fetchers
ChatGPT-User, Claude-User and Perplexity-User fetch one page because a person asked a question that needs it. They are not crawlers. OpenAI notes that because these actions are user-initiated, robots.txt rules may not apply. Perplexity says the same more directly: the fetcher generally ignores robots.txt. Anthropic frames Claude-User as the switch for user-initiated requests and warns that disabling it may reduce visibility.
The practical consequence: a site that wants to stop these fetchers has to do it at the network edge, and the same rule then removes it from every live answer. There is no version of blocking a user-fetch agent that keeps the citation.
"GET /pricing HTTP/1.1" 200 ... "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot"
"GET /pricing HTTP/1.1" 200 ... "Mozilla/5.0 ... compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot"
"GET /pricing HTTP/1.1" 200 ... "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot"Each operator publishes the IP ranges its bots use as JSON (for example openai.com/searchbot.json, openai.com/gptbot.json, perplexity.com/perplexitybot.json), so a log line can be verified rather than trusted on its user-agent string alone.
How this shows up in your score
Qomvia's robots.txt allows AI crawlers check under Machine access tests ten agents on /, mixing all three kinds. It also fetches your homepage as a declared, non-browser crawler (Serves content to a non-browser user agent) to catch the case where robots.txt says yes and the bot manager says no. A site can pass the first and fail the second; the second is the one that determines whether a fetcher gets a page or a challenge.
Questions
- Can I allow search agents and block training in one robots.txt?
- Yes. Give OAI-SearchBot, Claude-SearchBot and PerplexityBot their own Allow group and GPTBot, ClaudeBot and Google-Extended a Disallow group. The operators document these as independent settings.
- Why do I see ChatGPT-User on pages I disallowed?
- Because a user asked about that page. OpenAI states that robots.txt rules may not apply to user-initiated fetches. The same is true of Perplexity-User.
- Does one crawl serve both training and search?
- OpenAI says that if a site allows both bots, it may use the results of a single crawl for both purposes to avoid duplicate requests. The robots.txt settings remain independent.
Score your own site against the rubric this is written from.
Is your site agent-ready?
Free score against the same rubric, in under a minute.
Sign up free to keep the fixes and track the score.
AI monitor
PreviewHow often each model names your site across 11 tracked questions.