Published by Qomvia, , 3 min read
One file, three different decisions
Most robots.txt files were written for Googlebot. The AI platforms reuse the same file but split their traffic into separate user agents, and each agent answers a different question: may we train on this, may we show this in search answers, may we fetch this when a user asks. A blanket Disallow: / for one agent often removes a site from answers the owner wanted to be in.
What each agent does, from the operators' own documentation
- OpenAI.
GPTBotcrawls content that may be used to train foundation models.OAI-SearchBotsurfaces sites in ChatGPT search; sites that disallow it are not shown in search answers.ChatGPT-Userfetches a page when a user asks about it and, because the action is user-initiated, robots.txt rules may not apply. Changes take about 24 hours to propagate. Source: platform.openai.com/docs/bots. - Anthropic.
ClaudeBotcollects content for training.Claude-SearchBotindexes content for search quality.Claude-Userretrieves pages in response to a user question. Anthropic honours the non-standardCrawl-delaydirective and asks for opt-outs via robots.txt rather than IP blocks, because an IP block also stops them reading your robots.txt. Source: Anthropic Help Center. - Perplexity.
PerplexityBotsurfaces and links sites in results and is not used to train models.Perplexity-Userfetches pages for a live answer and, as a user-requested fetch, generally ignores robots.txt. Source: docs.perplexity.ai/guides/bots. - Google.
Google-Extendedis a robots.txt token, not a separate crawler: it controls whether content Googlebot already fetched may be used for Gemini training and grounding. It does not change how a site appears in Google Search. Source: Google crawler overview.
The pattern is consistent across all four operators: training and answering are separate switches. A site can refuse training and still be quoted, or allow training and still be invisible because a bot wall sits in front of the search agent.
A robots.txt that allows answers and refuses training
# Search and user-fetch agents: allowed, so the site can be quoted
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
# Training crawlers: your call. Disallow removes future training, not answers.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
# Everyone else
User-agent: *
Allow: /
Disallow: /account/
Disallow: /cart/
Sitemap: https://www.example.com/sitemap.xmlTwo details matter. Groups are matched by the most specific user agent, so a User-agent: * block does not apply to an agent that has its own group. And a disallowed private path (/account/, /cart/) should be disallowed for every group, including the AI agents you allow, otherwise a fetch on a user's behalf lands on a login page and reads as a broken site.
If you also want to state a policy that survives beyond the fetch (no training, yes to search, yes to grounding), add a Content-Signal line to the same file. The Content Signals article covers the syntax.
What Qomvia checks
The methodology awards up to 9 points under Machine access for robots.txt allows AI crawlers. It reads the live file and tests whether GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Claude-User, Google-Extended, Bingbot, Applebot-Extended and meta-externalagent are allowed on /. A missing file scores 7 with a note that crawl guidance is absent; blocking half or more of the list is a fail. The check is about answers, not policy: a site that blocks only the training crawlers keeps most of its points because search agents still get in.
Questions
- If I block GPTBot, will my site disappear from ChatGPT?
- No. GPTBot controls training. ChatGPT search uses OAI-SearchBot, and live page fetches use ChatGPT-User. Block those two and you disappear from answers; block GPTBot alone and you only opt out of future training.
- Do user-triggered fetchers obey robots.txt?
- Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a person asked for that page. Blocking them requires a WAF rule, which also removes you from the answer.
- Does Google-Extended affect my Google rankings?
- No. Google documents it as a control for Gemini training and grounding only. Search indexing is still governed by the Googlebot rules.
Score your own site against the rubric this is written from.
Is your site agent-ready?
Free score against the same rubric, in under a minute.
Sign up free to keep the fixes and track the score.
AI monitor
PreviewHow often each model names your site across 11 tracked questions.