Published by Qomvia, , 3 min read
robots.txt says yes, the edge says no
A robots.txt file is a request. A bot manager is an enforcement layer. When the two disagree, the enforcement layer wins, and the disagreement is invisible from a browser. The site looks fine to you, and a fetcher arriving without a browser fingerprint gets an interstitial that says 'Checking your browser' or a plain 403.
To an answer engine the challenge page is the page. It has no article, no product, no pricing, so the engine cites the competitor that returned content. Nothing in your analytics shows this: challenge pages are usually served before any tag fires.
How to tell if it is happening to you
Fetch your homepage without a browser and with a declared crawler user agent. If you get a shorter body than a browser does, a status of 403 or 503, or markup mentioning a challenge, verification or 'enable JavaScript', the edge is intercepting.
curl -sS -o /dev/null -w "%{http_code} %{size_download} bytes\n" \
-A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" \
https://www.example.com/
# Compare with a browser user agent
curl -sS -o /dev/null -w "%{http_code} %{size_download} bytes\n" \
-A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 Chrome/131.0 Safari/537.36" \
https://www.example.com/Qomvia runs this comparison on every scan: one fetch as QomviaBot, one as a browser, and the Serves content to a non-browser user agent check fails when the bot response is a challenge, a 4xx, or less than half the HTML a browser receives. The reason is listed on your site page with the server header, so you can see which layer answered.
Allowlist declared agents, keep the wall for everyone else
Every serious operator publishes the IP ranges its agents use, and the ranges are the reliable part of the identity; a user-agent string can be forged. The right rule combines both: user agent contains the declared name and source IP is in the operator's published list. Perplexity's own docs give the recipe for Cloudflare custom rules: field User Agent contains PerplexityBot or Perplexity-User, and IP source address is in the ranges from their JSON, action Allow.
- OpenAI:
openai.com/searchbot.json,openai.com/gptbot.json,openai.com/chatgpt-user.json - Perplexity:
perplexity.com/perplexitybot.json,perplexity.com/perplexity-user.json - Google:
developers.google.com/static/search/apis/ipranges/googlebot.jsonfor Googlebot; Google-Extended shares Googlebot's fetches - Anthropic: published in the Help Center article on crawling
If your CDN offers a 'verified bots' or 'known AI crawlers' category, use it: those lists are maintained against the operators' ranges. If the CDN only offers block or challenge for that category, choose Allow for the search and user-fetch agents and decide separately about training crawlers. And keep the wall on the paths that need it, /login, /account, /checkout, where no agent has business anyway.
A longer-term option is Web Bot Auth, where an agent signs its requests and you verify the signature against a published key directory instead of guessing from headers. It is still an IETF draft and only a few operators sign requests today, but a CDN that supports it lets you allow signed agents without maintaining IP lists.
The rendering trap behind the wall
Even when the bot manager lets an agent through, a site that renders in the browser gives it an empty shell. Bot walls and client-side rendering fail the same audience for different reasons, and both are scored under Machine access. The client-side rendering article covers the second half.
Questions
- Will allowing AI agents through my WAF invite scrapers?
- Not if the rule requires both the declared user agent and the operator's published IP range. A scraper can copy the string but not the source address.
- My CDN shows AI bots as 'blocked' by default. Is that a problem?
- For training crawlers it is a policy choice. For search indexers and user-triggered fetchers it removes you from AI answers while your robots.txt still says you are open.
- How do I know which layer served the challenge?
- The response's server and related headers usually name the CDN or bot manager. Qomvia records the status and server header on a failed bot-response check.
Score your own site against the rubric this is written from.
Is your site agent-ready?
Free score against the same rubric, in under a minute.
Sign up free to keep the fixes and track the score.
AI monitor
PreviewHow often each model names your site across 11 tracked questions.