Published by Qomvia, , 3 min read
Fetching and using are different permissions
robots.txt was designed in 1994 to answer one question: may this crawler fetch this path. It has no vocabulary for what happens next. A site that wants to appear in AI answers but not in training sets has had to express that by guessing which user agent does which, and the mapping shifts as operators add bots.
Content Signals, published by Cloudflare at contentsignals.org, add a Content-Signal directive to robots.txt that states permitted uses independently of who fetches. Three signals are defined: search, ai-input and ai-train, each set to yes or no. An omitted signal neither grants nor restricts.
What the three signals mean
- `search`: building a search index and providing search results, meaning hyperlinks and short excerpts. The definition explicitly excludes AI-generated search summaries.
- `ai-input`: inputting content into one or more AI models, for example retrieval-augmented generation, grounding, or other real-time use of content for generative answers. This is the switch for being quoted by an assistant.
- `ai-train`: training or fine-tuning AI models.
The definitions matter more than the names. A site that sets search=yes, ai-input=no is saying: list me, link me, but do not read me into an answer. For most businesses that is the opposite of what they want; ai-input is where visibility now happens. The being named versus being cited piece goes into why.
The syntax
# Content Signals apply per user-agent group.
User-Agent: *
Content-Signal: ai-train=no, search=yes, ai-input=yes
Allow: /
# A stricter group for a training crawler is still possible
User-Agent: GPTBot
Content-Signal: ai-train=no, search=yes, ai-input=yes
Disallow: /- The directive sits inside a user-agent group, like
Allow, so different bots can receive different signals. - Values are comma-separated
name=yes|nopairs. Order does not matter. - The same value can be sent as an HTTP response header,
Content-Signal: ai-train=no, search=yes, ai-input=yes, which is how Cloudflare's Markdown for Agents passes it along on converted pages. - The generator at contentsignals.org emits a comment block above the directive that states the terms and reserves rights under Article 4 of the EU Directive 2019/790. Keep it; it is what turns a preference into a stated reservation.
What it does and does not do
Content Signals are a declaration, not an enforcement mechanism. The specification says so plainly: robots.txt files express preferences and do not technically prevent anyone from taking content. Their value is that they are machine-readable, unambiguous and, through the rights-reservation text, legally legible. Cloudflare's April 2026 scan found 4 percent of top domains carrying them, and its readiness scanner counts them as a core signal.
Qomvia's Content Signals declared check under Agent protocols reads the live robots.txt for a Content-Signal: line and, failing that, the homepage response headers. Any value passes and is recorded; the point is that a stated policy exists. Pair it with the human-readable automated-access policy page, which is a separate check under Identity.
Questions
- Does ai-train=no stop my content being used for training?
- It states that you do not permit it, in a form crawler operators can parse and courts can read. Whether a given operator honours it is up to the operator; the signal itself does not block a fetch.
- Should I still block GPTBot if I set ai-train=no?
- You can do both. The Disallow prevents the fetch; the signal states the permitted use for anyone who fetches anyway or who obtained the content another way.
- Do I need search=yes for Google?
- Google reads robots.txt Allow and Disallow rules for indexing. The signal is an additional statement; omitting it neither grants nor restricts. Setting search=no is a stated request to be excluded from search indexes.
Score your own site against the rubric this is written from.
Is your site agent-ready?
Free score against the same rubric, in under a minute.
Sign up free to keep the fixes and track the score.
AI monitor
PreviewHow often each model names your site across 11 tracked questions.