Published by Qomvia, , 3 min read
Enumeration versus exploration
There are two ways for a machine to learn what pages a site has. It can follow links from the homepage, which finds what the navigation exposes and misses what it does not: old posts, deep product variants, help articles three clicks down. Or it can read a list. The sitemap is the list, and every crawler that matters, search or AI, starts by looking for one.
Cloudflare's April 2026 readiness analysis of the top 200,000 domains put robots.txt at 78 percent adoption and named the sitemap alongside it as the first place agents look. The two work together: robots.txt is where the sitemap is declared.
What Qomvia checks
Sitemap is discoverable and structured in the methodology fetches /sitemap.xml and every sitemap declared in robots.txt, follows index files one level down, and counts the URLs. A sitemap that returns 200 and is declared in robots.txt scores the full 5 points. Found at the default path but not declared: 3, because a reader that trusts robots.txt over convention will not find it. None: fail. The <lastmod> values also feed the Content carries a date check.
A sitemap that does its job
User-agent: *
Allow: /
Sitemap: https://www.brand.ch/sitemap.xml<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.brand.ch/sitemap-pages.xml</loc>
<lastmod>2026-08-11</lastmod>
</sitemap>
<sitemap>
<loc>https://www.brand.ch/sitemap-products.xml</loc>
<lastmod>2026-08-12</lastmod>
</sitemap>
</sitemapindex>- Declare it in robots.txt with an absolute URL. The
Sitemap:line is not tied to a user-agent group and can appear anywhere in the file. - Split by type once you pass a few thousand URLs: pages, products, articles. Each file stays under the protocol limits (50,000 URLs, 50 MB uncompressed) and an updated product file does not force a re-read of everything.
- Only canonical, indexable URLs. No redirects, no
noindexpages, no parameter variants. Every entry should return 200 with a self-referencing canonical. - Honest `lastmod`. Set it when the content changes, not on every build. Readers learn quickly whether your dates mean anything.
- No `priority`, no `changefreq`. Major consumers ignore both; they add bytes and false precision.
Sitemaps, feeds and llms.txt are different lists
It is easy to conflate three files that all list URLs. The sitemap is complete and neutral: everything indexable, no editorial view. A feed (RSS, Atom, JSON Feed) is chronological and carries content: what changed, most recent first, with the text or a summary. llms.txt is curated and short: what matters, for a reader with one question. An agent platform enumerating a site wants the first; one monitoring it wants the second; one answering about it wants the third. Publish all three from the same source of truth.
Questions
- Do AI crawlers read sitemaps?
- The operators describe their search crawlers as standard web crawlers that honour robots.txt, and the sitemap is the standard discovery mechanism those crawlers use. Cloudflare's readiness scanner checks for it as a core signal.
- Should the sitemap include images or videos?
- Only if they are content you want found on their own. For most sites the page URLs are enough; image and video extensions add size without changing what an agent can answer.
- How often should the sitemap be regenerated?
- Whenever content changes, ideally on publish. A daily rebuild is fine for most sites; a monthly one means agents see a stale list for weeks.
Score your own site against the rubric this is written from.
Is your site agent-ready?
Free score against the same rubric, in under a minute.
Sign up free to keep the fixes and track the score.
AI monitor
PreviewHow often each model names your site across 11 tracked questions.