Published by Qomvia, , 3 min read
Agents measure a different kind of slow
Web performance work optimises what a person sees: Largest Contentful Paint, layout shift, interaction delay. An agent sees none of it. A retrieval pipeline fetching twenty candidate pages for one answer has a per-request budget, often a few seconds, and it makes the request with a plain HTTP client. What it experiences is the time until the first byte of HTML arrives, and then how many bytes it has to download and parse before it finds the content. A page that scores 95 in Lighthouse because it streams a skeleton quickly and hydrates later can still be slow and heavy by these measures.
Time to first byte
TTFB is the interval from sending the request to receiving the first response byte: DNS, TLS, the server thinking, and the first packet back. For anonymous requests to a public page the server should not have to think at all: the HTML can be cached at the CDN edge and served in tens of milliseconds. The common reasons it is not are personalisation that forces every request to the origin (a cookie banner state, a currency switch, an A/B flag), a bot rule that routes non-browser user agents through a slower path, and origins that render each request from a database.
curl -o /dev/null -s -w "ttfb %{time_starttransfer}s total %{time_total}s bytes %{size_download}\n" \
-A "Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)" \
https://www.brand.ch/- Cache HTML for anonymous requests, including declared crawlers. Vary on the cookie that marks a logged-in session, not on every cookie.
- Do not send bots to a slower lane. Bot management rules that add a JavaScript challenge or a managed challenge double the TTFB at best and return a challenge page at worst; see bot protection.
- Watch the redirect chain.
httptohttpstowwwto a trailing slash is four round trips before the first byte of HTML. Agents count all of them.
HTML weight
After the first byte comes the rest of the document, and the agent has to read all of it to find the content. Modern frameworks inline a serialised copy of the page state into the HTML for hydration, sometimes hundreds of kilobytes of JSON that repeats what is already in the markup. Add inlined critical CSS, tracking snippets, and an SVG sprite, and a homepage reaches a megabyte of HTML before a single image. For a model-based pipeline that is also a cost problem: the HTML is converted to text and tokenised, and a bloated document either gets truncated or crowds out the passages that matter.
- Trim the hydration payload. Do not ship the full API response twice; most frameworks let you mark data as server-only.
- Move large inline scripts and styles to files. They are cached across pages and invisible to the text extraction.
- Compress. Brotli or gzip for
text/htmlis table stakes; check the response header rather than assuming the CDN does it. - Offer Markdown. Sites that support Markdown content negotiation hand agents a version that is a fraction of the size, with the structure intact.
What Qomvia checks
The Performance dimension is 8 points across two checks, both measured on the homepage request made with a declared crawler user agent. Time to first byte: 800 ms or less passes (5 points), up to 2,000 ms is partial (2), slower fails. HTML payload size: 500 KB or less passes (3 points), up to 1,200 KB is partial (1), larger fails. The dimension is deliberately light because access and legibility matter more, but it is also the one where a single CDN rule often moves a site from fail to pass.
Questions
- Our Core Web Vitals are green. Why does the TTFB check fail?
- Web Vitals are measured in real browsers by real visitors, mostly returning ones with warm caches. Qomvia measures a cold request from a declared crawler, which often takes a slower path through bot management and origin rendering.
- Is 500 KB of HTML really too much?
- For the document alone, before images and scripts, yes. Most well-built pages are under 100 KB of HTML. Above 500 KB almost always means duplicated state or inlined assets.
- Does the measurement location matter?
- Yes; a single measurement from one region is a sample, not a verdict. Use the trend across scans and your own curl from several locations before optimising.
Score your own site against the rubric this is written from.
Is your site agent-ready?
Free score against the same rubric, in under a minute.
Sign up free to keep the fixes and track the score.
AI monitor
PreviewHow often each model names your site across 11 tracked questions.