Published by Qomvia, , 2 min read
The step nobody sees
A page is mostly not content. Navigation, header, footer, consent banner, related links, share buttons and legal notices often outweigh the article by characters. Feeding all of that to a model is expensive and makes the answer worse, so every retrieval pipeline runs an extraction step first: find the main block, drop the rest.
Extractors use two kinds of evidence. Semantic markup (<main>, <article>, role="main") is a direct statement of where the content is. Heuristics (text density, link density, block length) are a guess. When the markup is there, the guess is unnecessary. When it is not, the guess sometimes picks the sidebar.
What Qomvia checks
Main content is extractable is the largest single legibility check in the methodology, worth 8 points. It looks for a <main> or <article> element in the raw HTML of a content page. Present: pass. Absent with enough text: partial, with the note that extraction has to guess. Under 400 characters of text: fail, because there is nothing to identify a main block in.
It is one of the cheapest checks to fix and one of the most common to fail, because most page builders and older themes wrap everything in <div>.
The markup that fixes it
<body>
<header>…site navigation…</header>
<main>
<article>
<h1>The one topic of this page</h1>
<p>Standfirst that states the claim.</p>
<h2>First section</h2>
<p>…</p>
</article>
</main>
<aside>…related links…</aside>
<footer>…legal, contact…</footer>
</body>- One
<main>per page. It is the document's main content, so a second one is a contradiction. <article>for a self-contained piece: a post, a product, a listing. Several<article>elements inside<main>are fine on an index page.- Keep chrome outside. A cookie banner or a newsletter box inside
<main>is inside the quote. role="main"on a<div>is accepted by extractors as a fallback, but the element is shorter and clearer.
Landmarks and headings work together
Once the extractor has the right block, the next pipeline stage splits it into passages, and it uses headings to do that. A landmark without headings gives a model one long blob; headings without a landmark give it well-structured navigation. The heading structure article covers the second half.
Questions
- Is this an accessibility thing or an AI thing?
- Both. Landmarks were introduced for screen readers, and extraction libraries adopted them for the same reason: they state structure instead of implying it.
- My CMS wraps content in a div with class 'content'. Is that enough?
- No. Class names are site-specific and extractors cannot rely on them. Change the wrapper element to <main> or add role="main" to it.
- Does a product page need <article>?
- A <main> is sufficient. If the product description and specifications are self-contained, wrapping them in <article> does no harm and helps some extractors.
Score your own site against the rubric this is written from.
Is your site agent-ready?
Free score against the same rubric, in under a minute.
Sign up free to keep the fixes and track the score.
AI monitor
PreviewHow often each model names your site across 11 tracked questions.