
Published by Qomvia, , 13 min read
Key takeaways
- A question can be rewritten: the search product may turn a request into targeted queries and follow-ups.
- Search partners can vary: OpenAI names Bing and Shopify among third-party providers; this is not a claim that one index supplies every answer.
- Citations are interface evidence: inline links and a Sources panel show presented sources, not every hidden retrieval step.
- Crawler roles differ: OAI-SearchBot is documented for search, GPTBot for potential training use, and ChatGPT-User for certain user actions.
- Access is the first gate: allowing a crawler removes an access barrier; source fit and visible answer evidence guide the next review.
What is ChatGPT search?
ChatGPT search is a web-enabled answer experience that can consult search providers, synthesize a response and present source links. OpenAI says it may rewrite a user's question into targeted queries, send them to providers and issue more specific follow-ups. The public description explains important parts of the path, but not a complete rule for choosing every result or citation.
A person can experience the result as a single answer, but the publishing and retrieval path may involve a sequence of actions. The system interprets the question, identifies information needs, obtains candidate material, composes a response and displays sources. If the query is ambiguous or asks for current information, a more targeted search may help. The specific path depends on the product mode and question, so an answer without visible web sources should not be treated as though the same retrieval occurred.
Call this sequence the query-to-source chain: question, rewrite, retrieval, answer and source display. The name is an explanatory model for a publisher's investigation, not an official OpenAI architecture diagram. Its value is that each link has a different possible observation. A query rewrite is described in product help, a fetch may be visible in a server log, a cited URL can appear in the interface, and the relationship between a passage and the final wording still requires a careful reading.
This article explains that mechanism rather than giving a complete page-improvement checklist. For the latter, see how to get cited by ChatGPT. For a broader view of model memory and retrieval, read LLM SEO and training, search and user-fetch bots.
How does ChatGPT search turn a question into searches?
Query rewriting is an observable step
OpenAI's search help says the product typically rewrites a question into one or more targeted queries sent to search providers. It may issue additional, more specific follow-up queries as it gathers information. That description matters for site owners because the words in a search request may not mirror the original prompt. A broad question about a product could be decomposed into separate needs for features, policies or comparisons.
The documentation also says that ChatGPT may partner with other search providers and names Bing and Shopify among the providers whose privacy policies may apply. That supports a narrow statement about partner use, not the stronger claim that Bing is the exclusive source for all web results. A site owner should avoid inferring a single upstream index from one citation or one successful crawl.
General location information can be shared with providers to improve search relevance, according to OpenAI, while the help page says the IP address itself is not shared. It also notes that when Memory is enabled, saved memories may inform rewritten queries. Those details remind marketers that a repeated prompt can still be shaped by account settings and context; they do not justify claiming that every answer is personalized in the same way.
For measurement, save both the original wording and the exact test conditions. Record whether search was visible, whether an answer displayed sources, what the query asked and which links appeared. A prompt about current pricing and a prompt about a general concept test different retrieval needs. A stable test set should preserve that intent instead of reducing all questions to one generic “brand visibility” count.
Do not try to predict every rewritten query and then force those phrases into a page. The public help material describes query rewriting but does not expose the rewritten request for every answer or prescribe a publisher-side phrase list. A durable page should answer the underlying need in accurate language. If a question has separate subtopics, use a useful structure that makes each answer and its limits easy for a human to follow.
As an editorial example, a request to compare two service plans can involve different subquestions: what each plan includes, which limits matter and how the terms differ. A useful page can make those comparisons explicit without pretending to know the hidden query sequence. Organize around the decision a reader needs to make, then keep each answer attached to its relevant source and qualification.
Separate provider statements from an individual test. Public help can describe that a product may rewrite a query or consult search providers, while a saved answer shows only what was visible in that run. A source panel may help a reader check the final response, but it does not expose every candidate page or internal decision. Keep a product description, a server log and an answer capture as different kinds of evidence.
When a team compares repeated answers, preserve the context that may affect the request. Record whether Memory was enabled, the market or general location available to the product, the wording of the prompt and whether a search state appeared. This does not make every variation predictable. It helps the reviewer distinguish a change in the test setup from a change in the cited sources.
Where do search partners and publishers fit?
A search provider may help retrieve pages, while the answer product controls how those pages are synthesized and presented. The relationship between these steps is important but should not be overstated. If a site has strong coverage in one search engine, that may be useful evidence for the publisher's broader discovery strategy; it is not proof that every answer will draw from that engine or that a ranking position transfers unchanged.
OpenAI's publisher guidance says any public website can appear in ChatGPT search. To be included in summaries and snippets, site owners should not block OAI-SearchBot. Treat this as an actionable inclusion control, then track citation, ranking and visit outcomes separately from crawl permission. Relevance, the question, available sources and the product's own selection choices still matter.
Publishers should check the entire request path, not just the robots file. A CDN rule, bot challenge, authentication wall or server error can prevent a permitted request from receiving usable content. Verify the canonical page and the actual response under the intended access policy. Keep private or customer-specific data restricted. A crawler allowance is not a reason to make protected account content public.
Keep crawler purposes distinct. OpenAI documents OAI-SearchBot for surfacing public websites in ChatGPT search, GPTBot for potential use of crawled content in model training, and ChatGPT-User for certain user actions. These are separate controls and should be evaluated according to the publisher's intent. Allowing search discovery does not require allowing training-related crawling, and a user-requested fetch is not the same as automatic search inclusion.
A noindex choice should be scoped to the intended URL and audience. The publisher guidance describes noindex as a way to prevent a title-only listing in Atlas when the crawler can read the directive. If the crawler is blocked before it can fetch the page, do not assume the directive has been observed. Check how the rule affects ordinary search eligibility and internal discovery before applying it to a directory or site-wide template.
For a public article that should be eligible for a search summary, inspect the response with the real access controls in place and confirm that the canonical content is not hidden behind a challenge. For a page that should not be public, protect it with authentication rather than relying on a title or summary control. The correct policy depends on the sensitivity of the content and the publisher's decision, not only on whether a page has appeared in an answer.
If a site opts out of OAI-SearchBot, OpenAI says a URL discovered from a third-party provider or other pages may still appear as a page title and link in ChatGPT Atlas when it seems relevant. The publisher FAQ says a noindex directive can prevent such a listing, provided the crawler is allowed to read the meta tag. This is a narrow, product-specific behavior. Evaluate the tradeoff with the page owner before adding a broad noindex rule that could affect ordinary search inclusion.

What do citations and the Sources panel tell a reader?
A source panel is evidence, not a reconstruction
OpenAI says search responses may include inline citations, and that a Sources panel can list cited sources alongside other relevant links. A citation makes the origin of some information easier to inspect. The panel can also provide a route to material not shown inline. These interface details improve auditability for readers, but they do not reveal the complete candidate set, every query issued or the contribution of each source to the generated text.
A citation should be evaluated against the sentence it accompanies. Open the destination and ask whether it supports the specific claim, whether the page is current and whether the relevant context is preserved. A page can be authoritative on one detail and irrelevant to another. A link to a company homepage may identify an entity while failing to substantiate a detailed product comparison.
Use citations as a reader's audit path, not as a complete explanation of the answer engine. An inline source may support a nearby clause without being the only material considered. A related link in the Sources panel may be useful to the person without being cited in the body. If a team treats every panel entry as a direct endorsement, it can overstate the evidence and miss whether the answer accurately represented the linked source.
A publisher can improve the audit path by making the destination specific. Link the page that states the policy, comparison or specification rather than a general homepage that requires another search. Keep the title and first paragraph consistent with the actual subject so that a person opening a citation can quickly confirm relevance. This makes the source easier to verify; measure whether its link appears in sampled responses separately.
For a publisher, being named and being cited are separate results. A model can describe an organization without a first-party URL, or link to an external page that discusses it. The distinction affects what to fix: identity may need clarification, a direct source page may be missing, or a response may select other evidence. The source-level playbook explains how to classify those cases without treating a citation as an endorsement.
| Visible result | What it establishes | What remains unknown |
|---|---|---|
| Inline source link | A link is attached to part of the response | The full internal retrieval sequence |
| Entry in Sources panel | The interface offers a source or related link | Its exact weight in answer composition |
| Brand name without link | The answer names the entity | Whether a first-party page was retrieved |
| Title-only Atlas link | A relevant page title may be surfaced | That the blocked content was summarized |
Keep the source URL and the passage that supports the statement in your archive. If the page later changes, a reviewer can understand what the answer may have drawn from at the time. Where the answer's attribution is unclear, classify it as uncertain. A precise uncertainty label is preferable to attributing a claim to a source that does not support it.
How should publishers interpret referral tagging?
OpenAI's publisher FAQ says ChatGPT automatically adds the UTM parameter utm_source=chatgpt.com to referral URLs. Analytics teams can use that parameter to identify visits that arrive through those tagged links. The tag is a useful convention for classifying observable sessions; it is not a record of all impressions, answer views or decisions influenced before a visit.
Preserve the original landing page and relevant campaign parameters in your analytics configuration. Check that redirects, canonicalization and consent controls do not erase the values before collection. If a visit comes through another product or a shared link, its referrer may look different. Do not assume that all answer interfaces use the same tag or that an absent tag means a person did not see an answer.
The UTM value has a narrow job: it helps classify a tagged referral session. It cannot tell the analyst whether someone saw a source card and came back later through a bookmark, or whether an answer changed a decision that ended offline. Keep the analytics definition literal. A report can say “sessions tagged with this source,” then add a separate answer-sampling result for links and mentions, without implying that either captures every exposure.
If a report includes a referral trend, show the exact filter and the landing pages included. Check that canonical redirects preserve the query string and that the analytics property records campaign parameters as expected. Treat a missing visit as missing visit evidence, not evidence that a citation was never displayed. For the broader implementation, see tracking AI traffic in GA4.
Use web analytics for the behavior it measures: visits and the events configured on the site. Use answer sampling for the text and sources displayed in a model response. These datasets can be compared at a strategic level, but they do not share a complete event trail. The broader guide to tracking AI traffic in GA4 distinguishes tagged referrals from answer visibility; AI search visibility measurement covers the question-sampling side.
How can site owners improve the source experience?
Give each important question a canonical source page. Explain the answer near the top, use headings that make the topic explicit and keep the relevant qualifications close to the claim. If the answer depends on a policy, make the policy easy to find. If it depends on a product specification, link the specification rather than hiding the detail in an image or an interactive configurator.
Make the page safe to summarize. Identify the publisher and subject, state the conditions for the guidance and link primary material. Avoid vague superlatives that cannot be checked. When a fact changes, update the explanation and its supporting reference instead of only changing the visible date. A source that is easy to quote but easy to misinterpret is not a good source.
Review the final link as a reader would. Does it open the page that supports the claim, or does it lead to a broad homepage that requires another search? Is the key qualification visible without logging in? Does the page retain the same offer or policy named in the response? These checks improve the source experience for people whether or not a particular ChatGPT search answer selects the page.
Then test how the intended crawler and the person experience the page. Confirm that public text is fetchable, that the title and canonical identify the correct page, and that the response does not require a user session. Protect sensitive information and test security rules. For a broader technical view, see Bing SEO for AI search and the AI agent readiness guide.
Sources and further reading
Questions
- How does ChatGPT search work?
- OpenAI says it may rewrite a question into targeted searches, consult search providers and issue more specific follow-ups. Responses may show inline citations and a Sources panel.
- How does ChatGPT search choose sources?
- Published help describes query rewriting and source presentation; OpenAI does not document every source-selection step. Review visible citations and check that OAI-SearchBot can access the relevant pages.
- Does ChatGPT search use Bing?
- OpenAI's help page says ChatGPT sometimes partners with other search providers and names Bing among them. That does not establish Bing as the exclusive source for every answer.
- What is OAI-SearchBot used for?
- OpenAI documents OAI-SearchBot as a crawler used to surface websites in ChatGPT search features. It is distinct from GPTBot and ChatGPT-User.
- Why does ChatGPT show a title-only link in Atlas?
- OpenAI says a relevant page that is disallowed to OAI-SearchBot may still surface as a title and link if its URL is found elsewhere. The publisher guidance describes noindex as a control for preventing that listing when the crawler can read the directive.
- How can I identify ChatGPT search traffic in analytics?
- OpenAI says it adds
utm_source=chatgpt.comto referral URLs. That can identify tagged visits, but it does not measure answer views or every influence on a later visit.
Score your own site against the rubric this is written from.
Is your site agent-ready?
Free score against the same rubric, in under a minute.
Sign up free to keep the fixes and track the score.
AI monitor
PreviewHow often each model names your site across 11 tracked questions.