
Published by Qomvia, , 13 min read
Key takeaways
- Separate the crawler roles: PerplexityBot is documented for search discovery, while Perplexity-User can fetch a page in response to a user's request.
- Robots policy is not the whole path: a permitted crawler may still meet a firewall challenge, rate limit or server error.
- Citations are inspectable evidence: compare the linked page with the exact claim instead of treating a source card as a ranking report.
- Freshness is an editorial responsibility: update material facts when they change and show the context that makes a claim current.
- Access is the first diagnostic: a successful request removes one crawl obstacle; assess citation evidence on the relevant answer surface.
What is Perplexity SEO?
Perplexity SEO is the practice of making a website's useful information accessible, clear and verifiable when Perplexity answers a question with sources. It aligns crawler-policy decisions, reliable page delivery, direct explanations and careful citation review with evidence that a page serves the searcher's task.
The publisher-facing distinction begins with two documented user agents. PerplexityBot is described as a crawler that surfaces and links websites in Perplexity search results, and Perplexity says it is not used to crawl content for AI foundation models. Perplexity-User supports user actions: it may visit a page to answer and link to it. The company says this user-requested fetch generally ignores robots.txt because the person asked for the page.
Those roles should not be collapsed into one category called “AI traffic.” A search crawler request, a user-requested page visit, a citation and a referral session are different evidence. The practical goal is to find which part of the path can be inspected and then fix the defect that belongs to that part. For broader access fundamentals, see robots.txt for AI crawlers and AI agent readiness.
How do PerplexityBot and Perplexity-User differ?
Discovery and user-requested fetching
PerplexityBot represents an automated discovery path. Its documented purpose is to surface and link websites in search results. Perplexity recommends allowing it and publishes IP ranges for verification. A site owner who wants pages considered in that path should inspect the bot's policy, the network response and the content that the request actually received. A robots allow rule alone cannot show whether a firewall or cache returned usable material.
Perplexity-User is tied to a user action rather than automatic search discovery. Perplexity says it may visit a site to answer a user and link to the destination, and generally ignores robots.txt for that requested fetch. That distinction is relevant to access governance: a public page can be available for a person-requested action even when automated discovery is restricted. It does not mean account pages or sensitive resources should be exposed without authentication.
Use a route-first diagnostic. If a page does not appear in search, examine the automated crawler path and the page's usefulness for the query. If a person asks the assistant to open a page and it fails, investigate the user-fetch request and the response. If a citation appears but the claim is wrong, inspect the source text and answer context. The same symptom, “Perplexity cannot see us,” can hide different causes.
Keep the two request roles separate in analytics and incident reports. Label an automated discovery request as a crawl event and a user-triggered visit as a fetch event when your evidence supports that classification. Do not merge both into a generic “AI bot” bucket. The first may relate to discovery; the second may follow a person's explicit request. Each carries different implications for access policy and measurement.
When attribution is uncertain, preserve the raw request details and classify it as unknown until a verification method supports a conclusion. A user-agent name is useful for triage but is not, by itself, proof of identity. Compare the request with the provider's published verification guidance and your edge logs. This protects against spoofed requests and prevents a false allow rule from becoming a security exception.
| Observation | What it can support | What it cannot prove |
|---|---|---|
| PerplexityBot request | A named crawler attempted or completed a request | That the page was selected in an answer |
| Perplexity-User request | A user action led to a page visit | Automatic search inclusion |
| Visible citation | A source link was presented | Every hidden retrieval step |
| Tagged referral session | A visit reached analytics | All prior answer exposure |
How should a site configure robots and a firewall?
Scope the rule to the public path
Begin with the publisher's intent. Identify the public pages that should be available for search discovery, pages that should remain private and any user-driven workflow that needs a separate policy decision. Read the current crawler documentation before editing robots.txt. Then check the CDN, firewall, origin server and authentication layer. A crawler may be permitted by the file but still receive a challenge page or an error at the edge.
If your infrastructure filters by network address, compare requests with the IP ranges Perplexity publishes. Treat the published source as a maintained input: confirm the retrieval path used by your security controls and assign someone to review it when the provider changes the list. Do not copy an IP range from an old log or assume that a user-agent string by itself proves request identity. Follow your organization's normal verification process.
Test a representative page through the same route the crawler uses. Record the requested URL, status, response type, canonical page and whether the intended article text is present. Check compressed and cached variants if the edge serves different responses. If the request fails, compare a direct origin test with the public edge response. A successful response from a development machine does not establish that a remote crawler can reach the public page.
Perplexity's documentation says changes may take up to 24 hours to be reflected. Use that as a window when evaluating a policy change, rather than expecting a crawler's behavior to update immediately. Continue to distinguish a delay in honoring a rule from a permanent access failure. Save the time of the change and the first later request that confirms the new behavior.
Build a small access test set from page types, not a random handful of URLs. Include a product page, a policy page, an article, a paginated view and any important localized variant. Check each through the same public hostname and edge path. If one template fails while the others work, compare its cache, redirect, authentication and application behavior rather than changing the crawler policy globally.
Inspect the body delivered to the remote request, not just the browser view seen by a logged-in employee. Personalized scripts, consent overlays, challenge pages and client-side rendering can create a different result. Save response headers and a text snapshot where your process permits. A successful status code can still accompany an interstitial or empty shell, so verify that the substantive page content arrived.
Treat an allowlist as a maintained security rule. Confirm the provider's published IP data is current, define how your edge vendor consumes it and establish an owner for updates. Do not bypass rate limits, authentication or bot controls on the assumption that all requests from an advertised range are safe. Protect write operations and account data even when a public article is intentionally available for crawling.
What makes a cited source useful to a reader?
A source is useful when a person can verify the claim it is attached to. Open the citation, locate the relevant sentence and check whether qualifications have survived the summary. A policy page may support a rule but not an exception. A product page may support a feature but not a comparative superiority claim. An article that is current in its title may still contain an outdated table or broken reference.
Write pages so that the subject and answer are explicit together. Use a descriptive heading, a direct opening explanation and enough context for the statement to remain true when quoted alone. Put caveats close to the claim. Identify the publisher, author or responsible team when relevant, and link to the primary source behind a technical or policy statement. Avoid forcing the reader to infer what “it,” “they” or “the solution” refers to from a paragraph several screens away.
Citation presence and source quality are different dimensions. A page can be linked but irrelevant, technically valid but unclear, or accurate but not selected for one question. Keep a record of the exact prompt, answer, source URL and supporting passage. If reviewers cannot agree whether the citation supports the sentence, preserve that ambiguity instead of scoring it as a win.
A source review should ask whether the page answers the task the person asked, not only whether it contains the same keywords. A product comparison that describes tradeoffs may be more useful for a choice than a category page with a repeated phrase. A help article may be the strongest source for a policy exception. Match the cited passage to the user's intent and the decision the answer is helping them make.
Separate the source's authority from the source's completeness. A first-party support page can establish a company's own terms, while an independent explanation may help a reader understand the broader issue. Make relationships and authorship visible, then let reviewers judge whether the source is appropriate for the claim. Avoid describing every citation to your domain as an endorsement or recommendation.

How should a publisher think about freshness?
Freshness begins with the facts that change. Identify prices, eligibility, policies, specifications, dates and availability that could make a page misleading when stale. Assign an owner to each material field, then update the claim where it is published and the source it depends on. A date label is useful only if someone reviewed the substance behind it.
Keep history legible. When a policy changes, say what changed and when it took effect. When a product is discontinued, mark the page clearly and point to the current option if appropriate. Do not silently replace the evidence in a way that makes an older citation impossible to understand. A reader needs to know whether a statement applied when an answer was generated or whether the source has since been revised.
Freshness matters when the underlying facts change. Use the substance of an update, not its timestamp alone, to assess whether the page is current. For a tested visibility method, compare saved answers and page snapshots with collection conditions held steady, then record response patterns before changing the content.
Use a change log for claims whose validity depends on a date or version. Record the previous statement, the new statement, the supporting source and the person who approved the update. If a cited answer reflects the old page, the record helps explain the difference without implying that the answer system failed to refresh. If the page was already current, the log also helps isolate a possible synthesis error.
Prioritize updates by reader consequence. A stale shipping promise, eligibility rule or safety instruction needs attention before a timeless explainer. Maintain source links and mark archived material clearly. The aim is not to change every publication date frequently; it is to ensure that statements a person may act on remain accurate and that historical context is still understandable.
How can teams diagnose a missing citation?
Work from the public page outward. Confirm that the exact URL resolves without an unintended challenge, that the important answer is present in the delivered content and that the page is the best destination for the tested task. Review the canonical tag, redirects and robots policy so the team is testing the page it intends to publish.
Then separate access evidence from selection evidence. A successful fetch documents that a request reached a response; record eligibility and citation selection from the relevant answer surface. A citation sample shows what appeared in that answer, not the full candidate set or internal selection decisions. Keep the two records distinct.
Work from observable layers. Confirm the page is public and returns useful content. Inspect whether the documented crawler can request it under the intended policy. Check that the topic and answer match the question. Review the pages that did receive citations and compare their evidence, specificity and currency. This sequence cannot reveal the platform's complete selection logic, but it can identify defects under the publisher's control.
Repeat a question only with a documented purpose. Save the wording, date, language, location context where available, answer and cited URLs. If the response changes, compare the full output instead of recording just whether the brand appeared. A changed citation could result from a changed query, a different available source or normal variation. Separate a stable pattern from a one-off response before changing the page.
Use a small issue record: symptom, evidence, likely layer, owner and retest. A blocked request belongs with infrastructure; an inaccurate source belongs with the page owner; an unclear test belongs with the measurement owner. If the source is correct and accessible but still absent, record that as an unresolved selection outcome rather than inventing a hidden ranking fix. For a general assessment, see how to measure AI search visibility.
Keep operational notes specific enough for another person to reproduce the check. Record the public URL, request path, response status, relevant edge action and the time of the test. For answer review, save the exact question, full response and cited destination separately. A single screenshot cannot show whether a crawler was blocked or whether a page later supported a citation.
Do not treat every missing citation as a reason to loosen access controls. First establish that the page is intended to be public, that the documented crawler can fetch useful content and that the page addresses the tested task. If each check passes, selection remains uncertain. Protecting private routes and customer data is more important than turning an unknown outcome into an unrestricted access rule.
Treat an IP allowlist as a separate deployment decision. Compare the provider's published list with current WAF controls and confirm that any rule is scoped to the intended crawler and public content. Review the effect on security controls before rollout, then test the requested path from the relevant edge. A valid IP match alone does not prove that the page delivered useful text.
Make an access change only when there is an access defect. If a crawler receives a successful response but an answer ignores the page, inspect task fit and the sources used. If a request is blocked, fix the exact response path while preserving controls for private routes. This ties crawler permission to a documented need instead of treating unrestricted access as a generic visibility tactic.
Choose the next test based on the suspected failure. If a page returned a challenge, reproduce the request through the same edge and inspect the response. If the page is available but the answer uses a stale fact, compare the cited page snapshot with the current canonical source. If the citation does not support the wording, record both passages and route the correction to the responsible editor. A test that cannot distinguish these causes will not guide a useful change.
Keep a record of unresolved cases as well as successful fixes. Repeatedly changing titles, robots rules or page copy without evidence can make the source less clear and obscure the original issue. A disciplined log lets the team identify when access is verified, when the content is corrected and when selection remains outside its control.
Sources and further reading
Questions
- How do I rank in Perplexity?
- No public formula guarantees a position. Make relevant pages accessible, specific and verifiable, then inspect repeated answers and their sources without treating crawler access as selection.
- What is PerplexityBot used for?
- Perplexity documents PerplexityBot as a crawler that surfaces and links websites in its search results. The company says it is not used to crawl content for AI foundation models.
- Does Perplexity-User follow robots.txt?
- Perplexity says Perplexity-User supports user actions and generally ignores robots.txt because a person requested the page fetch. This is distinct from automated search discovery.
- How do I allow PerplexityBot through a firewall?
- Review the published crawler guidance and IP ranges, then verify the public edge returns the intended page. Do not rely on a user-agent string alone as proof of request identity.
- How long do Perplexity crawler changes take?
- Perplexity says changes may take up to 24 hours to be reflected. Record the policy change and verify subsequent requests before treating it as complete.
- Why is my page cited by Perplexity?
- A visible citation shows that a source link was presented, but it does not expose the full selection process. Open the source and confirm that it supports the specific answer claim.
Score your own site against the rubric this is written from.
Is your site agent-ready?
Free score against the same rubric, in under a minute.
Sign up free to keep the fixes and track the score.
AI monitor
PreviewHow often each model names your site across 11 tracked questions.