How Perplexity picks its sources, and the two bots behind it

Perplexity sends two different bots to your site and cites sources on nearly every answer. What each bot does, why its citations track organic traffic more closely than other engines', and how to measure it.

· Updated 7 min read

Perplexity sends two different visitors to your site. PerplexityBot crawls and indexes pages so Perplexity can surface and link them in its answers. Perplexity-User fetches a page live when someone's question needs it, and Perplexity says it "generally ignores robots.txt rules" because a person triggered the visit. To be cited, PerplexityBot has to be allowed. The rest depends less on Perplexity-specific tricks than most guides claim, and more on whether your site already earns organic traffic.

PerplexityBot vs Perplexity-User

Perplexity's bot documentation describes them this way:

PerplexityBotPerplexity-User
Job"Designed to surface and link websites in search results on Perplexity""When users ask Perplexity a question, it might visit a web page to help provide an accurate answer"
robots.txtRespected. Perplexity recommends allowing it."Generally ignores robots.txt rules", since a user initiated the request
TrainingNot used to crawl content for AI foundation modelsNot used for crawling or collecting training data
If you block itYour pages drop out of the index Perplexity cites fromrobots.txt alone won't stop it; a CDN or firewall rule can

Two consequences. First, a robots.txt that disallows PerplexityBot is the one setting that reliably removes you from Perplexity's citations, and the 2024 "block AI crawlers" templates often included it alongside training bots. Neither Perplexity bot is a training crawler, so blocking PerplexityBot buys no training opt-out.

Second, edge-level bot rules can reach the fetcher that robots.txt can't. For new domains onboarding to Cloudflare from 15 September 2026, "Agent" bots, meaning bots acting in real time on a person's behalf, are blocked by default on ad-monetized pages (Cloudflare). Perplexity-User fits that description. Check your CDN settings and your logs for both user agents.

AI
Free tool · No signup
Free AI Crawler Checker
Paste a domain → see which AI crawlers your robots.txt allows and which it blocks, with training and retrieval bots told apart.
Check your crawlers →

PerplexityBot reads the HTML your server returns. If your content only appears after client-side JavaScript runs, check the raw response with view-source (not the browser inspector, which shows the page after scripts ran).

Why Perplexity is the easiest engine to measure

Because it cites sources on essentially every answer by design. ChatGPT, by contrast, attached a citation to 6.8% of US desktop answers in May 2026, according to Similarweb's 2026 Generative AI Landscape report, up from about 1.3% a year earlier. That average includes writing and coding prompts, but it still means many ChatGPT test runs return no sources at all, while a Perplexity run almost always returns a list you can check your domain against.

Perplexity also has an official API, Sonar, that returns the cited URLs with each answer. A short script can run the same questions every week and log the sources. API answers are not guaranteed to match perplexity.ai, so keep that data labeled as API data. citelity samples Perplexity this way and marks it with an "official API" badge.

Is Perplexity less dependent on Google rankings?

No, and this is the most repeated Perplexity myth. The reasoning goes: Perplexity runs its own crawler, so your Google performance doesn't matter. The first half is true. The conclusion is not.

Ahrefs measured how closely a site's organic traffic tracks how often AI engines mention it (June 2025 data). Perplexity had the strongest relationship of the engines tested:

Its own crawler decides whether Perplexity can find you. Whether it picks you depends heavily on the same authority that earns organic traffic. Keep funding your SEO.

What Perplexity tends to cite

Three engine-specific data points, each with its limits:

  • Older reference pages are tolerated. In Seer Interactive's 2026 study of 47,097 citations, 65% of the pages Perplexity cited had been updated within the past year, the lowest of the three engines measured (ChatGPT 73%, Gemini 78%). Seer describes Perplexity as more tolerant of older reference material. A well-maintained evergreen page has a better chance here than elsewhere.
  • Community sources are prominent. In Profound's dataset of 680 million citations (August 2024 to June 2025), Reddit was Perplexity's most-cited domain at 6.6% of its citations. That data predates several model changes, so treat it as a tendency, not a current ranking.
  • Attribution helped in the one controlled test. The GEO paper (Aggarwal et al., KDD 2024) tested page edits on Perplexity.ai and reported visibility gains of up to 37%, with citing sources, adding quotations and adding statistics among the strongest methods. The test is from 2023–24 and Perplexity has changed since.

The general page work (answer first, self-contained sections, named sources) is the same as for every engine and is in what is AEO.

What AI referral traffic is worth

Perplexity is one of the few engines that sends a clear referrer, so you can see its clicks in analytics. Expect them to be few. Across 13,770 domains, Conductor found AI referrals made up 1.08% of visits, and ChatGPT sent 87.4% of that. The quality numbers that circulate describe AI search traffic in general, not Perplexity: Semrush estimated the average AI-search visitor is worth 4.4× an organic visitor based on conversion rate, across about 500 digital-marketing topics (July 2025), and SE Ranking measured AI visitors spending 67.7% more time on site.

Perplexity is no longer the only place to see AI citation data, either. Bing Webmaster Tools reports citations for Copilot, and Search Console now shows impressions for Google's AI features. Perplexity's referrers remain the most direct view of clicks from Perplexity itself.

AI
Free tool · No signup
Free GEO ROI Calculator
Work out what AI search visibility is worth using your own traffic and conversion numbers, with no invented industry multipliers.
Run the numbers →

Measuring your Perplexity citation rate

  1. Confirm PerplexityBot requests in your server logs over the last 30 days. No hits means an access problem, and nothing else matters until it's fixed.
  2. Pick 10–20 questions your readers actually ask, phrased conversationally.
  3. Run each several times over a few days (or weekly through the Sonar API) and record whether your domain appears and which sentence was quoted.
  4. Compare against the same record after any change. How to track AI citations explains how many runs make a rate trustworthy.
How can I stop Perplexity-User from fetching my pages?
Perplexity says this fetcher generally ignores robots.txt because a user triggered the request, so a robots.txt rule may not stop it. A bot rule at your CDN or firewall can. Blocking it does not remove you from Perplexity's index, which PerplexityBot builds, but it stops Perplexity from reading your page live for a user's question.
Does Perplexity train AI models on my content?
Perplexity's documentation says neither PerplexityBot nor Perplexity-User is used to collect content for training AI foundation models. Blocking PerplexityBot therefore removes you from Perplexity's search results without any training benefit.

Sources

Updated October 1, 2026: corrected PerplexityBot's role (indexing, not answer-time fetching) and added Perplexity-User; added Ahrefs' correlation figures; corrected the Semrush conversion figure's scope; re-attributed the ChatGPT 6.8% figure to Similarweb; removed an invented example answer.