Perplexity picks its sources by running every question through retrieval first: it pulls candidate pages from its own crawler-built index plus live fetches, filters them through a ranking layer that weighs relevance, freshness, authority, and structure, then writes the answer from the surviving pages with numbered citations on every claim. There is no “answer from memory” mode. That single design decision, search first, always, is why Perplexity is the most winnable of the major AI engines for a brand willing to do the work, and the fastest to punish a brand that stops updating.
The contrast matters. In how Claude finds and cites sources we walked through an engine that leans on an external index. Perplexity is the opposite case: it owns its retrieval stack end to end, which changes both how it behaves and how you optimize for it.
The pipeline, end to end

Step one is the part people underestimate: every query triggers a search. Where ChatGPT decides per question whether to browse, and answers plenty of questions from training knowledge alone, Perplexity retrieves by design. That means your visibility there is almost entirely a function of the live web, the part you can actually change this quarter, rather than of frozen training data you cannot.
Step two is retrieval from Perplexity’s own index. The company runs its own crawler, PerplexityBot, and states plainly that the crawl is for surfacing and linking sites in results, not for training foundation models. On top of the index, fresh or specific questions trigger live fetches of individual pages at answer time.
Step three is where brands are won and lost. The candidate pages pass through a reranking layer that filters hard. Public analyses of Perplexity’s citation behavior consistently point at the same qualities surviving the filter: pages that match the question’s intent directly, pages updated recently, pages from domains with established authority, and pages structured so the answer is easy to extract. Only a handful of sources make it into a typical answer, so this is a competition for very few slots.
Step four, the model writes from the survivors with numbered citations. Worth knowing: Perplexity lets users swap the writing model, its own Sonar models or frontier models from other labs. The retrieval layer underneath is the same regardless. Optimizing for Perplexity means optimizing for retrieval, not for any particular LLM’s taste.
The two fetchers, and the different rules they play by
Perplexity touches your site through two documented user agents, and they do not follow the same rules.

PerplexityBot is the indexer. It crawls to build the index, it is documented as respecting robots.txt, and blocking it removes you from Perplexity’s view of the web. Perplexity-User is the live fetcher: it pulls your page at the moment a real user’s question needs it. By Perplexity’s own documentation it “generally ignores” robots.txt, on the argument that a human requested the fetch, so it acts as the user’s agent rather than a bot. Anthropic’s equivalent fetcher honors robots.txt; Perplexity’s, by its own admission, does not. A Perplexity-User hit in your logs is the strongest signal you can get from this engine, because it means a live answer pulled your page in. If you want to see what those log entries look like next to the other engines’ bots, is my website a source for ChatGPT and Perplexity walks through the full crawler trail.
There is also a dispute worth knowing about. In August 2025, Cloudflare published research accusing Perplexity of fetching pages with undeclared, stealth user agents to get around no-crawl directives, and Perplexity disputed the report publicly. We are not going to adjudicate that here. The practical takeaway is narrower: blocking Perplexity is unreliable in a way blocking a well-behaved crawler is not, so for most brands the productive question is not “how do I keep it out” but “is what it reads worth citing.” If you are weighing that trade-off across engines, see should you allow or block AI crawlers.
What survives the reranker
Four qualities keep showing up in pages Perplexity actually cites, and they compound.
Relevance to the literal question comes first: the reranker scores whether your page answers the query as asked, so a page targeting “best CRM for startups” will not surface for “CRM pricing comparison” no matter how good it is. Freshness comes second, and Perplexity’s recency bias is the strongest among the major engines, strong enough that an older, better page routinely loses to a newer, adequate one. Authority still counts, established domains and widely referenced pages surface more, but it buys less here than in classic Google SEO. And structure decides absorption: pages that front-load the answer, use question-shaped headings, and make claims in quotable, specific sentences get cited over pages that bury the same information in narrative.
Notice what is missing from that list: backlink volume, domain age, and the rest of the classic SEO scoreboard. They correlate with authority, but they are not the mechanism. We have seen small, fresh, well-structured pages outcite enterprise domains on Perplexity week after week.
What this means if you want Perplexity to cite you
Three consequences fall out of the mechanics. First, Perplexity is the fastest feedback loop in AI search: because everything is retrieval and retrieval favors freshness, a page you improve this month can be cited next month. Second, your technical floor has to be clean, PerplexityBot allowed and your content readable without JavaScript, because AI crawlers mostly do not execute JavaScript. Third, structure is not cosmetic, it is the difference between being retrieved and being quoted.
The full step-by-step is in how to get your brand recommended by Perplexity, and you can see how Perplexity specifically ranks and describes brands in your category on the Perplexity visibility tracker page.
One more practical note: Perplexity’s answers shift as its index refreshes, which is constantly. A one-time manual check tells you where you stood for one run on one day. Tracking the same buyer-intent prompts on a schedule is what shows your real position and its trend. Rankry does this for Perplexity alongside ChatGPT, Claude, Gemini, and Grok, keeping the raw answers and their cited sources as evidence.

FAQ
Where does Perplexity get its sources? From its own index, built by its crawler PerplexityBot, combined with live page fetches at answer time. A ranking layer filters candidates for relevance, freshness, authority, and structure before the model writes from the survivors.
Does Perplexity always cite sources? Yes, that is its defining behavior. Every answer carries numbered citations linking to the pages it drew from, typically a handful per answer. There is no uncited answer mode the way there is in chat assistants.
What are PerplexityBot and Perplexity-User? PerplexityBot crawls the web to build Perplexity’s index and is documented as respecting robots.txt. Perplexity-User fetches pages live when a user’s question needs them, and by Perplexity’s own documentation it generally ignores robots.txt because a human requested the fetch.
Does Perplexity favor fresh content? Yes, and more strongly than any other major AI engine. Recently updated pages with visible dates are consistently preferred, which makes a content refresh schedule one of the highest-leverage Perplexity tactics.
Why does Perplexity cite my competitor but not me? Your page is either failing retrieval, meaning it is not in the index, not fresh, or not readable to Perplexity’s fetchers, or failing absorption, meaning it was retrieved but had nothing quotable enough to shape the answer. Diagnose in that order.
See exactly which sources Perplexity cites in your category, and whether yours are among them, tracked weekly with raw answers kept as evidence. Start a free 7-day Rankry trial, no card, first report in two minutes.