HomeInsightsHow Does Perplexity Choose…

How Does Perplexity Choose Which Sources to Cite?

Perplexity has never published its ranking algorithm. Here's what's actually confirmed about how it picks sources — and what's guesswork dressed up as fact.

By Alex RiveraPublished October 10, 2026

**Perplexity has never published its ranking algorithm, but its own documentation and independent research show the real, confirmed picture: two separate bots do the work, a small cluster of domains — Reddit, LinkedIn, and a handful of others — dominates its citations, and the detailed 'reranking formula' guides circulating online cite no primary source at all.** That gap between what's confirmed and what's guessed is the whole story if you're trying to get a business cited.

How Does Perplexity Actually Find the Pages It Cites?

Perplexity runs two separate bots, and the distinction matters for any business trying to control what gets cited. PerplexityBot is the index-builder: Perplexity's own documentation states it "is designed to surface and link websites in search results on Perplexity" and explicitly "is not used to crawl content for AI foundation models" (Perplexity crawler docs, 2026). Perplexity-User is a second, separate agent that fires only when a live user's question needs a specific page checked in real time — Perplexity says it "might visit a web page to help provide an accurate answer and include a link to the page in its response," and because that visit is triggered by an actual person's question, it "generally ignores robots.txt rules" (Perplexity crawler docs, 2026). Block PerplexityBot and you fall out of the index Perplexity searches; Perplexity-User can still fetch your page live if a user's question lands on it directly, and Perplexity notes robots.txt changes can take up to 24 hours to take effect anyway.

What Domains Does Perplexity Actually Cite Most?

The most reliable public data on this comes from Semrush, which tracked over 100 million AI citations across 230,000 prompts on ChatGPT, Google AI Mode, and Perplexity over 13 weeks between July and October 2025. Perplexity's top cited sources were Reddit, LinkedIn, NIH, Microsoft, and Google (Semrush, 2026) — a different mix than the Wikipedia-heavy pattern a lot of people assume. Wikipedia actually accounted for only around 0.8% of Perplexity's citations in the study, far lower than its share on other engines (Semrush, 2026). Reddit's share drifted down slightly through September, but Semrush found the rest of Perplexity's top domains held steady across the full 13 weeks (Semrush, 2026).

Claim about Perplexity's sourcingStatusSource
Runs two separate bots — one builds the index, one fetches liveConfirmedPerplexity crawler docs
Reddit, LinkedIn, NIH, Microsoft, and Google dominate citationsConfirmed, 13-week studySemrush
Search API supports domain and recency filtersConfirmedPerplexity API docs
Leans heavily on news, media, and business-press domainsConfirmed, academic auditLi & Sinnamon, 2024
Runs a named 3-layer XGBoost reranker with exact % weights per factorUnverified — no primary source found—
Domain authority counts for an exact, published % of the ranking scoreUnverified — no primary source found—

Should You Trust the 'Perplexity Ranking Factors' Guides Online?

Search "how does Perplexity choose sources" and you'll find dozens of guides stating Perplexity runs a precise multi-layer reranking system with named percentage weights for domain authority, freshness, and citation frequency. None of the guides we checked while researching this piece linked to a Perplexity source, a published paper, or any verifiable methodology behind those numbers — and Perplexity's own documentation doesn't describe its ranking logic at that level of detail at all. That doesn't make every tactic in those guides wrong; freshness and relevance plainly matter, which lines up with the recency and date filters Perplexity actually exposes in its Search API (Perplexity API docs, 2026). But treat any precise percentage breakdown of Perplexity's "ranking formula" as marketing copy, not documented fact, until it's traceable back to Perplexity itself.

Does Perplexity Favor Big Media Over a Small Local Business?

There's a real limitation worth being honest about here. An academic audit of Perplexity, ChatGPT, and Bing Chat across 48 real queries over a week found all three systems relied heavily on "News and Media, Business and Digital Media websites" for their answers, with uneven source quality and commercial and geographic bias built into the results (Li & Sinnamon, 2024). If a query is broad and newsy, Perplexity is going to lean on a wire service or a trade publication before it leans on a single-location business's website — that's not something content tweaks fix.

Where a local business still wins is the narrow question nobody else answers directly: "who handles ___ in Kalispell" or "which companies in the Flathead Valley do ___" isn't a question Reddit, LinkedIn, or a national outlet has a current, specific answer to. That's the gap a crawlable, genuinely-answering local page fills — not by out-ranking Reddit on a broad topic, but by being the only accessible answer to a narrow one.

What Should a Montana or Northwest Business Actually Do About This?

  • Make sure PerplexityBot isn't blocked in robots.txt — Perplexity's own docs confirm blocking the bot removes a site from its index (Perplexity crawler docs, 2026).
  • Publish pages that answer the exact, narrow, local question — by city, by service — that no Reddit thread or national outlet covers.
  • Keep update dates visible and accurate: Perplexity's API exposes explicit recency and freshness filters, so a page with a real, current date has a mechanical advantage (Perplexity API docs, 2026).
  • Show up where Perplexity already over-indexes — relevant subreddits, LinkedIn posts, industry forums — instead of assuming a standalone site will out-rank Reddit on its own.
  • Re-test the exact questions buyers ask every few months; citation patterns shift, as Semrush's own data showed within its 13-week window (Semrush, 2026).

Everything above, a business can do without an agency — none of it requires a tool or a retainer, just the discipline to keep doing it across every engine, every month. Where it gets harder is at scale: tracking the same fixed set of buyer questions across ChatGPT, Perplexity, Gemini, and Google AI Overviews on a recurring basis, and knowing which gaps are actually worth fixing first. That's the part worth outsourcing if nobody in-house has the time to run it.

Skyline AI is the Northwest's AI automation agency — Montana-built, serving businesses across the Flathead Valley and the Pacific Northwest. We run a monthly AI-visibility sweep across ChatGPT, Perplexity, Gemini, and Google AI Overviews so you're working from what's confirmed, not from a guessed-at ranking formula. Book a free AI visibility audit.

Sources

  1. Perplexity crawler docs (2026)
  2. Perplexity API docs (2026)
  3. Semrush (2026)
  4. Li & Sinnamon (2024)
[ 05 ]Questions

Related questions

Clear answers to the questions operators ask most. Still not sure if AI fits your business? Talk to us — no pitch, just a straight read on where it pays off.

What's the difference between PerplexityBot and Perplexity-User?

PerplexityBot is the crawler that builds the index Perplexity searches, and it respects robots.txt. Perplexity-User is a separate agent that fetches a specific page live only when a user's question requires it, and it generally ignores robots.txt because a real person triggered the request (Perplexity crawler docs, 2026).

Does Perplexity cite Wikipedia as much as other AI engines?

No. A Semrush study tracking over 100 million AI citations found Wikipedia made up only around 0.8% of Perplexity's citations, a notably smaller share than on other AI search platforms in the same study (Semrush, 2026).

Can I stop Perplexity from citing my website?

Blocking PerplexityBot in robots.txt removes a site from the index it searches, though Perplexity notes changes can take up to 24 hours to take effect. Perplexity-User can still fetch a page live for a specific question even when PerplexityBot is blocked, since it operates separately (Perplexity crawler docs, 2026).

Is there a published 'ranking algorithm' for Perplexity I can optimize for?

No. Perplexity hasn't published the specifics of its ranking algorithm. Its public documentation confirms domain and recency filters exist in its Search API, but the detailed percentage-weighted 'ranking factor' breakdowns circulating in SEO guides aren't traceable to anything Perplexity has published (Perplexity API docs, 2026).

Why does Perplexity cite Reddit and LinkedIn so often?

A Semrush study tracking 230,000 prompts over 13 weeks found Reddit, LinkedIn, NIH, Microsoft, and Google were Perplexity's top cited domains overall — likely reflecting a weighting toward fresh, discussion-style and institutionally authoritative content (Semrush, 2026).

Related questions

Related services

More from Insights

No Pitch, No Obligation

See exactly where AI pays off in your business

Book a free AI audit. We'll map your biggest leak — missed calls, slow follow-up, manual admin — and show you the system that fixes it. No pitch, no obligation.

Free · no obligation~30 minutesYou own everything