Why Perplexity Is Not Citing Your Website
Perplexity cites a website only when it can reach the page, holds a current copy of it in its index and judges a passage on it to be the best available answer to the question. A website that never appears in the numbered sources under Perplexity’s answers is failing at one of those three points. The first is easy to overlook, because a firewall setting or a robots.txt rule can turn Perplexity away without anyone noticing. Earning citations is the aim of generative engine optimisation for B2B brands, but the work starts with technical checks that take minutes rather than months.
The sections below follow the same order Perplexity does. They start with crawler access and the two user agents Perplexity publishes, then move on to firewalls, content structure, freshness and mentions on other websites, with tracking at the end.
How Perplexity Finds and Cites Sources
Perplexity answers a question by searching the web, pulling the most relevant passages from the pages it finds and writing a summary with numbered citations that link back to those pages. Perplexity’s help centre article on how Perplexity works describes this as searching the internet in real time and adding numbered citations to each answer so that people can check the original sources.
The search runs on Perplexity’s own index rather than on Google’s results. In a research article on building the Perplexity Search API, Perplexity says its index tracks over 200 billion unique URLs. The same article explains that a machine learning model predicts which URLs need indexing and when, while documents from authoritative domains are among those it keeps ready for fast access.
Pages are not judged only as whole documents. Perplexity splits each page into smaller passages that are retrieved and ranked individually, using keyword and semantic matching together before a final reranking step. Early filters remove content judged stale or unresponsive to the question. A single passage wins the place in the answer, which is why a long page with the right information buried in the middle can lose to a shorter page that states the answer plainly.
Three conditions therefore have to hold before a page earns a citation. Perplexity has to be allowed to fetch it, the page has to be in the index in a current form and a passage on it has to answer the question better than the alternatives. A wider look at the platform itself is available in how Perplexity works as a search tool, which covers its growth and the way it differs from a traditional search engine.
Check Whether PerplexityBot Can Reach Your Website
PerplexityBot is the crawler that builds Perplexity’s search index, so a website that blocks it in robots.txt is unlikely to be cited for anything on its pages. Perplexity’s developer documentation on Perplexity crawlers says PerplexityBot is designed to surface and link websites in search results on Perplexity. It recommends allowing PerplexityBot in robots.txt and permitting requests from the IP ranges Perplexity publishes.
Perplexity’s help centre article on how Perplexity follows robots.txt says PerplexityBot will not index the full or partial text of a website that disallows it. A blocked page may still be indexed with its domain, headline and a brief factual summary, but Perplexity then has no text from the page to quote. The same article says the third party crawlers Perplexity works with to build its index have agreed to respect robots.txt as well.
Start by opening your own domain followed by /robots.txt in a browser and reading the groups line by line. The robots.txt standard, RFC 9309, says a crawler follows the group that names it and only falls back to the group for all crawlers, written as an asterisk, when no group names it. A rule for all crawlers that disallows everything therefore blocks PerplexityBot unless a separate PerplexityBot group allows it, while a PerplexityBot group with its own rules replaces the general group entirely for that crawler.
The common setups and what each one means for PerplexityBot are set out below. Each row assumes the file loads normally unless the row says otherwise.
| robots.txt setup | Effect on PerplexityBot | What to change |
|---|---|---|
User-agent: * with Disallow: / and no PerplexityBot group |
Blocked from every page | Add a PerplexityBot group that allows the pages you want cited |
A PerplexityBot group with Disallow: / |
Blocked from every page, whatever the general rules say | Remove the group or narrow the disallowed paths |
| Disallow rules on folders such as /insights/ or /services/ | Blocked from those folders only | Check that no page you want cited sits inside them |
| robots.txt returns a server error | The standard tells crawlers to assume everything is disallowed | Fix the error so the file loads with a 200 status |
| robots.txt returns a 404 | The standard lets crawlers access any page | Nothing for access, though a real file makes the rules clear |
Changes to robots.txt do not take effect straight away. Perplexity’s crawler documentation says each setting works independently and that it may take up to 24 hours for its systems to reflect a change, after which PerplexityBot still has to return to the pages before anything new can be cited.
On WordPress the robots.txt file is often generated by the software rather than stored as a file on the server. The WordPress code reference for do_robots() shows that the default output only disallows the admin area and that plugins can rewrite the file through the robots_txt filter, so an SEO or security plugin may be adding rules that nobody set deliberately. The same page notes that since WordPress 5.3 the setting that discourages search engines adds a robots meta tag instead of a disallow rule, so it will not show up in robots.txt at all.
Perplexity’s documentation does not say how PerplexityBot treats that meta tag. The safest course is to confirm the setting is switched off on the live website, since a staging copy can go live with it still turned on.
Why Perplexity Uses Two Separate User Agents
Perplexity publishes two user agents because it fetches pages in two different ways. Each one has a different job and a different relationship with robots.txt, so a website needs to treat them as separate decisions.
The crawler documentation describes Perplexity-User as the agent that supports user actions within Perplexity. When someone asks a question, Perplexity may visit a web page to help answer it and include a link to that page in its response, which makes this agent a direct route to a citation.
PerplexityBot
Crawls pages in advance to build the index that Perplexity searches when it answers a question. It obeys the rules in robots.txt. Perplexity also says it is not used to crawl content for AI foundation models.
Perplexity-User
Visits a page when a person’s question calls for it, so the answer can draw on that page and link to it. Because a person asked for the fetch, Perplexity says it generally ignores robots.txt rules.
The difference matters most when a firewall is involved. Blocking Perplexity-User at the server or the content delivery network stops Perplexity reading a page at the moment someone asks about it, even though a robots.txt rule would not have stopped that visit.
Perplexity says neither agent is used to collect content for training AI foundation models, which is the concern behind many blanket AI bot blocks. A business that wants to appear in Perplexity answers can usually allow both without handing its content over for model training, although that remains a commercial decision each organisation has to make for itself.
Firewalls and Bot Protection That Block Perplexity
A firewall or bot protection service can block Perplexity even when robots.txt allows it, because the request is refused before the crawler gets as far as reading the rules. Perplexity’s crawler documentation says websites behind a web application firewall may need to explicitly allow its bots. It recommends rules that match both the user agent string and Perplexity’s published IP address ranges.
Matching on IP ranges as well as the user agent matters because any request can claim to be PerplexityBot in its user agent string. Perplexity publishes the ranges for each agent as JSON files and advises fetching them regularly to update firewall rules, since the addresses change over time.
Cloudflare deserves a specific check because its AI bot settings can block crawlers without any change to the website itself. Cloudflare’s documentation on blocking AI bots groups crawlers by behaviour into Search, Agent and Training. Each group can be blocked on every page, blocked only on pages that show ads or allowed.
Cloudflare says that from 15 September 2026 new domains block Training and Agent bots on pages that display ads by default, with Agent covering chat fetch bots that act for a person in real time. Check the AI bot policies in the Cloudflare dashboard rather than assuming Perplexity can reach every page.
Cloudflare’s AI Crawl Control shows which AI services are requesting pages and lets website owners allow or block individual crawlers, which makes it a quick place to confirm whether PerplexityBot is getting through. The history between the two companies also explains why some Cloudflare settings treat Perplexity with suspicion. That dispute is covered in why Cloudflare blocked Perplexity AI.
Security plugins on WordPress and firewalls run by the hosting company can do the same thing with less visibility. Server access logs settle the question, because they show whether requests naming PerplexityBot are arriving and which status code each one received. A run of 403 responses points to a firewall rule, while no requests at all can mean Perplexity has not found the pages or is being turned away before the server sees anything.
Content Perplexity Can Quote
Once Perplexity can reach a page, the content has to give it a passage worth citing. Because Perplexity ranks individual passages rather than whole pages, each section should make sense when read on its own and answer its heading in the opening sentence.
Passages that state a fact plainly, name the thing they describe and give the condition or figure that answers the question are easier to lift into an answer than passages that build up to a point slowly. A section that answers in its first two sentences what its heading asks gives a retrieval system exactly the unit it is looking for.
The traits below follow from how Perplexity describes its own retrieval and from ordinary editorial practice. None of them is a trick. All of them help human readers too.
- Headings phrased the way people ask questions, with the answer straight after them
- One topic per section, so a passage does not depend on text elsewhere on the page
- Specific names, versions, dates and conditions rather than general statements
- Figures with a linked source, so the claim can be checked
- Tables and lists for comparisons and steps rather than long blocks of prose
- Important text in the page itself rather than only inside images
Thin pages give Perplexity very little to quote. A service page that describes what a business does in two short paragraphs offers less than a page that sets out who the service is for, how it works and what it involves, so the fuller page is far more likely to supply the passage that gets cited.
Perplexity’s help centre says its Research mode performs dozens of searches for a single question, so a page can be found through a related question rather than the exact one it targets. Covering the follow up questions a buyer would ask on the same page widens the number of searches it can answer, which is the idea behind query fan out in AI search.
Freshness and How Often Pages Are Indexed
Perplexity gives weight to current content and builds its index to refresh existing pages as well as add new ones. Its research article names freshness as one of three core requirements for its search system and says early filters remove stale content before ranking begins.
The same article says a model decides when to crawl each URL again based on how important it is and how often it is likely to change. Pages that are updated in a meaningful way on a regular basis give that model a reason to come back, while pages that never change are revisited less often.
Visible publication and update dates help readers judge how current a page is. A sitemap with accurate last modified dates gives crawlers a list of what has changed, although Perplexity does not document how it uses sitemaps, so dates should only change when the content does.
Refreshing a page means adding what has changed in the subject, correcting anything out of date and checking that every source still says what the page claims. A new date on unchanged content gives readers nothing and adds no value for any system that compares versions of a page.
Authority and Mentions Beyond Your Website
Perplexity draws on several websites for each answer, so the pages that describe a business are not limited to its own. Its help centre describes answers built from authoritative sources such as articles, websites and journals, which means industry publications, directories, review platforms and partner websites can all shape what Perplexity says about a company.
A business that is described consistently on credible third party websites gives Perplexity more sources that agree with each other. Inconsistent company names, old service descriptions or outdated addresses spread across directories work against that, because the answer may cite whichever version Perplexity finds first.
Original material earns mentions in a way that rewritten material does not. Survey results, clear definitions, detailed comparisons and documented processes give other writers a reason to link to a page. Each of those links is another route for Perplexity to find the page and another signal that other people trust it.
Mentions on other websites take longer to build than technical fixes. They are also the part of the work that a competitor finds hardest to copy, which is why they matter once access and content are in order.
Tracking Whether Perplexity Cites Your Website
Tracking Perplexity citations means asking the questions buyers ask and recording which websites appear in the numbered sources. Answers change from one run to the next, so a single check proves little and a fixed set of questions run on a schedule gives a far clearer picture.
Results also differ between Perplexity’s standard search, Pro Search and Research, since the deeper modes search more widely and read more sources. Record which mode was used alongside each result so that comparisons over time are fair.
A simple manual routine covers most of what a dashboard would show for a small set of priority questions. The steps below run from choosing the questions to acting on the gaps they reveal.
-
1
Choose the Questions
List the questions a buyer would ask before contacting a business like yours. Include comparison and problem questions as well as questions that name your services.
-
2
Run Them in Perplexity
Run each question in a new thread and note the mode used. Save the answer and every cited URL.
-
3
Record the Sources
Note whether your website is cited, which page is cited and which competitors appear. Count mentions without a citation separately.
-
4
Check the Gaps
Compare each cited competitor page with your own page on the same subject. Look at access, structure, freshness and outside mentions in that order.
-
5
Repeat Each Month
Run the same questions again every month. Change one thing at a time so the effect of each fix can be seen.
Tools that automate this testing can cover many more questions and AI platforms at once. The guide to checking LLM visibility covers the manual methods and the tools that help in more detail.
Priority Pixels tracks brand presence across ChatGPT, Gemini, Copilot and Perplexity using systematic prompt testing and response analysis as part of its LLM performance tracking and analytics service. Whoever runs the testing, the order of work stays the same, with access checked first, content improved second and authority built over time.
FAQs
Does blocking PerplexityBot remove a website from Perplexity completely?
Blocking PerplexityBot in robots.txt stops it indexing the text on your pages, but Perplexity says it may still index the domain, headline and a brief factual summary. Perplexity has no page text to quote in that case, so a citation becomes very unlikely.
How long does a robots.txt change take to reach Perplexity?
Perplexity’s crawler documentation says it can take up to 24 hours for its systems to reflect a change to robots.txt. PerplexityBot then has to crawl the affected pages again before they can be cited, which can take longer for pages it visits rarely.
Does Perplexity use website content to train AI models?
Perplexity says PerplexityBot is not used to crawl content for AI foundation models and that the company does not build foundation models itself. Its help centre answers the question of whether content allowed into Perplexity is used for AI training with a clear no.
Will ranking well on Google get a page cited by Perplexity?
Perplexity runs its own search index and ranking system, so a Google ranking does not carry over automatically. The qualities that help a page rank on Google, such as clear answers, accurate information and links from credible websites, tend to help with Perplexity as well.