How to Check Whether AI Crawlers Can Access Your Website
AI crawler access decides whether ChatGPT, Claude, Perplexity, Copilot and Google’s AI features can read your pages before they answer questions about your business. A single line in robots.txt, a firewall rule or a default setting at your content delivery network can shut those crawlers out without anyone on the marketing team noticing. Checking access is usually the first technical task in AI search optimisation for B2B websites, because no amount of content work helps if the crawlers are turned away before they reach a page.
Each AI company runs several crawlers with different jobs, so blocking the wrong one has very different consequences from blocking the right one. The sections below set out which crawlers matter, how to test each layer that can block them and how to decide whether training crawlers should be allowed at all.
Which AI Crawlers Need Access to Your Website
The crawlers that matter most belong to OpenAI, Anthropic, Perplexity, Google and Microsoft. Each company publishes the names its crawlers use, so website owners can allow or block them by name in robots.txt and find them in server logs.
| Crawler | Company | What it does | Effect of blocking it |
|---|---|---|---|
GPTBot |
OpenAI | Collects content that may be used to train OpenAI’s foundation models | Signals that content should not be used for training |
OAI-SearchBot |
OpenAI | Surfaces websites in ChatGPT search results | Pages are not shown in ChatGPT search answers |
ChatGPT-User |
OpenAI | Visits pages when a ChatGPT user asks a question | Robots.txt rules may not apply to it |
ClaudeBot |
Anthropic | Collects content that could contribute to model training | Future content is excluded from training datasets |
Claude-SearchBot |
Anthropic | Indexes content to improve Claude’s search results | Visibility in Claude’s search results may fall |
Claude-User |
Anthropic | Fetches pages when a Claude user asks a question | Claude cannot retrieve your pages for user questions |
PerplexityBot |
Perplexity | Surfaces and links websites in Perplexity search results | Pages may not appear in Perplexity search results |
Perplexity-User |
Perplexity | Visits pages when a Perplexity user asks a question | It generally ignores robots.txt rules |
Google-Extended |
A robots.txt token covering Gemini training and grounding | No effect on inclusion or ranking in Google Search | |
Bingbot |
Microsoft | Bing’s standard crawler | Pages that are not crawled are not indexed by Bing |
OpenAI’s documentation on its crawlers and user agents states that each setting is independent of the others. A website can allow OAI-SearchBot so its pages appear in ChatGPT search while disallowing GPTBot to keep its content out of model training. The same page says ChatGPT-User is not used to crawl the web automatically and that a robots.txt change can take around a day to reach ChatGPT’s search systems.
Anthropic splits its crawlers in the same way. Its help article on how Anthropic crawls the web describes ClaudeBot for training data, Claude-SearchBot for search quality and Claude-User for pages fetched when someone asks Claude a question. It also warns that blocking Anthropic’s IP addresses instead of using robots.txt may not work reliably, because the block stops the crawler reading the robots.txt file at all.
Perplexity’s crawler documentation says PerplexityBot surfaces and links websites in its search results and is not used to crawl content for AI foundation models. Its Perplexity-User fetcher visits pages in response to user questions and generally ignores robots.txt rules, because a person requested the fetch.
Google works differently because AI Overviews and AI Mode are part of Google Search. Google’s guidance on AI features and your website says robots.txt rules for Googlebot are the control for how a website is crawled for Search. A page must be indexed and eligible to show with a snippet before it can appear as a supporting link.
The separate Google-Extended token, listed among Google’s common crawlers, manages whether content is used to train future Gemini models and for grounding in Gemini Apps. Google states that it does not affect inclusion in Google Search and is not used as a ranking signal.
For Microsoft, the crawler to check is Bingbot, which Bing’s overview of its crawlers describes as its standard crawler. Bing Webmaster Tools now includes an AI Performance report that shows how often pages are cited in Microsoft Copilot and AI summaries in Bing. The announcement also confirms that Bing respects the preferences set in robots.txt.
Training Crawlers, Search Crawlers and User Fetches
AI crawlers fall into three groups and the group decides what blocking one of them costs. Mixing the groups up is how a website can disappear from AI answers when the only intention was to stop model training.
Model Training Crawlers
GPTBot, ClaudeBot and the Google-Extended token cover content that may be used to train models. Blocking them is a decision about training rather than about appearing in ChatGPT or Claude search.
Search Index Crawlers
OAI-SearchBot, Claude-SearchBot, PerplexityBot, Bingbot and Googlebot index pages so they can be cited in answers. Blocking them reduces or removes the chance of being cited.
User Triggered Fetchers
ChatGPT-User, Claude-User and Perplexity-User visit a page when a person asks a question. Some of them may not follow robots.txt, because a person requested the visit.
Cloudflare uses a similar set of categories in its own bot settings, which matters later when checking the firewall. The groups also explain why a website can be visible in one assistant and missing from another, since each company’s search crawler is controlled separately.
How to Check Whether AI Crawlers Can Access Your Website
Checking AI crawler access means testing every layer between the crawler and the page, starting with robots.txt and working outwards to the firewall and the server logs. A clean robots.txt file proves little on its own if a security setting returns an error to the same crawler.
-
1
Read the Live File
Open the robots.txt address on your domain in a browser. Read the version served to the public rather than the one in the CMS.
-
2
Find the Matching Group
Work out which group of rules applies to each crawler. A named group replaces the general rules for that crawler.
-
3
Check the Firewall
Review CDN, hosting and security plugin settings for AI bot blocking. These can refuse crawlers that robots.txt allows.
-
4
Read the Logs
Search the server logs for each crawler name and note the status codes returned. A 403 or a challenge page means the crawler is being refused.
-
5
Confirm With Platform Tools
Use the robots.txt report in Search Console and the reports in Bing Webmaster Tools. Check again after any change to plugins or security settings.
Each step is covered in more detail below. The order matters, because a robots.txt fix achieves nothing while a firewall rule in front of it is still refusing the same requests.
Reading Your Robots.txt File Correctly
Reading robots.txt correctly means finding the one group of rules that applies to each crawler, because crawlers do not combine a named group with the general one. Google’s robots.txt specification explains that a crawler follows the group with the most specific user agent that matches it and ignores the others. Groups for a specific user agent and the global group are not combined.
Bing applies the same logic. Its guide to creating a robots.txt file notes that Bingbot ignores the generic section once it finds instructions written for itself, so the general rules have to be repeated in its own group. A file that adds a short group for GPTBot to allow one folder can therefore leave that crawler free to reach areas the general group was meant to protect.
A file that keeps search crawlers in while opting out of training might look like the example below. The general WordPress rules are repeated inside the search group so they still apply to those crawlers.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
WordPress adds its own layer to the file. The core do_robots function, documented in the WordPress developer reference, builds a default robots.txt and lets plugins and themes change it through the robots_txt filter. An SEO or security plugin can therefore add AI crawler rules that nobody wrote by hand.
Errors on the file itself matter just as much as its rules. Google’s specification says a robots.txt address that returns a 4xx error is treated as if no file existed, while a server error makes Google stop crawling the website for the first 12 hours and then fall back to the last good copy. The robots.txt report in Search Console shows the files Google found, when they were last crawled and any errors, which is the quickest way to confirm what Googlebot sees.
How Cloudflare and Firewalls Block AI Crawlers
Cloudflare and similar security services can block AI crawlers before they ever read robots.txt. Some of these settings now apply by default. Cloudflare’s documentation on AI bot policies groups crawlers by behaviour into Search, Agent and Training. Each group can be blocked on every page, blocked only on pages that show ads or allowed.
The same page says that from 15 September 2026 new domains on Cloudflare block Training and Agent bots on pages that display ads by default, while Search stays allowed. Mixed purpose crawlers that combine search and training are also caught by every setting that blocks AI training. Any website behind Cloudflare should check these settings, even if nobody remembers changing them.
Cloudflare can also change robots.txt itself. With its managed robots.txt setting turned on, Cloudflare adds its own rules above the existing file. Its published example disallows crawlers such as ClaudeBot. The file served to crawlers can therefore differ from the one stored in WordPress, which is why the first check should always be the public address rather than the CMS.
Cloudflare’s AI Crawl Control lists the AI crawlers requesting your pages, how many of their requests were unsuccessful and how often each one broke your robots.txt rules. It is the fastest way to see whether the firewall and the file agree.
Other firewalls raise the same issue. Perplexity’s documentation recommends explicitly allowing its crawlers in web application firewalls by matching both the user agent and its published IP ranges. The same approach suits security plugins and hosting firewalls that block unfamiliar bots.
Checking Server Logs for Real Crawler Visits
Server logs are the only record of what AI crawlers requested and what your website sent back. Search the access logs for each crawler name and note the status codes, since a run of 403 responses or challenge pages shows a block that robots.txt testing will never reveal.
User agent names can be copied by anyone, so a log entry that says GPTBot is not proof of a visit from OpenAI. Bing’s guidance on verifying Bingbot uses a reverse DNS lookup followed by a forward lookup to confirm the address. OpenAI and Perplexity publish lists of the IP addresses their crawlers use, which serve the same purpose.
Logs also show whether access matters in practice. A crawler that requests your service pages regularly is reading them, while one that only ever fetches robots.txt may be obeying a disallow rule nobody knew about. Regular log file analysis for SEO can cover Googlebot and the AI crawlers in the same review.
Should You Block GPTBot and Other Training Crawlers
Blocking GPTBot and other training crawlers is a reasonable choice for many businesses, provided the search crawlers stay allowed. OpenAI states that disallowing GPTBot indicates content should not be used to train its foundation models, while OAI-SearchBot is the crawler that decides whether pages appear in ChatGPT search.
The case for blocking is strongest where content is the product, such as paid research, training material or original data a business sells. The case for allowing is stronger for most B2B service companies, because their pages exist to be read and there is little commercial reason to hold marketing pages back from training. Neither choice affects Google Search, since Google states that Google-Extended is not a ranking signal.
- Allow the search crawlers from every AI company
- Decide on each training crawler separately
- Repeat the general rules inside every named group
- Record each decision and the date it was made
- Block every AI crawler with one wildcard rule
- Rely on IP blocks to opt out of training
- Assume firewall defaults match your robots.txt
- Treat the Gemini training token as an AI Overviews switch
AI Overviews cannot be switched off through Google-Extended, because they are part of Google Search. Google’s guidance on AI features points to the nosnippet, max-snippet and noindex controls for limiting what is shown from a page. Each of those also changes how the page appears in standard search results.
UK law on training AI with copyright material is still under review. The government’s consultation on copyright and AI closed in February 2025 with aims that include giving rights holders more control over whether their work is used to train AI models. A decision on training crawlers is worth revisiting as that policy develops.
Keeping AI Crawler Access Under Review
AI crawler access needs checking whenever plugins, hosting or security settings change, because any of them can rewrite robots.txt or add a block without warning. New crawlers also appear as AI companies launch products. Each one needs a decision recorded alongside the existing rules.
A short quarterly check covers most of the risk. Read the public robots.txt, confirm the firewall settings still match it, look for each crawler in the logs and compare the result with the Search Console and Bing Webmaster Tools reports. A sudden fall in visits from AI assistants is a reason to run the same checks straight away, alongside the wider causes covered in why a website is not showing up on Google.
Crawler access is one part of a wider approach to AI visibility, which generative engine optimisation covers in more depth. Priority Pixels treats crawler access as part of the technical work behind its AI SEO service, alongside schema and the reporting that shows AI referral traffic.
Priority Pixels also configures robots.txt files and XML sitemaps as part of its technical SEO services, with server log analysis and Search Console data used to show current crawling patterns. Whoever does the work, the goal is a robots.txt file, a firewall and a set of logs that all tell the same story about which crawlers are welcome.
FAQs
Does robots.txt stop AI crawlers from reading a website?
A robots.txt file states your preferences and the main AI companies say their automatic crawlers follow it. It cannot physically stop a crawler, while the fetchers that visit a page when a person asks ChatGPT or Perplexity a question may not apply its rules. A firewall or bot management rule is needed where access has to be enforced rather than requested.
Does blocking GPTBot remove a website from ChatGPT search?
Blocking GPTBot only signals that your content should not be used to train OpenAI’s models. Appearing in ChatGPT search depends on OpenAI’s separate search crawler, which has its own robots.txt setting.
Can you stop Google using your content in AI Overviews?
AI Overviews are part of Google Search, so the separate token Google offers for Gemini training does not remove pages from them. Google’s guidance points to snippet controls and noindex for limiting what is shown. Those settings change how the page appears in standard results too.
How can you tell whether a visit really came from an AI crawler?
Check the IP address in the server log against the list the AI company publishes, or run a reverse DNS lookup followed by a forward lookup. User agent names are easy to copy, so the name alone does not prove who made the request.