The Complete Guide to robots.txt for AI Crawlers: Which Bots to Allow
GPTBot vs OAI-SearchBot, ClaudeBot vs Claude-User, Googlebot vs Google-Extended: which AI crawlers to allow, which to block, and copy-ready robots.txt templates.
By Crawlable · · 7 min read
On this page
Most robots.txt advice about AI treats every AI bot alike: block them all, or allow them all. Both miss how the operators actually work. OpenAI, Anthropic, Google and Perplexity each run separate tokens for separate jobs: one collects training data, another builds the index an answer engine cites from, a third fetches a page because a user just asked about it.
A blanket Disallow: / for "AI" removes you from ChatGPT search, Claude and Perplexity answers along with the training sets. This guide sorts the crawlers by job, compares the pairs people confuse, and gives you two copy-ready files.
Cheat sheet: the pairs people confuse
| Token | Operator | Job | Block it and… |
|---|---|---|---|
GPTBot | OpenAI | Collects training data | Your content is excluded from future training. ChatGPT answers are unaffected. |
OAI-SearchBot | OpenAI | Builds the ChatGPT search index | You stop appearing in ChatGPT search answers (navigational links may remain). |
ChatGPT-User | OpenAI | Fetches a page a user asked about | ChatGPT can't read the page for that user. OpenAI says robots.txt "may not apply". |
ClaudeBot | Anthropic | Collects training data | Excluded from future training. Claude answers are unaffected. |
Claude-SearchBot | Anthropic | Indexes pages for Claude's search | Anthropic says visibility in Claude's results may drop. |
Claude-User | Anthropic | Fetches a page a Claude user asked about | Claude can't retrieve your page when asked. |
Googlebot | Crawls for Google Search, including AI Overviews and AI Mode | You leave Google Search entirely. Don't. | |
Google-Extended | Control token, not a crawler: Gemini training and grounding | Gemini stops training and grounding on your content. Search rankings unchanged. | |
PerplexityBot | Perplexity | Indexes pages for Perplexity's answers | You stop being surfaced and linked in Perplexity. |
Perplexity-User | Perplexity | Fetches a page a user asked about | Perplexity says it "generally ignores robots.txt rules". |
The rule of thumb: tokens ending in Bot that aren't SearchBot usually mean training; SearchBot means the index you're cited from; -User means a live fetch for a person. Google-Extended is the exception that breaks every rule, covered below.
Every crawler Crawlable's scanner evaluates, with its job and whether it runs JavaScript:
| User-agent token | Operator | Job | Runs JavaScript | Blocking costs visibility |
|---|---|---|---|---|
| GPTBot | OpenAI | Model training | No | No |
| OAI-SearchBot | OpenAI | Search index | No | Yes |
| ChatGPT-User | OpenAI | Fetch on user request | No | Yes |
| ClaudeBot | Anthropic | Model training | No | No |
| Claude-SearchBot | Anthropic | Search index | No | Yes |
| Claude-User | Anthropic | Fetch on user request | No | Yes |
| PerplexityBot | Perplexity | Search index | No | Yes |
| Perplexity-User | Perplexity | Fetch on user request | No | Yes |
| Google-Extended | Model training | Yes | No | |
| Applebot-Extended | Apple | Model training | Yes | No |
| Amazonbot | Amazon | Search index | No | Yes |
| Bytespider | ByteDance | Model training | No | No |
| meta-externalagent | Meta | Model training | No | No |
| cohere-ai | Cohere | Model training | No | No |
| MistralAI-User | Mistral | Fetch on user request | No | Yes |
GPTBot vs OAI-SearchBot
This is the most searched pair, and the one most often blocked together by mistake. OpenAI's crawler documentation describes three bots:
- GPTBot crawls content that may be used for training. Disallowing it "indicates a site's content should not be used in training generative AI foundation models."
- OAI-SearchBot surfaces sites in ChatGPT search. Sites that opt out "will not be shown in ChatGPT search answers, though can still appear as navigational links."
- ChatGPT-User acts on user requests, and "because these actions are initiated by a user, robots.txt rules may not apply."
OpenAI says robots.txt changes take about 24 hours to take effect, and publishes IP ranges for each bot (gptbot.json, searchbot.json, chatgpt-user.json) so you can verify that a request claiming to be one of them really is.
# Opt out of OpenAI training, stay in ChatGPT search
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
Allow: /
User-agent: GPTBot
Disallow: /
ClaudeBot vs Claude-User (and Claude-SearchBot)
Anthropic describes three bots: ClaudeBot collects content that could contribute to model training, Claude-SearchBot indexes content to improve search results, and Claude-User fetches pages when a Claude user asks a question. Anthropic says blocking either of the last two "may reduce your site's visibility." All three honour robots.txt, and Anthropic also supports the non-standard Crawl-delay directive.
# Opt out of Anthropic training, stay visible in Claude
User-agent: Claude-SearchBot
User-agent: Claude-User
Allow: /
User-agent: ClaudeBot
Disallow: /
Googlebot vs Google-Extended
Google-Extended isn't a crawler. Google's documentation says it "doesn't have a separate HTTP request user agent string"; Googlebot does the fetching, and Google-Extended is a product token that decides what the fetched content may be used for. It governs whether content Google crawls is used to train Gemini models and to ground answers in Gemini Apps and the Vertex AI API for Gemini. Google is explicit that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal."
AI Overviews and AI Mode are part of Google Search, so they follow Googlebot and your snippet controls (nosnippet, max-snippet), not Google-Extended.
# Opt out of Gemini training and grounding; Google Search untouched
User-agent: Google-Extended
Disallow: /
User-agent: Googlebot
Allow: /
PerplexityBot vs Perplexity-User
Perplexity's documentation says PerplexityBot surfaces and links websites in its search results and is not used to crawl content for AI foundation models; allow it if you want to appear. Perplexity-User fetches pages for user requests and "generally ignores robots.txt rules." Perplexity documents no separate training crawler, so there's no split to make here: allow PerplexityBot or leave Perplexity.
The two other names worth knowing
CCBot builds Common Crawl's open web archive and respects robots.txt. That archive is a common source of language-model training data: filtered Common Crawl text was the largest part of GPT-3's training mix, according to the GPT-3 paper. It feeds no answer engine directly, so blocking it is purely a training decision.
Bytespider and meta-externalagent are training crawlers from ByteDance and Meta. See their pages, Bytespider and meta-externalagent, for what each operator says and what independent reports say about compliance.
Template 1: allow AI search, block training scrapers
This file is generated from the same registry Crawlable's scanner uses. It lets every answer-engine crawler in and opts out of every training crawler:
# Answer engines: allowed, so you can be found and cited
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Amazonbot
User-agent: MistralAI-User
Allow: /
Disallow: /admin/
Disallow: /api/
# Model training: opted out. No effect on Google Search, AI Overviews,
# or ChatGPT, Claude and Perplexity search results.
# Google-Extended also controls grounding in Gemini Apps and Vertex AI.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Bytespider
User-agent: meta-externalagent
User-agent: cohere-ai
Disallow: /
# Everyone else, including Googlebot and Bingbot
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /api/
Sitemap: https://example.com/sitemap.xml
Adapt before you ship it:
- Change the
Disallowpaths to your own private areas, and repeat them in every group (the group rule below explains why). - Change the
Sitemapline to your own absolute sitemap URL. - Add
User-agent: CCBotto the training group if you also want out of Common Crawl's archive.
Template 2: robots.txt allow all
The literal allow-all is two lines:
User-agent: *
Allow: /
An empty Disallow: line means the same thing. Both work, and both name no AI crawler, which has two costs. You can't opt one bot out of training later without restructuring the file, and you can't tell from the file whether "allowed" was a decision or a default. Crawlable's crawler-governance check scores a wildcard-only file as having no AI directives for exactly that reason.
The explicit version allows the same crawlers but names them, in two groups, so flipping training to Disallow: / later is a one-line change:
# Answer engines: allowed, so you can be found and cited
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Amazonbot
User-agent: MistralAI-User
Allow: /
Disallow: /admin/
Disallow: /api/
# Model training: allowed. Change Allow to Disallow: / here to opt out later.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Bytespider
User-agent: meta-externalagent
User-agent: cohere-ai
Allow: /
Disallow: /admin/
Disallow: /api/
# Everyone else, including Googlebot and Bingbot
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /api/
Sitemap: https://example.com/sitemap.xml
The group rule that catches everyone
Under the robots.txt standard, RFC 9309, a crawler follows the most specific group that names it and ignores the * group entirely. If you write User-agent: GPTBot with only Allow: /, GPTBot is now allowed into /admin/, even though your * group disallows it. Every named group needs its own copy of your private-path rules, which is why both templates repeat them.
The same standard lets several User-agent lines share one group, which is how the templates stay short. Don't put a search bot and a training bot in the same group if you ever want to treat them differently.
Why Disallow: / for every AI bot is a visibility mistake
Lists of "AI bots to block" circulate as copy-paste snippets, and most of them include OAI-SearchBot, Claude-SearchBot and PerplexityBot. Paste one in and you've removed yourself from ChatGPT search, Claude's search results and Perplexity, the products where a citation sends a click, while doing nothing about the -User fetchers that may not honour the file anyway.
If what you object to is training, block the training tokens. Leave the index and user fetchers alone.
Without compromising security
robots.txt is a set of instructions for well-behaved crawlers. It has no enforcement behind it.
- Never use it to hide anything. The file is public.
Disallow: /internal-reports-2026/advertises exactly the path you didn't want found. Protect private areas with authentication, and keep pages out of indexes with anoindexmeta tag or header, which only works if the crawler can fetch the page and see it. - Don't trust a user-agent string. Any scraper can call itself GPTBot. Google warns that its own user agent "is often spoofed" and documents verification by reverse DNS or published IP ranges. If you let a bot past a rate limit because of its name, check its IP against the operator's published list too.
- Enforce at the edge, not in the file. If a bot ignores your rules or hammers your server, block or rate-limit it in your CDN or firewall.
- Check your firewall isn't doing the opposite. One-click "block AI bots" or bot-fight settings in CDN and WAF products can block OAI-SearchBot and PerplexityBot even when robots.txt allows them. The file says yes, the network says no.
Common mistakes
User-agent: *withDisallow: /in production. Usually a staging config that shipped. It removes you from Google and every answer engine at once.- Expecting Google-Extended to hide you from AI Overviews. It doesn't; AI Overviews are Google Search.
- Relying on outdated tokens. Crawlers get renamed and split. Re-check the operators' documentation, not a two-year-old snippet.
- Expecting changes to be instant. OpenAI quotes about 24 hours. Other operators re-read the file on their own schedule.
Test it
Fetch your live file with a crawler's token, and check it returns 200 with text/plain rather than a bot challenge or a redirect. A WAF rule keyed on the user agent shows up here as a 403:
curl -sS -o /dev/null -w "%{http_code} %{content_type}\n" \
-A "OAI-SearchBot" \
https://yoursite.com/robots.txt
curl -sS https://yoursite.com/robots.txt
Then read each group against the cheat sheet above. Crawlable's free scan does this for you: it parses your robots.txt, evaluates it against every crawler in the table, and flags a search bot caught by a rule you inherited from a plugin or CMS template.
Frequently asked questions
- What does an allow-all robots.txt look like?
- The minimal allow-all file is two lines: User-agent: * followed by Allow: / (an empty Disallow: line means the same thing). Every crawler that honours robots.txt may fetch every URL. It names no AI crawler, though, so you cannot later opt one out of training without restructuring the file.
- Which AI crawlers should I allow in robots.txt?
- Allow the crawlers that put you in answers: OAI-SearchBot and ChatGPT-User (ChatGPT), Claude-SearchBot and Claude-User (Claude), PerplexityBot and Perplexity-User (Perplexity), Amazonbot and MistralAI-User. Whether to allow training crawlers such as GPTBot, ClaudeBot, Google-Extended or CCBot is a separate choice that does not change whether you are cited today.
- What is the difference between GPTBot and OAI-SearchBot?
- GPTBot collects content that may be used to train OpenAI's foundation models. OAI-SearchBot builds the index behind ChatGPT search. Blocking GPTBot opts you out of training only; blocking OAI-SearchBot keeps you out of ChatGPT search answers, though OpenAI says you can still appear as a navigational link.
- Does blocking GPTBot remove my site from ChatGPT?
- No. ChatGPT search draws on OAI-SearchBot's index, and ChatGPT-User fetches pages users ask about. GPTBot is the training crawler. You can disallow GPTBot and keep OAI-SearchBot allowed, which is what the first template in this guide does.
- Does blocking Google-Extended hurt my Google rankings?
- No. Google states that Google-Extended does not affect a site's inclusion in Google Search and is not a ranking signal. It controls whether content Google crawls is used for Gemini training and for grounding in Gemini Apps and the Vertex AI API. AI Overviews follow Googlebot, not Google-Extended.
- Can robots.txt stop ChatGPT-User or Perplexity-User?
- Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User because a person initiated the fetch, and Perplexity says Perplexity-User generally ignores robots.txt. If you must stop them, block their published IP ranges at your CDN or firewall.
Keep reading
- Why AI Search Engines Can’t Read Client-Side Rendered (CSR) Sites
Most AI crawlers fetch raw HTML and never run your JavaScript. Here's what they actually see on a client-rendered site, how to test it, and how to fix it.
- How to Make React Apps Crawlable for AI Search: SSR, Prerendering and Dynamic Rendering
Google can render React; ChatGPT, Claude and Perplexity crawlers can't. Code for Vite SSR, Next.js and Astro prerendering, dynamic rendering, and a curl check.
- What GPTBot Actually Sees When Your Site Relies on Client-Side JavaScript
AI crawlers like GPTBot and OAI-SearchBot fetch raw HTML without executing client-side JS. Here is what they actually extract from SPAs.