Skip to content
Crawlable

The Complete Guide to robots.txt for AI Crawlers: Which Bots to Allow

GPTBot vs OAI-SearchBot, ClaudeBot vs Claude-User, Googlebot vs Google-Extended: which AI crawlers to allow, which to block, and copy-ready robots.txt templates.

By Crawlable · · 7 min read

On this page

Most robots.txt advice about AI treats every AI bot alike: block them all, or allow them all. Both miss how the operators actually work. OpenAI, Anthropic, Google and Perplexity each run separate tokens for separate jobs: one collects training data, another builds the index an answer engine cites from, a third fetches a page because a user just asked about it.

A blanket Disallow: / for "AI" removes you from ChatGPT search, Claude and Perplexity answers along with the training sets. This guide sorts the crawlers by job, compares the pairs people confuse, and gives you two copy-ready files.

Cheat sheet: the pairs people confuse

TokenOperatorJobBlock it and…
GPTBotOpenAICollects training dataYour content is excluded from future training. ChatGPT answers are unaffected.
OAI-SearchBotOpenAIBuilds the ChatGPT search indexYou stop appearing in ChatGPT search answers (navigational links may remain).
ChatGPT-UserOpenAIFetches a page a user asked aboutChatGPT can't read the page for that user. OpenAI says robots.txt "may not apply".
ClaudeBotAnthropicCollects training dataExcluded from future training. Claude answers are unaffected.
Claude-SearchBotAnthropicIndexes pages for Claude's searchAnthropic says visibility in Claude's results may drop.
Claude-UserAnthropicFetches a page a Claude user asked aboutClaude can't retrieve your page when asked.
GooglebotGoogleCrawls for Google Search, including AI Overviews and AI ModeYou leave Google Search entirely. Don't.
Google-ExtendedGoogleControl token, not a crawler: Gemini training and groundingGemini stops training and grounding on your content. Search rankings unchanged.
PerplexityBotPerplexityIndexes pages for Perplexity's answersYou stop being surfaced and linked in Perplexity.
Perplexity-UserPerplexityFetches a page a user asked aboutPerplexity says it "generally ignores robots.txt rules".

The rule of thumb: tokens ending in Bot that aren't SearchBot usually mean training; SearchBot means the index you're cited from; -User means a live fetch for a person. Google-Extended is the exception that breaks every rule, covered below.

Every crawler Crawlable's scanner evaluates, with its job and whether it runs JavaScript:

User-agent tokenOperatorJobRuns JavaScriptBlocking costs visibility
GPTBotOpenAIModel trainingNoNo
OAI-SearchBotOpenAISearch indexNoYes
ChatGPT-UserOpenAIFetch on user requestNoYes
ClaudeBotAnthropicModel trainingNoNo
Claude-SearchBotAnthropicSearch indexNoYes
Claude-UserAnthropicFetch on user requestNoYes
PerplexityBotPerplexitySearch indexNoYes
Perplexity-UserPerplexityFetch on user requestNoYes
Google-ExtendedGoogleModel trainingYesNo
Applebot-ExtendedAppleModel trainingYesNo
AmazonbotAmazonSearch indexNoYes
BytespiderByteDanceModel trainingNoNo
meta-externalagentMetaModel trainingNoNo
cohere-aiCohereModel trainingNoNo
MistralAI-UserMistralFetch on user requestNoYes

GPTBot vs OAI-SearchBot

This is the most searched pair, and the one most often blocked together by mistake. OpenAI's crawler documentation describes three bots:

  • GPTBot crawls content that may be used for training. Disallowing it "indicates a site's content should not be used in training generative AI foundation models."
  • OAI-SearchBot surfaces sites in ChatGPT search. Sites that opt out "will not be shown in ChatGPT search answers, though can still appear as navigational links."
  • ChatGPT-User acts on user requests, and "because these actions are initiated by a user, robots.txt rules may not apply."

OpenAI says robots.txt changes take about 24 hours to take effect, and publishes IP ranges for each bot (gptbot.json, searchbot.json, chatgpt-user.json) so you can verify that a request claiming to be one of them really is.

# Opt out of OpenAI training, stay in ChatGPT search
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
Allow: /

User-agent: GPTBot
Disallow: /

ClaudeBot vs Claude-User (and Claude-SearchBot)

Anthropic describes three bots: ClaudeBot collects content that could contribute to model training, Claude-SearchBot indexes content to improve search results, and Claude-User fetches pages when a Claude user asks a question. Anthropic says blocking either of the last two "may reduce your site's visibility." All three honour robots.txt, and Anthropic also supports the non-standard Crawl-delay directive.

# Opt out of Anthropic training, stay visible in Claude
User-agent: Claude-SearchBot
User-agent: Claude-User
Allow: /

User-agent: ClaudeBot
Disallow: /

Googlebot vs Google-Extended

Google-Extended isn't a crawler. Google's documentation says it "doesn't have a separate HTTP request user agent string"; Googlebot does the fetching, and Google-Extended is a product token that decides what the fetched content may be used for. It governs whether content Google crawls is used to train Gemini models and to ground answers in Gemini Apps and the Vertex AI API for Gemini. Google is explicit that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal."

AI Overviews and AI Mode are part of Google Search, so they follow Googlebot and your snippet controls (nosnippet, max-snippet), not Google-Extended.

# Opt out of Gemini training and grounding; Google Search untouched
User-agent: Google-Extended
Disallow: /

User-agent: Googlebot
Allow: /

PerplexityBot vs Perplexity-User

Perplexity's documentation says PerplexityBot surfaces and links websites in its search results and is not used to crawl content for AI foundation models; allow it if you want to appear. Perplexity-User fetches pages for user requests and "generally ignores robots.txt rules." Perplexity documents no separate training crawler, so there's no split to make here: allow PerplexityBot or leave Perplexity.

The two other names worth knowing

CCBot builds Common Crawl's open web archive and respects robots.txt. That archive is a common source of language-model training data: filtered Common Crawl text was the largest part of GPT-3's training mix, according to the GPT-3 paper. It feeds no answer engine directly, so blocking it is purely a training decision.

Bytespider and meta-externalagent are training crawlers from ByteDance and Meta. See their pages, Bytespider and meta-externalagent, for what each operator says and what independent reports say about compliance.

Template 1: allow AI search, block training scrapers

This file is generated from the same registry Crawlable's scanner uses. It lets every answer-engine crawler in and opts out of every training crawler:

# Answer engines: allowed, so you can be found and cited
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Amazonbot
User-agent: MistralAI-User
Allow: /
Disallow: /admin/
Disallow: /api/

# Model training: opted out. No effect on Google Search, AI Overviews,
# or ChatGPT, Claude and Perplexity search results.
# Google-Extended also controls grounding in Gemini Apps and Vertex AI.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Bytespider
User-agent: meta-externalagent
User-agent: cohere-ai
Disallow: /

# Everyone else, including Googlebot and Bingbot
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /api/

Sitemap: https://example.com/sitemap.xml

Adapt before you ship it:

  • Change the Disallow paths to your own private areas, and repeat them in every group (the group rule below explains why).
  • Change the Sitemap line to your own absolute sitemap URL.
  • Add User-agent: CCBot to the training group if you also want out of Common Crawl's archive.

Template 2: robots.txt allow all

The literal allow-all is two lines:

User-agent: *
Allow: /

An empty Disallow: line means the same thing. Both work, and both name no AI crawler, which has two costs. You can't opt one bot out of training later without restructuring the file, and you can't tell from the file whether "allowed" was a decision or a default. Crawlable's crawler-governance check scores a wildcard-only file as having no AI directives for exactly that reason.

The explicit version allows the same crawlers but names them, in two groups, so flipping training to Disallow: / later is a one-line change:

# Answer engines: allowed, so you can be found and cited
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Amazonbot
User-agent: MistralAI-User
Allow: /
Disallow: /admin/
Disallow: /api/

# Model training: allowed. Change Allow to Disallow: / here to opt out later.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Bytespider
User-agent: meta-externalagent
User-agent: cohere-ai
Allow: /
Disallow: /admin/
Disallow: /api/

# Everyone else, including Googlebot and Bingbot
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /api/

Sitemap: https://example.com/sitemap.xml

The group rule that catches everyone

Under the robots.txt standard, RFC 9309, a crawler follows the most specific group that names it and ignores the * group entirely. If you write User-agent: GPTBot with only Allow: /, GPTBot is now allowed into /admin/, even though your * group disallows it. Every named group needs its own copy of your private-path rules, which is why both templates repeat them.

The same standard lets several User-agent lines share one group, which is how the templates stay short. Don't put a search bot and a training bot in the same group if you ever want to treat them differently.

Why Disallow: / for every AI bot is a visibility mistake

Lists of "AI bots to block" circulate as copy-paste snippets, and most of them include OAI-SearchBot, Claude-SearchBot and PerplexityBot. Paste one in and you've removed yourself from ChatGPT search, Claude's search results and Perplexity, the products where a citation sends a click, while doing nothing about the -User fetchers that may not honour the file anyway.

If what you object to is training, block the training tokens. Leave the index and user fetchers alone.

Without compromising security

robots.txt is a set of instructions for well-behaved crawlers. It has no enforcement behind it.

  1. Never use it to hide anything. The file is public. Disallow: /internal-reports-2026/ advertises exactly the path you didn't want found. Protect private areas with authentication, and keep pages out of indexes with a noindex meta tag or header, which only works if the crawler can fetch the page and see it.
  2. Don't trust a user-agent string. Any scraper can call itself GPTBot. Google warns that its own user agent "is often spoofed" and documents verification by reverse DNS or published IP ranges. If you let a bot past a rate limit because of its name, check its IP against the operator's published list too.
  3. Enforce at the edge, not in the file. If a bot ignores your rules or hammers your server, block or rate-limit it in your CDN or firewall.
  4. Check your firewall isn't doing the opposite. One-click "block AI bots" or bot-fight settings in CDN and WAF products can block OAI-SearchBot and PerplexityBot even when robots.txt allows them. The file says yes, the network says no.

Common mistakes

  • User-agent: * with Disallow: / in production. Usually a staging config that shipped. It removes you from Google and every answer engine at once.
  • Expecting Google-Extended to hide you from AI Overviews. It doesn't; AI Overviews are Google Search.
  • Relying on outdated tokens. Crawlers get renamed and split. Re-check the operators' documentation, not a two-year-old snippet.
  • Expecting changes to be instant. OpenAI quotes about 24 hours. Other operators re-read the file on their own schedule.

Test it

Fetch your live file with a crawler's token, and check it returns 200 with text/plain rather than a bot challenge or a redirect. A WAF rule keyed on the user agent shows up here as a 403:

curl -sS -o /dev/null -w "%{http_code} %{content_type}\n" \
  -A "OAI-SearchBot" \
  https://yoursite.com/robots.txt

curl -sS https://yoursite.com/robots.txt

Then read each group against the cheat sheet above. Crawlable's free scan does this for you: it parses your robots.txt, evaluates it against every crawler in the table, and flags a search bot caught by a rule you inherited from a plugin or CMS template.

Frequently asked questions

What does an allow-all robots.txt look like?
The minimal allow-all file is two lines: User-agent: * followed by Allow: / (an empty Disallow: line means the same thing). Every crawler that honours robots.txt may fetch every URL. It names no AI crawler, though, so you cannot later opt one out of training without restructuring the file.
Which AI crawlers should I allow in robots.txt?
Allow the crawlers that put you in answers: OAI-SearchBot and ChatGPT-User (ChatGPT), Claude-SearchBot and Claude-User (Claude), PerplexityBot and Perplexity-User (Perplexity), Amazonbot and MistralAI-User. Whether to allow training crawlers such as GPTBot, ClaudeBot, Google-Extended or CCBot is a separate choice that does not change whether you are cited today.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot collects content that may be used to train OpenAI's foundation models. OAI-SearchBot builds the index behind ChatGPT search. Blocking GPTBot opts you out of training only; blocking OAI-SearchBot keeps you out of ChatGPT search answers, though OpenAI says you can still appear as a navigational link.
Does blocking GPTBot remove my site from ChatGPT?
No. ChatGPT search draws on OAI-SearchBot's index, and ChatGPT-User fetches pages users ask about. GPTBot is the training crawler. You can disallow GPTBot and keep OAI-SearchBot allowed, which is what the first template in this guide does.
Does blocking Google-Extended hurt my Google rankings?
No. Google states that Google-Extended does not affect a site's inclusion in Google Search and is not a ranking signal. It controls whether content Google crawls is used for Gemini training and for grounding in Gemini Apps and the Vertex AI API. AI Overviews follow Googlebot, not Google-Extended.
Can robots.txt stop ChatGPT-User or Perplexity-User?
Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User because a person initiated the fetch, and Perplexity says Perplexity-User generally ignores robots.txt. If you must stop them, block their published IP ranges at your CDN or firewall.