Every AI crawler, and what blocking it costs
Not all AI crawlers do the same job. Some collect content to train models. Others build the index behind an answer engine, or fetch a page because a user asked about it right now. Blocking the first kind costs you nothing in visibility. Blocking the second kind removes you from answers people actually read.
Block these and you disappear
8 retrieval and user-action crawlers. Each one feeds a surface a person is looking at.
OAI-SearchBot
OpenAI · raw HTML onlyBuilds the index behind ChatGPT search results. Blocking it removes you from ChatGPT search.
Read more about OAI-SearchBotChatGPT-User
OpenAI · raw HTML onlyFetches a page when a ChatGPT user follows or asks about a specific link.
Read more about ChatGPT-UserClaude-SearchBot
Anthropic · raw HTML onlyIndexes pages to support Claude search results.
Read more about Claude-SearchBotClaude-User
Anthropic · raw HTML onlyFetches a page on behalf of a Claude user following a link.
Read more about Claude-UserPerplexityBot
Perplexity · raw HTML onlyBuilds the Perplexity index. Blocking it removes you from Perplexity citations.
Read more about PerplexityBotPerplexity-User
Perplexity · raw HTML onlyFetches a page when a Perplexity user opens or asks about a specific link.
Read more about Perplexity-UserAmazonbot
Amazon · raw HTML onlyFeeds Alexa and Amazon answer surfaces.
Read more about AmazonbotMistralAI-User
Mistral · raw HTML onlyFetches a page on behalf of a Le Chat user.
Read more about MistralAI-User
Blocking these is a licensing choice
7 training crawlers. Disallowing them keeps your content out of model training and has no effect on whether you appear in AI answers.
GPTBot
OpenAI · raw HTML onlyCollects content for OpenAI model training. Blocking it does not remove you from ChatGPT answers.
Read more about GPTBotClaudeBot
Anthropic · raw HTML onlyCollects content for Anthropic model training.
Read more about ClaudeBotGoogle-Extended
Google · renders JavaScriptControls use of your content for Gemini training. Does not affect Google Search ranking.
Read more about Google-ExtendedApplebot-Extended
Apple · renders JavaScriptControls use of your content for Apple foundation model training.
Read more about Applebot-ExtendedBytespider
ByteDance · raw HTML onlyCollects content for ByteDance model training. Widely blocked for aggressive crawl rates.
Read more about Bytespidermeta-externalagent
Meta · raw HTML onlyCollects content for Meta AI model training.
Read more about Meta External Agentcohere-ai
Cohere · raw HTML onlyCollects content for Cohere model training.
Read more about Cohere
Which of these can reach your site right now?
The free scan parses your robots.txt and evaluates every crawler above against it — including the rules you inherited from a plugin, a CDN setting or a template you forgot about.
Check my robots.txt