Skip to content
AI crawler list

AI crawler list

AI companies send different bots for different jobs: training models, indexing pages for AI search, and fetching a page when a user asks. Each has its own robots.txt token, so you can allow one and refuse another. Facts come from each operator's own documentation, last verified 2026-09-30.

AI training and dataset crawlers

Crawl the web for content that may be used to train models or to build a reusable dataset.

BotOperatorrobots.txt tokenPurpose
GPTBotOpenAIGPTBotCrawls content that may be used to train OpenAI's foundation models.
ClaudeBotAnthropicClaudeBotCollects web content that could contribute to Claude model training.
Meta-ExternalAgentMetameta-externalagentCrawls for training AI models or improving Meta products.
AmazonbotAmazonAmazonbotCrawls to improve Amazon products; may be used to train Amazon AI models.
CCBotCommon CrawlCCBotBuilds Common Crawl's open web archive, free for anyone to reuse.

AI search crawlers

Index pages so an AI assistant can find, cite and link them in its answers.

BotOperatorrobots.txt tokenPurpose
OAI-SearchBotOpenAIOAI-SearchBotSurfaces websites in ChatGPT's search features.
Claude-SearchBotAnthropicClaude-SearchBotIndexes pages to improve Claude's search results.
PerplexityBotPerplexityPerplexityBotSurfaces and links websites in Perplexity's search results.
DuckAssistBotDuckDuckGoDuckAssistBotCrawls in real time to support DuckDuckGo's AI-assisted answers.

AI user-triggered fetchers

Fetch a page because a person asked an assistant to. Several operators say robots.txt may not apply.

BotOperatorrobots.txt tokenPurpose
ChatGPT-UserOpenAIChatGPT-UserVisits a page when a ChatGPT or Custom GPT user's request needs it.
Claude-UserAnthropicClaude-UserFetches a page when a Claude user's question needs it.
Perplexity-UserPerplexityPerplexity-UserFetches a page when a Perplexity user's request needs it.
Meta-ExternalFetcherMetameta-externalfetcherFetches individual links on behalf of Meta product users.
MistralAI-UserMistral AIMistralAI-UserFetches a page when a Mistral user's action needs it.

AI robots.txt control tokens

Names you write in robots.txt to control AI use. No crawler sends them as a user agent.

BotOperatorrobots.txt tokenPurpose
Google-ExtendedGoogleGoogle-ExtendedA robots.txt token that controls Gemini training and grounding use.
Applebot-ExtendedAppleApplebot-ExtendedA robots.txt token that opts content out of Apple model training.

Opt out of AI training, stay in AI search

This group refuses the training crawlers and control tokens above and leaves the AI search crawlers alone. Check each bot's page before you copy it: some names also feed products you may want.

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Amazonbot
User-agent: CCBot
Disallow: /

robots.txt only asks. Several user-triggered fetchers say they may not follow it, and anyone can send a crawler name. Verify a crawler or browse the full bot directory.