AI crawler list
AI companies send different bots for different jobs: training models, indexing pages for AI search, and fetching a page when a user asks. Each has its own robots.txt token, so you can allow one and refuse another. Facts come from each operator's own documentation, last verified 2026-09-30.
AI training and dataset crawlers
Crawl the web for content that may be used to train models or to build a reusable dataset.
| Bot | Operator | robots.txt token | Purpose |
|---|---|---|---|
| GPTBot | OpenAI | GPTBot | Crawls content that may be used to train OpenAI's foundation models. |
| ClaudeBot | Anthropic | ClaudeBot | Collects web content that could contribute to Claude model training. |
| Meta-ExternalAgent | Meta | meta-externalagent | Crawls for training AI models or improving Meta products. |
| Amazonbot | Amazon | Amazonbot | Crawls to improve Amazon products; may be used to train Amazon AI models. |
| CCBot | Common Crawl | CCBot | Builds Common Crawl's open web archive, free for anyone to reuse. |
AI search crawlers
Index pages so an AI assistant can find, cite and link them in its answers.
| Bot | Operator | robots.txt token | Purpose |
|---|---|---|---|
| OAI-SearchBot | OpenAI | OAI-SearchBot | Surfaces websites in ChatGPT's search features. |
| Claude-SearchBot | Anthropic | Claude-SearchBot | Indexes pages to improve Claude's search results. |
| PerplexityBot | Perplexity | PerplexityBot | Surfaces and links websites in Perplexity's search results. |
| DuckAssistBot | DuckDuckGo | DuckAssistBot | Crawls in real time to support DuckDuckGo's AI-assisted answers. |
AI user-triggered fetchers
Fetch a page because a person asked an assistant to. Several operators say robots.txt may not apply.
| Bot | Operator | robots.txt token | Purpose |
|---|---|---|---|
| ChatGPT-User | OpenAI | ChatGPT-User | Visits a page when a ChatGPT or Custom GPT user's request needs it. |
| Claude-User | Anthropic | Claude-User | Fetches a page when a Claude user's question needs it. |
| Perplexity-User | Perplexity | Perplexity-User | Fetches a page when a Perplexity user's request needs it. |
| Meta-ExternalFetcher | Meta | meta-externalfetcher | Fetches individual links on behalf of Meta product users. |
| MistralAI-User | Mistral AI | MistralAI-User | Fetches a page when a Mistral user's action needs it. |
AI robots.txt control tokens
Names you write in robots.txt to control AI use. No crawler sends them as a user agent.
| Bot | Operator | robots.txt token | Purpose |
|---|---|---|---|
| Google-Extended | Google-Extended | A robots.txt token that controls Gemini training and grounding use. | |
| Applebot-Extended | Apple | Applebot-Extended | A robots.txt token that opts content out of Apple model training. |
Opt out of AI training, stay in AI search
This group refuses the training crawlers and control tokens above and leaves the AI search crawlers alone. Check each bot's page before you copy it: some names also feed products you may want.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Amazonbot
User-agent: CCBot
Disallow: /robots.txt only asks. Several user-triggered fetchers say they may not follow it, and anyone can send a crawler name. Verify a crawler or browse the full bot directory.