Skip to content
Bot directory

Bot directory

The crawlers and fetchers that most often reach a public website: who runs each one, what it does, its user agent and robots.txt token, and how to check that a request really comes from it. Facts come from each operator's own documentation, last verified 2026-09-30.

Only need the AI names? See the AI crawler list. Have an IP to check? Verify a crawler.

AI training and dataset crawlers

Crawl the web for content that may be used to train models or to build a reusable dataset.

  • GPTBot (OpenAI): Crawls content that may be used to train OpenAI's foundation models.
  • ClaudeBot (Anthropic): Collects web content that could contribute to Claude model training.
  • Meta-ExternalAgent (Meta): Crawls for training AI models or improving Meta products.
  • Amazonbot (Amazon): Crawls to improve Amazon products; may be used to train Amazon AI models.
  • CCBot (Common Crawl): Builds Common Crawl's open web archive, free for anyone to reuse.

AI search crawlers

Index pages so an AI assistant can find, cite and link them in its answers.

  • OAI-SearchBot (OpenAI): Surfaces websites in ChatGPT's search features.
  • Claude-SearchBot (Anthropic): Indexes pages to improve Claude's search results.
  • PerplexityBot (Perplexity): Surfaces and links websites in Perplexity's search results.
  • DuckAssistBot (DuckDuckGo): Crawls in real time to support DuckDuckGo's AI-assisted answers.

AI user-triggered fetchers

Fetch a page because a person asked an assistant to. Several operators say robots.txt may not apply.

  • ChatGPT-User (OpenAI): Visits a page when a ChatGPT or Custom GPT user's request needs it.
  • Claude-User (Anthropic): Fetches a page when a Claude user's question needs it.
  • Perplexity-User (Perplexity): Fetches a page when a Perplexity user's request needs it.
  • Meta-ExternalFetcher (Meta): Fetches individual links on behalf of Meta product users.
  • MistralAI-User (Mistral AI): Fetches a page when a Mistral user's action needs it.

AI robots.txt control tokens

Names you write in robots.txt to control AI use. No crawler sends them as a user agent.

  • Google-Extended (Google): A robots.txt token that controls Gemini training and grounding use.
  • Applebot-Extended (Apple): A robots.txt token that opts content out of Apple model training.

Search engine crawlers

Build the indexes behind web search results.

  • Googlebot (Google): Crawls pages for Google Search, Images, Video, News and Discover.
  • Googlebot-Image (Google): Crawls images for Google Images and image features in Search.
  • Bingbot (Microsoft): Crawls pages for Bing search.
  • Applebot (Apple): Crawls pages for Spotlight, Siri and Safari search features.
  • DuckDuckBot (DuckDuckGo): Crawls pages for DuckDuckGo search.
  • YandexBot (Yandex): Yandex's main search indexing crawler.
  • Baiduspider (Baidu): Crawls pages for Baidu search.
  • PetalBot (Huawei (Petal Search)): Crawls pages for Petal Search and Huawei Assistant.
  • SeznamBot (Seznam.cz): Crawls pages for Seznam, a Czech search engine.
  • Yeti (Naver): Crawls pages for Naver search in South Korea.

Other Google crawlers

Single-purpose Google crawlers for shopping, ads, testing tools and Vertex AI.

  • Storebot-Google (Google): Crawls product and store pages for Google Shopping.
  • AdsBot-Google (Google): Checks the quality of ad landing pages for Google Ads.
  • Google-InspectionTool (Google): Fetches pages for Search Console URL inspection and the Rich Results Test.
  • Google-CloudVertexBot (Google): Crawls sites whose owners asked for it when building Vertex AI Agents.

SEO and link-index crawlers

Build backlink and site-audit databases for SEO tools. None of them affects search rankings.

  • AhrefsBot (Ahrefs): Builds the link database behind Ahrefs and the Yep search engine.
  • SemrushBot (Semrush): Collects web data for Semrush's SEO tools.
  • MJ12bot (Majestic): Maps links between websites for Majestic's link index.
  • DotBot (Moz): Gathers data for the Moz Link Index.
  • DataForSeoBot (DataForSEO): Adds links to DataForSEO's backlink database.
  • Barkrowler (Babbar): Collects links and page metadata for Babbar.

Link preview bots

Fetch a page when someone shares its link, to build the preview card.

  • facebookexternalhit (Meta): Builds link previews when content is shared on Meta's apps.
  • Twitterbot (X): Fetches card markup when a link is shared on X.
  • Pinterestbot (Pinterest): Indexes pages behind Pins and keeps Pin details up to date.

Crawlers without operator documentation

The operator publishes no page describing the crawler, so its identity cannot be checked.

  • Bytespider (Unconfirmed (attributed to ByteDance)): A crawler with no operator documentation.