Skip to content
AI training and dataset crawlers

What is GPTBot?

GPTBot is OpenAI's web crawler for model training. OpenAI says it crawls content that may be used in training its generative AI foundation models, and that disallowing GPTBot indicates a site's content should not be used for that training.

Last verified 2026-09-30 against the operator's documentation. All bots

GPTBot at a glance

Operator
OpenAI
Type
AI training and dataset crawlers
Purpose
Crawls content that may be used to train OpenAI's foundation models.
robots.txt token
GPTBot
Example user agent
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
How to verify
Check that the source IP is in the ranges OpenAI publishes for GPTBot.

How often is GPTBot impersonated?

For the median organisation, 88% of requests claiming to be GPTBot failed verification.

Measured between 3 July 2026 and 28 September 2026 across organisations protected by Centinel. We only publish a share when at least five organisations each saw enough of these requests. Pooled across all of their requests, the share was 2%: a few large organisations carry most of the traffic, so the median organisation is the better guide to what yours will see.

This directory measures requests between 3 July 2026 and 28 September 2026. The fake crawler report covers its own, earlier window (late June to 21 September 2026), so its figures differ from the ones here. Read it for the method.

Have a request that claims to be GPTBot? Check its IP address.

Should you block GPTBot?

Allow it if
you are content for your pages to be used in OpenAI model training.
Block it if
you do not want your content used for training, while still allowing OAI-SearchBot so ChatGPT search can cite you.
What blocking changes
It tells OpenAI not to use your content for training. OpenAI documents GPTBot, OAI-SearchBot and ChatGPT-User as independent settings, so blocking GPTBot does not remove you from ChatGPT search.

To block it, add this group to your robots.txt. OpenAI uses the GPTBot robots.txt token to control training use.

User-agent: GPTBot
Disallow: /

robots.txt only asks. It does nothing against a client that ignores it or only pretends to be GPTBot. Stopping those takes verification at your edge.

Questions

GPTBot: common questions

See which bots reach your site, and which of them are who they claim to be.

What is GPTBot?

GPTBot is OpenAI's web crawler for model training. OpenAI says it crawls content that may be used in training its generative AI foundation models, and that disallowing GPTBot indicates a site's content should not be used for that training.

How do I verify that a request is really GPTBot?

Check that the source IP is in the ranges OpenAI publishes for GPTBot. The user agent alone proves nothing: any client can send it.

How do I block GPTBot in robots.txt?

Add a group for User-agent: GPTBot with Disallow: /. OpenAI uses the GPTBot robots.txt token to control training use. robots.txt only asks. A client that ignores it, or pretends to be GPTBot, has to be stopped at your edge.

What happens if I block GPTBot?

It tells OpenAI not to use your content for training. OpenAI documents GPTBot, OAI-SearchBot and ChatGPT-User as independent settings, so blocking GPTBot does not remove you from ChatGPT search.

Related bots

  • OAI-SearchBot (OpenAI): Surfaces websites in ChatGPT's search features.
  • ChatGPT-User (OpenAI): Visits a page when a ChatGPT or Custom GPT user's request needs it.
  • ClaudeBot (Anthropic): Collects web content that could contribute to Claude model training.
  • Meta-ExternalAgent (Meta): Crawls for training AI models or improving Meta products.
  • Amazonbot (Amazon): Crawls to improve Amazon products; may be used to train Amazon AI models.
  • CCBot (Common Crawl): Builds Common Crawl's open web archive, free for anyone to reuse.