Which AI bots does this robots.txt disallow?
Enter a site. The checker reads its robots.txt and reports what the file says to each of the 18 AI bots in the directory, and whether the site has an llms.txt.
How the result is worked out
The checker reads the file the way RFC 9309 describes. For each bot it looks for a group that names the bot's robots.txt token, without regard to upper or lower case. When no group names the bot, the * group applies. When there is neither, no rule applies and the bot is allowed.
It then tests one path, the home page /. The longest rule that matches decides, and an allow rule wins a tie. Disallowed means the home page is disallowed, which is nearly always a Disallow: / rule for the whole site. Partly disallowed means the home page is allowed and at least one other path is not.
Some operators have a bot use another token's group when its own is absent. The checker does not model that. It reads the bot's own group, then the * group.
What the checker requests from the site
Two files, /robots.txt and /llms.txt, over https. The requests carry this user agent:
CentinelRobotsCheck/1.0 (+https://www.centinelanalytica.com/robots-txt-checker)
The checker waits 5 seconds, reads the first 500 KiB of a file and follows up to 3 redirects, each to the same host with or without www. It can reuse an answer for up to 5 minutes, so repeating a check does not always send the site a new request. This page shows the verdicts and not the text of either file.
What a robots.txt rule does not do
A robots.txt rule states what the site owner asks for. Whether a crawler follows it is the crawler's choice. Each page in the bot directory records what the operator says its bot does with a disallow rule.
A rule also does nothing about a client that only uses a crawler's name. The crawler IP verifier checks whether a request that says it is GPTBot or Googlebot came from an address its operator publishes.
See which bots reach your site, and which of them are who they claim to be.
Request a site audit