What is CCBot?
CCBot is the crawler of Common Crawl, which publishes its web crawl data as open data for anyone to use. Common Crawl itself warns that other crawlers falsely identify themselves as CCBot.
Last verified 2026-09-30 against the operator's documentation. All bots
CCBot at a glance
- Operator
- Common Crawl
- Type
- AI training and dataset crawlers
- Purpose
- Builds Common Crawl's open web archive, free for anyone to reuse.
- robots.txt token
CCBot- Example user agent
CCBot/2.0 (https://commoncrawl.org/faq/)- How to verify
- Common Crawl says CCBot runs on dedicated IP ranges with reverse DNS (except IPv6). Check the source IP against its published list.
- Operator documentation
- https://commoncrawl.org/ccbot
How often is CCBot impersonated?
For the median organisation, 87% of requests claiming to be CCBot failed verification.
Measured between 3 July 2026 and 28 September 2026 across organisations protected by Centinel. We only publish a share when at least five organisations each saw enough of these requests. Pooled across all of their requests, the share was 46%: a few large organisations carry most of the traffic, so the median organisation is the better guide to what yours will see.
This directory measures requests between 3 July 2026 and 28 September 2026. The fake crawler report covers its own, earlier window (late June to 21 September 2026), so its figures differ from the ones here. Read it for the method.
Have a request that claims to be CCBot? Check its IP address.
Should you block CCBot?
- Allow it if
- you are content for your pages to be in a public dataset others reuse.
- Block it if
- you do not want your content in an open dataset that third parties can use for any purpose, including model training.
- What blocking changes
- Your pages stop being added to future Common Crawl archives. Copies already in past archives stay there.
To block it, add this group to your robots.txt. Common Crawl says CCBot obeys robots.txt and Crawl-delay.
User-agent: CCBot
Disallow: /robots.txt only asks. It does nothing against a client that ignores it or only pretends to be CCBot. Stopping those takes verification at your edge.
CCBot: common questions
See which bots reach your site, and which of them are who they claim to be.
What is CCBot?
CCBot is the crawler of Common Crawl, which publishes its web crawl data as open data for anyone to use. Common Crawl itself warns that other crawlers falsely identify themselves as CCBot.
How do I verify that a request is really CCBot?
Common Crawl says CCBot runs on dedicated IP ranges with reverse DNS (except IPv6). Check the source IP against its published list. The user agent alone proves nothing: any client can send it.
How do I block CCBot in robots.txt?
Add a group for User-agent: CCBot with Disallow: /. Common Crawl says CCBot obeys robots.txt and Crawl-delay. robots.txt only asks. A client that ignores it, or pretends to be CCBot, has to be stopped at your edge.
What happens if I block CCBot?
Your pages stop being added to future Common Crawl archives. Copies already in past archives stay there.
Related bots
- GPTBot (OpenAI): Crawls content that may be used to train OpenAI's foundation models.
- ClaudeBot (Anthropic): Collects web content that could contribute to Claude model training.
- Meta-ExternalAgent (Meta): Crawls for training AI models or improving Meta products.
- Amazonbot (Amazon): Crawls to improve Amazon products; may be used to train Amazon AI models.