What is Diffbot?
Diffbot is the crawler of the company with the same name. The company says it crawls the web to build a general search engine, so that websites can be discovered and cited from its Knowledge Graph and web search services, and that the crawl is not used for AI training.
- Operator
- Diffbot
- Purpose
- Crawls the web for the Diffbot Knowledge Graph and Diffbot's web search services.
- robots.txt token
Diffbot- How to verify
- Diffbot publishes no IP list or reverse DNS domain.
- Operator documentation
- https://www.diffbot.com/docs/crawl/faq/robots-txt
- Last verified
How often is Diffbot impersonated?
We have not measured requests claiming to be Diffbot yet, so we show no share.
This directory measures requests between 3 July 2026 and 28 September 2026. The fake crawler report covers its own, earlier window (late June to 21 September 2026), so its figures differ from the ones here. Read it for the method.
Have a request that claims to be Diffbot? Check its IP address.
Should you block Diffbot?
- Allow it if
- you want your pages to be found and cited through the Diffbot Knowledge Graph.
- Block it if
- you do not want your pages in Diffbot's data products. An IP block alone may not hold, because Diffbot uses proxies for its Extract API.
- What blocking changes
- Diffbot does not state the effect of a disallow rule beyond adhering to it. For its Extract API, Diffbot says it applies a proxy pool when a site blocks its default IP addresses.
To block it, add this group to your robots.txt. Diffbot says its web crawls adhere to robots.txt by default, including the disallow and crawl-delay directives. It says the instruction can be overridden in specific cases, typically a partnership or agreement with the site.
User-agent: Diffbot
Disallow: /robots.txt only asks. It does nothing against a client that ignores it or only pretends to be Diffbot. Stopping those takes verification at your edge.
Related bots
- 360Spider (360 Search): Crawls pages for the web search of 360 Search (so.com).
- Applebot (Apple): Crawls pages for Spotlight, Siri and Safari search features.
- Baiduspider (Baidu): Crawls pages for Baidu search.
- Bingbot (Microsoft): Crawls pages for Bing search.
- DuckDuckBot (DuckDuckGo): Crawls pages for DuckDuckGo search.
- Googlebot-Image (Google): Crawls images for Google Images and image features in Search.
See which bots reach your site, and which of them are who they claim to be.
Request a site audit