Skip to content

What is Diffbot?

Diffbot is the crawler of the company with the same name. The company says it crawls the web to build a general search engine, so that websites can be discovered and cited from its Knowledge Graph and web search services, and that the crawl is not used for AI training.

Operator
Diffbot
Purpose
Crawls the web for the Diffbot Knowledge Graph and Diffbot's web search services.
robots.txt token
Diffbot
How to verify
Diffbot publishes no IP list or reverse DNS domain.
Last verified

How often is Diffbot impersonated?

We have not measured requests claiming to be Diffbot yet, so we show no share.

This directory measures requests between 3 July 2026 and 28 September 2026. The fake crawler report covers its own, earlier window (late June to 21 September 2026), so its figures differ from the ones here. Read it for the method.

Have a request that claims to be Diffbot? Check its IP address.

Should you block Diffbot?

Allow it if
you want your pages to be found and cited through the Diffbot Knowledge Graph.
Block it if
you do not want your pages in Diffbot's data products. An IP block alone may not hold, because Diffbot uses proxies for its Extract API.
What blocking changes
Diffbot does not state the effect of a disallow rule beyond adhering to it. For its Extract API, Diffbot says it applies a proxy pool when a site blocks its default IP addresses.

To block it, add this group to your robots.txt. Diffbot says its web crawls adhere to robots.txt by default, including the disallow and crawl-delay directives. It says the instruction can be overridden in specific cases, typically a partnership or agreement with the site.

User-agent: Diffbot
Disallow: /

robots.txt only asks. It does nothing against a client that ignores it or only pretends to be Diffbot. Stopping those takes verification at your edge.

Related bots

  • 360Spider (360 Search): Crawls pages for the web search of 360 Search (so.com).
  • Applebot (Apple): Crawls pages for Spotlight, Siri and Safari search features.
  • Baiduspider (Baidu): Crawls pages for Baidu search.
  • Bingbot (Microsoft): Crawls pages for Bing search.
  • DuckDuckBot (DuckDuckGo): Crawls pages for DuckDuckGo search.
  • Googlebot-Image (Google): Crawls images for Google Images and image features in Search.

See which bots reach your site, and which of them are who they claim to be.

Request a site audit