Skip to content
AI training and dataset crawlers

What is Meta-ExternalAgent?

Meta-ExternalAgent is a Meta crawler. Meta says it crawls the web for use cases such as training AI models or improving products by indexing content directly.

Last verified 2026-09-30 against the operator's documentation. All bots

Meta-ExternalAgent at a glance

Operator
Meta
Type
AI training and dataset crawlers
Purpose
Crawls for training AI models or improving Meta products.
robots.txt token
meta-externalagent
Example user agent
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)
How to verify
Meta says a crawler request comes from Meta if its source IP is on the list this command returns: whois -h whois.radb.net -- '-i origin AS32934' | grep ^route. That list covers Meta's whole network (AS32934), not this crawler alone, and Meta notes the addresses change often.

How often is Meta-ExternalAgent impersonated?

Meta documents one check for its crawlers: the source IP must be in its network, AS32934. Our measurement does not yet apply that check, so we show no share for Meta-ExternalAgent.

This directory measures requests between 3 July 2026 and 28 September 2026. The fake crawler report covers its own, earlier window (late June to 21 September 2026), so its figures differ from the ones here. Read it for the method.

Have a request that claims to be Meta-ExternalAgent? Check its IP address.

Should you block Meta-ExternalAgent?

Allow it if
you are content for Meta to use your pages this way.
Block it if
you do not want your content used for Meta's AI training.
What blocking changes
Meta's crawler stops collecting your pages for AI training and product indexing. Link previews use a different crawler.

To block it, add this group to your robots.txt. Meta documents robots.txt control for this crawler.

User-agent: meta-externalagent
Disallow: /

robots.txt only asks. It does nothing against a client that ignores it or only pretends to be Meta-ExternalAgent. Stopping those takes verification at your edge.

Questions

Meta-ExternalAgent: common questions

See which bots reach your site, and which of them are who they claim to be.

What is Meta-ExternalAgent?

Meta-ExternalAgent is a Meta crawler. Meta says it crawls the web for use cases such as training AI models or improving products by indexing content directly.

How do I verify that a request is really Meta-ExternalAgent?

Meta says a crawler request comes from Meta if its source IP is on the list this command returns: whois -h whois.radb.net -- '-i origin AS32934' | grep ^route. That list covers Meta's whole network (AS32934), not this crawler alone, and Meta notes the addresses change often. The user agent alone proves nothing: any client can send it.

How do I block Meta-ExternalAgent in robots.txt?

Add a group for User-agent: meta-externalagent with Disallow: /. Meta documents robots.txt control for this crawler. robots.txt only asks. A client that ignores it, or pretends to be Meta-ExternalAgent, has to be stopped at your edge.

What happens if I block Meta-ExternalAgent?

Meta's crawler stops collecting your pages for AI training and product indexing. Link previews use a different crawler.

Related bots

  • Meta-ExternalFetcher (Meta): Fetches individual links on behalf of Meta product users.
  • facebookexternalhit (Meta): Builds link previews when content is shared on Meta's apps.
  • GPTBot (OpenAI): Crawls content that may be used to train OpenAI's foundation models.
  • ClaudeBot (Anthropic): Collects web content that could contribute to Claude model training.
  • Amazonbot (Amazon): Crawls to improve Amazon products; may be used to train Amazon AI models.
  • CCBot (Common Crawl): Builds Common Crawl's open web archive, free for anyone to reuse.