Meta crawlers
This directory covers 4 bots that Meta runs. Each has its own robots.txt token, so a site can allow one and disallow another. Last checked against Meta's documentation on .
| Bot | Type | What it does | robots.txt token |
|---|---|---|---|
| facebookexternalhit | Link preview bots | Builds link previews when content is shared on Meta's apps. | facebookexternalhit |
| Meta-ExternalAgent | AI training and dataset crawlers | Crawls for training AI models or improving Meta products. | meta-externalagent |
| Meta-ExternalFetcher | AI user-triggered fetchers | Fetches individual links on behalf of Meta product users. | meta-externalfetcher |
| Meta-WebIndexer | AI search crawlers | Crawls pages to improve the search results of Meta AI. | meta-webindexer |
How to verify a request from Meta
A user agent is a string any client can send, so the name alone proves nothing. Check an address from your logs, or apply the operator's own method:
- facebookexternalhit
- Meta's crawler page gives no IP list and no reverse DNS domain. It says to add the user agent strings or the IP addresses used by the crawler to your allow list, and publishes no list of those addresses.
- Meta-ExternalAgent
- Meta's crawler page gives no IP list and no reverse DNS domain for Meta-ExternalAgent.
- Meta-ExternalFetcher
- Meta's crawler page gives no IP list and no reverse DNS domain for Meta-ExternalFetcher.
- Meta-WebIndexer
- Meta's crawler page gives no IP list and no reverse DNS domain for Meta-WebIndexer.