Pay per crawl is a model in which a website charges an AI crawler for each page it fetches,
instead of only allowing or blocking it. A crawler without a payment arrangement gets an
HTTP 402 Payment Required response or a paywall; a crawler that has agreed to pay gets the
page, and a marketplace settles the charges. Every version assumes the site already knows which
crawler is asking.
This article describes how the main systems work and where they are still weak. Centinel does not sell, meter, or broker crawler access. Its part is identifying the crawler, covered at the end.
HTTP 402 was reserved for decades
HTTP/1.1 listed 402 in RFC 2068 in 1997 as "reserved for future use," and it was never given a meaning after that. The current HTTP standard, RFC 9110, still says only: "The 402 (Payment Required) status code is reserved for future use."
That leaves each pay-per-crawl system to define its own headers around the code. The status code tells a crawler that payment is possible. It does not say how much, in what currency, or how to pay. Those details are different in each system below.
How Cloudflare Pay per crawl works
Cloudflare announced pay per crawl in July 2025. Its documentation describes it as a feature of AI Crawl Control that is in closed beta, with Cloudflare acting as merchant of record. A site owner sets one price per zone and chooses, crawler by crawler, whether to allow, charge, or block.
The exchange uses custom headers. When a crawler requests a page it has not agreed to pay for, Cloudflare returns the price:
HTTP/2 402
crawler-price: USD 0.01
The crawler can retry with crawler-exact-price to accept that price, or send
crawler-max-price up front with the most it will pay. If the price fits, the page is served
with a crawler-charged header, and Cloudflare records a billing event. According to the
announcement, Cloudflare aggregates those events, charges the crawler, and pays the publisher.
Only a verified crawler can be charged. Cloudflare's crawler setup guide requires Web Bot Auth signatures and registration as a verified bot. Cloudflare's blog calls preventing anyone from spoofing a paying crawler "an incredibly important technical challenge."
Security rules come first: a block from Cloudflare's WAF or Bot Management overrides the charge setting. The crawler selection guide also warns that charging or blocking search engine crawlers can hurt indexing.
How TollBit works
TollBit uses a separate gateway instead of a status code on the main site. Its
architecture documentation describes the
setup: the publisher creates a tollbit subdomain, and edge logic at the CDN "checks each
request against a simple list of known bot user agents." Matching requests are redirected to
the subdomain.
There, TollBit checks for a valid token. Its platform documentation describes the token as a short-lived signed JWT that the AI company generates with its TollBit API key. A bot without a token sees a paywall page telling it how to get one. A bot with a token gets the content, fetched from the publisher, at the rates the publisher sets. TollBit's documentation describes the redirect being set up in a CDN or in a bot-management product.
Other approaches
RSL (Really Simple Licensing) is a machine-readable license format, defined in the
RSL 1.0 specification. A site links to its license file from robots.txt with a
License directive, from HTTP headers, or from the page. The license can state terms such as
pay per crawl or pay per inference, and a server may answer with 401 or 402 when access
depends on a license. The specification also defines a protocol for obtaining a license, and
lets a license name a payment protocol such as x402.
The x402 protocol is a general scheme for machine payments over HTTP. The server returns 402 with a PAYMENT-REQUIRED
header that describes accepted payment schemes and the price, and the client pays and retries
with a PAYMENT-SIGNATURE header. It is not specific to crawlers.
| System | How it is enforced | How the crawler is identified | Who settles payment |
|---|---|---|---|
| Cloudflare Pay per crawl | At Cloudflare's edge, on the site's own URLs | Web Bot Auth signature and verified-bot registration | Cloudflare, as merchant of record |
| TollBit | Redirect to a tollbit subdomain | User-agent match for the redirect, TollBit token for access | TollBit |
| RSL | Wherever the site enforces its license | Left to the site or an RSL license server | License server or the named payment protocol |
| x402 | In the server or its middleware | Payment payload, not crawler identity | The payment scheme named in the response |
The open problems
A price only applies to crawlers that ask
Every system above works on traffic that identifies itself. A scraper that sends a browser user agent from a residential address never matches a crawler list, so it never meets the paywall. Only bot detection can catch it.
User agents are easy to copy
A redirect or rule keyed on the user-agent string treats every request that uses the name the same way. That is a real problem at the edge: on the news and media sites in our fake crawler report, 1 in 7 requests calling themselves Googlebot on the typical site did not come from Google, and on the typical site 84% of requests using the GPTBot name were fake. A fake sent to a paywall does little harm. The same fake under a rule that allows crawlers by name gets the content for free.
Signed identity is new and not universal
Web Bot Auth builds on HTTP Message Signatures (RFC 9421). A crawler signs each request with a private key and publishes the public key in a directory. The IETF working group published the first version of its protocol draft on 1 September 2026, so it is not yet a finished standard. Crawlers that do not sign fall back to the older checks: the operator's published IP ranges or a reverse and forward DNS lookup.
A valid signature proves that the holder of a known key signed the request. It does not tell the site how the crawler will use the content. The verifier also has to enforce the signature's creation and expiry times to stop replays, and has to decide which key directories it trusts.
Search crawlers are a separate decision
Charging or blocking a search crawler can remove pages from search results, as Cloudflare's own documentation warns. Google also separates crawling from use: the Google-Extended robots.txt token controls whether content Google crawls may be used for Gemini training and grounding, but it has no user agent of its own. Crawling is done with Google's existing user agents, so a price on a user agent cannot separate those uses.
There is no common protocol yet
Cloudflare uses crawler-price headers, x402 uses PAYMENT-REQUIRED, TollBit uses a redirect
and a token, and RSL points to a license file. A crawler has to implement each one it wants to
pay through. Until 402 gets a standard meaning, a site that changes provider changes protocol
too.
Where crawler verification fits
Allowing, blocking, rate-limiting, and charging all need the same input: an identity that has been checked. Centinel identifies automated traffic and verifies declared crawlers with each operator's documented method: a signed request where the operator supports it, published IP ranges, or reverse and forward DNS. It also detects automation that does not declare itself. Centinel does not set prices, process payments, or negotiate access.
To check one address yourself, use the crawler verification tool. The AI agent verification guide covers each operator's method, and How to control AI agent access covers setting different responses by identity. For the scraping side of the problem, see scraping and AI crawler protection.
