Skip to content
Blog

Pay per crawl explained: HTTP 402 and TollBit

Pay per crawl charges AI crawlers per request. How Cloudflare, TollBit, RSL and x402 work, and why it depends on knowing which crawler is asking.

Frederick Jahn
Published
Pay per crawl explained: HTTP 402 and TollBit

Pay per crawl is a model in which a website charges an AI crawler for each page it fetches, instead of only allowing or blocking it. A crawler without a payment arrangement gets an HTTP 402 Payment Required response or a paywall; a crawler that has agreed to pay gets the page, and a marketplace settles the charges. Every version assumes the site already knows which crawler is asking.

This article describes how the main systems work and where they are still weak. Centinel does not sell, meter, or broker crawler access. Its part is identifying the crawler, covered at the end.

HTTP 402 was reserved for decades

HTTP/1.1 listed 402 in RFC 2068 in 1997 as "reserved for future use," and it was never given a meaning after that. The current HTTP standard, RFC 9110, still says only: "The 402 (Payment Required) status code is reserved for future use."

That leaves each pay-per-crawl system to define its own headers around the code. The status code tells a crawler that payment is possible. It does not say how much, in what currency, or how to pay. Those details are different in each system below.

How Cloudflare Pay per crawl works

Cloudflare announced pay per crawl in July 2025. Its documentation describes it as a feature of AI Crawl Control that is in closed beta, with Cloudflare acting as merchant of record. A site owner sets one price per zone and chooses, crawler by crawler, whether to allow, charge, or block.

The exchange uses custom headers. When a crawler requests a page it has not agreed to pay for, Cloudflare returns the price:

HTTP/2 402
crawler-price: USD 0.01

The crawler can retry with crawler-exact-price to accept that price, or send crawler-max-price up front with the most it will pay. If the price fits, the page is served with a crawler-charged header, and Cloudflare records a billing event. According to the announcement, Cloudflare aggregates those events, charges the crawler, and pays the publisher.

Only a verified crawler can be charged. Cloudflare's crawler setup guide requires Web Bot Auth signatures and registration as a verified bot. Cloudflare's blog calls preventing anyone from spoofing a paying crawler "an incredibly important technical challenge."

Security rules come first: a block from Cloudflare's WAF or Bot Management overrides the charge setting. The crawler selection guide also warns that charging or blocking search engine crawlers can hurt indexing.

How TollBit works

TollBit uses a separate gateway instead of a status code on the main site. Its architecture documentation describes the setup: the publisher creates a tollbit subdomain, and edge logic at the CDN "checks each request against a simple list of known bot user agents." Matching requests are redirected to the subdomain.

There, TollBit checks for a valid token. Its platform documentation describes the token as a short-lived signed JWT that the AI company generates with its TollBit API key. A bot without a token sees a paywall page telling it how to get one. A bot with a token gets the content, fetched from the publisher, at the rates the publisher sets. TollBit's documentation describes the redirect being set up in a CDN or in a bot-management product.

Other approaches

RSL (Really Simple Licensing) is a machine-readable license format, defined in the RSL 1.0 specification. A site links to its license file from robots.txt with a License directive, from HTTP headers, or from the page. The license can state terms such as pay per crawl or pay per inference, and a server may answer with 401 or 402 when access depends on a license. The specification also defines a protocol for obtaining a license, and lets a license name a payment protocol such as x402.

The x402 protocol is a general scheme for machine payments over HTTP. The server returns 402 with a PAYMENT-REQUIRED header that describes accepted payment schemes and the price, and the client pays and retries with a PAYMENT-SIGNATURE header. It is not specific to crawlers.

SystemHow it is enforcedHow the crawler is identifiedWho settles payment
Cloudflare Pay per crawlAt Cloudflare's edge, on the site's own URLsWeb Bot Auth signature and verified-bot registrationCloudflare, as merchant of record
TollBitRedirect to a tollbit subdomainUser-agent match for the redirect, TollBit token for accessTollBit
RSLWherever the site enforces its licenseLeft to the site or an RSL license serverLicense server or the named payment protocol
x402In the server or its middlewarePayment payload, not crawler identityThe payment scheme named in the response

The open problems

A price only applies to crawlers that ask

Every system above works on traffic that identifies itself. A scraper that sends a browser user agent from a residential address never matches a crawler list, so it never meets the paywall. Only bot detection can catch it.

User agents are easy to copy

A redirect or rule keyed on the user-agent string treats every request that uses the name the same way. That is a real problem at the edge: on the news and media sites in our fake crawler report, 1 in 7 requests calling themselves Googlebot on the typical site did not come from Google, and on the typical site 84% of requests using the GPTBot name were fake. A fake sent to a paywall does little harm. The same fake under a rule that allows crawlers by name gets the content for free.

Signed identity is new and not universal

Web Bot Auth builds on HTTP Message Signatures (RFC 9421). A crawler signs each request with a private key and publishes the public key in a directory. The IETF working group published the first version of its protocol draft on 1 September 2026, so it is not yet a finished standard. Crawlers that do not sign fall back to the older checks: the operator's published IP ranges or a reverse and forward DNS lookup.

A valid signature proves that the holder of a known key signed the request. It does not tell the site how the crawler will use the content. The verifier also has to enforce the signature's creation and expiry times to stop replays, and has to decide which key directories it trusts.

Search crawlers are a separate decision

Charging or blocking a search crawler can remove pages from search results, as Cloudflare's own documentation warns. Google also separates crawling from use: the Google-Extended robots.txt token controls whether content Google crawls may be used for Gemini training and grounding, but it has no user agent of its own. Crawling is done with Google's existing user agents, so a price on a user agent cannot separate those uses.

There is no common protocol yet

Cloudflare uses crawler-price headers, x402 uses PAYMENT-REQUIRED, TollBit uses a redirect and a token, and RSL points to a license file. A crawler has to implement each one it wants to pay through. Until 402 gets a standard meaning, a site that changes provider changes protocol too.

Where crawler verification fits

Allowing, blocking, rate-limiting, and charging all need the same input: an identity that has been checked. Centinel identifies automated traffic and verifies declared crawlers with each operator's documented method: a signed request where the operator supports it, published IP ranges, or reverse and forward DNS. It also detects automation that does not declare itself. Centinel does not set prices, process payments, or negotiate access.

To check one address yourself, use the crawler verification tool. The AI agent verification guide covers each operator's method, and How to control AI agent access covers setting different responses by identity. For the scraping side of the problem, see scraping and AI crawler protection.