Decide which AI crawlers and scrapers can read your content
Search crawlers, AI crawlers, assistants fetching a page for a user, and scrapers that hide behind a browser all read the same articles. Centinel identifies which is which, so publishers can keep search indexing and give every other client the access they choose.
Request a site auditrobots.txt is a request, not a control
RFC 9309, the robots.txt standard, says its rules are not a form of access authorization. Declared crawlers usually honor them. Scrapers that ignore them, or copy a crawler's name, do not. Some agents that fetch a page for a user do not apply robots.txt at all; OpenAI says its rules may not apply to ChatGPT-User.
Read the researchWhat the traffic looks like
AI crawlers, declared and not
Training and search crawlers from AI companies usually identify themselves. Others arrive with a browser user agent, or with a well-known crawler name they have no right to.
Archives and paywalls copied
Scrapers walk article archives and subscriber pages to copy full text, sometimes through a logged-in account and a real browser.
Agents reading for a user
AI assistants and browser agents load pages when a person asks them to, which makes them a different client from a bulk crawler and a different policy decision.
What runs against this traffic
Verify declared crawlers
Centinel can compare a declared crawler identity with the verification data its operator publishes, such as IP ranges or DNS records.
A real Googlebot keeps indexing, and a verified AI crawler gets the policy you set for it. A copied crawler name receives no automatic trust.
Read moreA policy for each crawler and section
Supported integrations apply the configured response for the client's identity and the requested resource: allow, rate-limit, challenge, or block.
Search indexing can stay open while AI training crawlers are refused, or allowed only on some sections, without waiting for every bot to read robots.txt.
Read moreEvidence against disguised scraping
In supported browser deployments, Centinel evaluates request and browser evidence on article and archive pages, not only the user agent.
A scraper posing as a reader can be rate-limited or challenged without putting subscribers through the same check.
Read more
Want to see what reaches your site?
Start with a free scraping audit: tested tools, exposed pages, and fixes to consider, reviewed by hand and emailed to you. This checks scraping exposure, not every abuse pattern or all production traffic. Discuss broader workflow coverage in a demo.
Media & publishing: common questions
What platform and security teams ask before they deploy.
How do publishers block AI crawlers without leaving Google Search?
Block the AI-specific crawler names and keep Googlebot. Google says its Google-Extended token controls use of content for Gemini and does not affect inclusion in Google Search. Other AI companies also publish separate names for their training and search crawlers.
Do AI crawlers respect robots.txt?
Declared crawlers from the large AI companies say they do. Scrapers that hide their identity do not, and agents that fetch a page for a user may not apply it. robots.txt states your preference; enforcing it needs a control on the server.
How can I tell a real GPTBot from a fake one?
Check the request against the verification data the operator publishes, such as OpenAI's IP ranges for GPTBot. A request that claims the name from an address outside those ranges is not GPTBot, whatever its user agent says.
Can paywalled articles still be scraped?
Yes. A scraper can log in with a paid account and read pages in a real browser, so the paywall alone does not stop it. The useful evidence is how the session behaves: how fast it reads, how many pages, and whether the browser is automated.
Does Centinel decide which AI uses of my content are allowed?
No. Centinel identifies and verifies the client and applies the policy you set for it. Which companies may use your content, and on what terms, is your decision.