Skip to content
Industry

Decide which AI crawlers and scrapers can read your content

Search crawlers, AI crawlers, assistants fetching a page for a user, and scrapers that hide behind a browser all read the same articles. Centinel identifies which is which, so publishers can keep search indexing and give every other client the access they choose.

Request a site audit
Why it matters

robots.txt is a request, not a control

RFC 9309, the robots.txt standard, says its rules are not a form of access authorization. Declared crawlers usually honor them. Scrapers that ignore them, or copy a crawler's name, do not. Some agents that fetch a page for a user do not apply robots.txt at all; OpenAI says its rules may not apply to ChatGPT-User.

Read the research
How it works against you

What the traffic looks like

  • AI crawlers, declared and not

    Training and search crawlers from AI companies usually identify themselves. Others arrive with a browser user agent, or with a well-known crawler name they have no right to.

  • Archives and paywalls copied

    Scrapers walk article archives and subscriber pages to copy full text, sometimes through a logged-in account and a real browser.

  • Agents reading for a user

    AI assistants and browser agents load pages when a person asks them to, which makes them a different client from a bulk crawler and a different policy decision.

How Centinel closes it

What runs against this traffic

  • Verify declared crawlers

    Centinel can compare a declared crawler identity with the verification data its operator publishes, such as IP ranges or DNS records.

    A real Googlebot keeps indexing, and a verified AI crawler gets the policy you set for it. A copied crawler name receives no automatic trust.

    Read more
  • A policy for each crawler and section

    Supported integrations apply the configured response for the client's identity and the requested resource: allow, rate-limit, challenge, or block.

    Search indexing can stay open while AI training crawlers are refused, or allowed only on some sections, without waiting for every bot to read robots.txt.

    Read more
  • Evidence against disguised scraping

    In supported browser deployments, Centinel evaluates request and browser evidence on article and archive pages, not only the user agent.

    A scraper posing as a reader can be rate-limited or challenged without putting subscribers through the same check.

    Read more
Related use cases
Start with your site

Want to see what reaches your site?

Start with a free scraping audit: tested tools, exposed pages, and fixes to consider, reviewed by hand and emailed to you. This checks scraping exposure, not every abuse pattern or all production traffic. Discuss broader workflow coverage in a demo.

Questions

Media & publishing: common questions

What platform and security teams ask before they deploy.

How do publishers block AI crawlers without leaving Google Search?

Block the AI-specific crawler names and keep Googlebot. Google says its Google-Extended token controls use of content for Gemini and does not affect inclusion in Google Search. Other AI companies also publish separate names for their training and search crawlers.

Do AI crawlers respect robots.txt?

Declared crawlers from the large AI companies say they do. Scrapers that hide their identity do not, and agents that fetch a page for a user may not apply it. robots.txt states your preference; enforcing it needs a control on the server.

How can I tell a real GPTBot from a fake one?

Check the request against the verification data the operator publishes, such as OpenAI's IP ranges for GPTBot. A request that claims the name from an address outside those ranges is not GPTBot, whatever its user agent says.

Can paywalled articles still be scraped?

Yes. A scraper can log in with a paid account and read pages in a real browser, so the paywall alone does not stop it. The useful evidence is how the session behaves: how fast it reads, how many pages, and whether the browser is automated.

Does Centinel decide which AI uses of my content are allowed?

No. Centinel identifies and verifies the client and applies the policy you set for it. Which companies may use your content, and on what terms, is your decision.