Skip to content
Blog

What is agentic traffic? AI agents explained

Agentic traffic comes from software that picks and performs steps toward a user's goal. How it differs from crawlers and fetchers, and how to measure it.

Frederick Jahn
Published Updated
What is agentic traffic? AI agents explained

Agentic traffic is web traffic created by software that chooses and performs steps toward a goal. A user might ask an agent to compare products, research a topic, or complete a workflow. The agent then decides which pages or APIs to visit and what to do next.

A crawler collecting pages on a schedule is automated, but it is not acting on one user's live task, and neither is a search bot building an index. Reports often count all three as "AI traffic". They need different access rules: a crawler reads public pages, while an agent may log in, fill a form, or place an order.

What agentic traffic means

There is no HTTP header that turns a request into agentic traffic. The useful definition is operational:

Agentic traffic is one or more web requests made by software that can select its next action while working toward a user or application goal.

The sequence may include page navigation, JavaScript execution, API calls, or state-changing actions such as submitting a form. It may also stop after one request if the task is simple. The test is whether the software chose the next step. Page count and the presence of a large language model do not decide it.

Sites also receive requests from AI products that are not agentic in this sense. OpenAI, Anthropic, and Perplexity each document separate clients for model training, for search indexing, and for retrieval started by a user. OpenAI states that GPTBot may collect training content, OAI-SearchBot supports ChatGPT search, and ChatGPT-User performs certain user-triggered actions (OpenAI crawler documentation). Anthropic separates ClaudeBot, Claude-SearchBot, and Claude-User the same way (Anthropic crawler documentation).

Classify the trigger before the client

Start with why the request happened. The client name comes second.

Was the request triggered by a live user or application task?
|
+-- No or not known
|   |
|   +-- Building a search index?            -> search crawler
|   +-- Collecting model-development data?  -> training crawler
|   +-- Purpose cannot be established?      -> unidentified automation
|
+-- Yes
    |
    +-- Fetching content for an answer?      -> user-directed retrieval
    +-- Choosing and performing steps?       -> agentic session

If the trigger cannot be established, record it as unknown.

The same distinction works as a traffic matrix:

Traffic classTriggerTypical request shapePublished examplesFirst policy question
Search crawlerOperator's indexing systemRepeated discovery and refresh requestsGooglebot, OAI-SearchBot, Claude-SearchBot, PerplexityBotDo we want this search product to discover and surface the content?
Training or model-development crawlerOperator's collection scheduleBroad or repeated content collectionGPTBot, ClaudeBotHas this use of the content been approved?
User-directed retrievalA user asks an AI product to fetch or answer somethingOne or more fetches tied to the requestChatGPT-User, Claude-User, Perplexity-UserShould this user-triggered fetch reach the requested resource?
Agentic sessionA user or application delegates a taskA sequence of chosen reads or actions, sometimes through a browserNo universal identifierWhich actions may this automated session perform?

The examples were checked against operator documentation on October 6, 2026. Names and stated purposes change, so recheck them before you rely on the table.

Perplexity says PerplexityBot supports search and indexing, and Perplexity-User may fetch a page after a user asks a question (Perplexity crawler documentation).

Google-Extended belongs in none of the rows. Google documents it as a robots.txt product token that controls certain uses of content Google has already crawled. It has no HTTP user agent of its own, so a log search for the string finds no real Google traffic (Google's crawler documentation).

What the classes look like in logs

A single log row can tell you the claimed client, source address, route, method, status, time, and response size. It usually cannot tell you whether a human prompt, a scheduler, or another application caused the request.

Collect enough context to classify a sequence later:

  • the source IP observed at a trusted network boundary;
  • the full user-agent header and any verified operator identity;
  • timestamp, method, host, path, status, and bytes transferred;
  • a request or trace identifier;
  • an authenticated account or first-party session identifier, when present;
  • connection and browser signals available at your edge;
  • the classification result, reason, and rule version.

Do not build sessions from IP addresses alone. Corporate gateways, mobile networks, privacy relays, and proxy services can place unrelated users behind one address. The reverse is also possible: one automated workflow may move between addresses.

A declared search or training crawler is the simplest case to label, once its identity has been verified. Named user fetchers provide evidence about the trigger too, because their operators publish their purpose. An ordinary browser session is harder. It might be a person, a browser agent, a test runner, or a scraper using browser automation. Agent frameworks such as Browser Use and Stagehand drive an ordinary Chromium browser, so the name of the framework never reaches your logs.

What a user agent cannot tell you

A user-agent header is a claim supplied by the client, not proof of who sent the request. Checking that claim against operator evidence is covered in How to verify AI agents. For classification, keep two separate fields in the traffic record:

FieldQuestionExample values
IdentityWho can we verify sent this?OpenAI, Anthropic, unknown
Traffic classWhat triggered the request and what is it doing?search crawl, training crawl, user retrieval, agentic session, unknown

A verified operator can run several types of client, so identity does not settle the class. The absence of a named AI user agent proves even less: agents can use direct HTTP clients or browser automation, and detection may establish that a session is automated without revealing who operates it.

How to recognize an agentic session

Look at the sequence and the actions it attempts. Agentic traffic often moves through a task-shaped path: discover a page, inspect details, follow a relevant link, then call an API or submit data. The exact path depends on the delegated goal.

None of these signals is conclusive on its own:

  • route transitions that follow content or application state;
  • a switch from read-only requests to state-changing methods;
  • browser execution followed by an API request;
  • repeated attempts after validation or permission errors;
  • timing and navigation patterns inconsistent with the site's human sessions.

Use connection, browser, request, account, and behavioral evidence to decide whether the client is automated. Use the request sequence to describe what the automation is doing. Leave the operator and trigger unknown when the evidence does not establish them.

Unidentified automation may be a price scraper, uptime monitor, SEO tool, fraud bot, test runner, or a custom agent. Label it agentic only when the sequence shows chosen steps.

Match the response to the traffic class

The class tells you which policy question to ask. The answer still depends on the route.

Traffic classSensible starting point
Search crawlerVerify the operator, then decide which content should be discoverable in that search product.
Training crawlerApply an explicit model-development access rule. Do not inherit the search-crawler rule.
User-directed retrievalVerify named clients where possible, protect restricted resources, and set bounded request limits.
Agentic sessionApply normal authorization to every action. Add abuse controls around login, checkout, forms, and other state changes.
Unidentified automationUse the site's default automated-traffic policy until stronger evidence changes the classification.

Verification does not override authorization. A genuine user-directed fetcher does not gain access to a subscriber-only article without the required credentials. An authenticated agent should not be able to submit an action its account is not permitted to perform.

The policy may allow, observe, rate-limit, challenge, or block a request. Make that decision using the resource, identity result, behavior, and business rule together. How to control AI agent access covers that layer in detail.

How to measure agentic traffic

One chart for every AI-related request mixes scheduled crawls with delegated tasks. Report traffic class and identity status separately. A report should include:

  • requests and sessions by traffic class;
  • verified, failed, unknown, and unidentified identity counts;
  • pages and bytes served by route group;
  • cache hits, cache misses, response status, and origin work where available;
  • state-changing attempts and their authorization result;
  • policy actions and reason codes;
  • the classifier version used for the report.

Keep raw observations separate from inferred labels. If a classification rule changes, versioning lets you reprocess historical data or at least explain why the chart moved. Otherwise a taxonomy update can look like a traffic surge.

For scheduled crawler mechanics, see What is an AI crawler?. For browser-controlled automation, see How to detect browser automation.

Finding the classes on your own site

Centinel helps teams identify and control automated traffic. When you evaluate any product for this, test whether it keeps operator identity, traffic class, and policy outcome separate, and what it reports when a request cannot be classified or verified. The bot management guide has the full evaluation checklist.

To see which classes reach your current pages before you set policy, request a site audit.

Sources