SEO and AI

AI Crawlers: Who Reads Your Site, What Each One Does and the Three Doors That Decide Access

AI crawlers read the web to train models, build search indexes and serve user actions. See the 2026 list and the three doors that decide their access to your site.

By , founder and lead strategist at Flowup

Anthropic, OpenAI, Perplexity, Google, Meta, Apple and Amazon each operate their own crawlers, with different purposes and different consequences. This guide provides the updated 2026 list, the traffic data showing who reads the web the most today, the real cost of blocking each operator and the three-door model: the declared door in robots.txt, the invisible one in the infrastructure and the outsourced one at the CDN. Two of them almost nobody audits.

Direct answer

AI crawlers are the bots that artificial intelligence companies use to read the web for three distinct purposes: training models, building search indexes and serving user actions in real time. Their access to your site is decided at three doors: the declared one in robots.txt, the invisible one in the infrastructure and the outsourced one at the CDN, and the last two are almost never audited. The scale justifies the care: according to Cloudflare Radar data for July 2026, Claude-User was the second most active bot on the web, behind only Googlebot, and ClaudeBot accounted for 16.28% of AI bot traffic. Controlling these doors is the foundation of any GEO and AEO strategy.

The three-door model: why this guide exists

Almost everything published about AI crawlers deals with a single question: what to write in robots.txt. It is a necessary question, and an insufficient one. The robots.txt file is only the declared door, the one the site owner chooses and publishes. There are two others, and they carry more weight.

The second door is the infrastructure: web application firewalls, commercial security rule sets and server configurations that stop crawlers before the request reaches the site, often as a factory default and without anyone knowing. The third is the CDN: since July 2025, the web's largest infrastructure provider has blocked AI crawlers by default on new domains, which means the decision may already have been made by a vendor, not by you.

The practical result: a site can have the most permissive robots.txt in the world and be completely invisible to an AI platform, with no alert in any tool. We documented a case like this firsthand, on a technically exemplary site, and it changed the way we treat the subject. This guide walks through the three doors in order, with the updated list of who knocks on them and an audit method that sees what the logs do not show.

What AI crawlers are and the three purposes that change the decision

AI crawlers are automated programs that travel the web collecting content for artificial intelligence systems. The mechanism is the same as any crawler's: HTTP requests, content extraction, link discovery. What changes is the destination of what was read, and it is the destination that defines the strategic decision.

The traditional split into two types, training and inference, is no longer enough for the 2026 ecosystem. There are three purposes, and the infrastructure providers themselves have started to classify them separately in their controls.

Training

GPTBot, ClaudeBot, Meta-ExternalAgent and Bytespider collect content to build the datasets for future versions of the models. The effect of blocking is long term and diffuse: the content stops shaping what the models will know, but nothing changes in today's answers. According to Cloudflare Radar data compiled for July 2026, training accounted for 44.54% of the declared purpose of AI crawling, up from 35.74% a year earlier.

Search and indexing

OAI-SearchBot, Claude-SearchBot and PerplexityBot build the indexes that AI search products query. The effect of blocking is direct: the site leaves that platform's results. It is the category closest to Googlebot's historical role, and the one that should least be blocked by anyone who depends on being found.

User action

Claude-User, ChatGPT-User and Perplexity-User do not crawl on their own: they fetch pages in real time when a person asks. Someone pastes your site's URL into Claude and asks what the company does; it is Claude-User that will read it. The effect of blocking is immediate and the most underestimated on the list: the AI cannot read your site even when your own customer asks it to, and citation in real-time answers disappears.

This third category changes the whole calculation. Blocking training is a negotiable position on intellectual property. Blocking user action is refusing to be read by a customer who is, at that very moment, researching you.

The size of it in 2026: who reads the web the most today

The 2026 numbers have retired two common intuitions: that AI crawlers are marginal traffic and that OpenAI dominates crawling.

According to Cloudflare Radar data for July 2026, in a documented independent compilation, ClaudeBot accounted for 16.28% of AI bot traffic, behind only Googlebot. A year earlier, in July 2025, ClaudeBot's share was 11.23%. The positions swing with each lab's training cycles, and any number in this paragraph should be read with the date beside it: these are monthly measurements of a volatile ecosystem.

The most important figure, however, is not about training. In the breakdown by individual bots, in the four weeks to August 1, 2026, Claude-User appeared as the second most active verified bot on the entire web, with 7.57% of requests, behind only Googlebot, with 12.87%, and ahead of Meta-ExternalAgent and GPTBot. A user action agent, driven by human demand in real time, overtook the classic search crawlers. What used to be a forecast about the agentic web became the second row of the table.

The honest counterpoint is the ratio between crawling and return. In the same July 2026 window, the ratio of pages crawled to visits referred was 2,237 to 1 at Anthropic and 217 to 1 at OpenAI, against ratios close to parity at traditional search engines. AI crawlers read a lot and send back few clicks, and it is exactly this asymmetry that feeds the block-by-default movement this guide covers at the third door. The decision to allow access is not obvious; it is strategic, and it depends on where your brand needs to exist: in the clicks, in the answers, or in both. The difference between these two worlds is the central theme of SEO versus GEO.

The 2026 list: user agents, purposes and the cost of blocking

The table below consolidates the relevant crawlers as of August 2026, with the column that lists usually leave out: what you lose, in practice, by closing the door on each one. Tokens that exist only as a robots.txt directive are identified, because treating them as crawlers is a common configuration mistake. The robots.txt compliance column reflects what each operator states in its documentation, plus public reports of noncompliance where they exist.

Bot or token Operator Purpose What you lose by blocking robots.txt
GPTBot OpenAI Training Presence in the datasets of future GPT models States compliance
OAI-SearchBot OpenAI Search and indexing Presence in ChatGPT search results States compliance
ChatGPT-User OpenAI User action Real-time reading and citation when the user asks The documentation says the rules may not apply, because the actions are initiated by a user
ClaudeBot Anthropic Training Presence in the datasets of future Claude models States compliance
Claude-SearchBot Anthropic Search and indexing Presence in Claude search results States compliance
Claude-User Anthropic User action Real-time reading by the second most active bot on the web in July 2026 States compliance
PerplexityBot Perplexity Search and indexing Presence in Perplexity's index and answers States compliance; publicly reported episodes of noncompliance
Perplexity-User Perplexity User action Real-time reading at the user's request The documentation notes that the user action fetcher generally ignores robots.txt rules
Googlebot Google Traditional search and the AI experiences in Search Removal from all of Google; never block States compliance
Google-Extended Google robots.txt-only token: Gemini training and grounding Presence in Gemini training; does not change Search or AI Overviews Token, makes no requests
Applebot Apple Search (Siri and Spotlight) Presence in searches on Apple devices States compliance
Applebot-Extended Apple robots.txt-only token: Apple Intelligence training Presence in Apple Intelligence training Token, makes no requests
Meta-ExternalAgent Meta Training Presence in the datasets of Meta AI and the Llama family States compliance
Amazonbot Amazon Search and answers (Alexa) Presence in the answers of Alexa and related services States compliance
Bytespider ByteDance Training Presence in ByteDance's models Recurring reports of noncompliance; requires enforcement by firewall
CCBot Common Crawl Public crawl archive Presence in Common Crawl and in the many open models trained on it States compliance
DuckAssistBot DuckDuckGo Assisted answers Presence in DuckAssist, an operator with a return ratio close to parity States compliance

Two notes on reading the table. First: full user agent strings vary from version to version, and firewall rules usually match the token, not the whole string; before configuring anything, check each operator's official documentation, listed in the references. Second: the Google-Extended row corrects a frequent mistake. According to Google's documentation, AI Overviews is a Search feature and follows Googlebot's index; Google-Extended controls the use of content in Gemini training and grounding. Blocking the token does not remove the site from AI Overviews, and confusing the two has already led many people to wrong decisions in both directions.

Operator by operator: what you lose at each closed door

The table summarizes; the decision calls for one more paragraph on the operators that concentrate the traffic.

Anthropic. It is the operator that grew the most over the year and the only one with all three types of bot at high volume at the same time. Blocking the whole trio in 2026 means being invisible to what was, in July 2026, the web's second-largest reader, and to an expanding search index. The company publishes documentation, official IP ranges and a contact channel for the crawler, which places it among the most verifiable operators.

OpenAI. It keeps the clearest separation between purposes, with one bot per function and documentation with IP ranges for origin verification. The classic granular decision, blocking GPTBot and allowing OAI-SearchBot and ChatGPT-User, protects content from training without leaving the answers.

Perplexity. It operates search and user action, with a public history of robots.txt noncompliance episodes that the company treated as fixed. The nuance almost nobody reads: the documentation for the user action agent notes that, on requests initiated by people, the fetcher generally ignores robots.txt rules, on the logic that it was a human who asked. For anyone who needs an effective block, this moves control from the declared door to the enforcement doors.

Google and Apple. The two cases in which confusing crawler and token costs the most. Googlebot and Applebot are search crawlers and should not be blocked by anyone who wants to be found; Google-Extended and Applebot-Extended are robots.txt tokens that control only training. No request reaches the server under the tokens' names, and putting them on a firewall allowlist has no effect at all.

Common Crawl. CCBot feeds a public archive used to train many open models. Blocking it has a cascading effect on an entire ecosystem, not on a single company, and the decision deserves that framing.

Door 1: robots.txt, the declared layer

The robots.txt file remains the central instrument of governance, and the only one of the three with behavior defined by a standard, RFC 9309. It works by cooperation: the main operators state compliance and, in general, comply. Three designs cover most cases.

Allow everything is the absence of a rule: with no directive for a user agent, access is open. It is the natural stance for anyone competing for visibility in answer engines.

Block training and allow search and user action is the granular design most used by those who want to protect editorial content without leaving the answers:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-User
Allow: /

Block by section combines commercial protection with informational visibility: an e-commerce site can close its price and inventory directories and keep the blog open, with Disallow rules by path instead of a total block.

Two notes on precision. The Crawl-delay directive is not part of the standard and has inconsistent support among AI crawlers; load control is done with rate limiting on the server or at the CDN, with limits tolerant enough not to interrupt legitimate crawls. And llms.txt does not belong to this door: it is a curation proposal that points models to relevant content, with no blocking power and still without declared adoption by the major operators, as we discuss in our analysis of whether llms.txt works. Anyone who wants to guide reading can publish it as a low-cost index, knowing that Google states that Search does not use the file and that publishing it neither helps nor hurts; anyone who wants to prevent access needs the next doors.

Door 2: the infrastructure, the layer that blocks without warning

Here lies the problem that user agent lists leave out. Between the crawler and your robots.txt there is an entire stack: network firewall, the server's security rule sets, web application firewall, cache. Any of these layers can drop the request before it reaches WordPress, robots.txt and your application logs, and the rules that do this usually ship preconfigured in commercial security packages, enabled by default on shared servers. On WordPress sites, checking these layers is part of the security routine of a WordPress specialist.

We documented a case like this in August 2026: a technically exemplary medical website, with a permissive robots.txt, on which a commercial ModSecurity rule classified ClaudeBot as a threat and closed the connections without returning any HTTP response. All the other crawlers got through with status 200; only one operator was shut out, and nobody had decided that. The server cache masked the problem on popular pages, and no traditional tool flagged anything. The full reconstruction, with the investigation, the two blocking layers found and the fix protocol, is in the case study on AI crawler blocking.

The normative consequence is what makes this door so severe. Under RFC 9309, when robots.txt is unreachable, with no usable response from the server, the crawler must assume that the entire site is disallowed. A silent infrastructure block does not deny one page: it turns the whole domain into off-limits territory, under the standard itself. And since there is no dashboard equivalent to Search Console for AI crawlers, no vendor warns the site owner. The block lasts until someone tests.

Door 3: the CDN, the layer that decides by default

The third door is not on your server: it is at the vendor in front of it. On July 1, 2025, Cloudflare, which by its own account handles about a fifth of web traffic, became the first major infrastructure provider to block AI crawlers by default on new domains, flipping the model from opt-out to opt-in and launching a marketplace for charging per crawl. In July 2026, it refined the system with separate controls by purpose: search, agent and training, plus specific protection for pages monetized by ads.

For those who publish content and want to be cited, the implication is direct: if the site sits behind a CDN with these controls, the AI crawler access policy may have been set by a vendor default, when the zone was created, without going through any editorial decision of yours. The audit of this door is a dashboard, not a file: the CDN's bot controls need to enter the same periodic review as robots.txt, with the explicit question of which categories are blocked, limited or allowed, and whether that reflects the strategy or just the default.

The movement also explains the general climate: with crawl-per-visit ratios in the thousands to one, blocking by default became a product, and the trend is toward more friction, not less. Whoever decides to allow access needs to really decide, at all three doors, because the entire ecosystem is deciding the opposite by inertia.

How to audit: from the outside in, not just through the logs

The traditional method of bot management starts in the server logs: filter by user agent, measure volume, verify authenticity by reverse DNS. It remains necessary, and it has a structural blind spot: the log shows who arrived. It does not show who was stopped before arriving, because the blocks at doors 2 and 3 happen before the application records anything. A site can have zero ClaudeBot requests in its logs for two opposite reasons, lack of interest or blocking, and the log does not tell the two apart.

The complete audit combines the two directions.

  1. From the outside in: access test. From a machine outside the server, requests with the user agent of each relevant crawler to three URLs: robots.txt, the home page and a rarely visited internal page, outside the cache. A control request with a browser user agent before and after the battery, to detect a ban on the test IP. A dropped connection confirmed with a second attempt after a pause. Status 200 is access; 403, 429 and a connection closed without a response are a block or a limit, each pointing to a different layer.
  2. From the inside out: log reading with origin verification. Filtering by user agent in the access logs, with validation by reverse DNS and by the official IP ranges that Anthropic and OpenAI publish, to separate authentic crawlers from imitations. This is where the classic log analysis tools and the CDN's bot dashboards come in.
  3. Dashboards for doors 1 and 3. A review of the published robots.txt, of what the SEO plugin generates dynamically and of the CDN's bot controls, recording what is blocked by choice and what is blocked by default.
  4. Two rounds and a calendar. Only what shows up in two batteries on different days is a confirmed block. And since commercial rule sets update without notice, the audit belongs in the maintenance routine, not on the list of one-off events.

The detailed protocol, with the commands, the table of symptoms by layer and the script for the hosting support ticket, is in the case study cited in Door 2. For those who manage WordPress, the review of the security layers ties directly into the work of WordPress security and cleanup: the same tools that protect the site are the ones that, poorly calibrated, make it invisible.

Decision matrix: business model x bot purpose

There is no universal answer to allowing or blocking; there is the intersection between what the site sells and what each category of bot does. The matrix below summarizes the starting recommendation by profile, always subject to the specific strategy.

Site profile Training Search and indexing User action Rationale
B2B brand, services, consulting Allow Allow Allow The asset is being found and cited; the content exists to circulate
Publisher with paid or exclusive content Block or negotiate Evaluate by platform Allow with measurement Protection of the editorial asset without disappearing from real-time answers
E-commerce Allow on content, evaluate on the catalog Allow Allow Price and inventory can be sensitive; product and content discovery are not
Corporate website with sensitive data by section Block by directory Allow on public areas Allow on public areas A rule by path works better than a total block
Knowledge base and documentation Allow Allow Allow It is the content with the highest citation rate by design

Two rules cut across every profile. The first: the decision declared at Door 1 needs to be checked at Doors 2 and 3, because there is no point in allowing in robots.txt what the firewall drops. The second: allowing access is a foundation, not a result. Citation depends on what comes after access, and that is the work of content, entity and authority, the territory of the complete GEO guide and of your official knowledge base for AIs (in Portuguese).

Common mistakes in managing AI crawlers

Editing robots.txt to fix a firewall block. If the request dies before reaching the site, the file is irrelevant. First identify the layer, then fix it there.

Concluding from the log that there is no block. An absence of requests in the log is compatible with a block in the earlier layers. The outside-in test is the only one that answers the question.

Testing only the home page. The cache makes the popular pages pass the test and hides the block exactly on the new and rarely visited URLs, the ones that most need to be discovered.

Treating Google-Extended and Applebot-Extended as crawlers. They are robots.txt tokens; no request arrives under those names, and a firewall allowlist for them has no effect.

Blocking Googlebot or Applebot when aiming at training. They are the search crawlers. The target for anyone who wants to restrict training is the Extended tokens and the training bots, never the search ones.

Trusting Crawl-delay. It is outside the standard and inconsistently supported. Load control is well-calibrated rate limiting, tolerant of legitimate crawlers.

Testing Googlebot from an ordinary IP and concluding there is a block. Firewalls verify Googlebot's origin by reverse DNS and stop imitations by design. That test measures the defense against spoofing, not the real Google's access.

Allowing by IP instead of user agent. Operators' IP ranges change; the stable way to allow access is by agent, with the official ranges serving authenticity verification, not the main rule.

Deciding once and never reviewing. Commercial rulesets update without notice, CDNs change defaults, operators launch new bots. Managing AI crawlers is a maintenance routine with an owner and a calendar, in line with the same foundation logic as the B.I.N.A. Method.

Frequently asked questions

What is the difference between training, search and user action crawlers?

They are three distinct purposes, and the classic split into just two is no longer enough. Training crawlers, such as GPTBot and ClaudeBot, collect content to build the datasets for future versions of the models: the effect of blocking is long term. Search crawlers, such as OAI-SearchBot, Claude-SearchBot and PerplexityBot, build the index that AI search products query: blocking takes the site out of those results. User action agents, such as Claude-User, ChatGPT-User and Perplexity-User, fetch pages in real time when a person asks: blocking has an immediate effect on citations and on the AI's ability to read your site at a customer's request.

How do I know whether AI crawlers can access my site?

By testing from the outside in, not just by reading logs. The server log shows who arrived; it does not show who was stopped before arriving, because firewall and CDN blocks happen before the application records anything. The right test uses curl from an external machine, sending the user agent of each crawler to three URLs: robots.txt, the home page and an internal page outside the cache. Status 200 indicates access; 403, 429 or a connection closed without a response indicate a block or a limit. Confirm dropped connections with a second attempt after a pause and repeat the battery on another day before drawing conclusions.

Does blocking AI crawlers in robots.txt guarantee they will not get in?

No, it does not. The robots.txt file is a cooperation protocol: the main operators state compliance and, in general, comply, but the file enforces nothing technically. Bots that ignore the protocol, with cases reported for Bytespider and episodes involving Perplexity, as well as user action agents that their own documentation describes as able to ignore the directive on requests initiated by people, are only contained by enforcement layers: firewall rules, bot management at the CDN or origin verification by reverse DNS and official IP ranges. Declaration and enforcement are different layers of the same control.

Does allowing AI crawlers increase my brand's citations?

Allowing access is a necessary condition, not a sufficient one. Without access, no citation is possible: the content does not enter an index, does not feed answers and cannot be read in real time. With access, citation comes to depend on what decides any editorial contest: answers that are clear and extractable, entity consistency, verifiable sources, accumulated authority and the competition on each question. A promise of guaranteed citation has no technical basis, because no vendor offers that control. The serious path treats access as a foundation, measures presence in answers over time and works on the content on top of that base.

Does Crawl-delay work for AI crawlers?

Do not count on it. Crawl-delay was never part of the standard formalized by RFC 9309, and support among AI crawlers is inconsistent: some interpret the directive, others ignore it completely, and most operators' documentation does not even mention it. To reduce crawl load without blocking, the reliable mechanisms are rate limiting configured on the server or at the CDN, with limits tolerant enough not to interrupt legitimate crawls, and the bot management controls that infrastructure providers offer, which make it possible to limit by category instead of dropping connections.

Do I need llms.txt in addition to robots.txt?

They are instruments of different natures, and neither replaces the other. The robots.txt file controls access: it says which URLs each user agent may crawl, and it is the only one of the two with behavior defined by a standard. The llms.txt file is an emerging curation proposal: a plain-text index that points models to the site's most relevant content, with no blocking power and still without declared adoption by the major operators. If the goal is to prevent access, the instrument is robots.txt combined with enforcement layers. If the goal is to guide reading, llms.txt can be published as a low-cost index, knowing that Google states that Search does not use the file and that publishing it neither helps nor hurts.

Next step

Are the three doors of your site the way you decided?

The B.I.N.A. Diagnosis audits the foundation before the content: AI crawler access at the three layers, robots.txt, firewall and CDN, structured data and readiness for answer engines. Ranking is not enough. Be the answer.

Guto Bertoncini is the founder of Flowup Agency, a Digital Authority Engineering company based in São Paulo, Brazil, and is responsible for the SEO, GEO and AEO strategy of the agency's projects. The case cited in Door 2 was conducted firsthand by the team in August 2026.

Transparency note: the traffic figures cited come from independent, documented compilations of Cloudflare Radar data and vary from month to month; each one is dated in its own paragraph. The robots.txt compliance column reflects what each operator states in official documentation, plus public reports of noncompliance when they exist, and does not constitute independent verification by Flowup. User agent strings should be checked in each operator's documentation before any configuration.

Sources and references

  1. Cloudflare. Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large. Official press release, July 1, 2025. cloudflare.com/press
  2. Cloudflare Blog. Your site, your rules: new AI traffic options for all customers. July 2026. blog.cloudflare.com
  3. Cloudflare Blog. A deeper look at AI crawlers: breaking down traffic by purpose and industry. August 2025. blog.cloudflare.com
  4. TechnologyChecker. Web Traffic Statistics 2026, independent compilation of Cloudflare Radar data, July 2026. technologychecker.io
  5. TechnologyChecker. Bot Traffic Statistics 2026, compilation of Cloudflare Radar data. technologychecker.io
  6. SEOmator. GEO Data Report 2026: crawl-to-refer ratio of AI crawlers, July 2026. seomator.com
  7. IETF. RFC 9309: Robots Exclusion Protocol. September 2022. datatracker.ietf.org
  8. Anthropic. Does Anthropic crawl data from the web, and how can site owners block the crawler? Official documentation. support.claude.com
  9. OpenAI. Overview of OpenAI crawlers. Official documentation. platform.openai.com
  10. Perplexity. Perplexity crawlers. Official documentation. docs.perplexity.ai
  11. Google Search Central. Google's common crawlers, including Google-Extended. Official documentation. developers.google.com

About the author

Portrait of Guto Bertoncini

Guto Bertoncini

Founder and lead strategist, Flowup Agency

Guto Bertoncini is the founder and lead strategist of Flowup Agency, which he has run since 2011. He is the author of the B.I.N.A. Method, Novo SEO and the Base Informacional Semântica (Semantic Information Base), and leads the agency's SEO for AI, GEO and AEO practice, preparing companies to be found on Google and cited by artificial intelligence platforms. He writes about search and AI on the Flowup blog and on his official website.

Keep reading

Marketing for Engineering and B2B Companies

In engineering and technical B2B, marketing has to prove competence before the first sales contact. An approach built on trust, digital authority, SEO, GEO and AEO.

Related content