SEO and AI

AI Crawler Blocking: The Technical Diligence That Decides Who Exists in the Answers

A firewall rule nobody knew about made an exemplary medical site invisible to ClaudeBot. See the case, the blocking layers and the protocol to audit AI crawler access.

By , founder and lead strategist at Flowup

A technically exemplary medical website, with signed content, academic sources and structured data, was invisible to Anthropic's crawler. No tool flagged the problem. The cause was a firewall rule that not even the site owner knew about. This article reconstructs the case firsthand, explains the layers where invisible blocking happens and delivers the audit protocol that every serious GEO and AEO engagement should include.

Direct answer

AI crawler blocking is the barrier, almost always invisible to the site owner, that keeps crawlers such as ClaudeBot, GPTBot and PerplexityBot from accessing a site's pages, and it cancels out any GEO and AEO strategy before the content even comes into play. In a case we diagnosed in August 2026, a commercial ModSecurity rule dropped ClaudeBot's connections without returning any HTTP response, while all the other crawlers received status 200. The context raises the risk: since July 2025 Cloudflare has blocked AI crawlers by default on new domains, and Cloudflare Radar data for July 2026 shows Claude-User as the second most active bot on the web. Access is the foundation: without it, no citation is possible.

The case: exemplary content, zero visibility

The site that prompted this article had everything the GEO and AEO literature recommends, and not as decoration. It is the website of a physician in São Paulo, Brazil, a Flowup client, with clinical condition pages that carry a direct answer block at the top, extensive frequently asked questions, seven academic sources with full references per page, international diagnostic criteria laid out in tables, visible medical council credentials and specialist registrations, signed authorship, an update date and a medical disclaimer. There was even an official knowledge base built specifically for AI agents to read.

And the AI agent could not reach it. Not that knowledge base, and not any other page on the domain.

During a working session on August 25, 2026, we asked an AI assistant to read a page of the site for a content audit. Every attempt returned the same verdict: automated access forbidden. The home page, the internal pages, robots.txt itself. For Anthropic's crawler, the entire domain did not exist.

The paradox is the thesis of this article in a single fact: the market sells GEO and AEO as a content discipline, but the first layer is infrastructure. All the editorial work was correct. The layer nobody audits tore down what the layer everybody audits had built.

The anatomy of a silent failure

Any professional's first reflex would be to look for a 403 error in the logs. There was no 403 error, and that absence is the most important piece of data in the case.

When we reproduced the crawler's request with curl, from outside the server, what came back was curl error 52: an empty reply. The TCP connection was accepted, the TLS handshake was completed, the request was sent, and the server closed the connection without returning a single byte. No status code, no HTML block page, no record at the application layer. On the following attempts the behavior changed to error 35, a connection reset during the handshake itself, because the first blocked request had put the test IP on a temporary ban list.

This is the signature of the silent failure: it does not show up as an error anywhere. The site worked perfectly for patients. Google kept indexing it normally. The analytics tools recorded nothing abnormal, because they measure visitors with a browser. The application firewall did not record the block, because the request died before reaching it. The problem only became visible because someone tried to use AI to read the site, and failed.

The investigation: ruling out hypotheses is also method

We record the wrong hypotheses here on purpose. A diagnosis that presents only the final answer looks elegant, but it does not teach anyone to retrace the path, and it is the path that separates technical diligence from guesswork.

Hypothesis 1: robots.txt

Ruled out with direct evidence. The file followed the most permissive pattern possible: user agent asterisk, disallow only on wp-admin. No directive against AI crawlers. The lesson built into it: when robots.txt is generated dynamically by the SEO plugin, as was the case here, the request for it travels through the entire PHP stack, and a block at a lower layer can prevent even the reading of the file that was supposed to declare the permissions.

Hypothesis 2: the WordPress security plugin

Plausible and wrong. The site ran an application firewall in extended mode, loaded before WordPress on every PHP request. It was the perfect suspect. But security plugins respond with HTML block pages and 403 codes; they do not close connections in silence. The curl test cleared the entire application layer: the request never reached PHP. Anyone who works with WordPress security knows this boundary: what happens before PHP is not WordPress's fault.

Hypotheses 3 and 4: geographic blocking and client fingerprinting

The geographic blocking hypothesis fell when the test run from a Brazilian residential IP failed in the same way as the US-based crawler. Detection by TLS fingerprint, a technique that identifies the tool by the cryptographic signature of the handshake and not by the user agent, was ruled out by the hosting provider itself, with a check of the server logs.

Four reasonable hypotheses, four dismissals backed by evidence. What was left pointed in a single direction: a server rule, matching the crawler's user agent specifically.

The real cause: two layers and a cache in between

The log analysis by the hosting provider's support team, a Brazilian company that conducted the diagnosis with precision, found the first answer: a rule from the commercial Malware.Expert rule set for ModSecurity, identified as 1200002, that classified the ClaudeBot user agent as a threat pattern. When the rule fired, the server dropped the connection without a response and flagged the source IP as a bot, refusing the connections that followed for a period. That explained the exact sequence of errors observed in the tests.

Nobody at that hosting company decided to block Anthropic. Nobody at the site did either. The rule came from the factory, inside a third-party security package, and it was active on a shared server that hosts dozens of other sites. This is the part of the case that generalizes: the antagonist is not a villain, it is a factory default that nobody reviewed.

The second layer

The first fix seemed to have solved it. The robots.txt file started responding 200 to the crawler's user agent. But when we tested internal pages in sequence, the server started responding 429, the code for too many requests, from the fifth request in three minutes onward. The hosting provider's investigation revealed that it was not a rate limit: it was a second user agent blocklist, at another layer of the firewall, that still contained the ClaudeBot entry. Two independent rules, with the same target, at different layers of the same stack.

The cache that hid the problem

And here is the most treacherous detail of the case. After the first fix, robots.txt and the home page responded 200 with a telling header: the hit came from the cache of the LiteSpeed server, which responds before the layer where the second rule lived. The popular pages, served from the cache, got through. The pages outside the cache, exactly the least visited and the most recent ones, hit the application server and fell under the rule.

The consequence is perverse and worth stating clearly: the newer and less visited the page, the greater the chance that it is blocked. The content you have just published, the content that most needs to be discovered, is the first to become invisible. A test that checks only the home page passes the site and gets the diagnosis wrong.

What the standard says: with no response, the crawler assumes complete disallow

There is a normative reason why the silent block is more serious than an ordinary 403. RFC 9309, which standardized the Robots Exclusion Protocol in 2022, defines the crawler's expected behavior according to the robots.txt response: if the file responds with a client error, such as 404, the crawler may treat the site as open; but if robots.txt is unreachable, with no usable response from the server, the crawler must assume that the entire site is disallowed.

That is exactly what happened in this case. Dropping the connection without a response made robots.txt unreachable, and the crawler, following the standard, treated the whole domain as off-limits. A perfectly permissive file that nobody could read produced the effect of a total disallow. The declared configuration said yes; the infrastructure answered with silence; and the standard says silence must be read as no.

Why nobody notices: GEO has no Search Console

SEO matured with the support of a free diagnostic dashboard. Search Console reports crawl errors, excluded pages and indexing problems, and it notifies the site owner. For AI crawlers, no such equivalent exists: neither Anthropic, nor OpenAI, nor Perplexity currently offers a dashboard that warns the site owner that the crawler was blocked. There is no alert, no history, no notification. A block can last for months without anyone knowing, because no instrument in the traditional stack was designed to see it.

The second reason is asymmetry. In the documented case, GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot and Bingbot accessed the site normally, with status 200, including from US IPs. Only ClaudeBot was shut out. Anyone who tested the site's presence by asking ChatGPT would conclude that everything was fine, and would be right about one vendor only. The intuitive heuristic of testing one crawler and assuming the result holds for the others does not survive this case: each platform has its own crawler, and the verdict is individual.

And the third reason is the cache effect already described, which makes the test itself lie when it is run only on the most visited pages.

The market context: blocking by default is growing

This case is not an isolated curiosity. It takes place within a structural movement of the web, and knowing that movement is part of technical diligence.

On July 1, 2025, Cloudflare, which by its own account handles about a fifth of web traffic, became the first major infrastructure provider to block AI crawlers by default on new domains, flipping the model from opt-out to opt-in and launching a marketplace for charging per crawl. In July 2026, the same company refined the system, separating the controls by purpose: search, agent and training. The direction is unmistakable: AI crawler blocking is no longer a handcrafted exception and has become a factory setting at scale.

At the same time, the cost of being blocked has never been so high. Independent compilations of Cloudflare Radar data for July 2026 show ClaudeBot with 16.28% of AI bot traffic, behind only Googlebot, and Claude-User, the agent that fetches pages in real time when a user asks Claude something, as the second most active verified bot on the entire web in the four weeks to August 1, 2026, with 7.57% of requests, behind only Googlebot, with 12.87%. Blocking Anthropic in 2026 does not mean losing a training crawler: it means closing the door on one of the internet's largest readers, driven by real user demand in real time.

There is also the side of legitimate choice. In the robots.txt sample analyzed by Radar on June 1, 2026, with 4,072 files, GPTBot appeared blocked, fully or partially, in 529 and ClaudeBot in 457, numbers that represent declared editorial decisions. This distinction needs to be sharp: the site owner who publishes a disallow is exercising governance over their own content, and that is legitimate. The problem in this article is a different one: the site that wants to be found, invests in content for that purpose and is blocked without knowing it, by a layer it never audited. Understanding the difference between SEO and GEO helps gauge what is at stake at each of these doors.

The technical protocol: how to audit AI crawler access

What follows is the protocol we apply at Flowup, refined by this case. It requires no paid tool; it requires rigor in running the tests and in reading the results.

  1. Test from outside the server, with each crawler's real user agent. Use curl from an external machine, sending the official user agent of each crawler, checked against each vendor's documentation in the week of the test. Test at least ClaudeBot, Claude-User, GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot and Bingbot. Tests from inside the server do not count: they do not go through the firewall.
  2. Three URLs per crawler, always including a page outside the cache. The robots.txt file, the home page and a rarely visited internal page. The case proved that the cache passes the first two and hides the block exactly where it hurts the most.
  3. A control request before and after the battery. Run the same test with a browser user agent before the bots and repeat it at the end. If the control passes before and fails after, your IP was banned during the test and the results in between are contaminated. That is what happened in this case, and it would have led to the wrong conclusion.
  4. Confirm every dropped connection with a second attempt after a pause. A single drop can be instability. A real block drops twice. Record the exact type of error: an empty reply, a connection reset, a timeout and a 403 point to different layers.
  5. Read the result layer by layer, not as a single verdict. The table below summarizes the correspondence between symptom, likely layer and means of confirmation that this case validated in practice.
  6. Ask the hosting provider for the specific rule, with evidence. An effective support ticket carries the source IP, the exact time, the user agent used and the type of error observed. Ask by name about ModSecurity rule sets, user agent lists in the firewall and geographic blocking. In this case, the hosting provider located the rule within hours because the ticket arrived with that data.
  7. Allow by user agent, never by IP, and revalidate after the service restarts. Vendors use cloud IP ranges that change; Anthropic publishes the ranges and the official documentation, but a stable allow rule is set by agent. And a WAF fix only counts after the web service has been restarted and the test has been run again on a page outside the cache.
  8. Two rounds on different days before closing the case. Only what shows up in both is a confirmed block. And since commercial rule sets update without notice, the audit is not a one-off event: it goes on the maintenance calendar, along with the review of robots.txt and of rate limiting.
Observed symptom Likely layer How to confirm
Connection closes without any HTTP response Network firewall or server rule with a drop action, before the application External curl returns an empty reply; the application logs record nothing; ask the hosting provider for the ModSecurity and firewall logs
Connection reset right after earlier attempts Temporary ban on the test IP, triggered by the first blocked request The control with a browser user agent also fails; wait and repeat from another IP
403 with an HTML block page Application WAF or security plugin The response body identifies the product; check the lists and rules in the plugin's dashboard
429 after a few requests in sequence Rate limiting, or a second blocklist disguised as a limit Ask the hosting provider which layer applies the limit and what the value is; in the case reported, no limit was configured and the 429 came from another user agent list
Home page responds 200 and internal pages fail Cache responding before the blocking layer Check the cache headers in the response; repeat the test on a URL outside the cache
Everything responds 200 and the content does not appear in any AI It is not access: it is content, entity or competition Editorial and structured data audit; the problem moves to another discipline

For rounds at scale, the same protocol can be automated: a script that goes through a list of domains, varies the user agents, applies the confirmations and classifies the verdicts. That is how we are measuring the extent of the problem on Brazilian websites, and the numbers from that measurement will be published when the two rounds are complete, with an open methodology.

Common mistakes when diagnosing blocks and allowing AI crawlers

Editing robots.txt to fix a firewall block. If the request dies before reaching the site, the file is irrelevant: the fix has to happen at the layer that blocks. In the case reported, robots.txt had been permissive from the start.

Testing only the home page and declaring the site open. The cache passes the home page and hides the rest. A page outside the cache is a mandatory item of the test, not a refinement.

Drawing a conclusion from one crawler and assuming it holds for all. Blocking is asymmetric by nature, because each rule set lists different agents. The verdict is individual, vendor by vendor.

Allowing by IP. The ranges change, the allow rule breaks in silence and the problem comes back without warning. The stable way to allow access is by user agent, with the additional option of validating the origin against the official published ranges.

Testing Googlebot from an ordinary IP and concluding that Google is blocked. Firewalls verify Googlebot's origin by reverse DNS, and stopping imitations is correct behavior on their part. That test measures the defense against spoofing, not the real Google's access.

Treating Google-Extended and Applebot-Extended as crawlers. They are tokens that exist only as a robots.txt directive, to control use in training; no request reaches the server under those names. Putting them on a firewall allowlist has no effect at all.

Blaming the hosting provider before bringing evidence. In this case, the hosting provider's support team diagnosed and fixed both layers in less than 24 hours, precisely because the ticket arrived with the IP, the times, the user agents and the error types. The factory default of a third-party rule set is nobody's bad faith; it is a layer that needs review.

What access does not guarantee

This article argues an infrastructure thesis, and honesty requires marking out what it does not cover. Allowing AI crawlers does not guarantee a citation in any answer. With the door open, the work that decides the rest begins: content that answers real questions with extractable clarity, entity consistency, verifiable sources, authority built over time and the competition on each question that matters. It is the difference between a necessary condition and a sufficient one, and anyone who promises guaranteed citation is selling what they do not control.

What access guarantees is the opposite: without it, everything else is irrelevant. No editorial quality gets through a connection that the server closes in silence. That is why, in the B.I.N.A. Method, verifying crawler accessibility is a foundation step, ahead of the content work, and the complete GEO guide covers the editorial layer that comes after it. The two layers complete each other; neither replaces the other.

There is also an invitation to independent verification: the protocol is described in full in this article, it is reproducible with free tools and it does not depend on taking our word for it. Run it on your site. If everything responds 200, you have gained a valuable confirmation. If it does not, you have just discovered something no dashboard was going to tell you.

Frequently asked questions

What is AI crawler blocking?

It is any barrier that prevents the crawlers of artificial intelligence systems, such as ClaudeBot, GPTBot and PerplexityBot, from accessing a site's pages. It can be intentional and visible, declared in robots.txt, or invisible, applied by firewalls, security rule sets and server configurations that the site owner is not even aware of. In the case documented in this article, a commercial ModSecurity rule dropped ClaudeBot's connections without returning any HTTP response, which made the entire site unreachable for that crawler without triggering an alert in any traditional SEO tool.

How do I know whether my site is blocking ClaudeBot or GPTBot?

Test from outside the server with curl, sending each crawler's user agent to three URLs: robots.txt, the home page and a rarely visited internal page, which tends to be outside the cache. Compare with a control request using a browser user agent. Status 200 indicates that access is open. Status 403, 429 or a connection closed without a response indicates a block or a limit. Repeat the confirmation test after a few minutes, because the first blocked request can temporarily ban your IP and contaminate the following attempts, producing false results.

Is blocking AI crawlers in robots.txt the same as blocking them at the firewall?

No, and the difference is the heart of the problem. The robots.txt file is a declared editorial choice: the site owner decides and publishes the decision, and well-behaved crawlers follow it. Blocking by a firewall, a WAF or a server rule happens before the request reaches the site, often as the factory default of a third-party rule set, without the owner's knowledge. The first is content governance. The second is a silent infrastructure failure that can contradict the intention declared in robots.txt itself, as in the case reported in this article.

Does allowing AI crawlers guarantee that my brand will be cited in the answers?

No. Access is a necessary condition, not a sufficient one. Once access is open, citation comes to depend on the quality of the content, the clarity of the answers, entity consistency, the sources, the authority built and the competition on each question. What blocking guarantees is the negative result: without access, no indexing or citation is possible, however good the content is. Be wary of any promise of guaranteed citation in AI systems. Serious GEO and AEO work treats access as the foundation and citation as a likely consequence, never as a contractual deliverable.

Why does my site appear on Google but not in AI answers?

Because access is evaluated crawler by crawler, and blocking tends to be asymmetric. In the case documented in this article, Googlebot, GPTBot, OAI-SearchBot, PerplexityBot and Bingbot accessed the site normally with status 200, while only ClaudeBot was dropped by a specific firewall rule. The site remained indexed on Google and triggered no alert in Search Console, which only reports Google's view. Each AI platform has its own crawler, and presence on one says nothing about presence on the others.

Does rate limiting get in the way of crawling by AI crawlers?

It can get in the way decisively. Crawlers do not request a single isolated page: on discovering a site, they crawl several URLs in sequence. A request limit that is set too low responds 429 on the very first pages, interrupts the crawl and delays the crawler's return, with a practical effect similar to that of a block. In the case reported, the 429 appeared from the fifth request in three minutes onward, but the cause was not a limit: it was a second user agent blocklist. The right adjustment preserves the protection against abuse and, at the same time, lets legitimate crawlers go through the site at a normal pace.

Next step

Does your site respond 200 to the ones that matter?

The B.I.N.A. Diagnosis checks the foundation before the content: AI crawler accessibility, blocking layers, structured data and readiness for answer engines. Ranking is not enough. Be the answer.

Guto Bertoncini is the founder of Flowup Agency, a Digital Authority Engineering company based in São Paulo, Brazil, and is responsible for the SEO, GEO and AEO strategy of the agency's projects. This article documents a real case conducted by the team in August 2026.

Transparency note: the case is reported firsthand, with dates, error codes and the sequence of events preserved from the original records. The client and the hosting provider are not named, as a matter of editorial standard; the hosting provider involved diagnosed and fixed the problem in less than 24 hours after receiving the ticket with the evidence. The market figures cited carry their source and date in the paragraph itself and reflect measurements for the periods indicated, subject to monthly variation.

Sources and references

  1. Cloudflare. Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large. Official press release, July 1, 2025. cloudflare.com/press
  2. Cloudflare Blog. Your site, your rules: new AI traffic options for all customers. July 2026. blog.cloudflare.com
  3. Cloudflare Blog. A deeper look at AI crawlers: breaking down traffic by purpose and industry. August 2025. blog.cloudflare.com
  4. TechnologyChecker. Web Traffic Statistics 2026, independent compilation of Cloudflare Radar data, July 2026. technologychecker.io
  5. TechnologyChecker. Bot Traffic Statistics 2026, compilation of Cloudflare Radar data. technologychecker.io
  6. TechnologyChecker. We Analyzed robots.txt Across Cloudflare’s Network, compilation of Cloudflare Radar data, 2026. technologychecker.io
  7. IETF. RFC 9309: Robots Exclusion Protocol. September 2022. datatracker.ietf.org
  8. Anthropic. Does Anthropic crawl data from the web, and how can site owners block the crawler? Official documentation. support.claude.com
  9. OpenAI. Overview of OpenAI crawlers. Official documentation. platform.openai.com
  10. Perplexity. Perplexity crawlers. Official documentation. docs.perplexity.ai

About the author

Portrait of Guto Bertoncini

Guto Bertoncini

Founder and lead strategist, Flowup Agency

Guto Bertoncini is the founder and lead strategist of Flowup Agency, which he has run since 2011. He is the author of the B.I.N.A. Method, Novo SEO and the Base Informacional Semântica (Semantic Information Base), and leads the agency's SEO for AI, GEO and AEO practice, preparing companies to be found on Google and cited by artificial intelligence platforms. He writes about search and AI on the Flowup blog and on his official website.

Keep reading

Marketing for Engineering and B2B Companies

In engineering and technical B2B, marketing has to prove competence before the first sales contact. An approach built on trust, digital authority, SEO, GEO and AEO.

Related content