SEO and AI

Why llms.txt Alone Won’t Make AI Recommend Your Brand

The promise was irresistible: place a text file on your server and artificial intelligences will start understanding — and citing — your website. Two years after the proposal, there is enough server data, official statements, and correlation research to answer with evidence: that is not how citation works. This article shows what llms.txt really is, what it is for, what it is not for — and what the data points to as the real predictors of AI recommendation.

Por

Pessoa digitando em um teclado sob um cérebro digital com a sigla LLM, representando a discussão sobre se o llms.txt funciona.
Direct answer

llms.txt won't make AI recommend your brand because no major AI platform documents consuming it — and telemetry confirms it: in an Ahrefs analysis of roughly 137,000 domains, 97% of llms.txt files never received a single request, and statistical models found no effect on citations. Google officially states that its Search, including generative AI features, does not use this type of file. What independent studies point to as the real predictor of citation is something else: distributed semantic authority — consistent brand mentions in sources the systems trust, with roughly three times the correlation of backlinks. Engineering without brand architecture is empty plumbing.

What llms.txt is — and what it promised

llms.txt is a proposed standard published in September 2024 by Jeremy Howard, co-founder of Answer.AI. The idea: a Markdown file at the site root (/llms.txt) working as a curated index for AI systems — the brand's essential content, stripped of navigation, ads, and interface noise, easy to consume for models with limited context windows.

The original intent was reasonable and the problem it describes is real: modern HTML pages carry a lot of noise for a system that only wants the content. The proposal, however, carries a built-in premise rarely said out loud: it only works if AI systems decide to read the file and trust it. It is a unilateral proposal — the site declares; nobody committed to listening.

The comparison with successful standards is instructive. robots.txt has decades of adoption and a clear enforcement mechanism. Schema.org was born from a coalition of Google, Bing, Yahoo, and Yandex — the standard's consumers participated in its creation. llms.txt was born on the publishers' side, with no major consumer committed. That difference of origin explains almost everything that came next.

What the platforms say — in the documentation, not on stage

Documented platform positions on llms.txt — verified July 2026
PlatformVerifiable position
Google The official AI features optimization guide, updated June 2026, states you do not need to create machine-readable files, AI text files, markup, or Markdown to appear in Search — including generative AI — because Search does not use them. John Mueller of the Search team stated in June 2025 that no AI system used llms.txt, compared it to the old keywords meta tag (abandoned for being easy to manipulate), and in 2026 called any benefit purely speculative.
OpenAI, Anthropic, Perplexity They document their crawlers with precision (what each collects, how to block it), but none documents consuming llms.txt as an input for search, training, or citation. The absence is eloquent: companies that specify their bots' behavior in detail do not mention the file.
Technical community Registered objections include the cloaking risk (showing bots a different version from what humans see — exactly the kind of manipulation that made search engines abandon self-declared signals) and the poor experience of citations pointing to raw text files, without navigation or context.

Mueller's parallel with the keywords meta tag deserves a paragraph. Search engines abandoned that tag because signals declared by the site itself, at zero cost and with no verification, are infinitely inflatable — and therefore useless for ranking trust. llms.txt repeats the structure of the problem: if every site can publish a flattering manifesto about itself, the manifesto differentiates no one. Modern answer systems select sources by signals the site does not control alone.

What the servers show: the file nobody reads

If official statements left any doubt, telemetry resolves it. Ahrefs analyzed real request behavior across a base of approximately 137,000 domains with llms.txt implemented. The findings:

  • 97% of the files never received a single request. It's not that bots read and ignored them — they never fetched them.
  • Of the requests that did exist, most came from SEO audit tools and non-AI crawlers — meaning the ecosystem that sells the tactic is its main consumer.
  • Statistical models applied to the base found no effect of the file on citations in AI answers.

Meanwhile, the same server logs show what AI bots actually request: the site's normal HTML pages, for training and for answer grounding. The operational conclusion is direct — the systems already read your site; what decides the citation is not the file format, it is what they find when they read.

Rigor note: Ahrefs data measures requests and correlations, not platforms' future intentions. Some consumption may emerge later — Mueller himself sums up the current state as speculation. Budget decisions, however, are made on observed behavior, not on tomorrow's lottery.

Where llms.txt has legitimate use — the part the hype leaves out

An honest article must register the other side: sophisticated tech companies — Stripe, Cloudflare, Vercel, Anthropic, among others — maintain llms.txt files. They are not naive. They are precise about the file's target audience: coding agents and agentic browsers that, in specific development workflows, choose to consume technical documentation in a clean format.

That is a real, narrow use case: API and technical product documentation, consumed by tools the developer points at the file. It has no relation to the scenario sold in the GEO market — that ChatGPT or AI Overviews would read the manifesto and start recommending the brand in commercial answers. Confusing the two scenarios is the central error of the discourse that overrates the tactic.

The distinction, summarized

llms.txt for technical documentation consumed by coding agents: legitimate, niche, verifiable. llms.txt as a citation tactic for brands in AI answers: no evidence of effect, no documented consumption, no requests in the logs. Implementing it does no harm — the cost is low and there is no penalty. The problem is one of priority: treating the deliverable with the least evidence of impact as the project's centerpiece.

Why the tactic sold so well

If the evidence is so weak, why did llms.txt become a line item in so many GEO proposals? Because it has the three properties of the perfect tactic to sell — and none of the ones that make a tactic work:

  • It is binary and screenshot-able. "We deployed the file" fits a checklist and a delivery screenshot. Semantic authority does not.
  • It is cheap to execute and expensive to look like. Generating a Markdown file takes minutes; presenting it as "preparing the site for the AI era" sustains an invoice line.
  • It cannot be refuted by the client. Since the promised effect is diffuse ("AIs will understand the site better"), the absence of results is never attributed to the file.

The pattern is not new. Every technology transition produces its layer of technical talismans — simple-to-implement objects sold as shortcuts to a result that actually depends on structural work. Recognizing the talisman is the first filter for evaluating any AI visibility vendor: ask what evidence supports each deliverable and what exactly it measures after delivery.

What actually predicts citation: the data

The right question was never "which file to put on the server" — it is "what do AI systems use to decide whom to cite". Here the evidence base is richer, and three independent studies, with different methodologies, converge:

Predictors of visibility in AI answers — independent studies, 2025–2026
SourceMethodCore finding
Ahrefs — 75,000-brand study (2025) Correlation between brand signals and AI Overview presence Web brand mentions: 0.664 correlation — versus 0.218 for backlinks, roughly 3x. YouTube mentions lead (0.737); branded anchors (0.527) and brand search volume (0.392) complete the top. Brands in the top quartile of mentions earned about 10x more AI Overview presence than the next quartile; the bottom half was practically absent from answers.
Muck Rack — 1M+ citations (2025–2026) Classification of the origin of citations in AI answers 82% of citations came from earned media and 94% from non-paid sources. Systems prefer what third parties say about the brand over what the brand publishes about itself.
Stacker / Scrunch (2025–2026) Controlled study across five LLMs Distributing content through third-party editorial outlets produced a 239% median lift in AI visibility, with cases above 300%.
Profound / Semrush (2025–2026) Mapping of the most-cited sources by platform Answers lean disproportionately on a small set of high-trust sources — communities, encyclopedias, video, professional networks — with different preferences per platform: what grounds ChatGPT is not what grounds AI Overviews or Perplexity.

Two method caveats, because rigor is part of the thesis. First, correlation is not causation — Ahrefs' own researchers stress that the data shows AI-visible brands also have broad web presence, not that manufacturing mentions mechanically generates citations. The underlying signal is brand authority as a composite, absorbed by the systems over time. Second, the AI measurement market is young and mostly operated by tool vendors; the numbers should be read as order of magnitude and direction, not physical constants.

Caveats made, the direction is unambiguous: all the strong signals live outside the brand's server. No independent study found, among the relevant predictors, any self-declared technical artifact.

The mechanism: how AI decides whom to cite

  1. The question triggers a search

    The major answer systems ground their responses with classic web search. Whoever is strong in organic enters the candidate set; whoever isn't indexable doesn't even exist for this stage.

  2. The system reads a small set of sources

    From the search, the model receives a handful of high-trust pages — frequently editorial outlets, communities, encyclopedias, reviews. This is where earned media weighs: the brand needs to be present inside the pages the system selects, not only on its own site.

  3. The model synthesizes — and cites what repeats with consistency

    If the brand appears consistently described across multiple independent sources, the model treats it as the probable answer. If it appears rarely, or described in contradictory ways, it is diluted in the synthesis. Recurrence + consistency + source trust: that is the tripod the correlation studies capture.

Seen through this pipeline, llms.txt's limit becomes obvious: it acts at a point that does not participate in the decision. The file doesn't improve the brand's position in the search that grounds the answer, doesn't insert it into the third-party sources the model reads, and doesn't increase the consistency with which the ecosystem describes it. It is plumbing — in an architecture problem.

The strategic reading: semantic authority is an asset, the file is an accessory

For the decision-maker, the conclusion matters more than the technicality. AI visibility is not a technical compliance problem solved with a file checklist — it is an asset-building problem: making the brand the entity the digital ecosystem describes with clarity, consistency, and proof, in the sources the systems trust.

  • Distrust GEO projects whose central deliverable is a self-declared technical artifact. If the proposal revolves around files, tags, and manifestos, it is optimizing the point of the pipeline that decides nothing.
  • Demand the authority layer in scope. Qualified mentions, editorial presence, publishable proprietary data, visible experts, entity consistency across channels — that is where the studies locate the effect.
  • Require measurement of what matters. Not "file deployed," but brand presence in real answers, per platform, compared against competitors, over time.

What to do instead, in order

  1. Secure the foundation the systems actually read. Crawlability, indexing, performance, and coherent structured data on the HTML pages — the true input of grounding.
  2. Structure content as answers. Real market questions, direct answer up front, depth and proof after it — the format that survives synthesis.
  3. Consolidate entity identity. The brand described the same way — what it does, for whom, with what differentiator — on the site, in the press, in reviews, directories, and profiles. Inconsistency dilutes citation.
  4. Build mentions where the systems read. Editorial presence, proprietary data that earns coverage, qualified participation in communities and review platforms relevant to the category.
  5. Measure citations continuously. Search Console (generative AI experiences), Bing Webmaster Tools (Copilot), and structured prompt tests with declared methodology.
  6. Only then, if you want, publish llms.txt. At the end, for completeness and at zero marginal cost — never as the center of the project.

Frequently asked questions

What is llms.txt?

A proposed standard created in September 2024 by Jeremy Howard (Answer.AI): a Markdown file at the site root offering AI systems a curated content index, free of navigation and ads. Unlike robots.txt, it is not an official standard, has no governing body, and no major platform documents consuming it.

Does Google use llms.txt?

No. The official documentation, updated in June 2026, states there is no need to create machine-readable files, AI text files, markup, or Markdown to appear in Search — including generative AI. John Mueller stated in 2025 that no AI system used the file, compared it to the keywords meta tag, and in 2026 called any benefit purely speculative.

Do AI bots read llms.txt?

As a rule, no. Ahrefs telemetry across ~137,000 domains found 97% of files with no requests at all; the few that existed came mostly from audit tools, not AI bots. Platform crawlers keep requesting normal HTML pages for training and grounding.

So is llms.txt useless?

It is irrelevant for the goal it is usually sold for — brand citations in AI answers. It has a legitimate niche in technical documentation consumed by coding agents (which is why companies like Stripe, Cloudflare, and Anthropic maintain it). Implementing it does no harm; the mistake is treating it as a priority or central deliverable.

What actually gets a brand cited by AIs?

Distributed authority: web brand mentions (0.664 correlation vs. 0.218 for backlinks, Ahrefs/75,000 brands), earned media (82% of citations, Muck Rack/1M+ citations), and presence in each platform's high-trust sources. A Stacker/Scrunch experiment measured +239% median visibility after editorial distribution through third-party outlets.

What should a GEO project prioritize?

In order: a real technical foundation; answer-format content with proof; entity consistency across channels; building qualified mentions where the systems read; and continuous citation measurement. A text file, if it enters at all, enters last — for completeness, not strategy.

Does your brand have a file — or authority?

Flowup's B.I.N.A. Diagnostic measures the four layers that decide whether search engines and AIs find, understand, and cite your company — from technical foundation to distributed authority. Immediate result, in-depth report by specialists.

Take the free diagnostic

About the author

Guto Bertoncini is the founder of Flowup Agency, where he leads Digital Authority Engineering projects — the integration of SEO, GEO, AEO, technology, and brand that prepares companies to be found, understood, and cited by Google and by artificial intelligence platforms alike.

Methodology note

Sources verified on July 30, 2026. Claims about platform positions rest on official documentation and named public statements by spokespeople; claims about bot behavior rest on published server telemetry with a declared sample base; claims about citation predictors rest on correlation studies and controlled experiments, with the limitations indicated in the body (correlation ≠ causation; a young measurement market mostly operated by tool vendors). Where the evidence is speculative — including possible future adoption of the format — the text identifies it as such.

References

  1. Google Search Central (2026). AI features optimization guide — June 2026 update on machine-readable files. developers.google.com
  2. John Mueller / Google (2025–2026). Public statements on the non-use of llms.txt by AI systems. Coverage: Search Engine Roundtable and Search Engine Journal.
  3. Ahrefs (2025–2026). llms.txt request telemetry across ~137,000 domains. ahrefs.com
  4. Ahrefs (2025). Correlation study between brand signals and AI Overview presence — 75,000 brands. ahrefs.com
  5. Muck Rack (2025–2026). Origin analysis of over 1 million citations in AI answers. muckrack.com
  6. Stacker / Scrunch (2025–2026). Controlled study on the effect of editorial distribution on visibility across five LLMs.
  7. Profound / Semrush (2025–2026). Mappings of the most-cited sources by answer platform.
  8. Jeremy Howard / Answer.AI (2024). Original llms.txt standard proposal. llmstxt.org
Tags:
AI visibilityGEOllms.txtSEO for AI

Leia também

Conteúdo relacionado

SEO and AI

The End of the Click: Why Your Brand May Become Invisible in the Age of Answers

For two decades, Google sent the customer to your website. Now, artificial intelligence reads your website — and delivers the final answer without anyone needing…
Ler artigo
SEO and AI

Generic Content Is a Liability: The Real Cost of Volume SEO

For a decade, publishing more was the winning strategy: more articles, more keywords, more traffic. Then the cost of producing generic text fell to zero…
Ler artigo
keyboard_arrow_up