What the data shows about how brands get cited by AI: the tactics from the Princeton study (up to 40% more visibility), the weight of external sources and a practical checklist.
By , founder and lead strategist at Flowup
What the data shows about how brands get cited by AI: the tactics from the Princeton study (up to 40% more visibility), the weight of external sources and a practical checklist.
By , founder and lead strategist at Flowup

Direct answer
To be cited by ChatGPT, Gemini and Perplexity, a brand has to win on three fronts at the same time: citable content (the Princeton GEO study, published at KDD 2024, showed that adding statistics, source citations and expert quotes raises visibility in generative answers by up to 40%), authority beyond its own website (Reddit, YouTube, LinkedIn and Wikipedia dominate the rankings of cited sources in the 2025-2026 studies) and a consistent entity identity, so that the systems know who the brand is on any surface. There is no such thing as a bought or guaranteed citation: what exists is probability, and it is built, as the GEO methodology details.
Before optimizing, you need to understand what is being optimized. When a user puts a question to a system with web access, three things happen in sequence: the system retrieves a set of candidate sources (searching the web or its own index), selects the passages that best answer the question and synthesizes an answer, attributing citations to some of the sources it used. On Google, this retrieval unfolds into several related searches, the so-called query fan-out.
A brand can be eliminated at any of the three stages: if it is not retrieved, it does not exist; if it is retrieved but its passages cannot be extracted, it is not selected; if it is selected but another source says the same thing more precisely, the citation goes to the competitor.
There is also a second path, with no retrieval: the model’s internal knowledge, formed in training. There, the brand “exists” to the extent that it appears consistently in the corpus that trained the system, which explains why a broad and lasting presence in public sources carries weight even when there is no real-time search. The two paths reinforce each other, and that is why serious AI visibility work tackles both: today’s retrievable content and the entity trail that feeds tomorrow’s training runs.
The field has a rare advantage: there is a peer-reviewed academic study at its foundation. The paper “GEO: Generative Engine Optimization” (Aggarwal et al., published at KDD 2024, with authors from Princeton and IIT Delhi) tested nine content modification strategies on a benchmark of 10,000 queries and measured the impact on visibility inside generative answers.
The central result: the winning tactics raised visibility by up to 40%, and the winners were precisely the ones that make a text more verifiable: adding statistics, citing sources and including expert quotes, with fluent writing also producing meaningful gains. “Keyword stuffing” tactics, inherited from the worst of SEO, ended up among the losers.
The second layer of evidence comes from the large-scale citation studies, which observe what the systems actually cite in production. The analysis by Semrush, with 230,000 prompts tracked for 13 weeks in 2025, and the Peec AI analysis of 30 million sources, summarized by Search Engine Land in 2026, converge on the same pattern: Reddit, YouTube, LinkedIn, Wikipedia and editorial outlets dominate the top of the cited sources: platforms where a brand can be present, but which it does not control.
The third layer calibrates ambition: Evertune’s analysis of 200 million prompts, compiled by Contently in 2026, showed that even the most-cited domain rarely exceeds 5% of total citations on any platform: the other ~95% are spread across thousands of domains. This is structurally different from classic SEO, where the top concentrates most of the clicks, and it is good news for mid-sized brands: the tail is long, and there is citable room in practically every niche.
The first front is your own website. The goal is not to “rank” in the classic sense: it is to produce passages that a synthesis system can extract, verify and attribute. The practices below derive directly from the winning tactics of the Princeton study and from the patterns observed in the citation studies.
Replace vague claims with sourced numbers. “Many companies have this problem” is invisible to a synthesizer; “93% of companies show errors in AI answers, according to study X (2025)” is a citable passage. Adding statistics was among the best-performing tactics in the Princeton benchmark. Cite quality external sources inside your content: paradoxically, pointing to trusted third parties increases the chance that your text is the one cited, because it signals verifiability.
Structure for extraction: each section should open with a paragraph that answers its implicit question in a self-contained way; use comparison tables, clear definitions and FAQs with complete answers. Write fluently: the study showed that clear prose, on its own, improves visibility: dense, rambling text works against the source. And eliminate generic content: AI synthesis collapses interchangeable texts into composite answers with no attribution; only the specific survives with the author’s name attached: the same math that doomed volume SEO.
The citation studies are unequivocal: third-party platforms weigh as much as (in many industries, more than) the brand’s own website. The strategic implication is uncomfortable for anyone who invests only in their own domain: a decisive part of your AI visibility is built on territory you do not control.
There is even quantitative evidence of the relationship: an SE Ranking study of 129,000 domains, cited in Contently’s compilation, found that domains with many brand mentions on Reddit received on average 3.9 times more ChatGPT citations than domains with minimal presence on the platform: a strong correlation, which does not prove causation, but lines up with what the source rankings suggest.
The work on this front has three rules. First: legitimate presence, not planted presence. Communities such as Reddit filter out self-promotion within hours, and the studies themselves offer signs that the systems favor threads of authentic discussion. Taking part for real (answering the industry’s questions, publishing your own data, contributing frameworks) is the only path that scales.
Second: choose your industry’s platforms. LinkedIn and review platforms such as G2 show up strongly in B2B queries (Perplexity emphasizes them, according to the Peec AI analysis); YouTube dominates where the answer is visual or a tutorial. Third: data-driven public relations. Editorial outlets are at the top of the citations on every platform, and the way into them is to give journalists what AI systems also prefer: original numbers, proprietary research, quotable experts.
The first two fronts fail if the systems cannot answer a prior question: which entity is this? A brand with an inconsistent name across platforms, without structured data and without an official factual layer forces the AI to guess, and AI systems that guess get things wrong or leave the brand out.
The minimum infrastructure: Organization structured data on the website, with the sameAs property pointing to every official profile; absolute consistency of name, description and factual data (founding, headquarters, leadership, public figures) across the website, LinkedIn, Wikipedia where applicable, directories and the press; factual reference pages that answer precisely what the company is, what it does and for whom, the instrument we call the official machine-readable knowledge base; and identifiable authorship, with named, verifiable experts signing the content.
A warning about priorities: this infrastructure is not a matter of fashionable files. The llms.txt file alone won’t make AI recommend your brand, and treating it as a shortcut diverts resources from what the data actually associates with citation.
The foundation is common to the three platforms, but the citation studies record differences in emphasis that are worth fine-tuning for, with the caveat that these patterns change within weeks, as the volatility documented by Semrush demonstrates.
| System | What the studies observe | Practical adjustment |
|---|---|---|
| ChatGPT | Heavy reliance on Wikipedia, Reddit and editorial outlets; citation patterns that are highly volatile from one period to the next (Semrush, 2025) | Prioritize the entity’s editorial and encyclopedic presence; monitor monthly, because the preferred sources change fast |
| Gemini / AI Overviews / AI Mode | Google ecosystem: heavy weight of YouTube and user-generated content; after September 2025, YouTube, Reddit and Facebook were the ones that grew the most in AI Mode (Semrush) | Treat video as a citation asset; keep the website strong in classic search, which feeds the retrieval behind the AI features |
| Perplexity | Emphasis on Reddit, LinkedIn and review platforms such as G2 in B2B queries (Peec AI, 2026) | For B2B: executive profiles and content on LinkedIn, presence in technical communities and active review management |
The sequence below organizes the three fronts into a plan you can execute. The order matters: measure before changing, foundations before expansion.
Honesty about the return: AI citation is not a traffic machine. The Pew Research Center measured, in a behavioral study of 68,879 real searches (data from March 2025, published in July 2025), that the links cited inside the AI summary get a click in only 1% of visits. The value lies elsewhere: in the presence inside the answer the user actually reads, the consideration layer where brands are included or excluded before any click.
And the clicks that do come are better ones. Seer Interactive found 35% more organic clicks for brands cited in the AI summary compared with those not cited on the same results page, and Semrush data indicates that a visitor coming from AI converts about 4.4 times more than one from traditional search. Less volume, more intent. Anyone who measures the strategy by sessions will underestimate it; anyone who measures it by presence in the answers and by the quality of the traffic sees the real return.
By working on the sources ChatGPT consults, not on ChatGPT itself. When the model browses, it searches the web and favors pages with specific data, cited sources and a clear structure, the tactics that the Princeton GEO study (KDD 2024) showed raise visibility by up to 40%. And it leans heavily on external platforms: Wikipedia, Reddit and editorial outlets appear among its most cited sources in the 2025-2026 studies. In practice: publish citable content on your own website, keep a legitimate presence in the communities and platforms of your industry and consolidate the entity’s identity so that the model knows who you are.
There is no guaranteed timeline, and be wary of anyone who promises one. Answers with web browsing can reflect new content within days or weeks, because they depend on real-time retrieval; the models’ internal knowledge, by contrast, changes in longer training cycles. The citation studies also show high volatility: Semrush documented Reddit falling from about 60% to 10% of ChatGPT’s answers in six weeks. That is why the reasonable working horizon is a matter of months, with continuous measurement, and the right expectation is to increase the probability of citation, never to guarantee it.
No. The llms.txt file is a proposed guide file for language models, but there is no evidence that the main systems use it as a citation criterion: Google has stated publicly that it does not use it. What the data associates with citation is something else: content with statistics and sources, an extractable structure, entity authority and presence on the platforms that AI systems consult. The file does no harm, but treating it as the main lever means investing in the wrong place.
It generates little direct traffic, but better traffic, and a lot of value beyond the click. The Pew Research Center measured that the links cited inside the AI summary get a click in only 1% of visits. On the other hand, Seer Interactive found 35% more organic clicks for brands cited in the summary compared with those not cited, and Semrush data indicates that a visitor coming from AI converts about 4.4 times more than one from traditional search. Citation works mainly as a consideration layer: the brand enters the answer the user reads, even when the user does not click.
It depends on where your audience is, but the data offers clues: ChatGPT concentrates the largest volume of use (in Brazil, Flowup’s home market, about 99% of the generative AI market, according to a survey by Cadastra with Similarweb), while Google’s AI Overviews and AI Mode reach Search’s base of billions of users. Perplexity has a smaller volume, but a qualified research audience and a strong reliance on Reddit, LinkedIn and review platforms in B2B queries. The good news: all three favor the same fundamentals (citable content, a clear entity and external authority), so the foundation of the work is shared, with fine-tuning by platform.
Not in the organic answers of the main systems. The citations in ChatGPT, Gemini, Perplexity and AI Overviews are selected by the retrieval and synthesis mechanisms, not sold. What does exist are ad formats around AI experiences (in Portuguese), which are labeled paid media: a different thing. The real route to organic citation is to build what the systems select: verifiable content, entity authority and presence in the sources they consult. Any offer of a “guaranteed citation” in exchange for payment deserves immediate skepticism.
Next step
The B.I.N.A. Diagnosis answers exactly that: it maps how your brand appears today in ChatGPT, Gemini, Perplexity and AI Overviews, identifies who occupies the space that should be yours and prioritizes the actions from the three fronts of this guide for your case. Ranking is not enough. Be the answer.
Founder and lead strategist, Flowup Agency
Guto Bertoncini is the founder and lead strategist of Flowup Agency, which he has run since 2011. He is the author of the B.I.N.A. Method, Novo SEO and the Base Informacional Semântica (Semantic Information Base), and leads the agency's SEO for AI, GEO and AEO practice, preparing companies to be found on Google and cited by artificial intelligence platforms. He writes about search and AI on the Flowup blog and on his official website.
Sources verified on July 30, 2026. The evidence in this guide is of different kinds: (1) the Princeton study (KDD 2024) is peer-reviewed academic research, conducted in a controlled environment that emulates generative engines: the “up to 40%” figure refers to the visibility metrics of the study’s benchmark, not to guarantees in production; (2) the citation studies (Semrush, Peec AI, Evertune, SE Ranking) measure specific time windows, with different methodologies and samples, and show documented high volatility: they even diverge on which domain leads; (3) correlations, such as the one between Reddit mentions and ChatGPT citations, do not establish causation; (4) the behavioral data (Pew, Seer) measures AI Overviews, not all platforms. This field changes within weeks: the recommendations rest on the patterns that converge across studies, not on isolated numbers.