SEO and AI

How to Measure Whether AI Recommends Your Company: Real Metrics and Tools

A four-layer framework to measure AI visibility: presence and citation, accuracy, referral traffic and conversion, with the metrics, the tools and a free manual protocol.

By , founder and lead strategist at Flowup

Direct answer

Measuring whether AI systems recommend your company takes four complementary layers: (1) presence and citation: a fixed set of 20 to 30 real questions run every month in ChatGPT, Gemini, Perplexity and AI Overviews, recording presence rate, citation rate and share of voice against competitors; (2) accuracy: is what they say correct?; (3) referral traffic: in GA4, filtering sources such as chatgpt.com and perplexity.ai (knowing that part of it arrives as direct traffic); and (4) conversion: this traffic converts about 4.4 times more than traditional organic, according to Semrush. No tool covers all four; the manual protocol covers the first two for free. The starting point is a baseline like the one in the B.I.N.A. Diagnosis.

Why measuring AI is different from measuring search

In SEO, measurement inherits two decades of infrastructure: Search Console delivers impressions, clicks and positions straight from the official source. In the AI layer, three characteristics break that model. First: answers are not deterministic. The same question produces different answers depending on the wording, the moment, the location and the model version: there is no such thing as “the position” of your brand; there is a probability of appearing. Second: there is no official data source. No AI platform offers a complete “Search Console for citations” today; Search Console itself includes AI Overviews and AI Mode data in the Performance report, within the Web search type, with no separate breakdown, according to Google’s documentation. Third: the dominant behavior is reading without clicking. Pew Research Center found that the sources cited inside the AI summary get a click in just 1% of visits, so the classic metric (sessions) sees only the edge of the phenomenon.

The consequence: anyone who tries to measure AI with search metrics concludes, wrongly, that “nothing is happening.” The right measurement is probabilistic, sample-based and trend-oriented, like market research, not like server analytics.

The framework: four layers of measurement

The complete program answers four questions in sequence, each with its own data source:

Layer Question it answers Data source Frequency
1. Presence and citation Does the brand appear in the answers? Is it cited as a source? In what share of the market’s answers (share of voice)? Standardized prompt set (manual or run by a tool) Monthly
2. Accuracy Is what AI systems say about the company correct and complete? Qualitative analysis of the answers from the same prompt set Monthly
3. Referral traffic How many visits do AI systems send, and to which pages? GA4 (referral sources) + server logs Continuous
4. Conversion and value What are this traffic and this presence worth to the business? GA4 (conversions by source) + brand indicators Monthly/quarterly

Layer 1. Presence and citation: the manual protocol

This is the most important layer and the only one that needs no tool. The protocol, step by step:

  1. Build the prompt set. Between 20 and 30 questions that real customers would ask, covering the funnel: category definitions (“what is X”), comparisons (“best X providers for Y”), recommendations (“which company should I hire for X in Brazil,” to use our home market as the example) and direct questions about the brand. Use the customer’s language, not internal jargon: real prompts are long and contextual.
  2. Standardize the conditions. Clean sessions (no history and no memory feature turned on), the same systems every month (ChatGPT, Gemini, Perplexity and Google Search, watching AI Overviews and AI Mode), recording the date and the model version when visible.
  3. Log it in a spreadsheet, by prompt and by system: was the brand mentioned? Was it cited as a source (with a link)? At what position in the answer? Which competitors appeared? Which of your pages was cited?
  4. Calculate the metrics: presence rate (% of answers that include the brand), citation rate (% with a link to your domains), share of voice (your share against competitors in the same prompt set).
  5. Repeat every month without changing the prompt set: changing the questions destroys comparability. Review the set once a quarter and version it.

Why standardization matters so much: citation patterns are volatile by nature. The Semrush study of 230,000 prompts documented Reddit plunging from about 60% to 10% of ChatGPT answers between early August and mid-September 2025, with no cause confirmed by the study. In an environment like this, only the comparison between identical measurements separates a real trend from system noise.

Layer 2. Accuracy: what AI says when it talks about you

Presence without accuracy can be a liability: an AI that recommends your company while describing what it does incorrectly, quoting outdated prices or confusing it with a namesake produces tainted consideration. In the same monthly prompt set, classify each answer that mentions the brand into one of three levels: correct and complete, correct with relevant omissions or with a factual error, and record the specific error and the source the AI cited when it got things wrong, because that source is the one that will need to be fixed.

The error rate feeds the action plan directly: recurring errors almost always point to inconsistencies in the company’s public factual layer, such as descriptions that differ between the website and profiles, outdated data in directories, or the absence of a structured official source. This is the problem the official machine-readable knowledge base exists to solve: giving the systems a canonical factual layer to consult before they improvise.

Layer 3. Referral traffic: setting up analytics

In GA4, create an exploration report or a dedicated channel group, filtering the referral sources of the main systems: chatgpt.com (and the legacy chat.openai.com), perplexity.ai, gemini.google.com and copilot.microsoft.com. Track sessions, landing pages (which reveal which of your content the AI systems are linking to), engagement and conversions by source. If the company also advertises on ChatGPT, keep the visits that come from the ad separate from the ones that come from the answer: they are different decisions, and the guide on ChatGPT ads (in Portuguese) shows how to read both sides.

State the limitations in the report itself, so that no one reads the number as a total: part of the clicks coming from AI experiences arrives without a referrer and falls into direct traffic (a behavior documented in Google’s own chat experiences), which means the measured traffic is a floor, not the ceiling. And AI Overviews do not generate a distinguishable source of their own: their clicks blend into Google organic, with Search Console aggregating the impressions of AI features without separating them. Reading this layer together with layer 1 corrects the myopia: high presence with low traffic is not failure, it is the pattern of the environment, as Pew’s 1% figure establishes.

Layer 4. Conversion and value: the numbers that justify the program

This is the layer that turns measurement into budget. Two patterns show up consistently in market data. Traffic coming from AI is small in volume and disproportionate in quality: Semrush data points to a conversion rate about 4.4 times higher than that of traditional organic search, as compiled by industry coverage. And citation protects performance on the results page itself: Seer Interactive, analyzing 25.1 million impressions, measured 35% more organic clicks for brands cited inside the AI Overview compared with those not cited, with the click studies gathered and detailed by specialized coverage.

Complete the layer with indirect indicators that capture value without a click: the trend in branded searches (Search Console), mentions and, when volume allows, a comparison of close rates between leads who name AI as the source in the “how did you hear about us” field and everyone else. The program has a single executive indicator: AI share of voice on the questions that define your market, with the accuracy rate beside it. It is the share of the answer that your brand occupies: the rest of the funnel derives from it.

The metrics, operationally defined

Metric Operational definition Source
Presence rate % of the answers in the prompt set in which the brand is mentioned (with or without a link) Prompt set
Citation rate % of the answers with a link to the brand’s domains Prompt set
AI share of voice The brand’s mentions as a share of the total mentions of competing brands in the same prompt set Prompt set
Accuracy rate % of the mentions classified as correct and complete Qualitative analysis
AI referral traffic Sessions whose source is a domain of an AI system (a floor, not the total) GA4
AI traffic conversion Conversion rate of the sessions from that source vs. traditional organic GA4
Cited pages Which of the brand’s URLs the AI systems link to (guides content prioritization) Prompt set + GA4 (landing pages)

Tools: the map of a market still taking shape

Between 2025 and 2026 an entire category of AI visibility monitoring tools was born, and several of them are the sources of the studies cited in this article: Semrush (with its set of features for AI search), Peec AI (the analysis of 30 million sources), Evertune (200 million prompts), plus names such as Profound, Scrunch, Otterly and Ahrefs’ Brand Radar. In essence, they all automate layer 1: they run prompt sets at scale, extract mentions and citations and calculate share of voice.

Three selection criteria, and one caution. The criteria: system coverage (does it include the systems that matter to your audience, ideally with support for prompts in your customers’ language, which for us in Brazil means Portuguese?), methodological transparency (does it show the prompts, the frequency and the collection conditions?) and exportability (can the data be exported to your executive report?). The caution: this is a market still taking shape, with diverging methodologies. The studies use different samples and methods, and the result changes with the platform: in Peec AI’s analysis Reddit leads, followed by YouTube and LinkedIn, while Wikipedia is the most cited domain on ChatGPT, as Contently’s roundup shows. The golden rule: choose one methodology and compare it with itself over time; never mix numbers from different sources in the same series. And remember: the tool automates collection, not strategy. What to do with the numbers is still the work described in our guide on how to get cited by AI.

Frequently asked questions

Can you measure AI visibility without a paid tool?

Yes. And every serious program should start that way. The manual protocol: define 20 to 30 real questions from your market (the ones customers would ask), run them every month in ChatGPT, Gemini, Perplexity and Google Search, watching the AI Overviews, always in clean sessions, and log in a spreadsheet: did the brand appear? Was it cited as a source? Is what was said correct? Which competitors appeared? This produces the three core metrics (presence rate, citation rate and accuracy) at zero cost. Paid tools automate and scale this process, but they do not change its nature.

Does Google Analytics show the traffic that comes from ChatGPT?

It shows part of it. Visits that start with a click on a link in ChatGPT arrive with chatgpt.com as the referrer, and the same goes for perplexity.ai and gemini.google.com: in GA4, a report or channel group filtering these sources reveals the volume, behavior and conversion of this traffic. The limitations: part of the clicks coming from AI experiences arrives without a referrer and falls into direct traffic, which underestimates the total; and traffic is only one layer of the value, because most users read the answer without clicking: Pew Research measured a click rate of only 1% on the sources cited inside AI Overviews.

What is AI share of voice?

It is the share of your market’s AI answers in which your brand appears, compared with competitors. Operationally: over a fixed set of relevant prompts, you measure in how many answers each brand is mentioned or cited; each brand’s proportion is its share of voice. It is the metric that turns “are we in AI?” into a number you can compare and track, and the closest thing to an executive indicator in this layer. Important: because citation patterns change fast, share of voice should be read as a trend across standardized measurements, not as a definitive snapshot.

How often should I measure AI visibility?

Monthly for the prompt set and share of voice; continuously (through analytics) for referral traffic. The monthly pace balances two facts: citation patterns change within weeks (Semrush documented Reddit falling from about 60% to 10% of ChatGPT answers in a month and a half), but measuring more often than you act produces noise without decisions. The practical rule: same set of questions, same conditions, same spreadsheet, every month; and a quarterly review to update the questions as the market changes.

Why do citation numbers vary so much between tools and studies?

Because each one measures a different slice of a non-deterministic system. AI answers vary with prompt wording, moment, location, history and model version; the studies use different prompt sets, time windows and methodologies: that is why Reddit appears as the number 1 source across platforms, while Wikipedia leads on ChatGPT. This does not invalidate measurement; it demands method: standardize your conditions, compare periods within the same methodology and read trends, not absolute values. Be wary of any AI visibility number presented without a methodology.

Is there value in being cited without getting the click?

Yes. And that is probably where most of the value lies. A citation places the brand inside the answer the user actually reads, at the moment consideration is being formed; Pew Research found that the sources in AI Overviews get a click in only 1% of visits, which means reading without clicking is the dominant behavior. The effects show up in indirect indicators: branded searches, mentions, direct traffic and the quality of the little traffic that does come: Seer Interactive measured 35% more clicks for cited brands, and Semrush points to conversion about 4.4 times higher for traffic coming from AI. Measuring sessions alone systematically underestimates this layer.

Next step

Do you know what AI systems answer when people ask about your market?

The B.I.N.A. Diagnosis carries out the complete baseline for you: it runs your market’s questions in the main systems, measures presence, citation and accuracy, compares you with your competitors and delivers the priorities: the starting point of any serious measurement program. Ranking is not enough. Be the answer.

About the author

Portrait of Guto Bertoncini

Guto Bertoncini

Founder and lead strategist, Flowup Agency

Guto Bertoncini is the founder and lead strategist of Flowup Agency, which he has run since 2011. He is the author of the B.I.N.A. Method, Novo SEO and the Base Informacional Semântica (Semantic Information Base), and leads the agency's SEO for AI, GEO and AEO practice, preparing companies to be found on Google and cited by artificial intelligence platforms. He writes about search and AI on the Flowup blog and on his official website.

Methodology note

Sources verified on July 30, 2026. This guide distinguishes: (1) data from measurement studies (Pew Research, Seer Interactive, Semrush, Peec AI, Evertune), which cover specific windows and samples, mostly from English-language markets, and show documented methodological differences among themselves; (2) technical patterns of traffic attribution (domain referrers, referrer loss in chat experiences), described according to the documentation and specialized coverage available, and subject to change by the platforms without notice; (3) protocol recommendations, which are Flowup’s professional practice, not an industry standard. The conversion multiplier (4.4x) and the click multiplier (+35%) are third-party measurements in specific contexts: use them as an order of magnitude, not as a projection for your case. The framework was designed to produce comparable series; no single metric from this layer should support a decision on its own.

References

  1. Tyneside Marketing (2026). Detailed compilation of the click and CTR studies: Pew Research (68,879 searches; 8% vs 15%; 1% click rate on the cited sources), Seer Interactive (25.1 million impressions; +35% for cited brands) and Ahrefs. tynesidemarketing.co.uk/blog/ai-overviews-click-through-rates
  2. Semrush (2025). The Most-Cited Domains in AI: a 3-month study of 230,000 prompts that documents the volatility of citation patterns. semrush.com/blog/most-cited-domains-ai
  3. Contently (2026). Roundup that combines several citation studies (Peec AI, Semrush, Profound, SE Ranking and others), with the differences in method noted. contently.com/2026/04/29/top-sources-llms-cite
  4. Search Engine Land / Peec AI (2026). Analysis of 30 million sources cited by AI systems. searchengineland.com/ai-search-engines-cite-reddit-youtube-and-linkedin-most-study-473138
  5. PikaSEO (2026). Compilation of the data on AI traffic conversion (Semrush: 4.4x) and on zero-click behavior. pikaseo.com/articles/zero-click-search-ai-overviews-2026
  6. SmartLinks (2025). Documentation, in Portuguese, on measuring Google’s AI features: the aggregated data in Search Console and the loss of the referrer on clicks coming from chat. smartlinks.pt/google-ai-mode
  7. Search Engine Journal (2026). Consolidation of the studies on CTR and on the visibility of pages cited by AI Overviews. searchenginejournal.com/ai-overview-ctr-fell-61-but-clicks-didnt-collapse
Tags:
GEOSEO for AI

Keep reading

Marketing for Engineering and B2B Companies

In engineering and technical B2B, marketing has to prove competence before the first sales contact. An approach built on trust, digital authority, SEO, GEO and AEO.

Related content

SEO and AI

SEO and Social Media Integrated: What Changes Now That Google Measures Social Search

Google now measures Instagram, TikTok, X and YouTube in Search Console. The timeline of how SEO and social media were integrated and how to act, with data and sources.
Read article
SEO and AI

Data Governance Applied to Digital Marketing: The Guide for the Age of AI Answers

What data governance applied to digital marketing is, why it became urgent in the age of AI answers and how to implement it in five steps, with market numbers.
Read article
SEO and AI

YMYL (Your Money or Your Life): What It Is, How Google Evaluates It and How to Build Trust

A guide to YMYL: where the concept comes from in Google’s guidelines, how it relates to E-E-A-T, what it changes in SEO, GEO and AEO, and a 24-item audit checklist.
Read article
SEO and AI

Flowup Method vs. Traditional SEO: From Traffic to Answer Governance

What changes between traditional SEO and a GEO and AEO operation with data governance? The market standard and the Flowup Method compared, with numbers measured at the source.
Read article