One study found wrong or missing facts in AI answers for 93% of companies. See what an Official AI Knowledge Base is, how to build one, and what it can and cannot control.
By , founder and lead strategist at Flowup
One study found wrong or missing facts in AI answers for 93% of companies. See what an Official AI Knowledge Base is, how to build one, and what it can and cannot control.
By , founder and lead strategist at Flowup

Direct answer
An Official AI Knowledge Base is the canonical page a company publishes on its own domain as the source of truth about itself: verifiable identity, positioning, services, disambiguation, and official channels, in clear text for humans and structured data for machines. It exists because AI systems do not know your brand: they approximate it from scattered fragments, and they get it wrong with documented frequency. A study of more than 13,000 queries to ChatGPT, Perplexity, and Gemini found at least one basic fact hallucinated or missing for 93% of the companies analyzed, with smaller companies suffering 2.5 times as many confident fabrications (5% against 2%). The knowledge base does not buy a recommendation (that still depends on external authority), but it changes what it can change: it hands over the factual narrative ready-made, instead of leaving the models to assemble it on their own.
Run the test right now: ask an AI assistant who your company is, what it does, and what sets it apart. You will get a fluent, articulate, confident answer. What you will not see is where it came from.
Language models do not hold a registration record for your brand. They infer the entity from the signals available: website pages, press mentions, directories, profiles, reviews, public databases, everything that entered the training data or that search retrieves at the moment of the answer. When those signals are complete and consistent, the approximation is good. When they are fragmented, outdated, or contradictory (the case for the vast majority of companies), the model fills the gaps the way models fill gaps: by generating the most plausible text, not the truest.
The result is a new category of reputational risk: a company can be described incorrectly, at scale, without knowing it, to an audience that, more and more, never visits the site to check. Unlike traditional misinformation, the error needs no published source: it is generated on demand, with every question, in the authoritative tone that makes people believe it.
The most extensive research available on the topic is the Searchable study (2026), which ran more than 13,000 queries on ChatGPT, Perplexity, and Gemini about real London-based companies, comparing SMEs with brands of 500+ employees, and checked each answer against verified sources such as official records and institutional profiles.
| Measurement | Result |
|---|---|
| Companies with at least one basic fact hallucinated or missing | 93%: nearly the entire sample |
| Companies that received at least one fabricated fact | 50% of SMEs vs. 32% of large companies |
| Brand name confused or misattributed | 4% for SMEs vs. 0.7% for large companies: five times more |
| Services described accurately | 87% for SMEs vs. 93% for large brands |
| Fabrications presented in confident language | 5% for SMEs vs. 2% for large companies: the error comes with a tone of certainty |
Two findings deserve a close reading. The first is the asymmetry by size: smaller companies fare worse on every metric, and the researchers’ explanation points straight to the thesis of this article. Large brands have a wide digital footprint and third-party validation; smaller brands, even with an accurate, up-to-date website, give the models fewer signals to triangulate. The model does not err out of bad faith; it errs for lack of reliable raw material.
The second is tone: fabrications come wrapped in the same assertive language as the truths. For the person asking, there is no visual cue that the founding date, the partner’s name, or the list of services has just been made up. The confidence of the text lends credibility to the error.
AI errors about brands have existed since the first chatbot. What changed is the weight of the answer in the decision:
In the old model, the corporate page existed for the visitor who arrived. In the current model, there is a new and more frequent reader: the systems that answer for you when nobody visits. The strategic question is no longer “what does the visitor find on our site” but “what does a machine that has to describe our company in three paragraphs find to lean on.” The AI Knowledge Base is the institutional answer to that question.
An Official AI Knowledge Base is the canonical reference a company publishes about itself, on its own domain, written in two layers:
@id, connects the official profiles via sameAs, and removes the ambiguity that forces models to guess.The design principle is simple to state and demanding to execute: every factual question that anyone (person or system) might ask about the company should have a verifiable answer on that page, and the page should state explicitly that, when sources diverge, the content published on the official domain prevails. Not as an imposition (no brand imposes anything on a model), but as an offer: the most complete, most current, and easiest-to-consume source on that specific subject.
Legal name, company registration number (in Brazil, the CNPJ), founding date, named leadership, addresses, phone, email, and official channels. This is the layer models get wrong most often about SMEs, and the easiest to shield, because each item can be checked against public records.
The official definition of the company in one paragraph: the sentence you want to see reproduced when someone asks “who is company X?” If the brand does not publish its own definition, each system synthesizes a different one.
The portfolio with official names and canonical URLs; methodologies and concepts with the house definition. This is what keeps the model from describing the services “by approximation,” in the generic vocabulary of the category.
What the company is not: brands with the same or a similar name, recurring mix-ups, old discontinued domains and where they point. Name confusion hits SMEs five times more often than large brands, and disambiguation is the only remedy that is under the company’s control.
A statement of which source prevails when sources diverge, and a visible last-updated date. Systems (and humans) need a criterion for choosing between conflicting versions; the page provides that criterion.
Registrations, certifications, published case studies, press coverage, with links. The base does not ask for trust; it points to where trust can be verified.
JSON-LD with Organization, Person, and related entities, each with a stable @id that is referenced (not duplicated) on the other pages of the site, plus sameAs for the official profiles. An engineering detail that avoids rework: the markup must coexist with what the CMS or SEO plugins already generate, referencing the existing nodes instead of redefining them. Conflicting structured data recreates, inside the site itself, the ambiguity the base exists to eliminate.
Honesty about limits is part of the instrument, and it is what separates serious practice from a technical talisman.
It is not an llms.txt. The difference is one of channel: llms.txt is a manifest that no major platform’s documentation says it consumes, and that telemetry shows is not being requested. The knowledge base is a normal HTML page: indexable, crawlable, and retrievable by the same search engines that ground AI answers. It operates through the channel that is proven to exist.
It is not a guarantee. Structured data does not guarantee display, and no brand controls a model’s output. What the base changes is the odds and the raw material: when a system looks for information about the brand, it finds a complete, structured source instead of fragments, and there is now a canonical record against which any error can be objectively compared and reported.
It does not buy a recommendation. This is the most important limit, and the one that shallow sales talk leaves out. The knowledge base governs the factual layer: what the company is, what it does, where it is. The preference layer (which brand the AI recommends when someone asks “what is the best company for X?”) is still dominated by external authority: about 82% of citations in AI answers come from earned media (Muck Rack), and brand mentions on the web outperform backlinks by 3x as a predictor of visibility (Ahrefs, 75,000 brands). The base is a necessary condition, not a sufficient one: it makes sure that, when external authority brings your brand into the answer, the description is right.
When someone asks directly about the company, systems that ground their answers search for and read the most relevant pages about that entity. A canonical, complete, up-to-date, structured page about the brand itself is the natural candidate for that retrieval: it is the query where the official domain has the greatest possible competitive advantage.
Models triangulate a brand’s identity by comparing sources. When name, description, address, and leadership match across the site, profiles, directories, and the press (with the base as the reference every channel replicates), the entity consolidates; when they diverge, the model chooses (or invents). The converging recommendation of the guides on fixing brand hallucinations is exactly that: align the core facts across the web, with a stable source of truth at the center.
Without a canonical record, correcting an AI error means arguing over versions. With one, the process becomes an operating routine: periodically audit how the platforms describe the brand, compare with the base, reinforce the diverging facts in the channels the systems read, and measure convergence over time.
@id, sameAs for official profiles, integrated without conflict into the site’s existing markup.Practicing what we describe: Flowup keeps its own official knowledge base published at flowup.agency/official-ai-knowledge-base/ (in Portuguese), with verifiable identity, disambiguation of legacy domains, a precedence policy, and a JSON-LD layer integrated into the site’s markup. The format described in this article is the one we applied to ourselves first.
An Official AI Knowledge Base is the canonical page a company publishes on its own domain as the source of truth about itself: verifiable identity, positioning, services, disambiguation, and channels, in clear text for humans and JSON-LD with identified entities for machines. Its job is to replace the approximation AIs make from fragments with a single, current, verifiable source.
AIs get companies wrong with documented frequency: in a study of 13,000+ queries to ChatGPT, Perplexity, and Gemini, 93% of companies had at least one basic fact hallucinated or missing. SMEs fare worse (50% received fabricated facts, and their names were confused five times more often than those of large brands), and the errors come in a confident tone, indistinguishable from a correct answer.
No. llms.txt is a manifest that no major platform’s documentation says it consumes. An Official AI Knowledge Base is a normal HTML page: indexable and retrievable by the search engines that ground AI answers. It operates through the channel that is proven to exist, not through a hypothetical one.
No. And a guarantee, here, is the sign of a weak pitch. An Official AI Knowledge Base changes the odds: it becomes the strongest source for queries about the brand, reduces entity ambiguity, and creates the record against which errors are compared. The preference layer (which brand the AI recommends) still depends on external authority: the base makes sure that, when the brand enters the answer, it enters correctly described.
An Official AI Knowledge Base has seven components: verifiable identity facts; a canonical description; services and method with official URLs; explicit disambiguation; a precedence policy with an update date; verifiable proof; and the JSON-LD layer with entities by @id and sameAs, integrated without conflicts into the site’s existing markup.
Audit: ask the main AI platforms who the company is, what it does, who leads it, and what sets it apart, repeating the questions. Record errors, omissions, and mix-ups. The inventory becomes the map for the knowledge base, and the audit repeated after publication becomes the progress metric.
Flowup’s B.I.N.A. Diagnosis assesses whether search engines and artificial intelligences find, understand, and correctly describe your brand: the information base is the first of the four layers measured. Immediate result, in-depth report by specialists.
The AI knowledge base is part of Flowup Solutions, alongside the AI Agent for WhatsApp, which uses the same official source to answer customers.
Founder and lead strategist, Flowup Agency
Guto Bertoncini is the founder and lead strategist of Flowup Agency, which he has run since 2011. He is the author of the B.I.N.A. Method, Novo SEO and the Base Informacional Semântica (Semantic Information Base), and leads the agency's SEO for AI, GEO and AEO practice, preparing companies to be found on Google and cited by artificial intelligence platforms. He writes about search and AI on the Flowup blog and on his official website.
Sources verified on July 30, 2026. The central error figure (93% of companies with a hallucinated or missing fact) comes from the Searchable study of more than 13,000 queries to three platforms, checked against verified sources; the sample is made up of London companies, and the percentages should be read as indicative of the pattern, not as universal constants. The statements about the limits of the instrument (structured data with no guarantee of display, the predominance of earned media in citations) follow the platforms’ documentation and the independent studies cited. This article describes a format that Flowup applies to itself and sells as a service; the reader should weigh that interest, and that is exactly why the limits of the instrument are stated in the body of the text.