There is a figure in current research on AI visibility that is almost always quoted with resignation: two thirds of what ChatGPT cites is out of reach for marketing. Wikipedia, other companies’ homepages, app stores — sources you cannot get at.
The figure is correct. The resignation is still wrong, and that is because of a detail that usually gets lost in the excitement over Wikipedia’s dominance.
Fig. 1 · Key figure
The second-largest source in ChatGPT belongs to the companies themselves
This building block belongs to the third visibility tier — Tier 3 · Cited. There, the point is no longer to be found or recommended, but to be the source an answer relies on.
What Ahrefs counted
In October 2025, Louise Linehan exported the 1,000 most-cited pages in ChatGPT and sorted them by content type. Published on 28 October 2025, with data as of September 2025.
| Content type | Share | Influenceable? |
|---|---|---|
| Wikipedia | 29.7% | no |
| Homepages and landing pages | 23.8% | your own: yes |
| Explainer and educational pages | 19.4% | yes |
| App stores | 6.6% | no |
| Reviews | 5.8% | yes |
| News and media | 5.2% | yes |
| Language and grammar sites | 4.0% | – |
| Dictionaries and reference works | 2.2% | – |
| Blog posts and articles | 1.9% | yes |
| Forums and communities | 0.9% | – |
| Company pages (“About us”, contact) | 0.5% | yes |
Ahrefs counts 32.3% as influenceable — explainer pages, reviews, news and blog posts. The rest are considered “dead” citations, a term coined by Ryan Law: sources that classic public relations cannot reach.
The flaw in this calculation
The classification is written from the perspective of an agency that earns placements for clients. From that perspective, a homepage really is out of reach — someone else’s.
For the company itself, the logic flips. The second-largest citation category in ChatGPT is a page over which it has complete control: text, structure, freshness, wording. No editor needs to be persuaded, no portal maintained, no notability hurdle cleared.
Put differently: of the roughly two thirds of “unreachable” citations, just under a quarter of the total — 23.8% — is, for every single company, precisely the category in which it alone decides. That is not a correction to the study, but a change of vantage point.
What this means for your own homepage
If homepages and landing pages are cited this often, it is worth asking which properties qualify a page for it in the first place. Two findings from the same study help.
First: authority sits with the domain, not the page. Of the cited pages that rank on Google at all, 65.3% have a Domain Rating of 81 or higher, with a median of 90. But 67.3% of these pages have a URL Rating between 0 and 10, with a median of 6. So ChatGPT cites pages from strong domains — but not their most heavily linked pages. A well-made subpage without backlinks of its own is not ruled out.
Second: rankings are not a prerequisite. 28.3% of the most-cited pages have zero organic keywords. Ahrefs names three possible reasons — fresh content, narrow niche topics and a different selection logic — and notes that all three probably work together. In practice, this means “Do we rank?” and “Are we cited?” are two different questions.
The five-minute test
Read your homepage and strike out every sentence that would apply just as well to a competitor. “Quality since 1978.” “We are rethinking customer focus.” “Your partner for demanding solutions.” Out they go.
What remains is the part from which a language model can form a statement about you. On many mid-sized companies’ sites, alarmingly little remains — not because these companies have nothing to say, but because the concrete detail has been moved into the product catalogue and the homepage sets a mood.
For a human, that works. They have the context in their head, they see the images, they may know the brand. A language model deciding whether your name belongs in an answer has only the text.
What belongs on a usable homepage can be said in one sentence: what you make or provide, for which industry, in which market, with which verifiable distinguishing feature — with exactly that clarity, without the reader having to work it out.
Fig. 2 · Before / after
The five-minute test for your homepage
Mood
Your partner for demanding solutions. Quality since 1978. We are rethinking customer focus.
Every sentence applies to a competitor too — nothing from which a model can form a statement
Statement
Beta Technik builds drivetrain test benches for automotive suppliers in Europe — since 1978, with 240 employees at two sites.
What, for whom, where, since when and how big — in one sentence
Six details that should be on your homepage
A short list has emerged from our audit work. It is unspectacular, and that is exactly the point — this is not about elegant phrasing, but about the information being available in text form at all.
- What you make or provide, in the language of your industry, not in a coinage of your own. “Drivetrain test benches” is usable, “solutions for the mobility of tomorrow” is not.
- For whom. Industry, company size, use case. Without this, a system cannot match you to a query that asks for exactly that.
- In which market. Region, countries, delivery area. A considerable share of buyer questions include a geographical restriction.
- What sets you apart, with evidence. A process, an approval, a depth of in-house manufacturing, a standard. Not “unique”, but the substance behind it.
- Since when and at what scale. Founding year, headcount, sites. These details are the anchors systems use to tell entities apart — especially when names are shared, which happens more often than you might think.
- How to reach you, as text, not just as a form. That sounds trivial, and it is surprisingly often missing in machine-readable form.
Two things are deliberately missing from this list. Keywords are missing because they add nothing that is not already covered by the six points. And structured markup alone is not enough: it helps with extraction, but it does not replace a sentence that is not there. A schema entry on a page that makes no statement remains a schema entry that makes no statement.
And the matter of freshness
From the same study: of the pages for which an update date could be determined, 89.7% had been updated in the current year. Excluding Wikipedia, around 82% remain.
The figure should be read with caution — the subset is small and Wikipedia makes up over half of it. But the direction matches another Ahrefs study, according to which AI assistants cite considerably fresher content than organic search does. For your homepage, this means it is not a document you write once. If it has listed the same product range for four years while two business areas have changed, that is not only wrong in substance but also a signal.
Wikipedia, Wikidata and the detour we do not recommend
That leaves the largest category: 29.7% Wikipedia. The obvious idea — “then we need a Wikipedia entry” — is the wrong one, and we consider it a mistake for three reasons.
Wikipedia has notability criteria that an average mid-sized company does not meet. An entry that a company initiates about itself breaches the conflict-of-interest rules and is regularly deleted — and the discussion about it stays publicly visible. And even an existing entry cannot be controlled; it belongs to the community.
The more viable route runs through structured open data sources, above all Wikidata. Different inclusion rules apply there, and the information is machine-readable in exactly the form language models use: entity, property, value. That is why our scope of services lists Wikidata, not Wikipedia.
Connected to this is the second half of this building block, which sounds unspectacular and decides a great deal: brand consistency. If your company name, category and description of services diverge across your website, your directory listings, your expert articles and your database entries, the evidence falls apart. A system checking whether you are the right answer then finds three half-pieces of evidence instead of one whole one.
A warning to close this section: the figures on Wikipedia’s dominance vary considerably depending on methodology — depending on the study, 8.9%, “under 20%” and 26 to 48% are all in circulation. After a change to Google’s search interface in September 2025, Wikipedia and Reddit lost a considerable part of their citation share within a few weeks. Here we deliberately rely on Ahrefs’ measurement of the citation corpus rather than on estimates based on the prompt corpus — and we would not build a strategy on any of these figures that only works at exactly that value.
Where this study is weak
In fairness, the limitations — and they are not small.
The categorisation comes from a language model. Ahrefs says so openly: Claude sorted the 1,000 URLs, the result is “not 100% certain” and provides a “directional understanding”. A page counted as a “landing page” could also be a product catalogue.
It is a snapshot of one system. Cut-off date September 2025, ChatGPT only, predominantly the English-language web. On the Google surfaces the distribution looks different — there, YouTube is the most-cited domain.
The freshness finding rests on little. A publication date could be found for only 12.6% of the pages. Among the update dates, which were available for just over half, Wikipedia alone accounts for 53% of the subset. The fact that 89.7% of the pages were updated in 2025 is therefore less meaningful than the figure suggests — even though the value remains at around 82% without Wikipedia.
How this connects to the other building blocks
This building block sits in Tier 3 · Cited and requires the two tiers below it. Work on your own homepage technically belongs to Tier 1 — if an AI crawler cannot read the page, its text is irrelevant. Consistency across third-party sources directly affects brand mentions on third-party sites: the same company under three names produces three weak entities instead of one strong one.
And the 32.3% that Ahrefs describes as influenceable is almost identical to what the article on comparison articles and vendor lists covers — third parties’ explainer pages, reviews and expert articles. How these building blocks together decide whether ChatGPT names a company as a supplier is summarised in the article Getting recommended by ChatGPT as a supplier.
Related terms in the GEO glossary: Consistent brand description, Model knowledge and Citation.
The honest conclusion
The message of this study is not that everything is out of reach. It is that the work is shifting: away from measures that knock on other people’s doors, towards two things a company has in its own hands — a homepage that says in plain words what the company does, and a description that reads the same everywhere.
Neither is new. Both are rarely done, because they are unspectacular and earn nobody applause. According to the available data, they are the part with the best prospects.
Whether ChatGPT and Google AI Overviews name your company for your buyers’ questions today — and which sources they draw on instead — is what the Digital Visibility Audit measures, using real questions from your industry.
How ChatGPT selects its sources in the first place is covered in why Bing matters; another source of your own besides the home page is YouTube.