86.5% of the non-branded prompts in AEO Copilot return no mention of the brand they belong to, in any of the 4 engines.
I expected a gap. I did not expect the gap to be the default state.
The dataset
AEO Copilot tracks how brands show up in AI answers. Users add prompts their buyers would actually ask ("best CRM for startups", "is Acme legit"), and the tool runs them against ChatGPT, Claude, Perplexity, and Google AI Overviews, logging whether the brand gets mentioned, in what position, and with what sentiment.
I pulled the data across all brands on the platform: 1,600+ prompts with results. For the cross-engine analysis I kept only the 239 prompts that ran on all 4 engines, because comparing engines is only fair when every engine answered the same question. 156 of those 239 are non-branded (the brand name is not in the prompt), and 83 are branded.
That is a small sample. I am publishing it anyway because the pattern is not subtle, and because most of what gets written about AI visibility cites no dataset at all.
Finding 1: only 5.1% of non-branded prompts show up in all 4 engines
Of the 156 non-branded prompts, 13.5% got a mention somewhere. Then it splits in a way I did not predict: 8.3% of prompts were mentioned in exactly 1 engine, 5.1% in all 4, and 0% in 2 or 3.
The engines either barely agree or fully agree. There is no middle.
The single-engine mentions are the majority of all mentions (62%), and that stops being surprising once you look at how differently these 4 systems find their sources:
- ChatGPT only cites when a query triggers web search, and its retrieval runs on Bing's index plus licensed publishers (OpenAI's own announcement).
- Claude searches through Brave, not Google or Bing, and cites fewer sources, each tied to an exact quoted passage (Anthropic's web search docs).
- Perplexity retrieves live for every query and leans on fresh pages and community content.
- Google AI Overviews pull from Google's own index via query fan-out, and Google states plainly there is no special markup to get in: you need to rank.
4 engines, 4 different indexes, 4 different citation rules. A brand that ranks on Google but not on Bing or Brave shows up in 1 engine and misses the other 3. That is exactly the 62% pattern in the data. Profound's analysis of 680 million citations found only 11% of domains get cited by both ChatGPT and Perplexity, so the low overlap is not an AEO Copilot quirk, it is how the ecosystem works.
!How ChatGPT, Claude, Perplexity, and Google AI Overviews cite differently: citation style, sources, and how to win each engine
The 5.1% that show up everywhere are the brands the category cannot be described without. They are in the listicles, the comparison pages, the Reddit threads, and the review sites, so every index finds them through its own path.
Finding 2: branded prompts hit 100%, and that is why they measure something else
Every branded prompt in the dataset (all 83) got the brand mentioned in all 4 engines. 100%.
Before you celebrate: when someone asks "is Acme legit", the engine has to say "Acme" in the answer. The mention is guaranteed by the question. What is not guaranteed is what comes after the mention: the sentiment, and whether the facts are right.
This is the cleanest split in the whole dataset. Branded prompts measure reputation: what AI says about you when asked directly. Non-branded prompts measure visibility: whether AI brings you up at all when a buyer asks a category question. 100% vs 13.5%.
Most tools and most audits blur these 2 numbers into 1 score. They should not be blurred. A brand can have a spotless reputation and be invisible in every category question that feeds a buying decision.
Finding 3: longer prompts change nothing
I assumed prompt specificity would show up in the data as a length effect. It does not.
Prompts with 6 to 9 words got a mention 25.9% of the time. Prompts with 10 or more words: 26.6%. Half the prompts on the platform are 10 words or longer, and the extra words buy almost nothing.
What actually moves the number is whether the answer to the prompt is a list your brand belongs to. "Best CRM for startups" returns a list; either you are on it or you are not. Adding "with good reporting and a free tier for a 5-person team" to the prompt does not put you on the list. Earning a place in the pages the engines pull lists from does.
What to do with this
3 moves, in order:
1. Split your tracking into visibility and reputation. Track category prompts and branded prompts as separate sets with separate targets. If your report shows 1 blended score, you cannot tell "we are invisible" apart from "AI describes us wrong", and the fixes for those 2 problems have nothing in common.
2. Treat each engine as its own channel. A mention in ChatGPT tells you nothing about Claude. Check where you show up and where you do not, then work the weakest engine's supply chain: Bing indexing for ChatGPT, Brave for Claude, fresh community and comparison content for Perplexity, classic rankings for Google AI Overviews.
3. Chase the list pages, not the prompt wording. The 5.1% that show up everywhere are on the roundups, comparison pages, and community threads that every index reads. Get named in the pages that answer "best X for Y" and the engines follow, each through its own path.
You can see where you stand in an afternoon. Run a free audit or track your first 50 prompts on the free tier, no credit card.
FAQ
How many prompts back these numbers?
1,600+ prompts with results across all brands tracked in AEO Copilot. The cross-engine findings use the 239 prompts that ran on all 4 engines (156 non-branded, 83 branded). Small sample, stated openly.
Why do engines mention a brand in 1 engine but not the others?
They read different indexes. ChatGPT retrieves through Bing plus licensed publishers, Claude searches through Brave, Perplexity runs its own live retrieval, and Google AI Overviews pull from Google's rankings. Ranking in 1 index and not the others produces exactly this pattern.
Is 100% branded visibility good news?
It is expected news. The brand name is in the question, so the mention is guaranteed. The useful signal in branded prompts is sentiment and factual accuracy, not the mention itself.
Do longer, more specific prompts improve visibility?
Not in this data. 6 to 9 word prompts and 10+ word prompts get mentioned at nearly the same rate (25.9% vs 26.6%). Visibility comes from being in the sources engines cite, not from how the prompt is phrased.
Where to go next