I think we’re measuring AI visibility the wrong way.
We’ve been analyzing thousands of shopping and recommendation responses across ChatGPT, Gemini, Claude, and Perplexity, and the biggest takeaway for me is:
Stop treating AI visibility as just a score. Start looking at what causes the score.
One of the strongest signals we found was website retrieval.
Across several brands:
- When the brand’s own website was retrieved, the brand was mentioned 89% of the time
- When the website wasn’t retrieved, the brand was mentioned only 24% of the time
So imagine this:
AI Visibility: 37%
Website Retrieval: 13%
Mention when Retrieved: 92%
That tells a very different story than “your visibility is 37%.”
The AI already seems comfortable mentioning the brand when it reaches the site. The real problem is retrieval.
But retrieval is only one part of it.
We also found brands that were strongly associated with one product category while being almost invisible for other categories they clearly sell.
So two brands can have exactly the same visibility score for completely different reasons:
- One isn’t being retrieved enough
- One gets retrieved but still isn’t recommended
- One is only understood in part of its catalog
- One is being measured against prompts where brands are rarely mentioned at all
The content being retrieved was also interesting.
For one brand, 116 of 165 own-site citations came from blog content, while only 3 came from product pages. One roundup article alone was cited 37 times.
That makes sense when you think about what users actually ask:
“Best [category] brands”
“Best [product] for [use case]”
“[Brand] vs [competitor]”
“Top alternatives to [brand]”
A PDP is often great at explaining a product.
It’s not necessarily built to answer those questions.
Another thing we learned: don’t overreact to a single visibility test.
In one dataset, roughly a quarter of identical prompt/model combinations changed between repeated runs.
So a move from 53% to 47% doesn’t automatically mean something broke. Trends, repeated runs, and confidence matter.
And the same applies off-site.
Instead of assuming “Reddit is good for GEO” or “YouTube is important,” it makes more sense to look at the actual prompts where competitors win and ask:
Which external sources are showing up in those answers?
Sometimes it’s Reddit. Sometimes a niche publisher, retailer, review site, YouTube video, or comparison page.
So I’m increasingly thinking the useful questions aren’t:
“What’s my AI visibility?”
But:
Why is my visibility what it is?
Is AI retrieving me?
Does it recommend me when it does?
Which categories does it associate me with?
Which sources are influencing the prompts I care about?
The score is the output.
The interesting part is diagnosing the inputs that created it.
Curious how others working on GEO/AEO are thinking about this - are you already separating retrieval, mentions, category association, and prompt quality, or mostly tracking one overall visibility metric?