u/Current-L

▲ 11 r/aeo+1 crossposts

How AEO platforms measure brand visibility - two paths

Hi everyone,

There seem to be two fundamentally different ways AEO platforms measure AI visibility. Here is my summary of how AEO platforms measure brand mentions, citations, and competitive positioning. If you have a different understanding or you have seen AEO platforms take a different approach that I may not be aware of, share it.

The focus is measuring brand mentions and citations, not other capabilities around additional AEO features - LLM bot traffic analysis, etc.

Method #1: Custom prompt measurement

There are a lot of AEO platforms that use this method. Peec, Otterly, Adobe Brand Visibility, LLM Pulse, etc. You create a set of prompts for your brand to measure your brand presence. For instance, "What is the best bike to buy for a 6 year old learning to ride a bike?" The platform executes that set of prompts against the AI platform it is measuring, typically daily, and then analyzes the resulting response for things such as:

  • Was the brand mentioned?
  • Which competitors were mentioned?
  • Where in the response did the brand appear?
  • What domains/pages were cited?
  • Was the brand's own website cited?

For ChatGPT in particular, the Responses API may be used by the AEO platform to execute these prompts and receive a response to store and analyze (though the AEO platforms typically don't reveal the exact LLM API they are using in their documentation).

Limitations: Because the AEO platform is executing a controlled prompt rather than observing a real user's session, it lacks some of the implicit context a real user may bring — precise location, time-sensitive context, prior conversation history, personalization, etc. Some platforms allow country/region/location to be configured explicitly.

Method #2: Clickstream / user opt-in data sets

Some larger AEO platforms like Semrush and Profound have invested in acquiring huge data sets from clickstream providers like Datos (owned by Semrush, now Adobe) or other methods. They are harvesting this data from real participants who allow these data gatherers to observe their browsing and anonymize their data.

For LLM interactions, they capture the users' real prompts and the LLM responses, then cluster / normalize the prompts to semantic topics or user intents. In this way, different prompts that users submit can be grouped together, and the responses can be captured, stored, and analyzed for brands, product names, citation links, etc.

This method is very different from the first method where brands curate a set of prompts that represent the space they want to measure their brand presence for.

Limitations: The clickstream data approach is valuable for directional market intelligence, but not ground truth for brand visibility.

  1. Panel bias - even though the data providers have a large pool of millions of users, their data may not represent your core audience, particularly if you are in a specialized field or a B2B space. e.g. It probably represents moms looking for their kid's first bike better than an AI architect looking for the NAND device with the highest storage density.
  2. Topic / prompt clustering - the process to distill clickstream data into measurable intelligence loses a lot of the nuance of individual prompt measurement.
  3. Observed demand ≠ business value. Clickstream popularity doesn't always equal business value. For instance, a highly specialized AI infrastructure purchasing question may have tiny observed volume but influence a multimillion-dollar purchase. Custom prompt measurement is much better suited to measuring a customer's journey from discovery to conversion.

Method 2.1: Search-demand-derived prompt modeling

This is an alternate on Method 2, specifically used by Ahrefs (and maybe others). Rather than relying on observed AI-user prompts, Ahrefs uses its traditional keyword database and People Also Ask data to identify real search demand, converts those questions into conversational prompts, executes them against AI platforms, and captures the responses. I place this as a variant on method 2 because the method is still creating a large database, just leveraging search demand data rather than user observation data.

Limitations: Real search demand doesn't necessarily equal real LLM prompt demand. The tradeoff is that this approach inherits assumptions from traditional search behavior. Real Google search demand can be a useful proxy for user interest, but it may not reflect how people naturally formulate questions in conversational AI.

Summary

Both methods 1 & 2 bring valuable insights to a brand as a part of a strong AI visibility program, and ideally a brand will use tools and either an internal team or an agency leveraging both methods. Some tools and agencies have capabilities from both visibility measurement methods, whereas others may only use one method.

reddit.com
u/Current-L — 7 days ago