How we measure brand visibility in AI

Every week we ask ChatGPT, Gemini and Perplexity the questions someone looking for a product in your category would ask, and check whether the answer names your brand and whether your site is among its sources. Below: how we score it, what the number is worth and what it does not tell you.

What we ask, and how

  • Engines: ChatGPT (OpenAI, gpt-4.1 with web search), Gemini (Google, gemini-3.5-flash with Google Search) and Perplexity (sonar). We ask through the API, at temperature 0, in the language set for the project. ChatGPT also gets the project’s country as an approximate location.
  • Cadence: each engine once a week, and the day after the questions change.
  • Three kinds of questions:
    • about the category, the way someone who does not know your brand would ask. Only these count toward the score;
    • about the brand (“what is X”, “X reviews”). They show what AI knows about you, but they do not count toward the score: a question that names the brand almost always gets it named back. In two projects such questions inflated the score 2 to 4 times before we took them out;
    • about a differentiator: whether AI links the brand with the trait that sets it apart. That is a separate number next to the score, not part of it.
  • We recognise a question about the brand by its text: the name, aliases and domain, not only by its label.

How the questions are written

We write new sets in two steps. First a model describes the market: the category, the offer, the audience, competitors and the brand’s differentiators. Then a second step writes customer needs blind. It sees only the category description, without the brand’s name, site or differentiators, so the questions do not repeat the words the brand uses about itself. Each need is written in two different phrasings.

An automatic check rejects questions that contain the brand’s name. A second model (Gemini) flags questions nobody would ask. We freeze the set before the first measurement and never pick questions by their results.

How we score

Each category question scores 0 to 100 points:

  • a brand mention: up to 75 points. 75 when the brand comes first; less when competitors are named before it (60, 45, at least 30);
  • a link to your site among the answer’s sources: 25 points.

The score is the average over all category questions, from 0 to 100%. We match the brand by its name and aliases as whole words, never as part of another word.

What the number is worth

  • Uncertainty range. The questions are a sample of what people might ask, so next to the score we show a 95% range (Wilson method). With 15 questions it runs from about ±13 to ±23 points, depending on the score. More questions narrow it; asking the same question twice does not.
  • Week-on-week change. We call it a change only when it passes a statistical test on pairs of the same questions (p < 0.05). Otherwise we say “no significant change”. Next to the latest score we show the average of the last four measurements.
  • Alerts. Engines answer a little differently every time: in our data, 16 of 31 weekly “drops” were gone a week later. So we report a drop only when the loss holds for two measurements in a row and more questions lost than gained.

When the questions change

We mark every change of the question set on the chart. A comparison across it shows no percentage change, because it would set answers to different questions against each other. We do not recompute past results.

What the engine searched for

For ChatGPT and Gemini we record the phrases the engine typed into the search engine, and for ChatGPT also the pages it looked at. That separates “AI did not find your site” from “it found it but did not cite it”.

What this measurement does not tell you

  • The API is not the app. The ChatGPT, Gemini and Perplexity apps can answer differently: they use other models, remember the conversation and know the user. The score is a measurement comparable over time, not a copy of what every customer sees.
  • We do not measure AI Overviews in Google search results.
  • The score depends on the questions. That is why we keep one set as long as we can and mark every change to it.