Home/Blog/METHODOLOGY
METHODOLOGY

How do you actually measure whether ChatGPT recommends your brand?

Most companies test their AI visibility by asking ChatGPT once. The result is worthless, for a reason that takes about five minutes to understand.

Felix Zeh·August 5, 2026·6 min read·AI-assisted
2NEED STATES6PERSONAS12OPENING QUESTIONS281AI ANSWERS
What you'll take from this article
  • Why a single question to ChatGPT tells you nothing about your visibility
  • How a test needs to be built for the numbers to hold up
  • The difference between ranking and being mentioned, and why it changes everything
  • What you can realistically measure, and what simply isn't measurable
281AI answers evaluated
12opening questions, each asked repeatedly
10brands compared
3surveys over four weeks

Ask ChatGPT once for the best providers in your category. If your brand shows up, the relief is real. If it doesn't, so is the alarm. Both reactions are premature.

We ran exactly this test, properly, for a manufacturer of battery-powered garden and forestry equipment in the DACH market: 281 individual AI answers, three surveys, ten brands compared. The client prefers to stay unnamed, but the numbers are real. This article series takes each finding apart one by one. This first piece is about how you measure something like this cleanly in the first place.

Why does asking ChatGPT once get you nowhere?

Because the same question can lead to a completely different conversation every time you ask it. We put twelve identical opening questions through six independent runs each, same model, word-for-word identical text. Only two of the twelve questions had the brand show up in every single run.

2 of 12Questions where the brand was named reliably in every run. For the other ten, it depended on how the conversation happened to unfold.

A single test is a sample of one. It tells you what happened in that one conversation on that one day. It says very little about the normal case. Just how large that swing can be is covered in detail in this article.

What is a simulated buying conversation?

Instead of isolated prompts, we recreate how people actually talk to an AI about a purchase. So not one question, but a conversation.

Six fictional people from two different need states ask their opening question, deliberately phrased the way people really type, without a brand name. One example from the case we analyzed: "I've ended up with a different battery for every single tool at this point, it's driving me crazy."

From there, the AI carries the conversation forward on its own. Every follow-up question grows out of the answer before it. If the AI happens to mention a particular competitor in passing, the simulated person follows up on exactly that. One opening question can branch into several completely different conversations, spanning the awareness, consideration and purchase phases.

That's the key difference from classic rank trackers: what gets measured isn't a search term, it's an entire advisory conversation.

What exactly gets measured?

Three things, and they belong together.

First: does the brand show up at all? Not at which position, just whether it appears. This is the central break from classic search. There are no ten blue links to rank fourth among anymore. There's one answer, and it either names you or it doesn't.

Second: where did the AI get its information? Every answer draws on sources. In the case we analyzed, that meant 5,387 citations across 585 different domains. Whoever shows up there has a say in what gets said about a brand.

Third: did the answer actually help the person? Every need state has a goal, something like "find and buy a suitable device." Whether the answer delivers on that gets scored separately, because visibility without usefulness is an empty number.

How to build a reliable test yourself

If you want to try this without outside help, here's what actually matters:

  1. Define need states, not target audiences

    Write down the situation someone is in when they google your category or ask an AI about it: what problem, what time pressure, what context. You don't need age or income for this.

  2. Write opening questions without any brand name

    Twelve questions are enough to start. The important part: type them the way customers actually type, casual language included. The moment you put your brand in the question, all you're measuring is whether the model recognizes your name.

  3. Ask every question at least three times, independently

    New chat, no history, no memory of previous runs. Only comparing multiple runs tells you whether a mention is stable or a fluke.

  4. Log the model and the date

    Every number needs to carry which model answered and when. Without that, you can't later tell a model switch apart from a real change.

  5. Record the sources, not just the mentions

    For every answer, note which websites the AI cites as evidence. That's the information you can actually act on, because it tells you where to intervene.

  6. Measure competitors the exact same way

    Otherwise you have a number with nothing to compare it to. 43 percent presence sounds mediocre, until you learn the best competitor sits at 16.

What this test doesn't answer

Two limits worth knowing before you reallocate budget.

The test doesn't measure how often these questions actually get asked in the market. The need states and questions are derived from expertise, not from search volume. So you know how the AI answers when someone asks this way. How many people actually ask this way is a separate question.

And it doesn't measure where in the answer the brand appears. Whether it's recommended in the first sentence or mentioned in a subordinate clause obviously matters to a reader. That wasn't captured in this case's data, so we're not claiming anything about it either.

Frequently asked questions

How many questions do I need for a meaningful test?

Twelve opening questions with three to four independent runs each has proven to be a workable floor. That gets you roughly 100 to 160 usable answers. Fewer than three runs per question doesn't make much sense, because you won't be able to see the swing between runs.

Can I do this with a free ChatGPT account?

For a first impression, yes. But make sure every run starts in a new chat with no history, and that personalized memory is switched off. Otherwise your earlier questions bleed into the answer, and you end up measuring your own usage pattern instead.

Why not just ask: do you know brand X?

Because that answers a different question. With your name in the prompt, you're testing whether the model knows you. Without your name, you're testing whether it recommends you. The second one is the commercially relevant question, and the results are usually a lot worse.

How often should you repeat a measurement like this?

A fixed cadence works well, roughly quarterly, plus a check after any major changes to your own website. What matters more than frequency is comparability: same model, same question wording, same number of runs.

How this was measured

Basis for every number in this series: 281 evaluated AI answers from three surveys (July 12, 2026 with gpt-4o-mini; July 26 and August 3, 2026 with gpt-5.6-terra), two need states, six personas, twelve opening questions with three to four independent runs each. Supplemented by a technical crawl of ten brands in the same competitive field.

Transparency note · AI-generated content

A note on this article: research, analysis and writing were produced with the support of AI systems (the Ex Tenebris agent team) and reviewed editorially by Felix Zeh. Every figure cited comes from a real LUX/GEO analysis; the analyzed brand is anonymized to protect the client relationship.

FZ
Felix Zeh
Founder, Ex Tenebris · Stuttgart, Germany

Has spent over a decade at the intersection of marketing, data and analytics, and uses LUX to measure how brands show up in the answers of ChatGPT, Claude, Gemini and Perplexity. Questions or pushback on this article? kontakt@extenebris.de

Get in touch

How visible is your brand in the AI answer today?

The methodology behind this article is the same one we use for every assessment we run, tailored to your category, your competitors and your need states.