Home/Blog/DATA QUALITY
DATA QUALITY

Same question, six runs, six different outcomes

Twelve identical questions, six independent runs each, same model, same wording. Only two had the brand show up every time. That has consequences for everything you report about AI visibility.

Felix Zeh·August 10, 2026·6 min read·AI-assisted
6 RUNS PER QUESTION2/6Q13/6Q24/6Q36/6Q41/6Q51/6Q65/6Q76/6Q84/6Q90/6Q105/6Q110/6Q12
What you'll take from this article
  • How much brand mentions swing between identical runs
  • Why that comes down to conversation mechanics, not a bug
  • How many runs you need at minimum for a number to hold up
  • How to tell a swing from real change
2/12questions mentioned in every run
6/12questions with a swinging result
4/12questions with practically no mentions

If you report a result from a single ChatGPT test to leadership, you might be reporting randomness. These numbers show how much randomness is actually in there.

How big is the swing, really?

We ran twelve opening questions through six independent runs each: three runs from one survey, three from a second survey three weeks later, both with the same model and text confirmed word-for-word identical.

The result sorts into three groups:

So for half of all questions, whether the brand gets mentioned visibly depends on how that particular conversation happened to develop, not on whether the brand actually fits the topic.

Why does the same question lead to different conversations?

It comes down to how these conversations form. After the opening question, nothing else is fixed. Every follow-up question grows out of the answer that came before it.

If the AI happens to mention a particular competitor in one run, the simulated person follows up on exactly that. From that point on, the conversation revolves around that provider. In a second run with the exact same opening question, a different name might come up, and the conversation heads in a completely different direction.

A single phrase in the first answer can end up shaping the entire rest of the conversation. That's not a model malfunction, it's normal behavior for a system that rephrases things fresh on every run.

Two questions, side by side

An example from the stable end: the question about the total cost of fully equipping a large orchard led to a brand mention in all six runs. Here the link between topic and brand is so tight that the conversation's path stops mattering.

The counter-example: the question about sensible contract terms when buying several chainsaws for a business led to zero brand mentions across all six runs. Not even close. Same category, same brand, opposite result.

The difference is telling, by the way: the first question is a product question, the second a commercial one. On contract terms, the AI apparently finds nothing from the manufacturer worth citing. That fits the pattern from the purchase-phase analysis.

What does that mean for your numbers?

Mainly this: a single test is not a basis for a decision. Not for relief, and not for alarm.

And it has a less comfortable consequence for reporting. If the swing between runs is large, a change from 47 to 43 percent between two measurement dates can fall entirely within that swing. Whether it does, we don't know from this data. To know, you'd have to fully repeat the same survey multiple times on the same day and calculate the spread.

Until that exists, every reported change should carry the caveat that it may sit within normal swing. We hold our own reports to that standard, even though it makes the claims less punchy.

How to handle the swing

  1. At least three runs per question, four is better

    Each in a fresh chat with no history. Three runs already show you whether a mention is stable; four noticeably calms the number down per question.

  2. Evaluate per question, not just as an overall average

    The overall figure hides the pattern. Only the breakdown shows which questions work reliably and which don't at all. That's the information you can actually act on.

  3. Sort into three tiers instead of refining percentages

    Stable, swinging, practically never. This grouping is more robust against randomness than a decimal point, and it's plenty for making decisions.

  4. Start with the stable zeros

    Questions where you show up in zero runs are the clearest gaps. They're also the easiest to interpret: there's simply no content there.

  5. Only report changes once the gap is clearly large

    If the swing between runs hasn't been determined, small differences between two measurement dates shouldn't be presented as a trend. Better honestly fuzzy than falsely precise.

Frequently asked questions

Does this mean AI visibility can't be measured in any meaningful way?

No. It means it can't be measured with a single query. With multiple runs per question and a per-question breakdown, you get a reliable picture. That's exactly how every other sample-based survey works too.

Why doesn't it swing at all on some questions?

Because there the link between topic and brand in the sources is so clear-cut that every conversation path arrives at it anyway. That's the state worth aiming for: not being mentioned as often as possible, but being mentioned regardless of how the conversation goes.

Does the swing get smaller with better models?

We don't have solid data on that. Between the two model generations we measured, the overall level of presence shifted noticeably, but that says nothing about the swing within the same model. That would be worth measuring on its own.

Is it fine to always use the same chat for testing?

No, that skews the result. Within one chat, the system remembers earlier questions and answers. You'd end up measuring your own conversation history instead of the normal case. Every run needs a fresh chat.

How this was measured

12 opening questions, run on July 26 and August 3, 2026 with text confirmed word-for-word identical and the same model (gpt-5.6-terra), three independent runs each time, six runs per question in total. Grouping: stable = 6 of 6, swinging = 1 to 5 of 6, practically never = 0 or 1 of 6.

Transparency note · AI-generated content

A note on this article: research, analysis and writing were produced with the support of AI systems (the Ex Tenebris agent team) and reviewed editorially by Felix Zeh. Every figure cited comes from a real LUX/GEO analysis; the analyzed brand is anonymized to protect the client relationship.

FZ
Felix Zeh
Founder, Ex Tenebris · Stuttgart, Germany

Has spent over a decade at the intersection of marketing, data and analytics, and uses LUX to measure how brands show up in the answers of ChatGPT, Claude, Gemini and Perplexity. Questions or pushback on this article? kontakt@extenebris.de

Get in touch

How visible is your brand in the AI answer today?

The methodology behind this article is the same one we use for every assessment we run, tailored to your category, your competitors and your need states.