If you report a result from a single ChatGPT test to leadership, you might be reporting randomness. These numbers show how much randomness is actually in there.
We ran twelve opening questions through six independent runs each: three runs from one survey, three from a second survey three weeks later, both with the same model and text confirmed word-for-word identical.
The result sorts into three groups:
So for half of all questions, whether the brand gets mentioned visibly depends on how that particular conversation happened to develop, not on whether the brand actually fits the topic.
It comes down to how these conversations form. After the opening question, nothing else is fixed. Every follow-up question grows out of the answer that came before it.
If the AI happens to mention a particular competitor in one run, the simulated person follows up on exactly that. From that point on, the conversation revolves around that provider. In a second run with the exact same opening question, a different name might come up, and the conversation heads in a completely different direction.
A single phrase in the first answer can end up shaping the entire rest of the conversation. That's not a model malfunction, it's normal behavior for a system that rephrases things fresh on every run.
An example from the stable end: the question about the total cost of fully equipping a large orchard led to a brand mention in all six runs. Here the link between topic and brand is so tight that the conversation's path stops mattering.
The counter-example: the question about sensible contract terms when buying several chainsaws for a business led to zero brand mentions across all six runs. Not even close. Same category, same brand, opposite result.
The difference is telling, by the way: the first question is a product question, the second a commercial one. On contract terms, the AI apparently finds nothing from the manufacturer worth citing. That fits the pattern from the purchase-phase analysis.
Mainly this: a single test is not a basis for a decision. Not for relief, and not for alarm.
And it has a less comfortable consequence for reporting. If the swing between runs is large, a change from 47 to 43 percent between two measurement dates can fall entirely within that swing. Whether it does, we don't know from this data. To know, you'd have to fully repeat the same survey multiple times on the same day and calculate the spread.
Until that exists, every reported change should carry the caveat that it may sit within normal swing. We hold our own reports to that standard, even though it makes the claims less punchy.
Each in a fresh chat with no history. Three runs already show you whether a mention is stable; four noticeably calms the number down per question.
The overall figure hides the pattern. Only the breakdown shows which questions work reliably and which don't at all. That's the information you can actually act on.
Stable, swinging, practically never. This grouping is more robust against randomness than a decimal point, and it's plenty for making decisions.
Questions where you show up in zero runs are the clearest gaps. They're also the easiest to interpret: there's simply no content there.
If the swing between runs hasn't been determined, small differences between two measurement dates shouldn't be presented as a trend. Better honestly fuzzy than falsely precise.
No. It means it can't be measured with a single query. With multiple runs per question and a per-question breakdown, you get a reliable picture. That's exactly how every other sample-based survey works too.
Because there the link between topic and brand in the sources is so clear-cut that every conversation path arrives at it anyway. That's the state worth aiming for: not being mentioned as often as possible, but being mentioned regardless of how the conversation goes.
We don't have solid data on that. Between the two model generations we measured, the overall level of presence shifted noticeably, but that says nothing about the swing within the same model. That would be worth measuring on its own.
No, that skews the result. Within one chat, the system remembers earlier questions and answers. You'd end up measuring your own conversation history instead of the normal case. Every run needs a fresh chat.
12 opening questions, run on July 26 and August 3, 2026 with text confirmed word-for-word identical and the same model (gpt-5.6-terra), three independent runs each time, six runs per question in total. Grouping: stable = 6 of 6, swinging = 1 to 5 of 6, practically never = 0 or 1 of 6.
A note on this article: research, analysis and writing were produced with the support of AI systems (the Ex Tenebris agent team) and reviewed editorially by Felix Zeh. Every figure cited comes from a real LUX/GEO analysis; the analyzed brand is anonymized to protect the client relationship.
The methodology behind this article is the same one we use for every assessment we run, tailored to your category, your competitors and your need states.