Picture reporting this number: our brand's presence in AI answers rose from 26.9 to 47.3 percent. Almost a doubling, in two weeks. The obvious question from the room would be: what did you do?
The honest answer in our case: nothing. The AI model doing the answering changed between the two measurements.
The first survey ran on an older model generation. The second, two weeks later, on a newer one. Same questions, same need states, same evaluation logic. Only the model was different.
So the jump from 26.9 to 47.3 percent mostly measures a difference between two model generations. The company's website, provably, hadn't changed at all in those two weeks: title, description and word count of the checked category page were identical.
We ran a third survey eight days after the second one. Same model as the second time, same question wording, same number of runs. It's the only comparison in the entire series that actually reflects a trend.
The direction is even the opposite of the first reading. Miss the model switch, and you celebrate a win, while the actually comparable numbers point slightly down.
Whether that four-point drop is itself meaningful can't be said either, by the way. It could sit within the normal swing between runs.
Because model switches are usually invisible. Providers update their systems continuously, often without users noticing anything. Testing through the web interface, you often don't even know exactly which version is answering.
There's also a habit of thought carried over from classic web analytics. You naturally compare two traffic numbers from different months, because the measuring instrument stayed the same. That assumption doesn't hold for AI answers: the measuring instrument is part of the system, and it changes.
A newer model can favor different sources, answer at greater length, name specific brands more or less often. All of that shifts your metric, with nothing about your brand having changed at all.
Model including version, survey date, exact question wording, number of runs. Miss one and the number becomes uninterpretable later.
Only through the API do you fix the model version yourself and hold it steady for months. In the web interface, you have no control over that.
When you move to a new model, measure once in parallel with both. That tells you how large the pure model effect is, and lets you clean your time series around it.
A vertical line in the chart with a label is enough. Anyone looking at the graph later without you should spot the break immediately.
A column showing change across a model switch implies a trend that doesn't exist. In those cases, show both values side by side and skip the difference column.
A line that's worked well in presentations: comparing two numbers from two different models is measuring with two different instruments and treating the gap as if the thing itself changed.
That's uncomfortable when the wrong reading currently looks like a win. But it's what keeps your numbers worth something six months from now.
Through the API, the model name comes back with the answer and you can log it. In the web interfaces of ChatGPT, Claude or Gemini, you usually only see a rough product name, not the exact version. For reliable time series, there's barely a way around the API.
That would be the wrong conclusion. Your customers use current models, so that's where you should measure too. The sensible move is switching with a double measurement: once in parallel with the old and new model, so you can quantify the jump and place your time series correctly.
Yes, even more so. Different providers pull from different sources and answer differently. Show those values side by side, but don't merge them into one combined figure. The Gemini app and Google AI Overviews are also two separate things, by the way, and belong apart.
Don't quietly fix them. A short follow-up explaining why the earlier reading doesn't hold up is more uncomfortable, but considerably more credible than a silently swapped number.
Survey 1: July 12, 2026, gpt-4o-mini, n=119. Survey 2: July 26, 2026, gpt-5.6-terra, n=162. Survey 3: August 3, 2026, gpt-5.6-terra, identical question wording. For the comparison of surveys 2 and 3, only the first three runs per question were used on both sides so the run count matches.
A note on this article: research, analysis and writing were produced with the support of AI systems (the Ex Tenebris agent team) and reviewed editorially by Felix Zeh. Every figure cited comes from a real LUX/GEO analysis; the analyzed brand is anonymized to protect the client relationship.
The methodology behind this article is the same one we use for every assessment we run, tailored to your category, your competitors and your need states.