Home/Blog/DATA QUALITY
DATA QUALITY

Your AI visibility nearly doubled? Check the model first

From 26.9 to 47.3 percent brand presence between two measurements. Almost a doubling, and nobody had done anything. A different AI model sat between the two measurements.

Felix Zeh·August 12, 2026·5 min read·AI-assisted
25507510026.9 %07/12MODEL A47.3 %07/26MODEL B43.3 %08/03MODEL BMODEL SWITCH BETWEEN THE FIRST AND SECOND SURVEY
What you'll take from this article
  • Why big jumps in AI metrics usually aren't wins
  • How to separate a model switch from real change
  • The four details every metric needs to stay interpretable later
  • How to present this cleanly to leadership or clients
26.9%with the older model
47.3%with the newer model, two weeks later
43.3%with the same newer model, eight days after that

Picture reporting this number: our brand's presence in AI answers rose from 26.9 to 47.3 percent. Almost a doubling, in two weeks. The obvious question from the room would be: what did you do?

The honest answer in our case: nothing. The AI model doing the answering changed between the two measurements.

What happened?

The first survey ran on an older model generation. The second, two weeks later, on a newer one. Same questions, same need states, same evaluation logic. Only the model was different.

So the jump from 26.9 to 47.3 percent mostly measures a difference between two model generations. The company's website, provably, hadn't changed at all in those two weeks: title, description and word count of the checked category page were identical.

What does the clean comparison look like?

We ran a third survey eight days after the second one. Same model as the second time, same question wording, same number of runs. It's the only comparison in the entire series that actually reflects a trend.

47.3 % → 43.3 %The clean comparison within the same model. If anything, a slight decline, not close to doubling.

The direction is even the opposite of the first reading. Miss the model switch, and you celebrate a win, while the actually comparable numbers point slightly down.

Whether that four-point drop is itself meaningful can't be said either, by the way. It could sit within the normal swing between runs.

Why does this mistake happen so often?

Because model switches are usually invisible. Providers update their systems continuously, often without users noticing anything. Testing through the web interface, you often don't even know exactly which version is answering.

There's also a habit of thought carried over from classic web analytics. You naturally compare two traffic numbers from different months, because the measuring instrument stayed the same. That assumption doesn't hold for AI answers: the measuring instrument is part of the system, and it changes.

A newer model can favor different sources, answer at greater length, name specific brands more or less often. All of that shifts your metric, with nothing about your brand having changed at all.

How to make your measurement comparable

  1. Log four details for every metric

    Model including version, survey date, exact question wording, number of runs. Miss one and the number becomes uninterpretable later.

  2. Measure through the API, not the web interface

    Only through the API do you fix the model version yourself and hold it steady for months. In the web interface, you have no control over that.

  3. Run a double measurement at a model switch

    When you move to a new model, measure once in parallel with both. That tells you how large the pure model effect is, and lets you clean your time series around it.

  4. Mark model boundaries visibly in the report

    A vertical line in the chart with a label is enough. Anyone looking at the graph later without you should spot the break immediately.

  5. Only calculate differences within the same model

    A column showing change across a model switch implies a trend that doesn't exist. In those cases, show both values side by side and skip the difference column.

How to explain this internally

A line that's worked well in presentations: comparing two numbers from two different models is measuring with two different instruments and treating the gap as if the thing itself changed.

That's uncomfortable when the wrong reading currently looks like a win. But it's what keeps your numbers worth something six months from now.

Frequently asked questions

How do I know which model just answered?

Through the API, the model name comes back with the answer and you can log it. In the web interfaces of ChatGPT, Claude or Gemini, you usually only see a rough product name, not the exact version. For reliable time series, there's barely a way around the API.

Should we then just stick with the old model?

That would be the wrong conclusion. Your customers use current models, so that's where you should measure too. The sensible move is switching with a double measurement: once in parallel with the old and new model, so you can quantify the jump and place your time series correctly.

Does this also apply to comparisons between ChatGPT and Gemini?

Yes, even more so. Different providers pull from different sources and answer differently. Show those values side by side, but don't merge them into one combined figure. The Gemini app and Google AI Overviews are also two separate things, by the way, and belong apart.

What do I do with old reports that contain this mistake?

Don't quietly fix them. A short follow-up explaining why the earlier reading doesn't hold up is more uncomfortable, but considerably more credible than a silently swapped number.

How this was measured

Survey 1: July 12, 2026, gpt-4o-mini, n=119. Survey 2: July 26, 2026, gpt-5.6-terra, n=162. Survey 3: August 3, 2026, gpt-5.6-terra, identical question wording. For the comparison of surveys 2 and 3, only the first three runs per question were used on both sides so the run count matches.

Transparency note · AI-generated content

A note on this article: research, analysis and writing were produced with the support of AI systems (the Ex Tenebris agent team) and reviewed editorially by Felix Zeh. Every figure cited comes from a real LUX/GEO analysis; the analyzed brand is anonymized to protect the client relationship.

FZ
Felix Zeh
Founder, Ex Tenebris · Stuttgart, Germany

Has spent over a decade at the intersection of marketing, data and analytics, and uses LUX to measure how brands show up in the answers of ChatGPT, Claude, Gemini and Perplexity. Questions or pushback on this article? kontakt@extenebris.de

Get in touch

How visible is your brand in the AI answer today?

The methodology behind this article is the same one we use for every assessment we run, tailored to your category, your competitors and your need states.