Home/Blog/IMPACT
IMPACT

Visibility isn't the goal. An answer that actually helps is.

Answers that named a brand actually helped the person asking 85.7 percent of the time. Answers without a brand, only 62.4 percent. The close of the series, and the finding that reframes all the others.

Felix Zeh·August 21, 2026·6 min read·AI-assisted
25507510085.7 %ANSWERS WITH BRANDN=7762.4 %ANSWERS WITHOUT BRANDN=85
What you'll take from this article
  • Why raw visibility numbers alone can mislead you
  • How to measure whether an AI answer actually helped
  • Why this relationship flipped sign between two measurements
  • What that means in practice for your content
85.7%goal reached when the brand was named (n=77)
62.4%goal reached with no brand named (n=85)
23.3point gap

For nine articles, this series has been about how often a brand shows up in AI answers. The final finding turns that framing on its head, and for good reason.

What was measured?

Every need state in the case we analyzed had a defined goal. For the private group, for instance: find at least one suitable device that fits the existing battery system. For the commercial group: a fleet-ready solution with a concrete recommendation.

For every single answer, we separately scored whether it met that goal, regardless of whether a brand was named in it.

The result: answers naming a brand met the goal 85.7 percent of the time. Answers naming no brand, 62.4 percent.

The finding that flips sign

Now comes the part nobody enjoys reporting. In the survey two weeks earlier, the ratio was reversed: answers with a brand met the goal 77.4 percent of the time, answers without a brand 81.6 percent.

No proven trendA model switch sat between the two surveys. The sign flip isn't a change over time, it's a difference between two model generations.

We're reporting this finding anyway, because leaving it out would be dishonest. The more recent survey has the larger sample and the more current model, so its result carries more weight. It's not a proven relationship.

What's cause, what's just correlation?

That can't be cleanly separated from this data either. Two explanations are equally plausible.

One: naming a specific brand makes the answer more concrete and therefore more useful. Someone asking which device to buy is further along with a product name than with a general category explanation.

The other: harder questions get a concrete brand recommendation less often, and they're also harder to answer well. In that case, the brand wouldn't be the cause, the difficulty of the question would be the shared cause of both.

Probably both are at play at once. Anyone selling you a causal relationship here is over-reading the data.

Why this reframes the whole series

The nine articles before this one largely treated visibility as the target metric: mentioned more often equals better. This last finding is a reminder that the mention itself isn't the point.

At the end of the chain is a person with a real question and a real goal. The AI is just the layer in between. Content optimized purely for citability, without actually helping, misses the exact mechanism that makes it citable in the first place.

Which is also why classic content quality and GEO don't contradict each other, by the way. A page that answers a question completely and with evidence is the right answer for both.

How to measure usefulness, not just mentions

  1. Set a verifiable goal per need state

    Not "get informed", but something like "receive a concrete product recommendation with a reason". The criterion has to be worded so two people independently reach the same judgment.

  2. Score every answer twice

    Once: does the brand show up? Once: was the goal met? Keep these two judgments strictly separate, or one bleeds into the other.

  3. Report both numbers side by side

    One metric alone is misleading. High presence with low goal completion means you're visible in answers that don't help anyone. That's not a win.

  4. When goal completion is low, read the actual answers

    Read the answers that missed the goal. Usually the missing piece of information jumps out immediately. That's the most concrete content to-do list you'll get.

  5. Avoid causal claims

    You're measuring a relationship, not an effect. Phrase it accordingly: "answers with a brand met the goal more often" is correct. "The brand mention improves the answer" is not.

Closing out the series

Ten articles, one case, no invented number. Everything in this series comes from real measurements with all the fuzziness that comes with them: model switches, swings between runs, open questions about cause.

If you take away one thing: don't ask whether ChatGPT knows you. Ask whether ChatGPT recommends you when someone describes a problem you can solve. That's a different question, and it almost always comes back less comfortable.

Frequently asked questions

How do you objectively judge whether an answer helped?

Through a predefined, verifiable criterion per need state. In the case we analyzed, for instance, whether a concrete product recommendation matching the stated need was included. What matters is that the criterion is fixed before the evaluation, or you end up unconsciously fitting it to the result.

Should we optimize less for visibility, then?

No, but not exclusively either. Visibility without usefulness is an empty metric, usefulness without visibility reaches no one. The practical path is building content that genuinely answers a question, then making sure it's findable and citable.

Why report a finding that flips sign like this?

Because leaving it out would flatter the data. Both measurements are real. That they point in different directions is itself a result: it shows how strongly model generations can move the numbers, and how careful you should be with any single measurement.

Does this relationship also hold for services?

We don't know. What was measured was a product category with clear purchase decisions. For services, even the goal is harder to define, because the close rarely happens within one conversation. The methodology transfers; the specific numbers don't.

How this was measured

Goal completion scored per answer against the predefined success criterion for that need state. Survey from July 26, 2026 (gpt-5.6-terra): with brand 85.7 percent (n=77), without brand 62.4 percent (n=85). Survey from July 12, 2026 (gpt-4o-mini): with brand 77.4 percent (n=31), without brand 81.6 percent (n=87). A model switch sits between the two surveys, which is why the difference doesn't represent a change over time.

Transparency note · AI-generated content

A note on this article: research, analysis and writing were produced with the support of AI systems (the Ex Tenebris agent team) and reviewed editorially by Felix Zeh. Every figure cited comes from a real LUX/GEO analysis; the analyzed brand is anonymized to protect the client relationship.

FZ
Felix Zeh
Founder, Ex Tenebris · Stuttgart, Germany

Has spent over a decade at the intersection of marketing, data and analytics, and uses LUX to measure how brands show up in the answers of ChatGPT, Claude, Gemini and Perplexity. Questions or pushback on this article? kontakt@extenebris.de

Get in touch

How visible is your brand in the AI answer today?

The methodology behind this article is the same one we use for every assessment we run, tailored to your category, your competitors and your need states.