Home/Blog/SOURCES
SOURCES

585 websites help decide what the AI says about your brand

An AI purchase recommendation doesn't come out of nowhere. It's distilled from sources. In the case we analyzed, that meant 5,387 citations from 585 different websites, and the distribution is anything but even.

Felix Zeh·August 19, 2026·6 min read·AI-assisted
25%50%75%100%OWN MANUFACTURER DOMAIN23.0 %REDDIT13.7 %583 OTHER DOMAINS63.3 %
What you'll take from this article
  • How many sources, and which ones, sit behind an AI answer
  • Why a handful of domains carry a disproportionate share
  • How to find the sources that actually matter for your category
  • What you can concretely do with that list
585distinct domains
5,387individual citations
33citations per answer on average
23.0%share of the single strongest source

There's a common assumption that AI answers come from some kind of diffuse language knowledge. For systems with web search, that's not true. They fetch pages, read them, and build an answer from what they find. Whoever shows up on those pages shapes the outcome.

How many sources sit inside one answer?

Across 162 evaluated answers from one survey, we counted 5,387 individual citations spread across 585 different domains. That's a good 33 citations per answer on average.

That number surprises most people. An AI answer made up of five paragraphs can draw on several dozen pages. It's one argument for why ranking first is losing importance: the AI doesn't read the top result, it reads many.

What kinds of sources are there?

585 raw domains aren't analyzable as they are. So we sort them into categories, applying the same scheme across all brands and all surveys:

Only that sorting makes the interesting question answerable: which category does my brand not appear in at all? Because that's an unclaimed channel where opinion about your category forms without you in the room.

How uneven is the distribution?

Very. The manufacturer domain of the brand under study was the single most-cited source at 23.0 percent of all citations, with Reddit following at 13.7 percent. Two domains cover more than a third of all citations, while the remaining 583 split what's left.

That's practically good news. You don't have to show up on 585 websites. You need to know the twenty that carry the biggest share, and be present, or at least correctly represented, there.

Why source lists go stale fast

One rule we've set for ourselves: every source percentage carries a survey date, and gets re-checked before every new project instead of carried over.

The reason is that the source mix depends on licensing and legal relationships between platforms, not only on relevance. After Reddit sued Perplexity in October 2025, Reddit citations there dropped sharply. Anyone working from an older number in a published study is planning against a reality that no longer exists.

It follows that no third-party source strategy should rest on a single platform. A contract change can shut that channel overnight.

How to build your own source map

  1. Ask questions and record the sources

    Twenty questions typical for your category, no brand name, each in a fresh chat, noting every cited source for every answer. It's grunt work, and it's the data foundation this whole approach rests on.

  2. Normalize domains

    Lowercase everything, strip the www prefix and paths. Otherwise you count the same website multiple times and distort the distribution.

  3. Sort by category

    Using the scheme above or your own. What matters is applying it consistently, including the next time you measure.

  4. Sort by frequency and cut at twenty

    Sources that show up once or twice are one-offs with no real signal. Your top twenty are your working list.

  5. Check per source how you're represented there

    Do you show up? Correctly? Current? Or is it only competitors? Sources that regularly feature competitors and never you are particularly telling.

  6. Prioritize by how much influence you actually have

    Your own website is fully in your hands. Trade press and associations are reachable through classic PR. Forums and communities need genuine participation over months. Review sites are reachable through good products and entering their tests.

What this list is not

It's not a playbook for planting mentions. Trying to seed paid recommendations in forums risks getting kicked out, and legal trouble for undisclosed advertising on top of that.

The value of the list is knowing where your category gets talked about, and whether what's said there is even accurate. An outdated product claim on a heavily cited trade portal does more damage than ten missing mentions.

Frequently asked questions

Can I influence which sources the AI draws on?

Only indirectly. You can make sure your own content is citable, that trade media and associations report on you accurately, and that factual errors on heavily cited sources get corrected. The system itself makes the selection.

Are forums really as important as often claimed?

They're a substantial share of the sources; in our measurement, Reddit was the second-strongest single source. The important caveat: these shares swing a lot between platforms and over time, partly for legal reasons. A strategy that leans on forums alone is risky.

How do I handle wrong information on third-party sources?

Document it first: which source, which claim, how often cited. Then contact the operator, factually, with evidence. In parallel, make sure the correct information is clearly available and citable on your own domain, so there's a solid counter-source.

Is Wikipedia work worth it for AI visibility?

Wikipedia shows up as a heavily cited source in many analyses, especially for ChatGPT. Important caveat: don't write about yourself there, it violates the guidelines and reliably gets caught. The legitimate path runs through solid coverage in independent media that an article can then cite.

How this was measured

585 distinct domains, normalized to lowercase with no www prefix, classified with a fixed category scheme applied identically across all surveys. Basis: 5,387 citations from 162 AI answers, survey from July 26, 2026 with gpt-5.6-terra. Citations are counted, not answers; a domain can appear more than once within the same answer.

Transparency note · AI-generated content

A note on this article: research, analysis and writing were produced with the support of AI systems (the Ex Tenebris agent team) and reviewed editorially by Felix Zeh. Every figure cited comes from a real LUX/GEO analysis; the analyzed brand is anonymized to protect the client relationship.

FZ
Felix Zeh
Founder, Ex Tenebris · Stuttgart, Germany

Has spent over a decade at the intersection of marketing, data and analytics, and uses LUX to measure how brands show up in the answers of ChatGPT, Claude, Gemini and Perplexity. Questions or pushback on this article? kontakt@extenebris.de

Get in touch

How visible is your brand in the AI answer today?

The methodology behind this article is the same one we use for every assessment we run, tailored to your category, your competitors and your need states.