652 sources measured: assistants don't read the same pages depending on the language of the question
Yesterday: the recommended brands change with the language. Today: the pages read to pick them. 90 % of the 652 cited domains appear in ONE language only.
In this article
- What we measured, exactly
- First result: the three engines are not playing the same game
- Second result: the reading lists barely overlap
- Third result: the two engines don't read the same pages either
- What these sources actually look like
- The number that changes what you should do
- What this means if you sell in several languages
- The limits of this measurement
Short answer: the sources an assistant consults to answer are almost entirely different from one language to the next. Out of 652 domains cited across ten languages, 589 - that is 90 % - appear in a single language. Exactly one domain out of 652 is cited in at least eight languages.
This follows directly from yesterday's measurement, which showed that the recommended list changes with the language, and left readers with the only question that matters next: so what do I actually do about it? The answer lies in the sources. An engine that searches does not recommend from memory - it reads pages, then pulls names out of them. Knowing which pages is worth more than knowing they exist.
Recommended brands overlap 43-74 % with English. The sources read to choose them: 2 to 9 %. Similar answers, produced from almost disjoint reading.
What we measured, exactly
The same design as yesterday, word for word, so the two measurements can be read together: five global sectors - CRM software, running shoes, online payments, hotel booking, web hosting -, ten languages, three assistants, and no country named anywhere in the question. One variable moves: the language. Whatever changes can only come from it.
What is new is what we record: no longer the text of the answer, but the URLs the engine declares it consulted. 150 answers, collected on 2026-08-24.
First result: the three engines are not playing the same game
| Assistant | Answers citing a source | Domains per answer |
|---|---|---|
| Perplexity | 50 / 50 | 16.7 |
| Gemini | 31 / 50, but URLs masked | titles only |
| ChatGPT (API, no browsing) | 0 / 50 | 0 |
ChatGPT queried through the API does not search: it answers from memory, without consulting a single page. That zero is not a measurement failure, it is the result - and it is a reminder that two very different mechanisms live under the word « assistant ». What you publish today cannot influence an answer that reads nothing.
Gemini does search, but hides its addresses: its links all point to an internal redirect domain. Counting them would make Google appear as the number-one source in every language - an artefact of the format, not a fact. So we set them aside and used the only usable field: the titles, which carry the domain names. That is an indirect measurement, and we treat it as one.
Second result: the reading lists barely overlap
For Perplexity, the source overlap between each language and English - the share of domains common to both, out of all domains across the two:
| Language | Domains cited | Overlap with English | National domains |
|---|---|---|---|
| English | 72 | - | 1 % |
| German | 82 | 5 % | 68 % |
| Polish | 84 | 3 % | 58 % |
| Portuguese | 83 | 9 % | 51 % |
| Japanese | 78 | 2 % | 50 % |
| French | 84 | 8 % | 38 % |
| Korean | 62 | 5 % | 29 % |
| Spanish | 84 | 8 % | 26 % |
| Chinese | 67 | 7 % | 15 % |
| Arabic | 79 | 7 % | 8 % |
Gemini, measured separately and by a different route, lands in the same range: 5 to 17 % overlap with English, and 82 % of its 279 domains in a single language. Two engines, two extraction methods, one conclusion - which is what allows us to publish it.
The right-hand column deserves a pause. English is the exception, not the norm: 1 % national sources, against 68 % in German and 58 % in Polish. Asking in English queries a web with no country; asking in German queries the German web.
Third result: the two engines don't read the same pages either
Same language, same question, same day: Perplexity and Gemini share only 9 to 18 % of their sources. So the gap between languages is not a special case - it is the general rule. There is no single list of pages that decides your presence in AI assistants. There is one per engine and per language.
What these sources actually look like
They are not the international comparison sites you would expect. On web hosting, the sector with the widest brand gap yesterday:
| Language | Sources cited |
|---|---|
| English | pcmag.com, hostingstep.com, websiteplanet.com, techradar.com |
| German | fuer-gruender.de, heise.de, websitewissen.com, hostinger.com |
| Polish | jakwybrachosting.pl, rankinghostingow24.pl, rankinghost.pl, hostingi.net |
| Japanese | lolipop.jp, value-domain.com, assirobo.com, crepas.co.jp |
| Korean | temkit.kr, ko.wix.com, webhosting.gabia.com |
| Arabic | mouqarin.com, adviserhost.com, afddal.com, hostingarabi.com |
Four families recur, and none of them is a global directory: national comparison sites built for a single market, the country's technical press (heise.de in Germany, clubic.com in France, xataka.com in Spain), local publishing platforms (blog.naver.com and brunch.co.kr in Korea, namu.wiki, cnblogs.com in China), and forums - reddit.com is the single most-cited domain in the whole measurement, across seven languages out of ten, and it shows up on Perplexity and Gemini alike.
One example that unsettles the received wisdom: in Polish, the assistant cites a CRM comparison published in the personal-finance section of a political news site. That is not noise - it is a genuine ranking article, on a site with a large national audience. The pages that decide your presence are not always the ones in your industry.
The number that changes what you should do
Out of 834 citations, 86 point to the website of a company named in the answer. Nine citations out of ten talk about you somewhere other than your own site.
You do not get recommended by writing on your own website. That is the fundamental break with classic SEO, where your domain is the primary asset. Here your site confirms what other pages already say - it does not trigger it.
What this means if you sell in several languages
An audit run in English does not measure your other markets.
It doesn't even measure your own if it isn't English-speaking: 2 to 9 % shared sources is a different question put to a different corpus.
Being cited on major English-language sites does not carry over.
A place in a US ranking counts for almost nothing in the Japanese answer, where half the sources carry a .jp extension.
The target is national, and it is short.
For one sector in one language, the assistant reads a few dozen domains. They are identifiable, and many are modest sites that no global brand-awareness strategy would ever reach.
Check before you translate.
On running shoes, the recommended brands were identical across all ten languages. Translating a visibility campaign for that sector would have been budget spent against a problem that does not exist.
An engine that doesn't search will never read you.
With an assistant answering from memory, the only variable is what was already written about you when it was trained. That is a slower game, and publishing does not catch it up.
The limits of this measurement
We state them because a number without its caveats is not a result. A single day, five sectors, ten languages: this is a snapshot, and assistants are not deterministic - we measured yesterday that two runs of the same English prompt overlap only 62-76 %. Gemini's sources are measured through their titles, for lack of readable URLs, and it produced none at all in Arabic. Finally, we measure what is READ, not what convinced: a cited page is not necessarily the one that put a name on the list. Those are two distinct questions, and we do not claim to have answered the second.
We promise nobody a place in ChatGPT's answers - nobody can do that honestly. What can be measured is where you stand, per language and per engine, and what the assistant read to get there.