The same need asked differently returns different companies: 31 % overlap
« The best X » and « which one for my 10-person company » overlap by only 31 %, against 77 % when the exact same question is asked twice.
In this article
Short answer: changing the buying context inside the question changes which companies get recommended - far more than the assistant's own randomness does. Ask exactly the same question twice: the lists overlap by 77 %. Ask « which one for my ten-person company » instead of « what are the best »: 31 %.
That is 46 points below the noise floor. In other words: how your buyer phrases their need weighs more on your presence than everything the assistant varies on its own from one run to the next.
A visibility audit that asks a single phrasing does not measure your market. It measures a quarter of it.
Why this question, now
We have shown that language changes the recommended list, that the sources read change with it, and that those sources genuinely carry the recommendation. What remained was the variable a company suspects least, because it does not control it at all: how its customer phrases the request.
Nobody types only « best CRM ». A buyer also says « I run a small outfit », « which has the best value for money », « I'm a beginner, which is simplest ». Those are four buying moments, not four ways of saying the same thing.
What we measured, exactly
Five global sectors, two languages (English, French), two assistants that search - Perplexity and Gemini. 120 runs, zero failures. The reference question is taken word for word from our earlier measurements, so this one can be read alongside them.
| Intent | Phrasing |
|---|---|
| Reference | « What are the best X? » |
| Small company | « I run a 10-person company. Which X would you recommend for us? » |
| Price | « Which X offer the best value for money? » |
| Beginner | « I am a complete beginner. Which X are the easiest to get started with? » |
The noise floor, without which this article would not exist
An assistant is not deterministic: ask it the same question again and it does not return the same list. Comparing two phrasings without knowing how much it contradicts itself would be publishing randomness and calling it a result.
So we ran the reference question three times per combination, which gives 46 pairs of « same question, two answers ». Their average overlap - 77 % - is the floor. Every gap between intents is read against it, never in the absolute.
The result
| Question asked | Overlap | Gap to noise |
|---|---|---|
| The same one, asked again (floor) | 77 % | - |
| Best value for money | 53 % | 24 points below |
| I'm a beginner | 40 % | 37 points below |
| 10-person company | 31 % | 46 points below |
Both engines move the same way, independently - which is what allows us to publish. Perplexity goes from an 84 % floor to 39 % on the « small company » question; Gemini, from 68 % to 23 %. The values differ; the direction and the magnitude do not.
The case that makes it concrete
On hotel booking, the assistant repeats itself very well: 86 % overlap between two runs of the same question. Add « I run a ten-person company » and the overlap falls to 0 %. Not 30, not 15: zero. Not one company in common.
Looking at the names explains it: the reference question returns the consumer platforms everyone knows. The « ten-person company » question surfaces TravelPerk, Engine, Booking.com for Business - business-travel players, absent from the first list. It is not a ranking that shifts, it is a different market answering.
It is not your position that changes with the question. It is the list of candidates.
The same pattern elsewhere: on English web hosting, Namecheap and Wix appear only on the « small company », « price » and « beginner » questions. On French CRM, HubSpot is absent from the reference list but present as soon as you specify « small company » or « I'm starting out ».
The rival hypothesis, tested and dismissed
A result this clean invites a serious objection, and we raised it against ourselves: our name extractor lets through advice fragments (« check the integrations », « minimum budget »). And a price question produces different advice vocabulary from a beginner question. That noise could have manufactured the whole gap.
So we redid the computation at four cleaning levels, from raw to strictest - keeping only names that recur across several answers, since a real company name repeats where an advice phrase is one-off. The noise floor rises, as expected, from 62 % to 79 %. The gap stays between 29 and 36 points in all four cases. The objection does not hold; we publish the strictest computation, not the most flattering.
What to do about it
List your customers' questions, not your category.
« Best invoicing software » is a journalist's question. Your buyers say « for a nonprofit », « without an accountant », « for invoicing abroad ». Those are the sentences to measure.
A good score on one phrasing does not transfer.
A 46-point gap means being recommended on the generic question predicts almost nothing about the question your real segment asks.
Buying context is leverage, not only risk.
If you are absent from the generic list but your segment has its own question, that is where to exist - and it is a far less crowded contest.
Make your pages say WHO you are for.
The sources that put a name into the « ten-person company » answer are the ones explicitly discussing that case. A page that does not say who it is for cannot be picked for a profile.
The limits
One sector removed from the measurement: running shoes.
The answers there cite numbered models (« Ghost 16 », « Clifton 9 ») that our extractor discards. The sector therefore came out at 0 % everywhere - an artefact of our filter, not a result. Publishing it would have been dishonest.
Two languages and two engines.
English and French; Perplexity and Gemini. ChatGPT through the API is excluded: we measured that it does not search at all, so intent does not play the same role there.
A single day.
These measurements are dated and assistants move. That is precisely why a noise floor accompanies every figure.
Our extractor remains imperfect
, even after cleaning: a few phrases (« technical support », « French SMEs ») survive in the sets. A fix is under way. We say so because a number whose cleaning is hidden is not a number.
We promise nobody a place in ChatGPT's answers - nobody can do that honestly. What can be measured is where you stand, per language, per engine, and now per question asked.