Reword the question, you change the list
Two ordinary ways of asking one engine the same thing, across 15 trades and 13 US cities. Only 11% of the names are shared, and 47% of comparisons share nothing at all. Yet at market level, 7 of the 10 most cited businesses stay the same.
11%
of names shared between two phrasings, same engine
In this article
Short answer: "what are the best plumbers in Chicago?" and "list the most recommended plumbers in Chicago and explain why" do not return the same list. Same engine, same city, same trade, two ordinary phrasings: 11% of the names are shared, and in 47% of comparisons no name comes back at all. The first business named is the same in 18% of cases.
47%
of comparisons share no name at all
7 of 10
most cited businesses survive the rewording
A one question audit measures a phrasing, not your visibility. What survives rewording is not a ranking, it is repeated presence.
What we measured
The data is the Waseit US study collected on 13 September 2026. Every cell (one trade, one city) received two questions, asked of both ChatGPT and Gemini. The first: "What are the best [trade] in [city]? Give me a list with your recommendations." The second: "List the most recommended [trade] in [city] and explain why." Two ordinary ways of asking the same thing.
We keep the 15 trades measured across the 13 cities here (hotels, restaurants and bakeries are set aside: their names often contain the city). That is 780 answers, 195 per engine per phrasing. For each cell and each engine we compare the list of businesses returned by the first question with the list returned by the second: 305 usable comparisons, the ones where both phrasings named at least one business.
The measure is simple: out of all the names the two phrasings cite together, what share is cited by both? Directories are counted separately and advice headings slipped into the lists are removed.
The phrasing moves almost the whole list
Across the 305 comparisons, the two phrasings cite 3,108 distinct business names between them. Only 370, or 12%, appear under both. Averaged per comparison, the overlap is 11%. And 143 of the 305 comparisons share no name whatsoever.
This could be a length effect: the second question asks for an explanation, so the engine spends text explaining rather than naming. We ran the measure again keeping only the first three names of each answer, before any cut-off: the overlap is still 13%, and 61% of comparisons share no name in their top three. This is not truncation, these are different businesses.
The engine still matters more than the phrasing
For comparison, we ran the same measure between the two engines at identical phrasing: the overlap falls to 2%. Switching engine changes the answer even more than rewording it. But the order of magnitude is the same: in both cases most of the list moves.
| Trade | Two phrasings, same engine | Comparisons sharing no name | Two engines, same phrasing |
|---|---|---|---|
| Staffing agency | 32% | 6 of 24 | 9% |
| Insurance broker | 17% | 7 of 25 | 3% |
| Driving school | 17% | 8 of 21 | 2% |
| Hair salon | 14% | 2 of 17 | 2% |
| Attorney | 13% | 6 of 15 | 0% |
| Plumber | 12% | 8 of 22 | 0% |
| Marketing agency | 11% | 11 of 25 | 3% |
| Auto repair shop | 11% | 7 of 25 | 0% |
| Accountant | 7% | 15 of 24 | 3% |
| Wedding photographer | 7% | 13 of 25 | 0% |
| Locksmith | 6% | 12 of 19 | 0% |
| HVAC contractor | 4% | 15 of 19 | 0% |
| Real estate agent | 3% | 12 of 16 | 1% |
| Dentist | 3% | 10 of 12 | 1% |
| Electrician | 3% | 11 of 16 | 2% |
Staffing is the most stable trade: a third of the names survive the rewording, and it is also the only one where the two engines agree at all (9%). At the other end, for dentists, electricians, real estate agents and HVAC contractors, asking the question differently returns what is practically another list.
What survives is the names that come back often
That would be a discouraging conclusion if it stopped there. It does not. At whole market level the rewording moves almost nothing: among the 10 most cited businesses across all cities and trades, 7 are the same under either question. Robert Half, Aerotek, Kforce, Insight Global, Randstad USA, HUB International and A1 Driving School appear in both rankings.
Another sign of stability: the share of citations going to a name present in at least three cities is 20% under both phrasings. The question changes who shows up in one answer, not who shows up often.
A single answer is a draw.
It depends on the phrasing, the engine and the cell.
Repeated presence is a signal.
A name that comes back across phrasings and cities owes nothing to how the question was typed.
There is no ranking.
There is no "the" AI answer for your trade in your city, there is a distribution.
The two questions do not make the engines talk the same way
The question that asks for an explanation shortens the lists. ChatGPT drops from 7.2 to 5.0 businesses named per answer, Gemini from 4.1 to 3.1. The share of answers naming at least one business barely moves: 94% then 90% for ChatGPT, 82% then 84% for Gemini.
With Gemini, the second question often produces a list of criteria rather than a list of businesses: headings such as "independent brokers" or "comprehensive service" appear where the first question gave names. Our cleaning removes most of them, not all.
What this does not prove
Two phrasings are not all phrasings.
These are the two we ask, not a sample of what your customers type. The real gap could be wider or narrower.
Answer length weighs on part of the figures.
We cap output at 600 tokens for ChatGPT and 1,400 for Gemini: 48 of 195 ChatGPT answers and 54 of 195 Gemini answers end mid sentence on the second question. The per answer name counts are therefore floors. The overlap finding, however, also holds on the first three names.
What an app user sees.
These answers come from the APIs; the consumer app may add web search and other settings.
A single collection.
One day, 13 September 2026. We are not measuring day to day variation here, nor variation between two runs of the identical question.
Perfect cleaning.
Advice headings are removed by rules; whatever escapes slightly inflates the distinct name count, under both phrasings.
What this changes for you
Do not judge your visibility on one question.
Ask several, in your customers' words, and look at what comes back.
Do not celebrate a single appearance, and do not panic at a single absence.
Both are noise.
In the unstable trades
(dentist, electrician, real estate, HVAC), count appearances across several phrasings before concluding anything.
In the stable trades
(staffing, insurance, driving schools), a repeated absence is a real diagnosis.
Our free audit asks ChatGPT and Gemini several questions and counts appearances rather than a rank. That is the only way to get a number that does not depend on how the question was typed.