Skip to content
Waseit
← All articles
AEOGEO

Reword the question, you change the list

Two ordinary ways of asking one engine the same thing, across 15 trades and 13 US cities. Only 11% of the names are shared, and 47% of comparisons share nothing at all. Yet at market level, 7 of the 10 most cited businesses stay the same.

11%

of names shared between two phrasings, same engine

September 16, 20268 min read
In this article
  1. What we measured
  2. The phrasing moves almost the whole list
  3. The engine still matters more than the phrasing
  4. What survives is the names that come back often
  5. The two questions do not make the engines talk the same way
  6. What this does not prove
  7. What this changes for you

Short answer: "what are the best plumbers in Chicago?" and "list the most recommended plumbers in Chicago and explain why" do not return the same list. Same engine, same city, same trade, two ordinary phrasings: 11% of the names are shared, and in 47% of comparisons no name comes back at all. The first business named is the same in 18% of cases.

47%

of comparisons share no name at all

7 of 10

most cited businesses survive the rewording

A one question audit measures a phrasing, not your visibility. What survives rewording is not a ranking, it is repeated presence.

What we measured

The data is the Waseit US study collected on 13 September 2026. Every cell (one trade, one city) received two questions, asked of both ChatGPT and Gemini. The first: "What are the best [trade] in [city]? Give me a list with your recommendations." The second: "List the most recommended [trade] in [city] and explain why." Two ordinary ways of asking the same thing.

We keep the 15 trades measured across the 13 cities here (hotels, restaurants and bakeries are set aside: their names often contain the city). That is 780 answers, 195 per engine per phrasing. For each cell and each engine we compare the list of businesses returned by the first question with the list returned by the second: 305 usable comparisons, the ones where both phrasings named at least one business.

The measure is simple: out of all the names the two phrasings cite together, what share is cited by both? Directories are counted separately and advice headings slipped into the lists are removed.

The phrasing moves almost the whole list

Across the 305 comparisons, the two phrasings cite 3,108 distinct business names between them. Only 370, or 12%, appear under both. Averaged per comparison, the overlap is 11%. And 143 of the 305 comparisons share no name whatsoever.

This could be a length effect: the second question asks for an explanation, so the engine spends text explaining rather than naming. We ran the measure again keeping only the first three names of each answer, before any cut-off: the overlap is still 13%, and 61% of comparisons share no name in their top three. This is not truncation, these are different businesses.

The engine still matters more than the phrasing

For comparison, we ran the same measure between the two engines at identical phrasing: the overlap falls to 2%. Switching engine changes the answer even more than rewording it. But the order of magnitude is the same: in both cases most of the list moves.

Waseit US study, collected 13 September 2026, 15 trades across 13 cities. "Two phrasings": average share of names shared between the two questions asked of the same engine. "Two engines": the same measure between ChatGPT and Gemini at identical phrasing.
TradeTwo phrasings, same engineComparisons sharing no nameTwo engines, same phrasing
Staffing agency32%6 of 249%
Insurance broker17%7 of 253%
Driving school17%8 of 212%
Hair salon14%2 of 172%
Attorney13%6 of 150%
Plumber12%8 of 220%
Marketing agency11%11 of 253%
Auto repair shop11%7 of 250%
Accountant7%15 of 243%
Wedding photographer7%13 of 250%
Locksmith6%12 of 190%
HVAC contractor4%15 of 190%
Real estate agent3%12 of 161%
Dentist3%10 of 121%
Electrician3%11 of 162%

Staffing is the most stable trade: a third of the names survive the rewording, and it is also the only one where the two engines agree at all (9%). At the other end, for dentists, electricians, real estate agents and HVAC contractors, asking the question differently returns what is practically another list.

What survives is the names that come back often

That would be a discouraging conclusion if it stopped there. It does not. At whole market level the rewording moves almost nothing: among the 10 most cited businesses across all cities and trades, 7 are the same under either question. Robert Half, Aerotek, Kforce, Insight Global, Randstad USA, HUB International and A1 Driving School appear in both rankings.

Another sign of stability: the share of citations going to a name present in at least three cities is 20% under both phrasings. The question changes who shows up in one answer, not who shows up often.

A single answer is a draw.

It depends on the phrasing, the engine and the cell.

Repeated presence is a signal.

A name that comes back across phrasings and cities owes nothing to how the question was typed.

There is no ranking.

There is no "the" AI answer for your trade in your city, there is a distribution.

The two questions do not make the engines talk the same way

The question that asks for an explanation shortens the lists. ChatGPT drops from 7.2 to 5.0 businesses named per answer, Gemini from 4.1 to 3.1. The share of answers naming at least one business barely moves: 94% then 90% for ChatGPT, 82% then 84% for Gemini.

With Gemini, the second question often produces a list of criteria rather than a list of businesses: headings such as "independent brokers" or "comprehensive service" appear where the first question gave names. Our cleaning removes most of them, not all.

What this does not prove

Two phrasings are not all phrasings.

These are the two we ask, not a sample of what your customers type. The real gap could be wider or narrower.

Answer length weighs on part of the figures.

We cap output at 600 tokens for ChatGPT and 1,400 for Gemini: 48 of 195 ChatGPT answers and 54 of 195 Gemini answers end mid sentence on the second question. The per answer name counts are therefore floors. The overlap finding, however, also holds on the first three names.

What an app user sees.

These answers come from the APIs; the consumer app may add web search and other settings.

A single collection.

One day, 13 September 2026. We are not measuring day to day variation here, nor variation between two runs of the identical question.

Perfect cleaning.

Advice headings are removed by rules; whatever escapes slightly inflates the distinct name count, under both phrasings.

What this changes for you

Do not judge your visibility on one question.

Ask several, in your customers' words, and look at what comes back.

Do not celebrate a single appearance, and do not panic at a single absence.

Both are noise.

In the unstable trades

(dentist, electrician, real estate, HVAC), count appearances across several phrasings before concluding anything.

In the stable trades

(staffing, insurance, driving schools), a repeated absence is a real diagnosis.

Our free audit asks ChatGPT and Gemini several questions and counts appearances rather than a rank. That is the only way to get a number that does not depend on how the question was typed.

Frequently asked questions

Do two ways of asking the same question return the same answer?
No. Across 15 trades in 13 US cities, two ordinary phrasings asked of the same engine share on average only 11% of the businesses cited, and 47% of comparisons share no name at all. The first name cited is the same in 18% of cases.
Is it only because the second question spends words explaining?
No. Keeping only the first three names of each answer, before any cut-off, the overlap is still 13% and 61% of comparisons share no name in their top three. These are different businesses, not simply fewer businesses.
What stays stable, then?
The frequently cited names. Among the 10 most cited businesses across all cities and trades, 7 are the same under either phrasing. And the share of citations going to a name present in at least three cities is 20% in both cases.
How many answers are these figures based on?
780 answers collected on 13 September 2026 through the ChatGPT and Gemini APIs, for 15 trades in 13 US cities, that is 195 answers per engine per phrasing, and 305 usable comparisons.