Skip to content
Waseit
← All articles
AEOMeasurement

Same engine, same question: one name in 32 survives when the AI is allowed to search

Measured this morning across four trades. The result forced us to fix our own score - an engine that did not search no longer votes « absent ».

August 10, 20267 min read
In this article
  1. What this means for your business
  2. We had to fix our own score
  3. The problem is far bigger than our case
  4. What should you do tomorrow morning?

Here is an experiment we had never run, and whose result cost us a week of fixes. Take a single assistant - Gemini. Ask it a single question: “Who are the best plumbers in Bordeaux?”. Change neither the model, nor the wording, nor the day. Change one thing only: allow it, or not, to consult Google before answering.

Run this morning, 10 August 2026, across four trades in four cities - a plumber in Bordeaux, a dentist in Nantes, a garage in Toulouse, a fine-dining restaurant in Strasbourg. Eight questions, one engine. The result: 32 business names cited from memory, 32 cited after searching, and a single name common to both lists. Three percent.

These are not two versions of the same answer. They are two different answers, wearing the same signature.

What this means for your business

Both answers carry exactly the same confident tone. Without search, Gemini announces “a selection of Bordeaux plumbers regularly recommended, based on customer reviews and local reputation”, then names businesses. With search, it announces “recommendations based on user feedback and specialist matching platforms”, then names different ones. Nothing in the phrasing tells the reader it just changed its source of information.

Your customer, though, uses the app - the one that searches. If you test your visibility with a tool that queries the model without search, you are measuring its memory: what it retained from training, frozen months ago. That is real and useful information - brand recall lives there - but it is not what your customer sees today.

We had to fix our own score

This distinction is not theoretical for us: it broke our product. Our free audit queries two engines - one searches the web, the other answers from memory - and our score averaged the two. The arithmetic consequence: a business cited in FIRST position by the searching engine was capped at 50 out of 100, dragged down by the silences of the one that does not search. In practice it landed at 15 or 30, and the report read “low visibility”.

Across the three preceding weeks, 56 of our 74 audits returned exactly 0. A user reported it in the simplest possible terms: his business appears in the ChatGPT app, and our tool gave him zero. He was right, and we were wrong.

The rule changed on 9 August: when at least one assistant actually searched the web, the score counts only those assistants. A model answering from memory does not testify to your visibility - its “I don't know them” is an unmeasured, not a zero. The same audit that returned a capped score now returns one that reflects what the engine actually found, and every report names which engines searched and which remembered.

The problem is far bigger than our case

More than twenty companies now sell AI-visibility measurement tools, with different methods that produce different answers for the same brand. The most common fault line is exactly the one we just measured: querying a model's API, or observing what the user sees in the app. Both approaches have flaws - an API understates what the app knows, a screen scrape overstates by taking one personalised session for the norm.

The practical consequence for anyone buying such a tool: ask your vendor whether its measurements involve a web search. If they cannot answer, or the question makes them uncomfortable, their number does not mean what you think. This is not a technical curiosity - it is the difference between 3% and 100% of the names cited.

What should you do tomorrow morning?

  • Check for yourself: ask the question in the ChatGPT or Gemini app, the one your customers use. It is the one measurement nobody can dispute.
  • Ask any AI-visibility tool - ours included - to tell you which engines searched and which answered from memory. A tool that does not say is blending two measurements.
  • Work both layers, they are not won the same way: search is won in the directories and reviews the engine reads at answer time; memory is won over the long run, through mentions the next training run will absorb.
  • Our free audit queries the assistants in 60 seconds, no account - and since yesterday, it tells you which one searched.

One closing note, because this is not the kind of article companies usually publish. Admitting that a measurement product was measuring badly is unpleasant, and probably bad for short-term sales. But what we sell is honest measurement: hiding a method correction would be exactly the behaviour we warn our readers about. The methodology is public, it is dated, and it now carries this rule in black and white.