Skip to content
Waseit
← All articles
AEOMeasurement

Do AI recommendations change from one day to the next?

Re-measured this morning: of 81 names AI assistants cited on 6 August, only 15 reappear on the 8th. What holds for 48 hours, what moves, and why.

August 8, 20266 min read
In this article
  1. What exactly did we measure this morning?
  2. Why do so few names hold?
  3. What should you do tomorrow morning?

Just over a year ago, on 7 August 2025, OpenAI swapped ChatGPT's default model for GPT-5 overnight. Millions of users watched the tone and content of answers change at once; the backlash was loud enough that the older model was restored for paying subscribers on 12 August. Then it happened again: on 5 May 2026, GPT-5.5 Instant became the new default. Twice in a year, the engine underneath the assistant changed - with no notice to the businesses whose names it cites, or stops citing.

Should businesses worry? We preferred to measure. This morning, 8 August 2026, we re-asked - word for word - the four questions from 6 August, about the best plumbers in Bordeaux and Lille, phrased the way a customer types them, to the same three assistants: ChatGPT, Gemini, Perplexity. Twelve answers re-collected, compared with Thursday's twelve. The answer to the question in the title: yes, massively. Of 81 distinct names cited on Thursday, only 15 reappear on Saturday - and two of those fifteen are section headings caught by the counter, not businesses. Roughly one cited name in six survived 48 hours. That instability is the first thing AEO (Answer Engine Optimization: existing in assistants' answers) has to look in the eye.

What exactly did we measure this morning?

The protocol fits in one sentence: the same questions, the same three engines, the same models, 48 hours apart, twelve answers on each side, names counted by the same parser as our reports - so any gap is the answers moving, not the method. The per-engine detail is the real finding. ChatGPT, queried via API without web search: zero names kept out of 26. Gemini: one out of 29 - Espace Aubade, a national bathroom retail chain. Perplexity, the only one of the three that reads the web while answering: 12 names kept out of 28.

The name confidently cited on Thursday is not cited again on Saturday - not even by the engine that put it forward.

The most telling case: on Thursday, ChatGPT recommended “Plomberie Lille Services”, complete with confident-sounding reasons - a name we could not match to any identifiable business, and flagged as unverified. On Saturday, it does not mention it at all. Even its author does not repeat it. Its refusals are not stable either: three questions out of four declined on Thursday, two out of four on Saturday.

Why do so few names hold?

Because stability is not a property of “the AI”: it is a property of its sources. ChatGPT answers from memory - it says so itself in Saturday's answer: “based on reviews and recommendations available up to 2023” (our translation). A memory sampled at every generation produces a different list at every draw. Perplexity re-reads directories at question time: its names hold because Bilik, ThreeBestRated or Travaux.com have not changed in 48 hours. And look at what survived on its side: first the directories and platforms themselves, then the tradespeople they rank. The stable layer of your AI visibility is the written sources - not the lottery of a single generation.

Now add the scale of the problem: when the default model changes - 7 August 2025, 5 May 2026 - the answers of tens of millions of users are reshuffled the same day. Your score can move without you, your competitors or your customers doing anything. A single snapshot, good or bad, therefore says almost nothing; what speaks is the series - how often your name comes back, measurement after measurement, and the sources that carry it.

What should you do tomorrow morning?

  • Draw no conclusion from a single measurement - no panic over one zero, no victory lap over one good score. Re-measure at a regular cadence: the trend means something, the point does not.
  • Invest in the layer that held: the directories and platforms the web-reading engine consults. In our measurement, the only tradespeople still cited 48 hours later came from there.
  • Note the model-switch dates (they are public). If your visibility moves on a default-change day, the cause is probably above you - not inside your business.
  • Measure the three engines separately, free, in 60 seconds - then measure again next month. The comparison is the product, not the photo.

The usual honesty to close: four questions, twelve answers, two cities, one trade - that is a sounding, not a census. But the order of magnitude matches what our three national studies already show (the engines agree with each other on only 2 to 3% of names), and it is dated, published and re-checkable. An AI recommendation is an event, not a state. You do not manage an event - you manage its frequency.