Skip to content
Waseit
← All articles
AEOGEO

Three days later, the list has barely moved - 6 points, against 46 for a different question

We sell repeated measurement. So we tested whether the list drifts from one day to the next. It barely does - and for one engine, not at all.

August 27, 20269 min read
In this article
  1. What we measured
  2. The result
  3. And here the two engines stop resembling each other
  4. Six against forty-six
  5. What we take from this for our own product
  6. The limits, and a data loss of our own making

Short answer: the day you ask changes almost nothing. How you ask changes everything. Three days apart, the same question returns 70 % overlap, against 76 % between two runs made minutes apart. Six points. Yesterday we measured that changing the buying context cost 46.

We sell repeated measurement. It would suit us to conclude that you must measure often, and that is exactly why the question had to be asked properly, accepting the inconvenient answer in advance.

What moves the list is not the calendar. It is the question, and it is the engine.

What we measured

The same reference question every measurement of ours has used since 23 August - identical to the character, verified rather than assumed. Four global sectors, two languages, two assistants that search. 60 runs today, zero failures, plus a 24 August collection kept as the three-day comparison point.

The noise floor is still the load-bearing part. An assistant already contradicts itself minutes apart. A difference between two days means nothing until it is read against what randomness produces in minutes. So we ran the question three times today to get the day's own floor, rather than importing the one from two days ago.

The result

Overlap of recommended lists. Four sectors, English and French, Perplexity and Gemini, 2026-08-27. The « cleaned » computation keeps only names that recur across several answers; the raw one is published beside it so the reader can see what cleaning changes.
Gap between the two runsCleanedRaw
Same day, minutes apart (floor)76 %61 %
Three days70 %53 %
Distance from the floor6 points8 points

Six points on the strictest computation, eight on the raw. Either way, the list from three days ago resembles today's about as much as two runs made in the same minute.

And here the two engines stop resembling each other

The same computation, engine by engine. An average across the two would lend one the stability of the other.
EngineSame dayThree daysDistance
Perplexity85 %71 %14 points
Gemini68 %69 %none

Gemini does not drift at all. Three days later, its list resembles today's exactly as much as two of its own consecutive answers do. That is not stability: it is that it is already so inconsistent with itself - 68 % between successive runs - that three days have nothing left to add.

Perplexity is the opposite. It repeats itself well in the short term, 85 %, and that is precisely why a real drift becomes visible: 14 points in three days. There, measuring a week apart means something, because the noise does not drown the signal.

On one engine, « how often should I measure? » has an answer. On the other the question does not arise, because a single run there already means nothing.

Six against forty-six

The figure only means something beside yesterday's, obtained on the same design, the same floor and the same cleaning.

What moves the recommended list, in points below the noise floor. Two measurements, two consecutive days, one method.
What you changeCost in points
The day you ask (3 days)6
Price as the criterion (« best value for money »)24
Level (« I'm a beginner »)37
Context (« my 10-person company »)46

Seven times more. An audit that repeats the same phrasing daily measures, very precisely, one eighth of the subject. A measurement budget is better spent on the variety of questions than on their frequency.

What we take from this for our own product

Daily measurement has no justification at this stage.

Over three days the gap stays inside the noise. We will not sell a cadence our own figures do not support.

A single measurement remains unusable

, which is the other face of the same result: at 68 % overlap between consecutive answers, one run on Gemini establishes nothing. Repetition is there to silence noise, not to follow the news.

The useful cadence depends on the engine.

On Perplexity a real drift shows in three days. On Gemini you first need several runs to know what you are looking at.

The lever is the variety of questions.

That is where 46 points sit, against 6 for the calendar.

The limits, and a data loss of our own making

We destroyed our own 26 August data. An analysis script imported a function from the previous day's collector; that collector ran its measurement at module load and rewrote its output file. 120 runs cut to 12, with no copy. The one-day and two-day comparisons this article was meant to carry are therefore not measurable: a past day's measurement cannot be redone. Only the day's floor and the three-day gap remain. A guard has been added so that importing triggers nothing.

Two points in time, not a curve.

Six points at three days says nothing about how drift behaves at a week or a month. We do not extrapolate it.

Four sectors, two languages, two engines.

Running shoes are excluded: the answers there cite numbered models our extractor discards, which made the sector come out at 0 % everywhere - a filter artefact, not a result.

Our extractor remains imperfect.

That is why both cleaning levels are published side by side: the conclusion holds in both, or there is no conclusion.

A calm period stays a calm period.

These three days saw no major launch and no announced model change in the sectors tested. Weak drift in calm conditions promises nothing in turbulent ones.

We promise nobody a place in ChatGPT's answers - nobody can do that honestly. What can be measured is where you stand: per language, per engine, per question asked. And now, with some idea of what time does to it.

Frequently asked questions

How often should I measure my AI visibility?
Less often than you would think, according to our 27 August 2026 measurement. Three days apart, the recommended list overlaps 70 %, against 76 % between two runs made minutes apart: a gap of only six points. A measurement budget is better spent on the variety of questions asked than on their frequency.
Does the list of companies an AI recommends change from one day to the next?
Very little over three days, and it depends on the engine. Perplexity, which repeats itself well in the short term (85 %), shows a real drift of 14 points in three days. Gemini shows none - not from stability, but because it is already so inconsistent with itself (68 % between consecutive answers) that three days add nothing.
What moves an AI's recommendations the most?
The question, not the calendar. Changing the buying context - « what are the best X » against « which one for my ten-person company » - costs 46 points of overlap. Waiting three days costs 6. That is seven to one, measured on the same design on two consecutive days.
Is a single audit enough to know whether an AI recommends me?
No. Two consecutive answers to the same question overlap only 68 % on Gemini and 85 % on Perplexity. A single run cannot tell a real change from the day's randomness. Repetition is there first to silence that noise, not to follow the news.
Why publish a result that argues against daily measurement?
Because we sell repeated measurement, and a figure chosen to suit the seller is worth nothing. The question was asked properly, with a noise floor and two cleaning levels, and the answer is published as it came.