Three days later, the list has barely moved - 6 points, against 46 for a different question
We sell repeated measurement. So we tested whether the list drifts from one day to the next. It barely does - and for one engine, not at all.
In this article
Short answer: the day you ask changes almost nothing. How you ask changes everything. Three days apart, the same question returns 70 % overlap, against 76 % between two runs made minutes apart. Six points. Yesterday we measured that changing the buying context cost 46.
We sell repeated measurement. It would suit us to conclude that you must measure often, and that is exactly why the question had to be asked properly, accepting the inconvenient answer in advance.
What moves the list is not the calendar. It is the question, and it is the engine.
What we measured
The same reference question every measurement of ours has used since 23 August - identical to the character, verified rather than assumed. Four global sectors, two languages, two assistants that search. 60 runs today, zero failures, plus a 24 August collection kept as the three-day comparison point.
The noise floor is still the load-bearing part. An assistant already contradicts itself minutes apart. A difference between two days means nothing until it is read against what randomness produces in minutes. So we ran the question three times today to get the day's own floor, rather than importing the one from two days ago.
The result
| Gap between the two runs | Cleaned | Raw |
|---|---|---|
| Same day, minutes apart (floor) | 76 % | 61 % |
| Three days | 70 % | 53 % |
| Distance from the floor | 6 points | 8 points |
Six points on the strictest computation, eight on the raw. Either way, the list from three days ago resembles today's about as much as two runs made in the same minute.
And here the two engines stop resembling each other
| Engine | Same day | Three days | Distance |
|---|---|---|---|
| Perplexity | 85 % | 71 % | 14 points |
| Gemini | 68 % | 69 % | none |
Gemini does not drift at all. Three days later, its list resembles today's exactly as much as two of its own consecutive answers do. That is not stability: it is that it is already so inconsistent with itself - 68 % between successive runs - that three days have nothing left to add.
Perplexity is the opposite. It repeats itself well in the short term, 85 %, and that is precisely why a real drift becomes visible: 14 points in three days. There, measuring a week apart means something, because the noise does not drown the signal.
On one engine, « how often should I measure? » has an answer. On the other the question does not arise, because a single run there already means nothing.
Six against forty-six
The figure only means something beside yesterday's, obtained on the same design, the same floor and the same cleaning.
| What you change | Cost in points |
|---|---|
| The day you ask (3 days) | 6 |
| Price as the criterion (« best value for money ») | 24 |
| Level (« I'm a beginner ») | 37 |
| Context (« my 10-person company ») | 46 |
Seven times more. An audit that repeats the same phrasing daily measures, very precisely, one eighth of the subject. A measurement budget is better spent on the variety of questions than on their frequency.
What we take from this for our own product
Daily measurement has no justification at this stage.
Over three days the gap stays inside the noise. We will not sell a cadence our own figures do not support.
A single measurement remains unusable
, which is the other face of the same result: at 68 % overlap between consecutive answers, one run on Gemini establishes nothing. Repetition is there to silence noise, not to follow the news.
The useful cadence depends on the engine.
On Perplexity a real drift shows in three days. On Gemini you first need several runs to know what you are looking at.
The lever is the variety of questions.
That is where 46 points sit, against 6 for the calendar.
The limits, and a data loss of our own making
We destroyed our own 26 August data. An analysis script imported a function from the previous day's collector; that collector ran its measurement at module load and rewrote its output file. 120 runs cut to 12, with no copy. The one-day and two-day comparisons this article was meant to carry are therefore not measurable: a past day's measurement cannot be redone. Only the day's floor and the three-day gap remain. A guard has been added so that importing triggers nothing.
Two points in time, not a curve.
Six points at three days says nothing about how drift behaves at a week or a month. We do not extrapolate it.
Four sectors, two languages, two engines.
Running shoes are excluded: the answers there cite numbered models our extractor discards, which made the sector come out at 0 % everywhere - a filter artefact, not a result.
Our extractor remains imperfect.
That is why both cleaning levels are published side by side: the conclusion holds in both, or there is no conclusion.
A calm period stays a calm period.
These three days saw no major launch and no announced model change in the sectors tested. Weak drift in calm conditions promises nothing in turbulent ones.
We promise nobody a place in ChatGPT's answers - nobody can do that honestly. What can be measured is where you stand: per language, per engine, per question asked. And now, with some idea of what time does to it.