134 wrong answers out of 200, 15 doubts expressed: being cited is not being cited correctly
Two major studies measured AI engine citations. More than 60% are incorrect, 31% have a sourcing problem, and the confidence never drops. What that changes for a brand.
more than 60%
of citations incorrect across 1,600 queries, eight engines tested (Tow Center, March 2025)
In this article
Short answer: an AI citation can name you and get you wrong. The published measurements are not about appearing, but about the accuracy of what is attributed: more than 60% incorrect citations in a study of 1,600 queries, 31% with sourcing problems in another covering more than 3,000 responses. The useful question is therefore not only « am I cited », but « is what I am made to say true, and does the link go anywhere ».
31%
of responses with a sourcing problem - missing, misleading or incorrect attribution (EBU and BBC, October 2025)
134 against 15
articles misidentified by ChatGPT, against the number of times it flagged a doubt, across 200 responses
An engine that is wrong and says so leaves you a chance. An engine that is wrong with confidence leaves none: the reader has no reason to go and check.
What two major studies measured
The first comes from the Tow Center for Digital Journalism at Columbia, published on 6 March 2025. The protocol is simple and reproducible: twenty publishers, ten articles drawn at random from each, excerpts copied verbatim, and eight answer engines asked to identify the headline, the original publisher, the date and the URL. That is 1,600 queries. Collectively, the engines returned an incorrect answer to more than 60% of them.
| Measure | Result |
|---|---|
| All eight engines together | more than 60% incorrect answers |
| Perplexity, best of the panel | 37% incorrect |
| Grok-3, worst of the panel | 94% incorrect |
| Grok-3, citations leading to an error page | 154 out of 200 |
The second is broader and comes from public service media. Published on 22 October 2025 by the European Broadcasting Union with the BBC, it had professional journalists evaluate more than 3,000 responses from ChatGPT, Copilot, Gemini and Perplexity, across 22 organisations, 18 countries and 14 languages. Its central conclusion is that the failure is systemic: it is tied neither to language, nor to market, nor to any one assistant.
| Type of fault | Share of responses |
|---|---|
| At least one significant issue | 45% |
| Sourcing problem: missing, misleading or incorrect attribution | 31% |
| Major accuracy issue: invented details, outdated information | 20% |
| Gemini, responses with a significant issue | 76% |
The worst part is not the error, it is the confidence
One figure from the Tow Center deserves isolating, because it decides everything else. ChatGPT misidentified 134 articles out of 200, and signalled a lack of confidence only fifteen times. Error and doubt are not correlated: the engine is wrong in two thirds of cases and says so in one case out of thirteen.
That is the difference between a fallible source and a misleading one. An engine that hesitates leaves the reader a signal, and therefore a chance to check. An engine that asserts leaves nothing. For a company it means an erroneous citation does not correct itself: nothing in the answer invites anyone to doubt it.
Links that lead nowhere
The other finding is more concrete still. More than half of the responses from Gemini and Grok-3 pointed to fabricated or broken URLs. For Grok-3, 154 citations out of 200 landed on an error page. The engine does not merely misattribute: it invents an address that has the appearance of proof.
The EBU study describes the same mechanism on the sourcing side: claims attributed to newsrooms that had never made them, links leading to an article on a different subject or to a page that no longer exists, and sometimes a complete citation - outlet name, headline, date - for an article that was never published. A well-formed reference is not a true one.
What this changes for a brand
The reasoning transfers directly. If an engine can attribute a sentence to a newsroom that did not write it, it can attribute your argument, your price or your speciality to a competitor - or the reverse, credit you with a promise you never made. In both cases you are « cited », and in both cases a counter that merely counts citations reports good news.
That is the limit of a visibility score taken alone, ours included. Appearing is a necessary condition, it is not the measure of the outcome. A citation is read with three questions: are we named, is what is attributed to us accurate, and does the offered link lead to a page that exists.
These studies are about news, not about business recommendations.
The Tow Center protocol is to trace the origin of a press excerpt; the EBU one has journalists evaluate news answers. The attribution mechanism is the same, but the transfer to commercial recommendations is ours and was not measured by this work.
They date from 2025, and the models have changed since.
The Tow Center published in March 2025, the EBU and the BBC in October 2025. The versions tested are no longer the ones served today, and neither measurement says what the same protocol would give this month. We have not redone it.
An average across eight engines hides enormous gaps.
Between 37% and 94% errors depending on the engine, the global « more than 60% » is an average that the worst performer lifts. Quoting that single figure to describe one particular engine would be a reading error.
We reproduced neither of them.
These are works published by a university research centre and by a consortium of public service media, read at source. They are not our measurements, and we do not present them as such.
What this changes
Read the citation, not just the counter.
Being named and being well described are two different outcomes. A dashboard that counts only appearances cannot tell good news from an error about you.
Check that the link exists.
An invented address on your domain looks like proof and is not. Click what the engine offers: half the responses from two tested engines led to a dead page.
Watch what is attributed to your competitors.
The error runs both ways: your argument can be credited to another name, and that is invisible from monitoring that only watches yours.
Do not count on the engine's doubt.
It is wrong far more often than it hesitates - 134 errors against 15 reservations across 200 responses. A confident tone says nothing about the accuracy underneath.