When the Machine Answers Wrong
Measured error rates in AI-generated answers about sources and organisations are high enough to be treated as an operating condition rather than an anomaly.
The failures of AI assistants are usually discussed as accidents: a bad answer, a fabricated citation, an embarrassing screenshot circulated for a day. Treated that way they are anecdotes, and anecdotes cannot be planned around. There are now measurements, taken by parties with no product to sell, and they support a different reading. Wrong answers about sources and organisations occur at rates that make them an ordinary output of the system rather than a malfunction of it.
What has been measured
The Tow Center for Digital Journalism published a study in March 2025 in which eight AI search tools were given 1,600 queries asking them to identify the source of excerpts from news articles. The tools collectively answered incorrectly more than 60% of the time. The spread was wide: Perplexity was wrong on 37% of queries, ChatGPT Search on roughly 67%, and Grok-3 on 94%, with 154 of its 200 citations resolving to dead links. The finding that matters more than the rates themselves is the behavioural one. With one exception, every tool was more likely to produce a confident wrong answer than to state that it could not determine the source.
A larger study reached a compatible result on different ground. The European Broadcasting Union and the BBC coordinated an examination of more than 3,000 AI assistant responses across 22 public broadcasters, 18 countries and 14 languages, published in October 2025. It found that 45% of responses contained at least one significant issue, and 31% had serious sourcing problems. The authors’ emphasis was on consistency: the failure pattern did not vary meaningfully by language or territory, which is what distinguishes a property of the systems from a property of any particular market.
These two studies examined news, which is the easiest domain to audit because the correct answer is documented and someone owns it. There is no reason to expect the systems to perform better on the harder, thinner, less contested material that describes a mid-sized company, and no equivalent body has audited that case.
Why the errors are not random
An error rate above 60% on source attribution is not noise around a correct answer. It indicates that the operation being performed is not the operation the reader assumes.
A citation, as a reader understands it, is a claim that a specific statement came from a specific place, checkable by going there. What these systems produce is a plausible-looking attribution generated alongside the answer, sometimes derived from a retrieved document and sometimes not. The dead links in the Tow Center’s Grok-3 results are the clearest illustration: a citation that points nowhere was never a record of consultation.
The design pressure runs the same way. A system that declines to answer is judged unhelpful, and a system that answers is judged helpful until someone checks. Since almost nobody checks, the incentive is stable. The Tow Center’s observation that the tools preferred a wrong answer to an admission of limitation is not a quirk of implementation; it is what optimising for apparent helpfulness produces.
The auditing problem
The practical difficulty for any organisation is not that the description may be wrong. It is that the description is generated per query, per user, and leaves no artefact.
Published material can be checked. A directory entry can be read, corrected and re-read. A generated answer exists once, is seen by one person, and is gone. Two people asking the same question in the same week may receive materially different accounts, and neither account is retained anywhere the subject can reach. There is no version, no timestamp, no correction procedure and no counterparty.
Nor is the citation layer a way back to the source. Similarweb’s measurement of US ChatGPT responses found that 6.8% carried a source citation in May 2026, up from 1.6% a year earlier — a fourfold rise that still leaves more than nine in ten answers with nothing attached to follow. An improving trend and a mostly unattributed corpus are the same fact seen from two directions.
Even where citations appear, they are not much used. The Reuters Institute’s Digital News Report 2026, drawing on 97,520 respondents across 48 markets, found that 42% of people who use AI chatbots for news click through to the underlying sources. The majority take the answer as given.
What follows for an organisation
Three responses do not work, and it is worth saying why before describing the one that does.
Monitoring the assistants directly by asking them questions about oneself produces a sample of one user’s results at one moment, generated by a system that is not deterministic and is retrained without notice. It is not nothing, but it should not be mistaken for measurement, and it will not detect a description that appears reliably to a different population of askers.
Correction at the point of output is unavailable. No mechanism exists for disputing an output that was never disclosed to the party it concerned, and the small number of feedback channels that do exist act on individual responses rather than on the underlying material.
Publishing more first-party material addresses only part of the problem, because the failures documented above are failures of attribution and synthesis rather than failures of supply. A system that misattributes an accurate statement is not short of source text.
What remains is the unglamorous option. The systems are wrong at measurable rates, but they are not wrong arbitrarily; they assemble from what is reachable, and the weight they give a version of events tracks how often and how consistently that version appears. An organisation whose public record is thin, contradictory, or supported only by its own domain supplies the conditions under which a confident wrong answer is the likely output. One whose record is consistent across several independent sources narrows the space in which the error can occur.
That is a reduction in exposure, not a guarantee, and it should be described as such. The honest summary of the current evidence is that the description of an organisation in circulation is partly outside its control, changes without notice, and is wrong often enough that treating accuracy as the default case is not supported by the data.