In October 2025, KPMG published a report called Total Experience: Redefining Excellence in the Age of Agentic AI. In June 2026 the firm pulled it from its websites. The report cited 45 sources. According to an analysis by the detection company GPTZero, only five of those citations pointed to real, intact sources, and 40 of the 45 titles were fabricated.
What makes this worth writing about is not that a large firm had a bad day with a language model. It is how the problem was found. The report made claims about the AI programmes of named organisations, and those organisations objected. UBS, the UK's National Health Service, Swiss Federal Railways and Transport for London all disputed what the report said about them. The errors were not caught by review, or by an editor, or by the tooling. They were caught by the subjects of the claims, in public, after publication.
That is the part every professional services firm should sit with. The detection mechanism was the client.
The model did the expected thing
A language model asked to support an argument will produce citations that look exactly like citations. Plausible authors, plausible journals, plausible years. Fabrication is not a malfunction here, it is the predictable output of a system optimised to produce text that reads correctly. Anyone who has used these tools for research knows the shape of it.
So the interesting question is not "why did the model invent sources". It is "what was supposed to catch that, and why didn't it". Forty fabricated citations is not a subtle failure. Checking whether a reference resolves is mechanical. GPTZero did it from the outside, without access to the drafting process, apparently without much difficulty. Any check at all, applied once, before publication, would have found it.
There wasn't one. Or there was one on paper and it did not run.
The guidelines existed
KPMG's own response is the most instructive document in this story. A spokesperson said the firm removed the report while it conducted an investigation, and added: "We expect all our people to follow our guidelines on the responsible use of AI, including human oversight to validate content and verify independent sources."
Read that carefully. The control was already written down. Human oversight to validate content and verify independent sources is precisely the control that would have caught this, and it existed as an expectation before the report shipped. The gap was not knowing what to do. The gap was that nothing in the path from draft to publication required it to have been done.
This is the distinction we keep returning to. A policy is a statement about what people should do. A control is something that happens whether or not anyone remembers. When the only thing standing between a fabricated citation and a client-facing PDF is somebody's diligence on a deadline, the outcome is a matter of luck, and luck is not a governance posture. We have written about this before in Governance is code, not policy, and this is the most expensive worked example we have seen.
Provenance is not a feature you add at the end
The reason the citation check is hard to retrofit is that by the time you have a finished document, the link between each claim and its source has already been lost. You are left doing forensics on your own output, which is exactly what GPTZero had to do.
Systems that get this right do not verify at the end. They carry the source with the claim from the moment the claim is made, so that "where did this come from" is a lookup rather than an investigation. That is an architectural decision taken early, and it is unglamorous. It shows up as slightly slower drafting and slightly more plumbing, and it pays out exactly once, on the day something would otherwise have gone out with 40 invented references attached to your name.
The firms most exposed to this are the ones whose product is credibility. If you sell audit, advisory or assurance, the document is not a by-product of the work, it is the work. A withdrawn report is not an embarrassing news cycle, it is a direct hit on the thing being sold.
We nearly did the same thing writing this
We should be straight about how close this piece came to being an example of its own subject.
The KPMG story reached us as an internal content suggestion, raised automatically by one of our own research processes, with a single source link and a confident summary. The summary said the story broke in the week of 21 July 2026. When we opened the source, it was a short aggregator article with no report title, no named outlet behind it, and no link to any KPMG statement or to any other coverage. On its own it was not sufficient evidence that a Big Four firm had withdrawn a publication.
So we went looking, and the story held up: TechCrunch, The Register, Inc. and others reported it, the report title is specific, the GPTZero citation analysis is specific, and KPMG's statement is on the record. But two details from our internal summary did not survive contact with the sources. The date was wrong by about six weeks, and the specifics that make the story actually useful, the 45 citations and the disputing organisations, were absent from the summary entirely.
Had we published from the summary, we would have written a piece about fabricated sourcing that was itself built on unverified sourcing, and got the date wrong in public while doing it. The check took a few minutes. It is the same few minutes that were missing upstream of the report we are writing about, and we do not think we deserve much credit for it. That is the point. The check is cheap. It only has to be mandatory.
What this is and isn't
This is not a claim that KPMG is careless with AI, or that this could not happen to a firm with good practices. The opposite is closer to the truth: this happened to an organisation with published AI guidelines, a large assurance practice and every commercial reason to get it right. That is what makes it worth studying rather than enjoying.
Nor is it a prediction that the firm has a systemic problem. KPMG withdrew the report, said publicly that it was investigating, and named the control it expects. On the available evidence that is a reasonable response to a bad outcome.
It is also not a case for using less AI. A rule that says "do not use models for research" will be ignored under deadline, without anyone announcing it, which is how you end up with unreviewed model output in a client deliverable and no record of where it came from. The workable position is that model-assisted drafting is fine and unverifiable claims are not, and the difference between those two things is a control that runs every time.
Where we sit
ORCAHQ builds and operates governed AI systems, so we have an obvious interest in the conclusion that provenance matters. Read this with that in mind. We have tried to keep to what is on the record and to link it, and where we are describing our own process we have described the near-miss rather than the save.
The practical question we would ask of any organisation publishing AI-assisted analysis is narrow, and it is not about models at all. Between the draft and the client, what runs? If the answer is a person's judgement on the day, that is the same posture KPMG had, and the same expectation, written down, that did not execute. If the answer is that unresolvable citations cannot leave the building, you have a control.
The organisations named in that report were not doing KPMG a favour when they objected. They were doing the last check in the chain, from outside, in public, after publication. It is a slow and expensive place to find out.
Sources
- TechCrunch, KPMG pulls report on AI usage due to apparent hallucinations, 13 June 2026
- The Register, KPMG's AI report becomes an accidental demo of AI hallucinations, 12 June 2026
- Inc., KPMG sells 'AI trust' to clients, but it just pulled its own report over alleged AI hallucinations
- People Matters, KPMG withdraws AI report after UBS, NHS and Swiss Railways dispute claims