The correction arrived. It just arrived from outside, after release, once the cost had already been incurred.
In October 2025, Deloitte Australia agreed to repay the final instalment of an A$440,000 contract with the Department of Employment and Workplace Relations. Chris Rudge, a University of Sydney researcher in health and welfare law, had identified references to nonexistent academic papers and a fabricated quote from a federal court judgment in the 237-page report. A revised version was published, disclosing that a generative AI system had been used in writing it. The account is here.
One detail is worth stating precisely, because it changes where the argument has to sit. Deloitte did not say that AI caused the errors. Asked directly, the firm declined to answer. What is established is that the errors were there, that they were corrected, and that the revised report disclosed the use of a model.
The immediate explanation was predictable anyway: someone trusted AI-generated content that should have been checked by qualified professionals.
That explanation may well be accurate. It is certainly incomplete.
Wherever the errors came from, the failure did not end there. It continued through an operating process that allowed unsupported claims to survive review and enter a client deliverable.
What was missed
At the time, I wrote that Deloitte had treated AI as an answer machine rather than a diagnostic instrument.
Instead of asking only for a completed report, the process needed to ask:
- What evidence supports each conclusion?
- Which citations have been independently verified?
- Where do the models disagree?
- Which assumptions entered the analysis?
- What remains unresolved?
- Who has the authority to approve publication?
Telling people to ask better questions helps. Training them to become sophisticated AI users helps. Requiring human review helps.
None of those measures proves that meaningful challenge occurred.
A reviewer can read a polished report and miss the same unsupported claim as the person who produced it. Several reviewers can agree because they inherited the same frame, relied on the same sources, or assumed someone else had verified the evidence.
That is correlated error disguised as assurance.
The real failure was the absence of a loop
The feedback that corrected the report existed. It came from Chris Rudge, a University of Sydney researcher, reading the published document and telling the media it was full of fabricated references. The process worked — from outside the process, after release, once the commercial and reputational cost had already been incurred.
That shape is familiar to anyone who has watched a system satisfy its gate and miss its purpose.
In 1998 I built a supplier ranking system for a defence maintenance operation measured on next-day parts delivery. After it went into production it took on an attribute nobody had specified at design: capturing defective and wrong-part shipments. That mattered more than it sounds. On-time delivery of the wrong item had been counting as success. The metric was satisfied. The technician still could not close the call.
Operating the system revealed that delivery had been the wrong measure — and the architecture absorbed the new one rather than being rebuilt around it. A system that improves its score on a fixed set of measures is optimizing. A system that reveals a missing measure and takes it on is discovering. Only one of those was ever the hard part.
I have been making that argument in public for a long time. In June 2005, keynoting the PTDA Canadian Conference in Calgary, I described what the algorithms were doing this way:*
Each and every transaction that happens, the system gets more and more intelligent. I won’t call it artificial intelligence, but it learns. *
Someone in the room pushed on the obvious weakness — how does the data stay current, and how long before any of it is worth anything. The answer is on the recording:
The transition of these systems and the history starts building from day one, and the system’s value becomes more valuable — six months, nine months, twelve months — because you never lose that data and never lose that intelligence. But you’ve got to start somewhere. *
I also described, that morning, what the loop produced: a bid broadcast to hundreds of suppliers where the cheapest supplier ranked last, because the system weighted delivery and quality history against price.* And I described why the record mattered — the information was captured immediately, reflected in the rankings, and created, in the language I used then, an audit trail to prove its value.*
That was 2005. The mechanism was built in 1998, and I have presented it many times since.
Virginia’s eVA is the same mechanism at institutional scale. It has been adjusted monthly for two decades — not because it was built wrong, but because a system in contact with real operating conditions keeps finding out what it was wrong about.
There is no learning without a continuous loop. A single pass produces a result. A loop produces a system capable of discovering that it was measuring the wrong thing.
Deloitte’s report passed its gate without an effective pre-release loop. The gate was satisfied: reviewed, approved, delivered, published. Whatever internal review occurred, it did not stop the difference between a completed review and a correct one from reaching publication. The difference was found by a reader, in public, two months later.
What a governed process does instead
The report would not have moved directly from model output to human editing to client release.
ARA™ challenges the frame first. What was the report required to establish? Which conclusions carried legal, financial or policy consequences? What evidence standard applied? Which claims required primary-source confirmation? What conditions would make the emerging conclusion wrong?
RAM 2025™ then submits the work to independent positions and progressive challenge. Agreement does not automatically count as verification. Each model’s objections, qualifications and unresolved concerns stay visible rather than dissolving into a consolidated draft.
SLAP OS™ preserves the decision record — which model raised each issue, what evidence supported or contradicted it, which correction was accepted, which objection was declined, who changed the reasoning frame, what remained unresolved at publication, and who authorized release.
Under that structure a fabricated citation becomes a failed evidence condition. An invented legal quotation requires confirmation against the authoritative source. A factual inconsistency stays visible until it is corrected, explicitly accepted as a limitation, or escalated to the authorized human decision-maker.
With those verification and release conditions properly configured and enforced, the unresolved items are inside the process rather than outside it. That is the whole of the claim: not that every error is caught, but that the loop runs before release rather than after.
We have now demonstrated the difference
On 27 September 2026, a multimodel panel reviewed one Procurement Insights post through five levels of challenge. Each round was recorded on the SLAP OS™ Multimodel Litmus Test™ — five dimensions, one column per reviewer, one cell per position.
At Level 5, all four reviewing models approved publication — and one of them identified a must-fix defect in the same breath: an unsourced superlative sitting inside a post that criticized unsupported claims.
The defect had survived four rounds of review, including every pass by the Anchor Model that produced the consolidated draft.
Level 5. Four reviewers on the Level 4 draft, the round that returns decision authority to the Human Orchestrator. Every verdict cell is clear. One accuracy cell is not. (SLAP OS™ Multimodel Litmus Test™, from ARA™ RAM 2025™ powered by SLAP OS™.)
What the panel caught on 27 September was an unsupported ranking claim rather than a fabricated source. So it is fair to ask whether the same governing principle has caught the defect class that appeared in the DEWR report. It has, and the instance is public. In October 2025 a model in the panel inserted the word “Amen” into a quoted comment. It was not there. It fit the passage well enough that it read as authentic, which is what makes that class of error dangerous. It was caught by a single question — where did you read that — and it never left the room. I published the case at the time, here.
A fabricated word inside a quotation and a fabricated quotation inside a court judgment are the same defect at different scales: a confident claim with no traceable source, wearing the appearance of something verified.
I am not going to tell you the panel would have caught Deloitte’s. I did not run it, and I have no standing to say what a process I did not apply would have produced. What I can say is that the defect class is one this process is built to raise before release rather than after — and that the published record shows it doing so.
Under a conventional workflow, the approval count can carry the decision. The document was polished. The reviewers agreed. Publication appeared justified.
Under this structure, the unresolved defect stayed visible.
That single rust-colored cell mattered more than the field of approvals around it. Consensus and correctness are not the same condition, and the matrix is built so the difference cannot be averaged away.
The difference between advice and an operating system
My October 2025 response identified three principles: challenge the norm; get it right rather than be right; document what others will not.
Those principles remain valid. What has changed is that they no longer depend on people remembering to follow them.
It does not merely advise users to question AI. It creates independent challenge. It does not merely recommend verification. It makes evidence discipline visible. It does not merely call for human oversight. It identifies the human authority responsible for the decision. It does not merely preserve the final answer. It preserves the reasoning, the disagreement, the corrections and the authorization that produced it.
The contemporaneous decision record behind the exercise has been preserved. It substantiates every cell in the matrix above. It is not being distributed, because it also contains the operating logic behind the instrument.
Would this make mistakes impossible?
No system can promise that.
People can disregard evidence. Decision-makers can override warnings. Organizations can weaken controls to meet a deadline or protect a preferred conclusion.
The difference is that the failure can no longer hide behind the AI made a mistake.
What the structure changes is that the unresolved items are visible before publication. If a report is released anyway, the decision transcript shows what was known, what remained unresolved, and who authorized the release.
That is the difference between having humans somewhere in the loop and governing the decision.
The Deloitte incident did not demonstrate that generative AI cannot produce professional work.
It demonstrated what happens when polished output passes a gate without an effective pre-release loop — a process that cannot reliably distinguish a good answer from a smooth one, and has nothing inside it that would find out.
A correction found by the client is not governance. The governing question is whether your loop finds the defect before the client does.
-30-
* Source for the 2005 passages. PTDA Canadian Conference keynote, Calgary, June 2005. The recording is public, in three parts:
All four passages quoted above are in the Q&A segment, at approximately 12:42 (the audit trail), 17:44 (it learns), 18:33 (the cheapest supplier ranked last) and 26:20 (value at six, nine and twelve months).
Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™
Related
No Loop, No Learning: What the Deloitte Report Reveals About Governed Review
Posted on September 28, 2026
0
The correction arrived. It just arrived from outside, after release, once the cost had already been incurred.
In October 2025, Deloitte Australia agreed to repay the final instalment of an A$440,000 contract with the Department of Employment and Workplace Relations. Chris Rudge, a University of Sydney researcher in health and welfare law, had identified references to nonexistent academic papers and a fabricated quote from a federal court judgment in the 237-page report. A revised version was published, disclosing that a generative AI system had been used in writing it. The account is here.
One detail is worth stating precisely, because it changes where the argument has to sit. Deloitte did not say that AI caused the errors. Asked directly, the firm declined to answer. What is established is that the errors were there, that they were corrected, and that the revised report disclosed the use of a model.
The immediate explanation was predictable anyway: someone trusted AI-generated content that should have been checked by qualified professionals.
That explanation may well be accurate. It is certainly incomplete.
Wherever the errors came from, the failure did not end there. It continued through an operating process that allowed unsupported claims to survive review and enter a client deliverable.
What was missed
At the time, I wrote that Deloitte had treated AI as an answer machine rather than a diagnostic instrument.
Instead of asking only for a completed report, the process needed to ask:
Telling people to ask better questions helps. Training them to become sophisticated AI users helps. Requiring human review helps.
None of those measures proves that meaningful challenge occurred.
A reviewer can read a polished report and miss the same unsupported claim as the person who produced it. Several reviewers can agree because they inherited the same frame, relied on the same sources, or assumed someone else had verified the evidence.
That is correlated error disguised as assurance.
The real failure was the absence of a loop
The feedback that corrected the report existed. It came from Chris Rudge, a University of Sydney researcher, reading the published document and telling the media it was full of fabricated references. The process worked — from outside the process, after release, once the commercial and reputational cost had already been incurred.
That shape is familiar to anyone who has watched a system satisfy its gate and miss its purpose.
In 1998 I built a supplier ranking system for a defence maintenance operation measured on next-day parts delivery. After it went into production it took on an attribute nobody had specified at design: capturing defective and wrong-part shipments. That mattered more than it sounds. On-time delivery of the wrong item had been counting as success. The metric was satisfied. The technician still could not close the call.
Operating the system revealed that delivery had been the wrong measure — and the architecture absorbed the new one rather than being rebuilt around it. A system that improves its score on a fixed set of measures is optimizing. A system that reveals a missing measure and takes it on is discovering. Only one of those was ever the hard part.
I have been making that argument in public for a long time. In June 2005, keynoting the PTDA Canadian Conference in Calgary, I described what the algorithms were doing this way:*
Someone in the room pushed on the obvious weakness — how does the data stay current, and how long before any of it is worth anything. The answer is on the recording:
I also described, that morning, what the loop produced: a bid broadcast to hundreds of suppliers where the cheapest supplier ranked last, because the system weighted delivery and quality history against price.* And I described why the record mattered — the information was captured immediately, reflected in the rankings, and created, in the language I used then, an audit trail to prove its value.*
That was 2005. The mechanism was built in 1998, and I have presented it many times since.
Virginia’s eVA is the same mechanism at institutional scale. It has been adjusted monthly for two decades — not because it was built wrong, but because a system in contact with real operating conditions keeps finding out what it was wrong about.
There is no learning without a continuous loop. A single pass produces a result. A loop produces a system capable of discovering that it was measuring the wrong thing.
Deloitte’s report passed its gate without an effective pre-release loop. The gate was satisfied: reviewed, approved, delivered, published. Whatever internal review occurred, it did not stop the difference between a completed review and a correct one from reaching publication. The difference was found by a reader, in public, two months later.
What a governed process does instead
The report would not have moved directly from model output to human editing to client release.
ARA™ challenges the frame first. What was the report required to establish? Which conclusions carried legal, financial or policy consequences? What evidence standard applied? Which claims required primary-source confirmation? What conditions would make the emerging conclusion wrong?
RAM 2025™ then submits the work to independent positions and progressive challenge. Agreement does not automatically count as verification. Each model’s objections, qualifications and unresolved concerns stay visible rather than dissolving into a consolidated draft.
SLAP OS™ preserves the decision record — which model raised each issue, what evidence supported or contradicted it, which correction was accepted, which objection was declined, who changed the reasoning frame, what remained unresolved at publication, and who authorized release.
Under that structure a fabricated citation becomes a failed evidence condition. An invented legal quotation requires confirmation against the authoritative source. A factual inconsistency stays visible until it is corrected, explicitly accepted as a limitation, or escalated to the authorized human decision-maker.
With those verification and release conditions properly configured and enforced, the unresolved items are inside the process rather than outside it. That is the whole of the claim: not that every error is caught, but that the loop runs before release rather than after.
We have now demonstrated the difference
On 27 September 2026, a multimodel panel reviewed one Procurement Insights post through five levels of challenge. Each round was recorded on the SLAP OS™ Multimodel Litmus Test™ — five dimensions, one column per reviewer, one cell per position.
At Level 5, all four reviewing models approved publication — and one of them identified a must-fix defect in the same breath: an unsourced superlative sitting inside a post that criticized unsupported claims.
The defect had survived four rounds of review, including every pass by the Anchor Model that produced the consolidated draft.
Level 5. Four reviewers on the Level 4 draft, the round that returns decision authority to the Human Orchestrator. Every verdict cell is clear. One accuracy cell is not. (SLAP OS™ Multimodel Litmus Test™, from ARA™ RAM 2025™ powered by SLAP OS™.)
What the panel caught on 27 September was an unsupported ranking claim rather than a fabricated source. So it is fair to ask whether the same governing principle has caught the defect class that appeared in the DEWR report. It has, and the instance is public. In October 2025 a model in the panel inserted the word “Amen” into a quoted comment. It was not there. It fit the passage well enough that it read as authentic, which is what makes that class of error dangerous. It was caught by a single question — where did you read that — and it never left the room. I published the case at the time, here.
A fabricated word inside a quotation and a fabricated quotation inside a court judgment are the same defect at different scales: a confident claim with no traceable source, wearing the appearance of something verified.
I am not going to tell you the panel would have caught Deloitte’s. I did not run it, and I have no standing to say what a process I did not apply would have produced. What I can say is that the defect class is one this process is built to raise before release rather than after — and that the published record shows it doing so.
Under a conventional workflow, the approval count can carry the decision. The document was polished. The reviewers agreed. Publication appeared justified.
Under this structure, the unresolved defect stayed visible.
That single rust-colored cell mattered more than the field of approvals around it. Consensus and correctness are not the same condition, and the matrix is built so the difference cannot be averaged away.
The difference between advice and an operating system
My October 2025 response identified three principles: challenge the norm; get it right rather than be right; document what others will not.
Those principles remain valid. What has changed is that they no longer depend on people remembering to follow them.
It does not merely advise users to question AI. It creates independent challenge. It does not merely recommend verification. It makes evidence discipline visible. It does not merely call for human oversight. It identifies the human authority responsible for the decision. It does not merely preserve the final answer. It preserves the reasoning, the disagreement, the corrections and the authorization that produced it.
The contemporaneous decision record behind the exercise has been preserved. It substantiates every cell in the matrix above. It is not being distributed, because it also contains the operating logic behind the instrument.
Would this make mistakes impossible?
No system can promise that.
People can disregard evidence. Decision-makers can override warnings. Organizations can weaken controls to meet a deadline or protect a preferred conclusion.
The difference is that the failure can no longer hide behind the AI made a mistake.
What the structure changes is that the unresolved items are visible before publication. If a report is released anyway, the decision transcript shows what was known, what remained unresolved, and who authorized the release.
That is the difference between having humans somewhere in the loop and governing the decision.
The Deloitte incident did not demonstrate that generative AI cannot produce professional work.
It demonstrated what happens when polished output passes a gate without an effective pre-release loop — a process that cannot reliably distinguish a good answer from a smooth one, and has nothing inside it that would find out.
A correction found by the client is not governance. The governing question is whether your loop finds the defect before the client does.
-30-
* Source for the 2005 passages. PTDA Canadian Conference keynote, Calgary, June 2005. The recording is public, in three parts:
All four passages quoted above are in the Q&A segment, at approximately 12:42 (the audit trail), 17:44 (it learns), 18:33 (the cheapest supplier ranked last) and 26:20 (value at six, nine and twelve months).
Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™
Share this:
Like this:
Related