Somewhere in Denmark right now, an AI pilot is summarising compliance documentation. Drafting control descriptions. Answering an auditor’s pre-questions. And it sounds excellent doing it — because it always sounds excellent. That’s the property nobody in the room has fully priced in.
In an unregulated context, an AI that invents a plausible answer is a quality problem. Annoying, costly, fixable. In a GxP, NIS2 or DORA environment, the same behaviour is something else entirely: it’s fabricated evidence. A control description that reads beautifully but doesn’t describe the actual control. A summary of a validation document the model never opened. A status that was reconstructed from patterns rather than read from the system.
A hallucination in a regulated environment isn’t a wrong answer. It’s an audit finding with a delay on it.
I run an AI factory, and I’ve watched this failure mode from the inside: early on, an agent delivered a polished analysis of a dataset it had never had access to. Well-structured, specific, professional — and fiction. It survived two days of review, because linguistic confidence in a language model is not a quality signal. It’s a stylistic constant. The model sounds the same whether it read the source or invented it.
Now put that behaviour in front of a regulator.
The question an auditor will eventually ask
Regulated organisations have spent decades building evidence discipline for human work: signatures, audit trails, data integrity principles — who did what, when, based on which records. Then an AI arrives, gets wired into the compliance workflow, and produces output that carries none of that lineage. The auditor’s question is entirely predictable: “How do you know the AI read the source?” Most organisations currently deploying AI cannot answer it.
The answer that works is the same one that works for human evidence — provenance, made mechanical:
- Every operational claim carries an evidence classification — source-based, mixed, heuristic or unverified. In writing, attached to the output.
- Source access is logged, not assumed. Which document was opened, when, and which observations the conclusion rests on. No readback, no credit.
- “Unverified” is a legitimate, visible state — not something fluent text is allowed to paper over. In my factory the rule is mechanical: unknown = blocked.
- Live state beats the model’s description of state. An AI’s account of a system status is a claim, not a record. The system itself is the record.
None of this is exotic. It’s ALCOA-era thinking — attributable, legible, contemporaneous, original, accurate — applied to a new kind of worker. The organisations that already live by those principles for human work are, ironically, best placed to govern AI well. They just haven’t connected the two yet.
One procurement question worth stealing
If a vendor is selling you AI for anything compliance-adjacent, skip the capability deck and ask for their failure catalogue: the actual ways their system has failed, and the mechanical countermeasure for each. I publish mine — ten patterns, from live operation. A vendor who can’t show you one hasn’t run long enough to have one, or won’t admit what’s in it. Either answer tells you what you need to know.
Speed is not the enemy here — false speed is. Every hour spent on unverified output is slower than any hour spent on gates. In regulated environments, that sentence isn’t philosophy. It’s the difference between an AI capability and a finding.
Adapted from a post first published on LinkedIn.