AI Compliance Evidence: 4 Proven Records Regulators Want

A few years ago, AI governance meant an ethics committee, a set of principles, and a slide deck the board saw once.

That will not survive an examination now.

The distinction auditors draw is between intent and enforcement. A policy document states what your organisation intends to do. It says nothing about whether that happened on any particular day, to any particular decision, involving any particular person.

Practitioners have a name for the failure mode: governance theatre. Static PDF policies and annual reviews look substantial on a shelf and produce nothing when a regulator asks to see a specific decision reconstructed.

The number that frames the problem: roughly 78% of enterprises have deployed AI, while only about 25% have formal governance documentation in place.

And the asymmetry that makes it urgent — building compliance evidence after the fact is exponentially harder than generating it continuously. Logs that were never captured cannot be recreated. A human review that was never recorded did not happen, evidentially speaking, even if it did happen in reality.

Key Takeaways
  • 78% of enterprises have deployed AI. Only about 25% have formal governance documentation. That gap is an open audit finding waiting to be written up.
  • A policy proves intent. Auditors ask for technical proof the policy was enforced, and audit trails are the largest evidence gap in most organisations.
  • If your log shows a service account rather than a named human, the record does not substantiate the assertion. Attribution failure voids the evidence entirely.
  • The EU AI Act sets two different clocks: logs for at least six months, technical documentation for ten years — covering every version, not just the current one.
  • ISO 42001 certification can accelerate EU AI Act readiness by an estimated 30–40%, and leaves real gaps in conformity assessment, logging retention and post-market surveillance.

Quick Navigation


The Four Tiers of AI Compliance Evidence

Not all evidence carries equal weight. Sorting it into tiers explains why well-resourced governance programmes still fail audits.

Tier 1 — Assertions. Policies, principles, acceptable-use documents, ethics statements. These establish that a standard exists. They prove intent and nothing else. Weakest tier, and the one organisations invest in most heavily.

Tier 2 — Attestations. Someone signed something. A sign-off form, a completed checklist, a manager’s approval. Better than an assertion because it names a responsible party, but self-reported and generated by the party being audited.

Tier 3 — Artefacts. Model cards, risk assessments, technical documentation, data lineage records, test results. Substantive and specific. Their limitation is that they are point-in-time snapshots, produced deliberately, and they drift from production reality between updates.

Tier 4 — Traces. Tamper-evident logs generated automatically at the moment an event occurs. Who accessed what data, under what authorization, with what outcome, at what time. Strongest tier because it is contemporaneous and not produced for the auditor.

The pattern that sinks audits: heavy investment in tiers 1 and 3, near-total absence at tier 4. Audit trails represent the largest evidence gap for most enterprises, and tier 4 is the one that answers “show me.”

Organisations that perform well under regulatory scrutiny are not those that spent most on governance tooling. They are those that can answer “show me” with records rather than policy documents.


What Regulators Ask For First

Five documentation categories are emerging consistently across financial services, healthcare and critical infrastructure examinations.

AI inventory. A complete, current list of every AI system in use — including systems a team stood up without telling anyone. You cannot evidence governance over a population you cannot enumerate, and shadow AI is the most common first finding.

Risk classifications. Each system mapped to applicable regulatory categories, with the reasoning recorded. Not just “we assessed this,” but what the assessment concluded and on what basis.

Policy and control documentation. Written policies paired with the technical controls that enforce them. The pairing is the point; a policy without a matching control is a tier 1 assertion.

Audit trails. Per-decision records with inputs, model version, output, and any human review. This is where most examinations break down.

Incident history. What went wrong, when it was detected, what was done, and how it resolved. An empty incident log is not a good sign — it usually indicates you are not detecting incidents rather than not having them.

The most common substantive failure is describable in one sentence: an output exists with no reasoning, no context and no trail, so the decision cannot be reconstructed.


The AI Compliance Evidence the EU AI Act Requires

The EU AI Act is the most prescriptive regime, so it is worth using as the reference standard even outside Europe.

For high-risk systems, the operative articles are:

ArticleRequirementEvidence produced
9Risk management system per systemRisk register, treatment plans, review records
10Data quality and governanceProof training and validation data was reviewed for bias and coverage gaps, with remediation documented
11 + Annex IVTechnical documentationStructured pack covering design through post-market monitoring
12Automatic event loggingOperational audit trail enabling traceability across the lifecycle
13Transparency toward deployersInformation packs enabling deployer oversight
14Human oversightDocumented oversight design, monitoring capability, override paths
15Accuracy and robustnessPre-market testing per harmonised standards

Two structural points matter more than the list.

These are requirements on the system and its lifecycle, not on the organisation surrounding it. An organisational management system does not discharge them.

The articles chain. Article 12 requires logging capability. Annex IV then requires your technical documentation to describe that capability — what is logged, how logs are retained, who has access. You cannot write the documentation before building the logging, which is why teams that start with documentation stall.

Note the current timing. Standalone Annex III high-risk obligations moved to 2 December 2027 under the Digital Omnibus. Article 50 transparency duties and GPAI obligations did not move.


Retention: How Long AI Compliance Evidence Must Live

Two clocks run simultaneously, and conflating them is a common planning error.

Logs: at least six months. Article 19 sets the floor for automatically generated logs, with deployers retaining them under Article 12.

Technical documentation: ten years. Article 18 requires retention for ten years from the date the system is placed on the market, and national authorities may request access at any point in that window.

AI compliance evidence

The ten-year requirement has an implication most teams miss on first reading. It applies to all versions, not just the current one. Documentation for superseded model versions must remain accessible.

For an organisation retraining quarterly, that is forty documentation versions to keep coherent over a decade — each linked to its specific training dataset, model weights and deployment configuration. This extends well beyond typical experiment-tracking retention policies.

What the ten-year pack includes: every documentation version, change logs and modification impact assessments, test results and validation reports, risk management records, post-market monitoring data, correspondence with notified bodies, and the EU declaration of conformity.

US state requirements are shorter but structurally similar. Colorado’s successor statute requires three years of records demonstrating compliance, a duty falling on both developers and deployers as set out in state AI laws and what builders and deployers each owe.


The Identity Problem That Voids AI Compliance Evidence

This is the failure mode least discussed and most likely to invalidate an otherwise complete evidence package.

An auditor asks for log records showing that authorization was evaluated and enforced for a specific access event. You produce them. If those records show a service account identity rather than a human user identity, the evidence does not substantiate the assertion.

The log exists. It is well-formed, timestamped and tamper-evident. And it proves nothing, because it cannot answer who acted.

The required standard is that AI data access logs contain individual user attribution, per-request policy enforcement decisions, and data-asset specificity equivalent to what is required for human data access. Most AI deployments currently generate logs meeting none of these.

This is where a security problem becomes an evidentiary one. When agents share credentials, attribution collapses — and attribution is the foundation every compliance claim rests on. The mechanics of that collapse are set out in why shared credentials are the real exposure.

The practical consequence for planning: distinct agent identity is not only a security control. It is a precondition for producing admissible evidence. Teams treating it as a security backlog item are deferring a compliance dependency.


Why ISO 42001 Is Not Enough on Its Own

ISO 42001 certification is worth having, and it is not a substitute for regulatory compliance. Both statements are true and the distinction matters commercially.

Analysis suggests certified organisations can accelerate EU AI Act readiness by roughly 30–40%. The governance infrastructure transfers, particularly for the organisational sections of Annex IV.

The gaps are specific:

Conformity assessment procedures. ISO certification is not an EU conformity assessment. Self-assessment or notified body assessment per the Act’s requirements remains separate work.

EU database registration. No ISO equivalent exists.

Logging retention specifics. ISO 42001 leaves retention to organisational policy. The Act sets floors.

Post-market surveillance. Ongoing monitoring obligations with defined content.

Annex IV format. ISO establishes documentation practices without prescribing the Act’s specific sections. Certified organisations typically cover the organisational sections and still need to produce detailed architecture, data and validation documentation.

The useful framing: ISO 42001 evidences that a management system exists. The AI Act requires evidence about specific systems and their lifecycles. One is about the organisation, the other about the artefact.

NIST AI RMF sits differently again — it earns an explicit enforcement safe harbour under Texas TRAIGA, which neither ISO nor the EU framework provides.


The AI Compliance Evidence Deployers Owe

Most coverage addresses providers. The heavier operational burden usually falls on deployers, and it is worth separating.

A provider produces evidence once per system version. A deployer produces evidence per decision, and the volume difference is enormous.

Three obligations drive it.

Pre-use notice records. Proof that a consumer was told before an AI system influenced a decision about them. Not a copy of your notice template — evidence that this particular person received it, when, and through what channel.

Adverse-outcome explanations. Where a decision went against someone, a plain-language account of the system’s role and the principal factors it used, typically within 30 days. This is the single heaviest lift in the current US frameworks, because it requires per-decision explainability that your model vendor may not supply and your contract may not oblige them to.

Human review records. Evidence that a meaningful review path existed and, where used, that a human with authority actually revisited the outcome.

The dependency worth flagging early: you cannot produce a deployer explanation from a provider who has not given you the underlying documentation. Intended uses, known limitations, the categories of data used in training — these flow from provider to deployer, and without them the deployer’s own obligation is unsatisfiable.

That makes evidence a procurement question. If your vendor contract does not name documentation as a deliverable, you have accepted an obligation you cannot discharge. Enterprises processing millions of AI-driven actions also find that manual audit preparation simply does not scale at this volume, which is what pushes evidence generation into the pipeline rather than a quarterly exercise.

One structural point closes the loop. The Act’s articles chain from provider to deployer the same way Article 12 chains to Annex IV. Provider documentation feeds deployer notices; deployer logs feed regulator inspections. A break anywhere in that chain surfaces at the far end, usually during an examination.


AI Compliance Evidence for Autonomous Agents

Every framework above was written for systems that make decisions about people. Agents that chain tool calls, modify records and take actions fit awkwardly.

Singapore moved first. Its Model Governance Framework for agentic AI, unveiled 22 January 2026, is the first framework specifically addressing autonomous system documentation. It requires organisations to define agent authority boundaries, autonomy classifications, and decision audit trails.

Those three requirements are a reasonable evidence template regardless of jurisdiction.

Authority boundaries. What is this agent permitted to do, expressed as scope rather than capability. Documented before deployment, not inferred afterwards.

Autonomy classification. Which actions proceed automatically, which require approval, which are prohibited. This is where human-in-the-loop checkpoints become an evidence artefact rather than only a safety control — an approval record is tier 2 evidence with a named actor and timestamp.

Decision audit trails. The full chain: which agent, spawned by which agent, on whose authority, touching which resource.

That third item is where multi-agent systems break most evidence architectures. Each delegation hop erodes attribution, and a trail that cannot name the originating human authority will not substantiate a compliance assertion.

Testing records deserve a specific note here too. Where a framework requires adversarial testing, a single passing result is weak evidence — an argument developed in why one passing red team test proves nothing. Record attempt counts and confidence bounds, not verdicts.


Generating AI Compliance Evidence Continuously

Six practices, ordered by dependency rather than difficulty.

Instrument before you document. Article 12 logging feeds Annex IV Section 3. Build the capability, then describe it. Teams that start with the document write fiction.

Attribute every action to a human or a distinctly identified agent. Service account logs do not substantiate assertions. This is the highest-leverage single fix.

Make storage tamper-evident. Standard database logs can be altered by a compromised admin account or a determined insider. Tamper-evident architecture is increasingly named in the technical annexes of governance frameworks.

Version documentation against system versions. Link each documentation version to the specific training dataset, model weights and deployment configuration it describes. Retention covers all versions.

Maintain a traceability matrix. Map each regulatory requirement to the specific document, section or record that addresses it. This is what makes an audit tractable rather than an archaeology project.

Capture human review as structured data. Who reviewed, when, what they saw, what they decided. Reviews recorded only in email threads or meeting notes are difficult to produce and easy to dispute.

The test to apply to your own programme: pick a decision your AI system made last quarter and try to reconstruct it. What data, which model version, what output, who reviewed it. If that takes more than an afternoon, you do not have an evidence layer — you have a policy binder.


Primary sources

Retention periods and article references reflect the AI Act as amended by the Digital Omnibus. This article is general information, not legal advice.


Frequently Asked Questions

Is a policy document enough to pass an AI governance audit?

No. A policy demonstrates intent. Audits increasingly require technical proof the policy was enforced — a current system inventory, documented risk classifications, testing records, data lineage and audit trails showing controls actually operated.

How long must AI logs and documentation be retained?

Under the EU AI Act, automatically generated logs must be kept for at least six months, and technical documentation for ten years from market placement, covering all versions. US state requirements are typically three years.

Does ISO 42001 certification satisfy the EU AI Act?

No. It can accelerate readiness by an estimated 30–40%, but gaps remain in conformity assessment, EU database registration, specific logging retention and post-market surveillance. ISO evidences a management system; the Act requires evidence about specific systems.

What is the most common AI compliance evidence gap?

Audit trails. Most organisations have policies and some documentation, and cannot produce per-decision records showing inputs, model version, output and human review.

Can I build compliance evidence retroactively?

Only partially, and at disproportionate cost. Documentation can be reconstructed; contemporaneous logs cannot. Events that were never captured leave no admissible record regardless of what actually happened.


Keep reading

AI compliance evidence

AI Compliance Evidence: 4 Proven Records Regulators Want

A few years ago, AI governance meant an ethics committee, a set of principles, and a slide deck the board saw once. That will not …

Read more

EU AI Act GPAI

GPAI Obligations: 4 Critical Gaps in the US Patchwork

A general-purpose AI model under the EU AI Act is a model capable of performing a wide range of distinct tasks. The obligations attach to …

Read more

Agent skills security

Agent Skills Security: 4 Hidden Gaps in Every Registry

An agent skill is a folder of instructions, scripts and resources that an AI agent discovers and loads on demand. Anthropic introduced the concept in …

Read more

AI Red Teaming: 4 Hidden Flaws in a Passing Test

Vendor datasheets lead with FLOPS. For most language model serving, FLOPS is the wrong number. Here is the physical reality of generating one token. The …

Read more

Advertisement

Leave a Comment