Back to Blog
AI Compliance

One Control Plane, Three Rulebooks: Building Agents That Can Prove What They Did

ISO/IEC 42001, the EU AI Act, and AIUC-1 overlap by roughly 70 percent, yet most programmes run three separate workstreams. Eight controls, built once, satisfy all three.

The blocker for stalled agent pilots is evidence, not model quality. Here is the one control set that satisfies ISO/IEC 42001, the EU AI Act, and AIUC-1 at once, and the twelve-week sequence to build it.

September 4, 202614 min read
ISO 42001EU AI ActAIUC-1Agent GovernanceAudit Trail
AgentTrustOSAI COMPLIANCEISO 42001 · EU AI ACT · AIUC-1One Control Plane,Three RulebooksBuilding agents that can prove what they didTHREE RULEBOOKS, ONE SETISO/IEC 42001EU AI Act, Art. 9-15AIUC-1, 6 domainsONE CONTROL SET8 controls, 3 renderingsSources: ISO/IEC 42001:2023 · Regulation (EU) 2024/1689 · AIUC-1 · AgentTrust OS 2026agent-trust.tech
Three instruments, one origin question: if the evidence is not produced by the system that does the work, it is a story, not a record.
AI Compliance · Agent Governance · Enterprise Risk
Key Facts
  • ISO/IEC 42001, published December 2023, is the first certifiable management system standard for artificial intelligence — structurally a sibling of ISO/IEC 27001.
  • Regulation (EU) 2024/1689, the EU AI Act, entered into force 1 August 2024. Obligations on high-risk systems listed in Annex III were legislated to apply from 2 August 2026. Penalties run to €35M or 7% of global turnover for prohibited practices.
  • AIUC-1, introduced in 2025 by the Artificial Intelligence Underwriting Company, is built with insurers and independent auditors so the risk of deploying a third-party agent can be underwritten — it requires independent audit and adversarial-testing evidence, not self-attestation.
  • The three rulebooks overlap by roughly 70% in what they actually require. Eight controls, built once and mapped three ways, satisfy all three audiences.
  • The Commission's Digital Omnibus package has proposed targeted adjustments to some AI Act application dates — confirm current status with counsel before building a plan around a specific month. The substance of Articles 9–15 has not been in question.
TL;DR
  • The blocker for most stalled agent pilots is evidence, not model quality — nobody can produce a defensible record of how the agent behaves, who approved it, and what happens when it is wrong.
  • Three rulebooks now govern that record — ISO/IEC 42001 for the management system, the EU AI Act for legal obligation, AIUC-1 for independent audit — and they overlap by roughly 70%.
  • Compliance should be an output of the runtime, not a parallel documentation exercise. Eight controls, enforced in the request path, satisfy all three at once.
  • The run log carries the weight. Seven of the other eight controls cannot be evidenced without a trustworthy, tamper-evident log underneath them.
  • Build the log first. Every team that started with the governance framework instead needed a restart; every team that started with the log did not.
Keep reading → the eight controls, the twelve-week build sequence, and where AgentTrust OS fits.

Ask an engineering team to demo their agent and you will be impressed. Ask them to prove what it did last Tuesday, on whose behalf, and who approved it, and the room goes quiet. That silence, not model quality, is what keeps most enterprise agents stuck in pilot.

A team builds an agent that genuinely works. It triages alerts, reconciles ledgers, drafts adjudication notes, answers customers. Stakeholders see the demo and want it live in six weeks. Then it meets the second line of defence, and a handful of very ordinary questions turn out to be unanswerable: which model version produced this output, and can you reproduce it; what data did it retrieve, and was that the customer's own; which tool calls can it make without a human, and who set that boundary; show me the last thirty escalations and what the human decided; if a regulator asks for this in eighteen months, where is it stored.

Definition — The Evidence Gap

The distance between an agent that behaves well and an organization that can prove it behaves well. None of the questions above are model questions — they are systems questions, and the honest answer in most pilots is that the information exists in fragments across an application log, a wiki page, a chat thread, and one engineer's head. The pilot does not fail. It simply never leaves the pilot.

Three rulebooks pointing the same way

Three instruments now define what "provable" means for an AI agent. They came from very different places — a standards body, a legislature, and an insurance market — and that is exactly why reading them together is useful.

ISO/IEC 42001 — the management system
The same Annex SL clause structure as ISO/IEC 27001 — context, leadership, planning, support, operation, performance evaluation, improvement — wrapped around a Plan-Do-Check-Act cycle. Certifiable by an accredited body, which makes it the first AI artefact a procurement team can treat the way it treats an ISO 27001 certificate.
EU AI Act — the legal obligation
For a high-risk agent, Articles 9–15 read like an engineering specification: lifecycle risk management (Art. 9), data governance (Art. 10), technical documentation to Annex IV (Art. 11), automatic event logging (Art. 12), transparency (Art. 13), human oversight designed in (Art. 14), and accuracy, robustness and cybersecurity (Art. 15).
AIUC-1 — the view from someone taking the risk

Structured around six domains — data and privacy, security, safety, reliability, accountability, societal impact — and requires independent audit and evidence of adversarial testing rather than self-attestation. It publishes crosswalks to the NIST AI RMF, the EU AI Act, ISO/IEC 42001, and MITRE ATLAS. Put the three side by side: ISO tells you how to manage, the AI Act tells you what is mandatory, AIUC-1 tells you what somebody willing to take financial risk on your agent actually wants to see. An insurer has no incentive to accept a well-written policy document in place of test results.

Around those three sit the regional expectations most likely to knock on your door first: NIST AI RMF 1.0 and its Generative AI Profile in North America, OSFI Guideline E-23 in Canada, the FCA/PRA's technology-neutral posture in the UK, MAS FEAT and AI Verify in Singapore. Long list, real convergence — the same eight demands keep appearing across all of them.

What most teams get wrong

Compliance for AI agents should be an output of the runtime, not a parallel documentation exercise running alongside it. Three things follow.

01
One control set, three dialects
Define the controls once. Map each to its ISO Annex A objective, its AI Act article, and its AIUC-1 domain. When an auditor arrives, render the same underlying evidence in whichever vocabulary they use. Never maintain three descriptions of one system.
02
Enforcement in the path, not the policy
A rule that lives in a document is a hope. A rule that lives in a gate the agent cannot bypass is a control. Every control should have a runtime enforcement point and a log line, or it does not count.
03
Evidence generated, not assembled
The run log is written as the agent works. The evidence pack, the thing you hand an auditor, is a rendering job over that log. It should be a report, not a project.
If I could change one question in the room

Stop asking "are we compliant with the EU AI Act?" and start asking "can our runtime produce, unprompted, a tamper-evident record of every decision this agent made last quarter?" If the answer is yes, compliance becomes a mapping exercise. If the answer is no, no amount of documentation will save the audit.

The clause mapping, control by control

Eight controls, built once, satisfy all three rulebooks. What differs between audiences is only the rendering — a Statement of Applicability for the certification body, an Annex IV technical file for conformity assessment, a domain-by-domain bundle for the AIUC-1 auditor, all generated from the same underlying log.

ControlISO/IEC 42001EU AI ActAIUC-1 domainEvidence artefact
1 · Lifecycle risk & impact assessmentA.5, Cl. 6.1Art. 9ReliabilityRisk register and sign-off
2 · Data governance and provenanceA.7Art. 10Data & PrivacyDataset sheet and lineage
3 · Technical documentationA.6Art. 11, Annex IVAccountabilityVersioned agent card
4 · Automatic logging & traceabilityA.6, A.9Art. 12AccountabilityTamper-evident run log
5 · Human oversightA.9Art. 14SafetyEscalation & decision records
6 · Accuracy, robustness, cybersecurityA.6Art. 15Security & SafetyScorecard & red-team report
7 · Transparency and disclosureA.8Art. 13, Art. 50SocietyDisclosure copy & UX proof
8 · Third-party and supply chainA.10Art. 25AccountabilityAttestations & model BOM

Control 4 carries the weight — seven of the other controls cannot be evidenced without a trustworthy log underneath them. Clause references are indicative for planning; confirm against published texts before certification.

Inside the eight controls

Policy gate: bind the purpose before anything executes

The policy gate answers, before any model call, three questions: who is asking, for what declared purpose, and what is the risk tier of the most consequential action this request could trigger. Purpose binding is the piece that gets skipped — a request declared as "customer balance enquiry" should not be able to end in a funds transfer, even if the model reasons its way there. Encoding purpose at the gate and re-checking it at the tool broker turns a whole family of prompt-injection outcomes from a security incident into a denied call with a log line.

Agent identity: stop agents borrowing human credentials

The most common architectural flaw in pilots is an agent running with a service account that inherits the permissions of the most privileged user it serves. It makes the demo easy and the audit impossible, because you can no longer distinguish an action taken by a person from an action taken by software on that person's behalf. Each agent needs its own workload identity, scoped credentials, and a delegation record when it acts for a user — short-lived tokens, scopes narrowed to the job, an on_behalf_of claim carried through the whole call chain.

Tool broker: the real security boundary

Prompts are advisory. Tools are what actually change the world, so the tool broker is where enforcement belongs. Every tool is registered with a declared risk tier, an input/output schema, a rate limit, a quota, and an idempotency requirement for anything that writes. Deny by default — an agent gets a tool allowlist tied to its job, not the full catalogue — and the broker re-validates purpose and tier at the point of action, because the plan may have changed since the gate.

Guardrails, and the abstention path everyone forgets

Input screening covers prompt injection, jailbreak patterns, and indirect injection carried inside retrieved documents — the vector most teams have never tested. The piece left out is abstention: an agent needs a first-class way to say "I am not confident, escalate this," and that path has to be cheap and well instrumented. Agents without an abstention path do not become more accurate, they become more confident, which is worse.

Human oversight, measured honestly

Article 14 asks for oversight a human can meaningfully exercise — a higher bar than a rubber-stamp queue. Two numbers tell you the truth: median review time and override rate. Four seconds on a complex adjudication is not oversight, it is a queue being cleared. An override rate near zero over months means either an excellent agent or a disengaged reviewer, and you need eval data to know which.

The run log and the evidence pack

Article 12 requires automatic recording of events over the lifetime of a high-risk system. Per run, the log needs: a run identifier, agent identity and version, model identifier and version, the full prompt assembly including retrieved context, every tool call with arguments and results, guardrail verdicts, escalations and human decisions, final output, latency, tokens, and cost — append-only with hash chaining, retained to the regulated process rather than an observability vendor's 30-day default. The evidence pack is a renderer over that log, producing an Annex IV technical file, a Statement of Applicability, or a domain-by-domain audit bundle on demand. Build it once and the marginal cost of the next audit collapses.

The Agent Control PlaneTriggerPolicy Gatepurpose + risk tierAgent Runtimemodel + plannerTool Brokerscopes + quotasAgent Identityown scoped credentialI/O Guardrailsinjection, PII, abstainHuman Oversightapprove, amend, rejectIMMUTABLE RUN LOG, APPEND-ONLY AND TAMPER-EVIDENTprompt · retrieved context · every tool call · model & version · approvals · latency · tokens · costEvidence Pack, rendered on demand
Figure 1: The agent control plane. Every obligation in the three rulebooks has a physical enforcement point in the request path and a matching log line — the evidence pack is a query over that log, not a project.

Building it in twelve weeks

This sequence assumes one product team, one risk partner, and an agent already in pilot. The order matters more than the durations.

Weeks 1

Classify. Risk classification per agent and per tool, purpose statements, a four-tier action model agreed with risk.

How you know it worked: second line signs the tier table.
Weeks 2–3

Log first. Run-log schema, append-only store, hash chaining, retention tiering, instrument the existing agent.

How you know it worked: you can replay any run from the log alone.
Weeks 3–4

Identity. Workload identity per agent, scoped credentials, delegation claim carried to systems of record.

How you know it worked: no agent uses a shared human account.
Weeks 4–6

Broker. Tool registry with tiers, schemas, quotas, idempotency; deny-by-default allowlists.

How you know it worked: an off-allowlist call is denied and logged.
Weeks 6–8

Guardrails. Input and output screening, indirect-injection tests on retrieval, abstention wired to escalation.

How you know it worked: red-team suite runs in CI.
Weeks 8–9

Oversight. Review console with reasoning trace and sources, amend capability, review-time and override telemetry.

How you know it worked: both oversight metrics on a dashboard.
Weeks 9–11

Evidence. Evidence renderer — Annex IV set, Statement of Applicability, audit bundle.

How you know it worked: pack generated in under a day, unassisted.
Week 12

Dry run. Internal audit against the pack; gaps become backlog items with owners.

How you know it worked: internal audit raises no critical finding.
Five ways I have watched this go wrong

Documentation first — the policy set is written before the runtime exists, and twelve weeks later the runtime does something different. The shared service account — the hardest item on this list to retrofit, which is why identity belongs at week three, not week nine. Guardrails on the prompt only — teams screen user input, ship, then find the injection arrived inside a retrieved PDF. Oversight theatre — an approval queue with a four-second median review time passes a walkthrough and fails the first real incident. Retention set by the observability tool — thirty-day retention on a process with a seven-year record-keeping obligation, discovered during audit, unfixable in retrospect.

What it is worth

The safety argument lands with risk and rarely with finance. These numbers tend to move a budget conversation, because they are about what the organization keeps.

2Q
of delay on a $1.2M/yr agent costs roughly $600K — usually several times what the control plane took to build
200–400
person-hours spent assembling evidence by hand for one regulated use case, every year, per audit
~0.7×
per-agent governance cost after the first, once the control plane exists — the fifth agent is a configuration exercise
AgentTrust OS: The Eight Controls, Built Once

The control plane in Figure 1 is not a diagram to build from scratch. AgentTrust OS implements it across the agent lifecycle, so the evidence for ISO 42001, the EU AI Act, and AIUC-1 is a rendering job over a runtime you were already going to need.

Trust Certify
Certifies the policy gate, agent identity boundary, and guardrail layer before production — Agent Discovery maps the real tool surface, the Attack Engine runs prompt injection and tool abuse against it, and the Compliance Engine checks the result against all three rulebooks in one pass.
Trust Runtime
Is the tool broker and human-oversight tier, enforced live. Every tool call is scored, routed to auto-approve, escalated, or blocked at the moment it is attempted, with the specific rule that fired attached to the result.
Trust Audit
Is the run log and evidence pack, generated as a by-product of running the agent. Every decision carries who authorized it, what was decided, the confidence score, and whether it fell inside policy.

Frequently Asked Questions

Most of it, yes. ISO/IEC 42001 and AIUC-1 apply regardless of your AI Act risk tier — they are a management system standard and an insurance-grade certification, not EU legislation. And the underlying design principle, that compliance should be a query over a runtime log rather than a parallel documentation exercise, is true for any regulated or reputationally sensitive agent, high-risk classification or not. What changes if you are not high-risk is urgency, not architecture.
SOC 2 covers your infrastructure and organizational controls — access management, change management, availability. It says nothing about whether a specific agent decision was authorized, what data it retrieved, or whether a human meaningfully reviewed a specific escalation. Those are agent-level, decision-level questions, and they need a decision-level log. SOC 2 and ISO/IEC 42001 are complementary, not overlapping — most enterprises running agents in regulated workflows end up needing both.
For one agent already in pilot, yes — that is the scope the timeline in this post assumes. The honest caveat is the log: if you build it first, as recommended, the other seven controls become verification work rather than construction work, because most of them are really just log queries with a policy attached. Skip the log and try to build documentation-first, and twelve weeks becomes a moving target, because the documentation keeps needing to be rewritten as the runtime changes underneath it.
Letting an agent run under a shared service account. It is the fastest way to ship a demo and the single hardest thing to retrofit once a portfolio of agents has grown up around it, because by then every downstream system has learned to trust that one identity. If you take one thing from this post before you take anything else, give every agent its own scoped credential from day one.
No single platform closes every gap here, and AgentTrust OS is no exception. Trust Certify and Trust Runtime give you the policy gate, the tool broker, and the run log as a built system rather than something you assemble from scratch — but you still need to do the actual risk classification and purpose-binding work for your specific agents, because nobody outside your organization can tell you which of your workflows are Annex III high-risk. What the platform removes is the twelve weeks of undifferentiated plumbing between deciding you need this and having an evidence pack an auditor will accept.
Ready to Close the Evidence Gap?

No AI Agent Enters Production Without AgentTrust

Certify before production, enforce the policy gate live, and render the evidence pack on demand — one control set, three rulebooks satisfied.

Start Free →

More from the blog

AI ComplianceJuly 22, 2026AI ComplianceJuly 22, 2026AI GovernanceJuly 22, 2026AI GovernanceJuly 22, 2026AI ArchitectureJuly 22, 2026AI SecurityJuly 16, 2026MLOpsJuly 10, 2026AI ImplementationJuly 8, 2026AI Agent ArchitectureJuly 5, 2026AI Agent ArchitectureJune 30, 2026EngineeringJune 23, 2026AI Agent ArchitectureJuly 2, 2026AI Agent ArchitectureJuly 1, 2026IntegrationsJuly 1, 2026IntegrationsJuly 1, 2026IntegrationsJuly 2, 2026AI SecurityJuly 3, 2026AI SecurityJuly 1, 2026AI ComplianceJuly 2, 2026AI ComplianceJuly 3, 2026AI StrategyJuly 2, 2026AI StrategyJuly 3, 2026AI StrategyJuly 3, 2026AI StrategyJuly 3, 2026AI GovernanceJuly 28, 2026Healthcare AIJuly 28, 2026ArchitectureJuly 29, 2026ArchitectureJuly 29, 2026AI StrategyJuly 29, 2026AI StrategyJuly 30, 2026AI SecurityJuly 30, 2026AI ComplianceJuly 30, 2026EngineeringJuly 30, 2026AI GovernanceAugust 4, 2026EngineeringAugust 4, 2026EngineeringAugust 4, 2026Agent SecurityAugust 19, 2026EngineeringSeptember 4, 2026AI StrategySeptember 4, 2026AI GovernanceSeptember 4, 2026