ISO/IEC 42001, the EU AI Act, and AIUC-1 overlap by roughly 70 percent, yet most programmes run three separate workstreams. Eight controls, built once, satisfy all three.
The blocker for stalled agent pilots is evidence, not model quality. Here is the one control set that satisfies ISO/IEC 42001, the EU AI Act, and AIUC-1 at once, and the twelve-week sequence to build it.
Ask an engineering team to demo their agent and you will be impressed. Ask them to prove what it did last Tuesday, on whose behalf, and who approved it, and the room goes quiet. That silence, not model quality, is what keeps most enterprise agents stuck in pilot.
A team builds an agent that genuinely works. It triages alerts, reconciles ledgers, drafts adjudication notes, answers customers. Stakeholders see the demo and want it live in six weeks. Then it meets the second line of defence, and a handful of very ordinary questions turn out to be unanswerable: which model version produced this output, and can you reproduce it; what data did it retrieve, and was that the customer's own; which tool calls can it make without a human, and who set that boundary; show me the last thirty escalations and what the human decided; if a regulator asks for this in eighteen months, where is it stored.
The distance between an agent that behaves well and an organization that can prove it behaves well. None of the questions above are model questions — they are systems questions, and the honest answer in most pilots is that the information exists in fragments across an application log, a wiki page, a chat thread, and one engineer's head. The pilot does not fail. It simply never leaves the pilot.
Three instruments now define what "provable" means for an AI agent. They came from very different places — a standards body, a legislature, and an insurance market — and that is exactly why reading them together is useful.
Structured around six domains — data and privacy, security, safety, reliability, accountability, societal impact — and requires independent audit and evidence of adversarial testing rather than self-attestation. It publishes crosswalks to the NIST AI RMF, the EU AI Act, ISO/IEC 42001, and MITRE ATLAS. Put the three side by side: ISO tells you how to manage, the AI Act tells you what is mandatory, AIUC-1 tells you what somebody willing to take financial risk on your agent actually wants to see. An insurer has no incentive to accept a well-written policy document in place of test results.
Around those three sit the regional expectations most likely to knock on your door first: NIST AI RMF 1.0 and its Generative AI Profile in North America, OSFI Guideline E-23 in Canada, the FCA/PRA's technology-neutral posture in the UK, MAS FEAT and AI Verify in Singapore. Long list, real convergence — the same eight demands keep appearing across all of them.
Compliance for AI agents should be an output of the runtime, not a parallel documentation exercise running alongside it. Three things follow.
Stop asking "are we compliant with the EU AI Act?" and start asking "can our runtime produce, unprompted, a tamper-evident record of every decision this agent made last quarter?" If the answer is yes, compliance becomes a mapping exercise. If the answer is no, no amount of documentation will save the audit.
Eight controls, built once, satisfy all three rulebooks. What differs between audiences is only the rendering — a Statement of Applicability for the certification body, an Annex IV technical file for conformity assessment, a domain-by-domain bundle for the AIUC-1 auditor, all generated from the same underlying log.
| Control | ISO/IEC 42001 | EU AI Act | AIUC-1 domain | Evidence artefact |
|---|---|---|---|---|
| 1 · Lifecycle risk & impact assessment | A.5, Cl. 6.1 | Art. 9 | Reliability | Risk register and sign-off |
| 2 · Data governance and provenance | A.7 | Art. 10 | Data & Privacy | Dataset sheet and lineage |
| 3 · Technical documentation | A.6 | Art. 11, Annex IV | Accountability | Versioned agent card |
| 4 · Automatic logging & traceability | A.6, A.9 | Art. 12 | Accountability | Tamper-evident run log |
| 5 · Human oversight | A.9 | Art. 14 | Safety | Escalation & decision records |
| 6 · Accuracy, robustness, cybersecurity | A.6 | Art. 15 | Security & Safety | Scorecard & red-team report |
| 7 · Transparency and disclosure | A.8 | Art. 13, Art. 50 | Society | Disclosure copy & UX proof |
| 8 · Third-party and supply chain | A.10 | Art. 25 | Accountability | Attestations & model BOM |
Control 4 carries the weight — seven of the other controls cannot be evidenced without a trustworthy log underneath them. Clause references are indicative for planning; confirm against published texts before certification.
The policy gate answers, before any model call, three questions: who is asking, for what declared purpose, and what is the risk tier of the most consequential action this request could trigger. Purpose binding is the piece that gets skipped — a request declared as "customer balance enquiry" should not be able to end in a funds transfer, even if the model reasons its way there. Encoding purpose at the gate and re-checking it at the tool broker turns a whole family of prompt-injection outcomes from a security incident into a denied call with a log line.
The most common architectural flaw in pilots is an agent running with a service account that inherits the permissions of the most privileged user it serves. It makes the demo easy and the audit impossible, because you can no longer distinguish an action taken by a person from an action taken by software on that person's behalf. Each agent needs its own workload identity, scoped credentials, and a delegation record when it acts for a user — short-lived tokens, scopes narrowed to the job, an on_behalf_of claim carried through the whole call chain.
Prompts are advisory. Tools are what actually change the world, so the tool broker is where enforcement belongs. Every tool is registered with a declared risk tier, an input/output schema, a rate limit, a quota, and an idempotency requirement for anything that writes. Deny by default — an agent gets a tool allowlist tied to its job, not the full catalogue — and the broker re-validates purpose and tier at the point of action, because the plan may have changed since the gate.
Input screening covers prompt injection, jailbreak patterns, and indirect injection carried inside retrieved documents — the vector most teams have never tested. The piece left out is abstention: an agent needs a first-class way to say "I am not confident, escalate this," and that path has to be cheap and well instrumented. Agents without an abstention path do not become more accurate, they become more confident, which is worse.
Article 14 asks for oversight a human can meaningfully exercise — a higher bar than a rubber-stamp queue. Two numbers tell you the truth: median review time and override rate. Four seconds on a complex adjudication is not oversight, it is a queue being cleared. An override rate near zero over months means either an excellent agent or a disengaged reviewer, and you need eval data to know which.
Article 12 requires automatic recording of events over the lifetime of a high-risk system. Per run, the log needs: a run identifier, agent identity and version, model identifier and version, the full prompt assembly including retrieved context, every tool call with arguments and results, guardrail verdicts, escalations and human decisions, final output, latency, tokens, and cost — append-only with hash chaining, retained to the regulated process rather than an observability vendor's 30-day default. The evidence pack is a renderer over that log, producing an Annex IV technical file, a Statement of Applicability, or a domain-by-domain audit bundle on demand. Build it once and the marginal cost of the next audit collapses.
This sequence assumes one product team, one risk partner, and an agent already in pilot. The order matters more than the durations.
Classify. Risk classification per agent and per tool, purpose statements, a four-tier action model agreed with risk.
How you know it worked: second line signs the tier table.Log first. Run-log schema, append-only store, hash chaining, retention tiering, instrument the existing agent.
How you know it worked: you can replay any run from the log alone.Identity. Workload identity per agent, scoped credentials, delegation claim carried to systems of record.
How you know it worked: no agent uses a shared human account.Broker. Tool registry with tiers, schemas, quotas, idempotency; deny-by-default allowlists.
How you know it worked: an off-allowlist call is denied and logged.Guardrails. Input and output screening, indirect-injection tests on retrieval, abstention wired to escalation.
How you know it worked: red-team suite runs in CI.Oversight. Review console with reasoning trace and sources, amend capability, review-time and override telemetry.
How you know it worked: both oversight metrics on a dashboard.Evidence. Evidence renderer — Annex IV set, Statement of Applicability, audit bundle.
How you know it worked: pack generated in under a day, unassisted.Dry run. Internal audit against the pack; gaps become backlog items with owners.
How you know it worked: internal audit raises no critical finding.Documentation first — the policy set is written before the runtime exists, and twelve weeks later the runtime does something different. The shared service account — the hardest item on this list to retrofit, which is why identity belongs at week three, not week nine. Guardrails on the prompt only — teams screen user input, ship, then find the injection arrived inside a retrieved PDF. Oversight theatre — an approval queue with a four-second median review time passes a walkthrough and fails the first real incident. Retention set by the observability tool — thirty-day retention on a process with a seven-year record-keeping obligation, discovered during audit, unfixable in retrospect.
The safety argument lands with risk and rarely with finance. These numbers tend to move a budget conversation, because they are about what the organization keeps.
The control plane in Figure 1 is not a diagram to build from scratch. AgentTrust OS implements it across the agent lifecycle, so the evidence for ISO 42001, the EU AI Act, and AIUC-1 is a rendering job over a runtime you were already going to need.
Certify before production, enforce the policy gate live, and render the evidence pack on demand — one control set, three rulebooks satisfied.
Start Free →