Assurance framework

AI agent assurance

50 requirements in six domains that the internal auditor uses to test AI agents: those of the business and those of the audit function itself. Under whose authorisation did the agent act, and where is the evidence that it stayed within that mandate?

Everyone sells agents for internal audit, nobody tests them

Diligent launched AuditAI at the IIA GAM 2026 conference. AuditBoard is now called Optro and acquired Midship, with its own claim that agents automate up to 87 percent of routine SOX tasks. TeamMate applies AI across the audit lifecycle. The public product information of these vendors emphasises automation and efficiency above all. Far less is made concrete about how mandate, authorisation, error measurement, human intervention and full reconstruction are demonstrably controlled. That is precisely the question a CAE has to ask, about every agent: the vendor's, the business's, and the one inside the audit function itself.

Audit files are not innocent input

An agent searching audit files may process names, email addresses, log data, HR information, fraud reports and special or criminal-offence personal data. Including when that data was never deliberately handed to the agent, but rode along with an attachment, an export or a mailbox.

In August 2026 the framework gained five requirements that test exactly this: whether the purpose for which the agent processes that data is compatible with the purpose it was collected for, whether it sees no more than its task requires and whether the organisation recognises the personal data riding along in attachments and exports, what remains in prompts, memory and vector indexes and whether deletion actually reaches those places, whether the vendor uses the input to train its models, and who else is in the chain, where that data sits and whether a data subject can exercise their rights there. Alongside that, existing requirements already covered which data categories the agent may access and whether that boundary is technically enforced, and what remains in the action log and for how long.

Most agents have no worked-out identity or authorisation

In February 2026 NIST opened a track on identity and authorisation for software and AI agents. That concept paper puts exactly the questions that matter for assurance at its centre: how do you identify an agent, how does it prove its authority for a specific action, and how do you bind its identity back to the human who authorised it.

Mandate has to be traceable

Singapore's IMDA framework for agentic AI, published in January 2026 and updated in May, emphasises identity and access management as a precondition for traceability and accountability. It is guidance, not law. The question it raises is unavoidable all the same: which agent acted, with what permissions, on whose behalf.

For most agents the kill switch has never been exercised

A stop mechanism exists on paper, but nobody has tested whether it works once the agent actually acts outside its limits, or whether an action interrupted halfway is left in a recoverable state.

The agent's error rate is unknown

As long as failed, aborted and human-corrected actions are not recorded as consistently as successful ones, any statement about an agent's reliability is an estimate. An agent that confidently completes something incorrect counts too.

The reverse question: we test the agent

A fixed set of requirements instead of a questionnaire reinvented for every engagement, and without the sales pitch about what the agent automates.

Testing instead of selling

The framework does not assess what an agent promises to automate, but whether its actions are demonstrably authorised, bounded, controlled and recorded. Those are questions vendor material tends to answer less prominently.

Applies to the audit function itself

The same requirements apply to the agent handling HR processes and to the agent searching audit files. An audit function that does not put its own agent through this framework will struggle to ask the rest of the organisation that question.

23 substantively validated mappings

Every mapping to a Unified Control was assessed individually on its merits and not generated automatically. The remaining 27 requirements deliberately carry no mapping: an ISO 27001-based control set has no fair counterpart for human intervention or an agent's error rate.

Part of the CRAFT library

The framework sits in CRAFT alongside ISO/IEC 42001, NIST AI RMF and OWASP LLM. Not a standalone document, but a framework that evolves with the rest of the library.

How the framework covers the agent, from inventory to board reporting

The framework works through the agent from the outside in: first what is running, then under what mandate and with access to which data, then who can intervene, what is evidenced, how reliable the behaviour is, and what has to be accounted for towards the board and the regulator.

A

Inventory and identity

Which agents are running, under which own technical identity, who owns them, and how you detect shadow agents set up outside that overview.

B

Mandate and boundaries

Under whose authorisation the agent acts, which systems, data sources and data categories it may access (and whether that boundary is technically enforced rather than merely instructed), whether the processing purpose is compatible with that of the source, whether it receives no more data than the task requires, which tools are prohibited, and which limits apply.

C

Human intervention

Where prior approval is required, whether that was a real assessment or routine click-through, who can stop the agent, whether that kill switch has ever been exercised, and whether an action can be reversed.

D

Evidence and traceability

An action log the agent cannot alter itself, traceability to the instruction and model version behind it, recording of failures, a retention period derived from the limitation period of the decision rather than a generic IT standard, and visibility of what remains outside the source system in prompts, memory and vector indexes.

E

Reliability and testing

Whether the agent was tested against a predefined standard, whether there is a measured error rate, whether silent degradation is noticed, and how it behaves when input tries to steer it off course.

F

Compliance and reporting

Disclosure towards the people the agent communicates with, AI literacy of those who steer and review it, accountability to the board and audit committee, and the regime for purchased agents: who is responsible, whether the vendor uses your input to train its models, which sub-processors are in the chain and where that data sits, and whether you may audit them. Whether a specific agent falls under the transparency obligation of Article 50 of the AI Act is established per use case in this domain.

Testing an AI agent against the 50 requirements, with a judgement and test approach per requirement
The testing itself: a judgement with explanation per requirement, the test approach alongside it, and visible whether a requirement maps to a shared control or is agent-specific.

What a requirement actually says

One requirement per domain, exactly as it appears in the framework.

A05

Own technical identity per agent

Every agent authenticates with its own technical identity and does not run under a shared or generic service account together with other agents or processes. Where a shared account is unavoidable, that is explicitly recorded as a deviation with a compensating measure.

B10

Data minimisation and personal data riding along

The agent sees no more data than its task requires. The organisation has recognised that audit files, mailboxes, exports and attachments carry personal data nobody deliberately handed over, including special-category and criminal-offence data, and has established how it detects that. An arrangement where the agent can reach everything the user can does not count as minimisation.

C05

Exercising the kill switch

The kill switch has been genuinely exercised at least once rather than only described on paper, and the result is recorded. It is determined what happens to an action the agent is halfway through: complete safely, cancel, or leave in a recoverable intermediate state.

D02

Traceability to instruction and model version

For every statement, recommendation or decision of an agent it can be established which instruction and which model version produced it. Without that link the organisation cannot reproduce a result, nor determine whether a fault lay with the model, the instruction or the input.

E02

Known error rate and measurement method

The organisation can state an error rate for every agent: how often it fails to complete a task, completes it partially or completes it incorrectly, measured on a representative set of tasks with a recorded definition of what counts as an error.

F05

Allocation of responsibility for purchased agents

For an agent supplied by a vendor or embedded in a purchased service, it is contractually recorded who is responsible for the agent's behaviour, what assurance the vendor gives about training, steering and limits, and how the organisation continues to meet its own obligations.

Where a requirement aligns with one of the 23 mapped Unified Controls, the testing uses that existing evidence. For the remaining requirements, including the exercised kill switch and the measured error rate, the auditor gathers their own evidence: logs, test results, mandate documents, escalation paths. That makes the outcome reproducible, regardless of which vendor supplied the agent or which department procured it. The testing feeds into the regular audit programme and into reporting to management and the audit committee.

The AI agent assurance framework in the CRAFT library, alongside the AI Act with ISO/IEC 42001 and ISO/IEC 27001
The framework sits in the library as a requirements set of its own, alongside the AI Act with ISO/IEC 42001 and ISO/IEC 27001, with 50 requirements and 23 cross-references.

Fixed requirements, not a loose questionnaire

Domains

6 domains

From inventory and identity to compliance and reporting.

Requirements

50 requirements

Spread across the six domains, each separately testable.

Mapped

23 mappings

To the shared Unified Controls, each mapping assessed individually on its merits.

Unmapped

27 requirements

Deliberately unmapped: the Unified Controls are ISO 27001-based and have no fair counterpart for, say, an exercised kill switch.

What the testing delivers

The testing gives the CAE a substantiated answer to the question the board, the external auditor or the regulator will ask sooner or later: which agents are running, who authorised them, which personal data did they touch, and where is the evidence that they stayed within that authorisation.

Existing obligations can be relevant too, depending on the use case. Article 50 of the AI Act has applied since 2 August 2026 to, among others, AI systems that interact directly with natural persons; not every agent falls under it, and the qualification has to be established per use case. Article 4 requires providers and deployers to take measures to support the development of AI literacy. The high-risk obligations have been postponed to 2 December 2027 and 2 August 2028. Domain F establishes per agent which of these regimes applies, instead of applying one rule to all agents.

The AI system register showing the systems flagged as agents
The agents come from the AI system register, so there remains one register instead of a second list.

Test the agents already running

Browse the full framework in the CRAFT library, or let us show you in a demonstration what the testing would deliver for your agents.