Governing agentic AI models in anti-financial crime

How governance and validation approaches need to evolve
By Stefano Boezio and David Choi
Home  // . //  //  Governing agentic AI models in anti-financial crime

Financial institutions are beginning to put agentic artificial intelligence to work in anti-financial crime (AFC) programs. These systems can gather evidence, sequence investigative steps, identify typologies, and recommend escalations. Our latest findings, from our joint report with Hawk AI, Governing AFC Agents: How Model Governance and Validation Must Evolve for Agentic AI in Financial Crime Compliance, show that the opportunity is significant, but so is the governance challenge.

Previous Oliver Wyman research on AI-based solutions in AFC compliance highlighted the need to rethink governance and operating structures as AI adoption grows. Agentic AI presses that need further. Unlike a traditional model or generative AI tool, an agent can plan, use tools, interact with multiple data sources, and exercise delegated authority across a workflow.

Existing supervisory guidance and industry frameworks offer useful principles, but they do not yet form a complete governance standard for agentic AI. Institutions need an approach that connects materiality, validation, human authority, and enterprise assurance.

Agentic AI expands the unit of governance beyond the model

Traditional model governance focuses on a model’s inputs, methodology, outputs, performance, and limitations. For an agentic system, that perimeter is too narrow.

An AFC agent combines a foundational model with orchestration, prompts, tools, data access, permissions, human oversight, and an audit trail. Its behavior is shaped by how those components interact, which means governance needs to address the workflow as a whole.

Exhibit 1: From model to workflow: the expanded unit of governance
Two-panel governance diagram comparing a classic model with an agentic system centered on a foundational model.

Institutions should be able to demonstrate that AI-supported activity is bounded, evidence-based, auditable, monitored, and aligned with the institution’s AFC risk profile. Established principles such as accountability, independent challenge, validation, and clear intended use remain essential, but they need to extend to the way an agent actually operates.

A risk-based agentic AI framework can match controls to materiality

A practical approach starts with the inherent materiality of business activity. Financial value at risk, process criticality, regulatory requirements, and the reversibility of an action establish the baseline.

Institutions can then use engineering choices to constrain an agent’s autonomy and action, helping shape its residual risk. The final report makes clear that these controls can change deployment conditions without changing the underlying materiality of the use case.

Consider an anti-money laundering (AML) investigation agent that gathers know-your-customer information, identifies typologies, and drafts a Suspicious Activity Report narrative. If it can only recommend an outcome and requires human approval for every disposition, its risk profile differs from that of an agent permitted to take consequential action.

Agent design therefore becomes a governance lever, allowing institutions to align controls and validation efforts with the risk that remains.

Why agentic AI validation must follow the workflow, not just the model

Validation needs a wider perimeter. Institutions should assess the agent’s intended use and design alongside how it accesses information, uses tools, interacts with humans, and produces outcomes. Pre-deployment testing must examine end-to-end behavior using representative, edge, and adversarial cases to see whether an agent remains within its charter when faced with ambiguous information, attempted prompt injection, or manipulated source data.

Oversight should continue in production. Performance, human overrides, drift, and guardrail breaches can provide early warning of problems. Institutions also need sufficient tracing to reconstruct the agent’s execution path, including its sources, tool use, outputs, human interventions, and system configuration. Together, these practices shift validation toward continuous assurance.

Human authority and accountability in agentic AI should be explicit by design

AFC introduces a question that traditional models rarely pose: What is the system authorized to do?

Every material agent needs a written charter defining its objective, scope, delegated authority, prohibited actions, and required human decision points. Technical permissions should reflect those boundaries.

Human decision rights should scale with materiality. Assistive and recommending agents may draft or propose dispositions while humans retain decision authority. For investigative AML use cases, a qualified human should approve any consequential action in the target system. AFC use cases that permit bounded autonomous action should be treated as critical-tier systems and subject to the strongest authorization, monitoring, containment, and intervention controls.

Exhibit 2: Governance operating model and decision-rights ladder
Matrix mapping AI governance tiers by materiality and design-controlled exposure, from Tier 1 assistive low risk to Tier 4 autonomous critical.

A three-stage approach to agentic AI governance

Institutions do not need to reach the target state at once.

The first stage is to establish visibility and ownership by taking inventory of agentic AFC activity, assigning accountable owners, classifying use cases by materiality, and defining clear operating boundaries.

Second is to standardize controls and evidence. That includes establishing reusable control patterns, validation requirements, monitoring thresholds, and revalidation triggers. Governance responsibility remains with the institution whether the technology is built internally or sourced from a third party, so external solutions must provide the testability, evidence access, and intervention rights needed for effective oversight.

The third stage involves embedding assurance into the technology environment. An enterprise agent assurance layer can apply common lifecycle controls across internal and third-party platforms, making governance more scalable as agent use expands.

The priority now is to turn established governance principles into controls that can be demonstrated in practice. Institutions that can demonstrate that their agents are bounded, testable, observable, challengeable, and interruptible will be better positioned to capture the benefits of agentic AI while maintaining the control and accountability that effective financial crime compliance demands.

Authors