MLCommons Agent Reliability Profile Named Finalist in Global Agentic Regulator Hackathon
Addressing the authorization and supervision gap Financial institutions and regulators still lack a reliable way to determine whether an AI agent will operate within its authori...
By AI Engineering Team
Addressing the authorization and supervision gap
Financial institutions and regulators still lack a reliable way to determine whether an AI agent will operate within its authorized boundaries. The MLCommons Financial Services Working Group’s Agent Reliability Profile addresses this challenge and has been selected as a finalist in the C:>DIR Global “Agentic Regulator” Hackathon.
AI agents do more than answer questions: they can take actions. This capability makes them useful, but it also creates a need for institutions and supervisors to verify what an agent is permitted to do and whether it has remained within that scope.
Existing agent-identity standards can establish who an agent is, but they do not necessarily show what the agent should be authorized to do after connecting to financial digital infrastructure. This unresolved issue is known as the authorization and supervision gap.
The Agent Reliability Profile is the working group’s flagship initiative. It is intended to provide a standardized, regulator-compatible framework for describing, validating, and benchmarking the reliability of agentic deployments in financial services.
The C:>DIR hackathon
The C:>DIR hackathon is run by the University of Cambridge’s Digital Innovation and Regulation Initiative. It is supported by the BIS Innovation Hub, the Global Financial Innovation Network, the Digital Regulation Cooperation Forum, and more than 35 other organizations.
The competition brings regulators and industry participants together to develop practical, deployable prototypes designed to strengthen trust and accountability in agentic AI.
Organizers received 336 submissions from more than 65 countries. A total of 36 teams, six in each problem space, advanced to the final round.
The MLCommons submission
Working group co-chairs Mike Hsu and Medha Bankhwal entered the submission in the Know Your Agent (KY-A), Digital Verification & Digital Public Infrastructure problem space. The submission demonstrates two tools built around the Agent Reliability Profile:
- Profile Builder processes an institution’s existing evidence, including design documents, configuration exports, and policies. It compiles that material into a Level 1 Asserted Profile and identifies differences between stated intent and the system’s as-built configuration.
- Profile Validator uses the Asserted Profile as a test specification and attempts to falsify it against the system’s actual behavior. The result is a Level 2 Validated Profile.
The demonstration applies both tools to an open-banking data-sharing and consent scenario. It tests whether an agent remains within the data-sharing scope to which a customer actually consented, providing a practical example of the authorization and supervision gap.
Schedule
Final-round development runs from September 1 through September 8. Live demonstrations and judging by regulatory fellows from the Cambridge Regulator Visiting Fellowship Program are scheduled for September 15 and 16.
The winners will be announced on September 18, 2026, at the C:>DIR Summit.
The Agent Reliability Profile is being developed as a working group effort involving financial institutions, regulators, and technical contributors. The Financial Services Working Group meets on Wednesdays from 12:00 to 1:00 PM ET.