Assessing and Managing AI Risk in Business Deployments
Organizations responsible for AI risk face two simultaneous demands: protect the business from harm while enabling broader and faster AI adoption, often without additional resou...
By AI Engineering Team
Organizations responsible for AI risk face two simultaneous demands: protect the business from harm while enabling broader and faster AI adoption, often without additional resources. AI governance helps reconcile these priorities by connecting risk management with business objectives and organizational values.
AI risk is already part of enterprise risk
Mature organizations generally follow an established risk-management cycle: identify exposures, estimate potential losses, apply controls proportionate to the risk, and report residual risk to an accountable decision-maker. Insurers price risk, auditors evaluate controls, and boards approve the resulting decisions.
Organizations have long understood the consequences of security failures, and AI can amplify those costs. IBM’s 2025 Cost of a Data Breach study reports a global average breach cost of USD 4.44 million and a record USD 10.22 million average in the United States. The study also found that AI adoption was outpacing oversight: 63% of breached organizations had no AI governance policy or were still developing one. Organizations with extensive “shadow AI” activity experienced breaches that cost approximately USD 670,000 more on average.
Regulation adds further consequences for poor outcomes. Frameworks established before generative AI, including GDPR and CCPA, already impose strict requirements on certain automated decisions. The EU AI Act includes penalties of up to EUR 35 million or 7% of global turnover, creating a strong incentive for proactive AI risk management.
Managing AI risk in deployed systems does not require an entirely new philosophy. Organizations can apply the same standard of evidence used elsewhere in operations, supported by measurement tools designed for AI systems.
What the airline industry demonstrates
Effective management of a new risk produces more than avoided losses. It can improve the resilience of an entire industry.
In 1959, U.S. commercial aviation experienced approximately 40 fatal accidents per million departures. Within a decade, that figure had fallen below 2, and it continued to decline by roughly half every decade for the following half-century. Aviation became safer as the industry developed systems for measuring, managing, and continuously improving risk.
Measurement made air travel safe enough to trust and insure at a large scale. Several billion passengers travel each year, while lenders and insurers support more than 30,000 commercial aircraft.
Independent measurement is becoming essential
AI risk management depends on trustworthy measurement. Organizations seeking reliable assurance should use measurement practices that are independent and transparent. Pharmaceutical companies do not adjudicate their own clinical trials, and public companies do not audit their own financial statements.
In AI, however, vendors have often run their own benchmarks and reported their own results. In one AI code-review assessment, a vendor’s self-run benchmark reported an 82% bug-catch rate. A rival that reran the benchmark on the same repositories recorded 45%, while an independent laboratory’s benchmark found that no vendor exceeded 63%. The products had not changed. The party conducting the evaluation had.
AILuminate and independent AI evaluation
MLCommons’ AILuminate addresses this measurement gap. Launched in December 2024, AILuminate is a family of independent, consensus-governed benchmarks that evaluate how reliably AI systems avoid harmful responses.
The benchmarks are developed by working groups that include participants from AI companies, academia, and civil society. Private test sets and independent evaluators help protect the process. Vendors do not write their own examinations, see the questions in advance, or control the grading.
Governance teams can use the results in three main ways:
- Identify concentrated exposure. Per-hazard grades for a model show where risks may be concentrated for a particular use case. These results provide evidence for an AI risk register.
- Inform specific controls. If a system’s grade declines sharply under adversarial pressure, the result supports a concrete control requirement at the system layer rather than a general instruction to exercise caution.
- Evaluate the deployed configuration. Buyers should test the configuration they actually deploy, using the same evaluation applied to the vendor’s system. The resulting score should reflect real-world deployment, not only the model as delivered by its developer.
The difference resembles the relationship between a vehicle’s official mileage rating and the mileage achieved during an actual commute. The rating describes the vehicle as shipped, while the second measure reflects how it performs in a particular operating environment.
Before-and-after grades can be recorded in a risk register as quantified, comparable, and auditable evidence. They also provide feedback for iterative improvement and for detecting regressions over time.
The need for measurable assurance is also visible in insurance. In late 2025, major insurers asked U.S. regulators for permission to exclude AI-related liabilities from corporate policies. The issue was not necessarily that AI risk could not be managed, but that insufficient measurement made the risk difficult to underwrite. Risk that cannot be measured is difficult to insure.
Evaluate what you actually deploy
A grade for a foundation model describes how that model performed when it was released by its developer. It does not establish how the model will behave in a deployed system.
Most organizations do not deploy models in isolation. They wrap them in prompts, connect them to internal data, provide tools, and sometimes configure them to operate as agents. These integrations make models useful, but they can also weaken or change safeguards provided by the model vendor.
One example of a more deployment-specific assurance approach is the Agent Reliability Profile, a project of the MLCommons financial services working group and a finalist at the 2026 CDIR hackathon. The profile is intended to document a bounded and falsifiable claim about an individual agent. It records whether a system reliably operates at a specified autonomy tier, within a defined operational design domain, and under a stated control envelope.
This type of artifact does not express a general impression about whether an agent is trustworthy. Instead, it defines a specific claim that an auditor or insurer can evaluate.
Apply established risk-management practices
AI risk can be identified, priced, managed proportionately, and reported using approaches similar to those applied to credit risk, currency exposure, and workplace safety. Applying these practices does not inherently prevent AI adoption. It provides a basis for adopting AI while understanding and managing its potential exposure.
The relevant risk is already part of many organizations’ operations. The central question is whether the organization will measure it.
Further reading
- McKinsey, The State of AI (2025) and The AI Reckoning: How Boards Can Evolve (2026)
- PwC, 2026 Annual Corporate Directors Survey
- IBM, Cost of a Data Breach (2025)
- Center for Democracy & Technology, Getting Third-Party AI Assessment Right (2026)
- NIST, AI Risk Management Framework
- MLCommons, AILuminate v1.1 and Jailbreak Benchmark v1.0, arXiv:2610.02827 (June 2026)