Skip to main content
Back to Blog
AI/MLSecurityEnterprise
1 October 20266 min readUpdated 1 October 2026

A New Taxonomy Maps Privacy Risks for AI Agents Handling Sensitive Data

MLCommons Agent Privacy Risk Taxonomy The MLCommons AI Risk & Reliability (AIRR) working group uses a straightforward approach to evaluating AI systems: define the behavior a sy...

By AI Engineering Team

MLCommons Agent Privacy Risk Taxonomy

The MLCommons AI Risk & Reliability (AIRR) working group uses a straightforward approach to evaluating AI systems: define the behavior a system should exhibit, then measure how consistently it delivers that behavior. Such measurements help deployers manage risk and cost while supporting higher reliability standards across the industry.

In April, MLCommons published the AI Reliability Map. It describes the rules an AI system should follow, including functionality, data protections, product safety, frontier safety, and psychosocial limits. It also identifies the conditions under which those rules should be tested, including normal use and adversarial scenarios.

The Reliability Map emphasizes the importance of data protection in overall system reliability. As the Privacy Working Group began its work, it focused on evaluating how reliably an AI agent handles personal data. Before that reliability can be measured, however, the privacy risks specific to agents must be defined.

A standalone chatbot generally creates data protection risks concentrated around the individual user interacting with it. An agent can create a much broader range of risks because it may continuously observe its environment, use external tools, retain information, and interact with other agents or services.

To organize these risks, MLCommons followed the first step in its community-based benchmark development process and created the v0.1 Agent Privacy Risk Taxonomy.

Five areas of data protection risk

The taxonomy covers five areas of data protection risk in agentic systems. It draws on expertise from industry, academia, and civil society, as well as reviews of known privacy and security incidents and existing risk-management frameworks.

  1. Data ingestion and processing: what the agent takes in and retains

    Agents may observe activity in the background, record their actions, and carry memory between sessions. As a result, they can retain more information than a task requires. Removing names and account numbers may not be enough when free-text passages contain clues that can reveal a person’s identity.

  2. Aggregation, use, and sharing: what the agent does with collected data

    Information that appears harmless in isolation can become sensitive when combined with other data. An agent might share or delete information without authorization, rely on third-party tools and services, or write and execute code dynamically.

    In some cases, a model may memorize fragments of personal data from user-consented training datasets. It could then recall or infer sensitive details about individuals and organizations.

  3. Inconsistent privacy practices across agents: what happens during handoffs

    One agent may operate under a different privacy policy from another agent with which it interacts. Information can spill into logs during these exchanges. When a user deletes or corrects information, or withdraws consent, downstream agents may not receive the update and may continue retaining data the user intended to remove.

  4. The limits of static consent models for runtime agent behavior

    Traditional consent is often treated as a decision made once, while agents make data-related decisions continuously. An agent may try to infer a user’s privacy expectations in a particular context and get them wrong. Alternatively, it may ask for permission repeatedly, encouraging users to approve requests without reading them because of alert fatigue.

    Agents also do not fully expose how they reach decisions, which can make it difficult for users to understand what they have agreed to.

  5. Accountability and governance: what happens after a failure

    When something goes wrong, it may be difficult to determine whether responsibility lies with the user, prompt, model, tool, or another agent. There is also no standard privacy-preserving method for logging agent interactions across organizations.

    Existing logs and memory stores can themselves become valuable targets for attackers seeking information about people or access to sensitive systems.

How the taxonomy can be used

MLCommons recommends incorporating the taxonomy into AI development and deployment processes, from pretraining through deployment monitoring. The intended benefits include reducing privacy violations and related risks, strengthening privacy and security protections, limiting potential liabilities, and reducing harm to individuals and organizations.

Next steps

Version 0.1 does not rank the risks. Their importance depends heavily on an agent’s capabilities and operating context. A bank customer-service agent and a personal assistant with access to an inbox may face different risks, with different priorities.

The next stage is to work with deployers to determine which risks matter most in specific deployment settings and which risks are broadly applicable. The aim is to develop measurements that work across the widest possible range of use cases.

MLCommons plans to engage large-scale deployers, in consultation with AIRR participants, to prioritize critical risks. It will then identify indicators for detecting those risks, pilot benchmarks for measuring them, and develop mitigation practices. The taxonomy will also continue to evolve, with the effort planned for completion by Q1 2027.

Some risks can be tested before deployment. For example, evaluations can determine whether an agent collects more data than a task requires or forwards a deletion request to downstream agents. Other risks, particularly those involving accountability and governance, concern the organization surrounding the agent and may not be captured by a benchmark alone.

A key objective is to create measurement tools that show how agents perform overall and how they perform in the areas most relevant to a particular deployment context.

Feedback on the taxonomy

The Agent Privacy Risk Taxonomy is an initial version and is expected to change. MLCommons is seeking feedback from industry practitioners, academic researchers, civil society organizations, and regulators.

The full taxonomy is available in the v0.1 document. The Privacy Working Group meets every other Thursday from 11:30 to 12:30 ET.