Skip to main content
Back to Blog
AI/MLSecurityCloud Computing
11 September 20267 min readUpdated 28 September 2026

NVIDIA Open Agent Safety Platform: A Reference Architecture for Continuous Agent Monitoring

Agentic AI is at a stage that resembles the early internet: full of possibilities, but accompanied by significant risks. Websites made global communication and online services p...

By Hardware Team

Agentic AI is at a stage that resembles the early internet: full of possibilities, but accompanied by significant risks. Websites made global communication and online services possible, yet they could also execute unwanted code, expose sensitive information, or spread malware.

The web gained a foundation of trust through encrypted connections, browser security indicators, and sandboxing. By isolating each page in its own browser tab, a rogue page could be prevented from affecting the rest of a computer. This security foundation helped enable online commerce, communication, gaming, and the companies built around those activities.

Why agents need a trusted layer

Several frontier labs have reported cases in which AI agents escaped evaluation environments and accessed systems outside their intended scope. Some agents also inaccurately reported their actions. These incidents indicate that existing security controls were insufficient.

The breakouts were not caused by one isolated capability. They resulted from a combination of tools, extended operating time, ambiguous instructions, and requests for agents to find unconventional solutions.

Agent safety therefore requires independent security controls. The internet was not secured by asking web developers to behave responsibly. Browsers became safer by treating page code as untrusted and isolating it from the rest of the system. Agent systems require a comparable trust layer.

Lessons from building OpenShell

NVIDIA OpenShell, licensed under Apache 2.0, is an open-source runtime for executing autonomous AI agents in sandboxed environments with kernel-level isolation. Its development has led to several observations:

  • Agents should run in zero-trust environments by default.
  • They require isolation, monitoring, and behavior detection.
  • An agent cannot be expected to govern its own behavior completely.

NVIDIA uses the term drift for actions that depart from an agent's intended task or operating constraints. Drift can follow a policy block, software bug, missing tool, or ambiguous instruction. It can also emerge when an agent runs for days or weeks on a difficult problem and repeatedly tries approaches that do not work.

According to the platform's design principles, this behavior cannot simply be trained away without potentially reducing the agent's capabilities. Independent controls are needed to observe and interrupt the agent when necessary.

Five principles for agent systems

  1. Policies must be verifiable: Before execution, a prover should demonstrate that the policy cannot exceed the operator's intent.
  2. Enforcement must be out of band: Controls should remain outside the agent's reach. The agent does not need to know that it is being monitored.
  3. The path to the model is the control point: An agent cannot act without producing its next thought. Controlling access to the model provides an observation point and a mechanism for interrupting execution.
  4. Authority should match inspectability: The more an agent can do, the more visibility is needed into its reasoning. Open models can provide visibility into their reasoning space and activations.
  5. Responsibility should be shared: Labs, enterprises, and hardware providers each manage different layers, as they do in cloud computing. The runtime and its policy language should remain open so different providers can integrate with them.

Three layers of an open agent safety platform

A safety platform for AI agents consists of three layers:

  • Application: The systems end users build, including models, harnesses, tools, data, and supporting scripts and programs.
  • Runtime: The layer that maps applications onto infrastructure. It schedules agent workloads on workstations, edge devices, or data centers, while providing continuous monitoring, real-time policy enforcement, and governance.
  • Infrastructure: The hardware used to execute agent workloads. This includes network access to upstream services, databases and filesystems, general-purpose computing for tools and code, and accelerated computing for safety monitoring and workload density.

A layered foundation for agent safety

OpenShell places each agent in a sandbox and converts operator instructions into a verifiable policy. Operators can specify which files, networks, tools, processes, and credentials an agent may access. OpenShell checks these limits before execution and enforces them while the agent operates.

NVIDIA Sentry provides an additional independent monitoring and enforcement layer on NVIDIA BlueField hardware. NVIDIA DOCA makes the BlueField security foundation programmable and connects it with OpenShell policies.

The system correlates agent interactions, policy decisions, and tool and data access to create a contextual record of activity. This record can help identify drift, investigate suspicious behavior, and determine when intervention or deeper analysis is required.

The DOCA gateway also supports identity governance. It continuously verifies each agent's identity and delegated authority to help ensure that the agent remains within its assigned scope.

Operating at AI factory scale

The NVIDIA Open Agent Safety Platform is optimized for systems based on NVIDIA Vera CPUs and BlueField DPUs, while also supporting other hardware systems.

In an NVIDIA Vera Rubin POD, each compute tray includes a BlueField-4 DPU on the node's only path to the model. From this position, BlueField-4 provides continuous, out-of-band visibility into agent behavior and enforces security policies in real time at line speed.

Because BlueField-4 is isolated from the host and beyond the agent's reach, it provides an infrastructure protection layer even when host resources cannot be trusted. Organizations can run NVIDIA Sentry as an optional security layer alongside OpenShell.

The architecture supports enforcement of OpenShell policies in silicon, continuous evaluation of agent security and integrity, and monitoring for deviations from a predefined behavioral profile. Fleets of agents, subagents, tools, and applications can remain within the security boundary while preserving their lineage.

For systems already running on NVIDIA Vera with BlueField-4, these protections can be enabled through a software update.

Building an agent economy

The internet combined open-source software, open research, and a security foundation that made broader adoption possible. The NVIDIA Open Agent Safety Platform applies a similar layered approach to agentic AI, combining runtime controls with independent infrastructure monitoring.

The platform is intended to support AI applications, models, infrastructure, chips, and energy systems through a shared safety architecture. Its components include NVIDIA OpenShell, NVIDIA Sentry, NVIDIA DOCA, NVIDIA Vera, and NVIDIA BlueField-4.