Add Runtime Controls to AI Agents with NVIDIA OpenShell
AI agents can receive a goal, write code, use tools, and continue working as new information becomes available. This enables applications that investigate software failures, run...
By Hardware Team
AI agents can receive a goal, write code, use tools, and continue working as new information becomes available. This enables applications that investigate software failures, run experiments, perform business-critical actions, and conduct research over days or weeks.
Useful agents often need access to workspaces, compute resources, data, credentials, and external services. Broader access also introduces more serious failure modes, including modifying production data, exposing confidential information, or acting beyond the assigned task.
NVIDIA OpenShell 0.1.0 is an open-source runtime for defining and enforcing which systems and data an agent can access. It combines sandboxed execution, controlled service access, credential management, and formal policy analysis. Teams can grant agents the capabilities required for a task while OpenShell enforces those permissions outside the workload.
OpenShell supports Codex, Claude Code, Pi, Hermes, and future frameworks across enterprise applications, frontier research, and physical AI. Its uses include internal agent fleets, long-horizon research, robotics, and edge systems.
This article explains how NVIDIA OpenShell 0.1.0 adds enforceable runtime controls around an existing AI agent without requiring the agent itself to be rewritten. The controls can restrict API operations, protect credentials, and allow permission changes to be reviewed outside the agent workload.
OpenShell provides the runtime layer of the broader NVIDIA Open Agent Safety Platform, which extends protection across application, runtime, and infrastructure layers.
How organizations are adopting OpenShell
OpenShell is open source and available for enterprise adoption across a broad partner ecosystem. Organizations are using it in applications involving chip design, enterprise automation, accelerated computing, and physical AI.
- Cadence uses OpenShell for chip design with its ChipStack Autonomous RTL Design Engineer.
- Slack is building an on-demand agent platform on OpenShell to automate tasks.
- Gecko Robotics uses OpenShell to govern agents that make decisions on physical robots.
OpenShell capabilities
OpenShell 0.1.0 supports sandbox operations, policy verification, governance integration, credential protection, and flexible compute.
| Capability | How it helps |
|---|---|
| Multi-tenant platform support | Operates agent services for multiple teams or customers with separate workspaces, permissions, and service access on shared infrastructure. |
| Formal policy verification | Shows human and AI reviewers whether requested permissions remain within defined security boundaries and identifies where they exceed them. |
| Extensible security and governance | Connects third-party security services, governance systems, and custom checks to enforcement outside the agent workload. |
| Credential-protected service access | Uses authenticated services while keeping real credentials outside the agent workload and binding them to authorized requests. |
| CPU and GPU execution | Runs experiments and data processing on CPUs or GPUs across containers, virtual machines, and Kubernetes environments. |
Table 1. Capabilities introduced in OpenShell 0.1.0
Enforce permissions outside the agent
An agent can interpret instructions, select tools, and change its approach over time. OpenShell preserves that flexibility while enforcing permissions outside the agent workload.
OpenShell can manage fleets of agents and their sandboxes. Each sandbox has its own permissions, while governance can be applied across groups. Three components provide this control:
- OpenShell Gateway: Manages the lifecycles and policies of multiple sandboxes.
- OpenShell Supervisor: Runs alongside each sandbox, outside the agent workload, and checks outbound requests against policy.
- OpenShell Sandbox: Runs the workload with kernel-level controls over its filesystem and processes. It has no network path except through the supervisor.
The sandbox runtime uses operating-system kernel controls to limit which files a workload can read or change and to prevent it from acquiring additional system privileges. Network rules can be more specific than simply allowing a connection to a service. The supervisor can inspect configured HTTP, GraphQL, and Model Context Protocol (MCP) traffic, allowing a data query while blocking a write through the same API.
These controls remain active when an agent starts a shell, runs generated code, launches child processes, or proposes delegating work to sub-agents. OpenShell records policy decisions in an Open Cybersecurity Schema Framework (OCSF) audit trail. When it blocks an inspected request, it can return a descriptive error that helps the agent determine what to do next.
Watch a policy decision happen
The following example uses curl and an unauthenticated endpoint in the GitHub REST API. This makes each policy decision visible without requiring an API key or a language model. The same controls apply when an agent makes the request.
Install and start OpenShell 0.1.0 using the installation guide. Then download the accompanying no-network.yaml and github-readonly.yaml policy files into an examples directory.
First, create a sandbox with no outbound network access:
openshell sandbox create --name policy-demo \\
--no-auto-providers \\
--policy examples/no-network.yaml
The command opens a shell inside the sandbox. Try reading a public endpoint:
curl -sS --max-time 10 https://api.github.com/zen
The request fails because the sandbox has no outbound network permission. Open a second terminal on the host and inspect the logs to see which program made the request and why it was blocked:
openshell logs policy-demo --since 5m
Next, replace the sandbox policy with one that permits read-only access to the GitHub REST API. Policies are authored in YAML and compiled to OPA/Rego, which OpenShell evaluates for each outbound request.
network_policies:
github_api:
name: github-api-readonly
endpoints:
- host: api.github.com
port: 443
protocol: rest
enforcement: enforce
access: read-only
binaries:
- path: /usr/bin/curl
This rule allows /usr/bin/curl to reach the GitHub API on port 443. With protocol: rest, OpenShell inspects HTTP requests, allowing reads while blocking writes.
In the host terminal, apply the complete replacement policy without restarting the sandbox:
openshell policy set policy-demo \\
--policy examples/github-readonly.yaml --wait
Return to the sandbox shell and try both requests:
## Read: allowed
curl -sS --max-time 10 https://api.github.com/zen
## Write: blocked
curl -sS --max-time 10 -X POST https://api.github.com/zen
The host logs show that OpenShell blocked the POST request. An agent running these commands encounters the same restrictions.
Access services without exposing credentials
Many agents need model APIs or private services to complete their tasks. OpenShell authorizes this access while keeping the real credentials outside the agent workload.
A provider profile defines the credentials, endpoints, and permitted programs for a service. If a GitHub provider named github is already configured, it can be attached to a new sandbox while launching Codex:
openshell sandbox create \\
--provider github \\
-- codex
Authorization for one service does not make its credential available to another. If an agent sends the placeholder credential to a destination outside the credential's approved endpoints, OpenShell rejects the request.
The receiving service continues to enforce the permissions attached to the real credential. OpenShell adds a separate control over how the agent can use that credential. For example, an inspected read-only API policy can block write requests even when the credential itself has write access.
Adjust network access while the agent runs
An agent may discover that it needs a service or data source that was not known when the task began. When policy blocks the request, OpenShell records the denial so an operator or an AI agent approver can review it.
With policy advisor enabled, the agent can propose a narrowly scoped network or file policy change. The proposal remains pending for human review by default, and the agent cannot approve its own request. After approval, OpenShell loads the new rule into the running sandbox, allowing the agent to retry without restarting its work.
Filesystem and process restrictions are established when the sandbox starts. Changing those controls requires a new sandbox.
Enable policy advisor for a sandbox with this command:
openshell settings set <sandbox-name> \\
--key agent_policy_proposals_enabled \\
--value true
Verify the permissions a policy grants
A policy can appear restrictive while leaving another route open. Blocking a GitHub write through one tool does not help if another permitted tool or generated code can use a credential to make the same write.
The OpenShell policy prover uses formal logic to check the permissions granted by a policy, including provider-contributed access. It can prove that modeled permissions remain within an operator-defined boundary or identify a concrete action that crosses it. Because the result comes from the policy model, an agent's explanation cannot change the result. The formal-methods primer describes this process.
In long-horizon adversarial experiments, frontier agents with reduced safeguards spent up to two hours trying to persuade an AI reviewer to grant permissions that would allow them to modify a protected GitHub repository. OpenShell's formal policy analysis provided evidence of what those permissions allowed, even when agents attempted to manipulate the reviewer. The combined review and runtime controls allowed useful access, and no protected repository writes occurred in these tests.
Ongoing work extends policy analysis across multiple agents, where one agent's access can combine with another's. The goal is to check the permissions of the system they form together. The prover documentation lists the supported checks.
Build locally and deploy into shared infrastructure
A local sandbox can be used while building an application and defining its permissions. To serve multiple users, the workspaces and access guide describes how to use the SDK to create and manage sandboxes. Each workload has its own policy and attached providers.
Trusted middleware outside the sandbox can connect identity services and add application-specific checks to the request path. Compute drivers connect OpenShell to Docker, Podman, MicroVM, and Kubernetes. The support matrix lists current requirements.
The 0.1.0 migration notes cover changes for existing deployments. The quickstart explains how to run an agent with OpenShell and configure the services it can access.