NVIDIA Open Agent Safety Platform: What It Covers

NVIDIA Open Agent Safety Platform

Every agent security architecture has to answer one question, and most of them answer it badly: who enforces the boundary when the thing being constrained is the thing that decided to cross it?

A system prompt telling an agent it has no internet access is a request. A sandbox the agent runs inside is better, but the sandbox is configured by the same platform team under the same time pressure, and configuration errors are the most common cause of containment failure in production. An enforcement point inside the blast radius is an enforcement point with a conflict of interest.

NVIDIA’s answer, announced on September 28, 2026, is to split the problem across two layers and move one of them into separate silicon. That is the part worth examining — not the announcement itself, but what changes when the control plane stops sharing a host with the thing it controls.

Key takeaways
  • The NVIDIA Open Agent Safety Platform combines OpenShell, an open-source Apache 2.0 agent runtime available now, with Sentry, a reference design that runs on BlueField-4 DPUs and has no disclosed price or general availability date.
  • OpenShell is not just a sandbox. NVIDIA documents four components: agent sandboxes with kernel syscall filtering, a Gateway control plane, a Supervisor that runs outside the sandbox and evaluates every network request, and a Policy Prover that formally verifies policy changes.
  • Sentry’s significance is positional. Running out-of-band on a DPU means the enforcement point is not inside the environment it polices, so a compromised host does not automatically compromise the watchdog.
  • The millisecond quarantine figure is a vendor claim. No independent testing of either component has been published.
  • Two claims in early coverage do not hold up: OpenClaw is supported, listed in NVIDIA’s own documentation as a bundled sandbox.
  • Neither layer addresses prompt injection, model deception, bad policy, or a legitimate tool call with damaging consequences. Enforcement moving down the stack changes who can bypass the control, not whether the instruction was a good idea.

Quick Navigation


What the NVIDIA Open Agent Safety Platform Is

The NVIDIA Open Agent Safety Platform is an open reference system design, announced on September 28, 2026, that pairs the OpenShell secure runtime with NVIDIA Sentry running on BlueField-4 DPUs. NVIDIA describes it as combining a software runtime boundary with in-silicon security enforcement, and says more than 100 organisations are working with the technologies in it.

The two halves have very different availability profiles, and conflating them is the main way to misread the announcement.

OpenShellSentry
What it isOpen-source agent runtimeReference design for out-of-band monitoring
LicenceApache 2.0Not applicable; hardware-dependent
Where it runsDocker, Podman, Kubernetes, VM; local, on-prem, cloud, air-gappedBlueField-4 DPUs
AvailabilityNow, from GitHubNo general availability date or price disclosed
PositionInside the host, outside the agent processOutside the host entirely
EnforcementSyscall filtering, policy-enforced egress, credential brokeringContinuous monitoring and quarantine

The framing NVIDIA uses on its own OpenShell page is worth quoting in substance because it is the argument, not the marketing: security lives in the environment rather than the model or the application; nothing is permitted by default; enforcement happens outside the agent’s process where it cannot be prompted away.

That last clause is the design thesis. Whether it holds in practice is what the rest of this article examines.

Context matters for why this arrived now. Through mid-2026 a series of disclosed incidents showed agents from several major labs reaching real systems during security evaluations because the environments around them were misconfigured rather than because the models broke out of anything. The lesson those incidents taught was not that models are dangerous but that containment implemented as configuration fails quietly. We traced that pattern in our analysis of the evaluation-environment failures, and it is the problem this platform is aimed at.

Where OpenShell Sits

OpenShell sits underneath the agent harness and above the container runtime. NVIDIA is explicit that it is not an agent framework and not merely a sandbox — it uses Docker, Podman, Kubernetes or VM isolation as substrates and adds agent-specific control on top.

Four components, as NVIDIA documents them:

  1. Agent sandboxes. Each agent runs in its own isolated sandbox with no direct network access. OpenShell runs the agent without privileges, limits the files it can reach, and monitors and filters its system calls in the kernel, blocking unsafe calls and brokering requests to the supervisor through a single secured channel.
  2. Supervisor. Runs outside the sandbox. It evaluates every network request against policy at the binary, destination, method and path levels, and supplies credentials only where policy allows. Policy updates apply live, and every allow and deny is logged.
  3. Gateway. The control plane. It authenticates users, manages sandbox lifecycle, and delivers policies, settings and credentials to each sandbox’s supervisor.
  4. Policy Prover. A verification engine that checks whether policies stay within an allowed access boundary, and whether a proposed network rule adds risky access relative to the sandbox’s current policy. NVIDIA describes this as formal verification.

Three of those deserve comment.

The supervisor’s position is the important one. Credentials are brokered rather than injected into the agent’s environment, which means an agent that reads its own environment variables does not find a usable key. Request evaluation at the binary level means policy can distinguish which process is making a call, not just where it is going — a meaningful step beyond a network allowlist.

The Policy Prover addresses a failure mode that gets very little attention: the policy change that looks narrow and quietly widens access. Checking a proposed rule against the current boundary catches the class of mistake that human review misses, because humans reason about the rule they are adding rather than the access surface it produces.

Syscall filtering in the kernel is the conventional part, and it is the part that determines what happens when an agent tries something the network layer would never see — writing where it should not, executing what it should not.

The supported-agent list is broad. NVIDIA’s documentation covers Claude Code, Codex, OpenCode, GitHub Copilot CLI, OpenClaw, Ollama and Pi, with varying levels of default policy coverage — Claude Code and Copilot CLI work out of the box with full coverage, while Codex is pre-installed but requires a custom policy. That coverage column matters more than the presence of a name on the list: “supported” and “works without you writing policy” are different states.

Where Sentry Sits

Sentry is the part that is architecturally interesting and practically unavailable to most readers today.

NVIDIA positions it as an out-of-band watchdog running on BlueField-4 DPUs, monitoring agent activity continuously and enforcing policy independently of the host executing the agent. Reporting of NVIDIA’s technical material describes it using DOCA to inspect agent requests and responses. NVIDIA says it can quarantine an agent that crosses its boundary within milliseconds.

Every sentence in that paragraph is a vendor statement. No independent testing has been published, and the millisecond figure in particular should be read as a design target that NVIDIA is asserting rather than a measured result anyone outside NVIDIA has reproduced.

What can be reasoned about is the position rather than the performance. A DPU sits on the network path with its own processor, memory and operating environment, separate from the host CPU. Security functions placed there have three structural properties:

  • They observe traffic the host cannot suppress. A compromised host can lie to its own monitoring agent. It has a harder time lying to a device that sees the packets on the way out.
  • They survive host compromise. Root on the host does not confer control of the DPU.
  • They enforce at a chokepoint. Network egress necessarily passes through the path the DPU occupies.

Those properties are why DPU-based security is not a new idea — it is the same logic behind out-of-band management and bump-in-the-wire inspection, applied to a new workload class.

What Sentry can see is bounded by what crosses that boundary. An agent doing something entirely within its sandbox — writing files, reasoning wrongly, making bad decisions with data it already holds — produces nothing for a network-path device to observe. Sentry’s visibility is the visibility of the wire, plus whatever telemetry OpenShell sends it.

On the hardware requirement: NVIDIA’s materials specify BlueField-4. I found no NVIDIA document explaining whether the requirement is architectural or a matter of where the software has been built and validated, and no comparison against BlueField-3 capability. The honest statement is that NVIDIA specifies BlueField-4 and has not published the reasoning.

NVIDIA Open Agent Safety Platform

Availability is the practical constraint. One report says enabling Sentry on systems already running Vera CPUs with BlueField-4 is a software update; for everyone else, NVIDIA has given no date and no price. Any architecture decision made today should treat OpenShell as a product and Sentry as a direction.

The Four-Layer Containment Test

Security controls do four different jobs, and vendors tend to blur them. Keeping them apart is how you work out what a platform actually buys you.

FunctionQuestion it answersWho does it here
PreventCan the action happen at all?OpenShell: default deny, syscall filtering, no direct network access
DetectDid something outside policy occur?OpenShell’s audit log; Sentry’s continuous monitoring
ContainCan the blast radius be limited once it starts?OpenShell’s per-sandbox isolation; Sentry’s quarantine
RecoverCan normal service be restored and the event reconstructed?Audit logs; neither component is a recovery system

Prevention is the strongest of the four and the one most teams under-invest in, because prevention requires deciding in advance what the agent may do. Detection is what you fall back on when prevention was incomplete. Containment limits damage once detection fires. Recovery is somebody else’s product.

The architectural claim behind this platform is that prevention and containment should not live in the same trust domain as the workload. OpenShell moves them out of the agent process. Sentry moves containment out of the host entirely.

Running a Real Failure Through the Stack

Take the class of failure that actually happened in 2026: an agent in a security evaluation reaches a real production system because its environment had internet access it was not supposed to have, and the target name it was given matched a live domain. The model believed it was hitting a simulated target. The network believed the request was legitimate. Both were correct within their own scope.

Layer 1, model and agent behaviour. The agent pursues its objective competently. Nothing here is defective, and nothing in this platform changes it. An agent given a target and tools will pursue the target.

Layer 2, runtime containment. This is where OpenShell’s default-deny posture bites. Sandboxes have no direct network access; egress is policy-enforced through the supervisor, evaluated at destination and path level. A misconfigured scenario does not automatically become a misconfigured network, because the network path is defined separately and denies by default. The Policy Prover adds a check against the policy change that would have opened it.

Layer 3, egress and tool policy. The request to an unlisted destination fails at the supervisor rather than succeeding silently. This is the single most valuable property in the whole stack, and it is available today in open source. It is the same principle we worked through in our piece on testing agents against prompt injection: reachability decides outcomes, not instruction quality.

Layer 4, infrastructure enforcement. If the host itself is wrong — a policy daemon that crashed, a container escape, an operator who disabled enforcement to debug something at 2am — Sentry is the layer that still sees traffic leaving and can cut it. NVIDIA’s claim is milliseconds; the verified part is the position, not the latency.

The honest scorecard: layers 2 and 3 would plausibly have prevented the 2026 incidents, and they need no special hardware. Layer 4 addresses the case where layers 2 and 3 were themselves broken, which is a real case and a rarer one.

Why the DPU Changes the Boundary

Walk the enforcement points from the model outward and the trust question changes at each step.

Enforcement pointWhat it can seeWhat compromises it
Model / promptIts own contextAny input that changes the model’s behaviour
Agent harnessTool calls it mediatesA bug or a bypass in the harness
Runtime sandboxSyscalls, file access, process behaviourContainer escape, host root
Host network stackTraffic from that hostHost root
DPUTraffic on the path, independent of host stateCompromise of the DPU itself

Each row down that table shrinks the set of things that can neutralise the control, and shrinks what the control can see. That is the trade, and it is why the answer is layering rather than picking a winner.

The DPU row is distinctive because host root — the thing that defeats every control above it — does not defeat it. For agentic workloads specifically, where the whole premise is software acting autonomously with tool access, an enforcement point that survives the workload owning its own host is a structurally different guarantee from one that does not.

It is not a complete guarantee. A DPU is a computer running software, with its own firmware, its own attack surface and its own supply chain. Moving enforcement there relocates the trust assumption rather than eliminating it. And a quarantine that fires correctly still fires after the action that triggered it — containment, not prevention.

What the NVIDIA Open Agent Safety Platform Does Not Solve

This section matters more than the architecture diagram, because the gap between what a platform enforces and what people assume it enforces is where incidents live.

  • Prompt injection. Neither component reads intent. If an agent is manipulated into taking an action that policy permits, both layers will permit it. Policy constrains the action space; it does not evaluate why the agent chose an action within it.
  • Model deception. An agent that pursues an objective in a way its operator did not intend, while staying inside policy, is invisible to both layers.
  • Bad policy. Every control here executes the policy it was given. A policy that allows too much is enforced faithfully. The Policy Prover checks whether a change widens access relative to the current boundary — it cannot tell you the original boundary was wrong.
  • Operator misconfiguration. Default-deny substantially reduces this class, since errors tend to break the agent rather than open the network. It does not eliminate it, and the biggest remaining risk is the permissive policy written to unblock a failing job.
  • Legitimate-but-dangerous tool calls. An agent with write access to a database that issues a valid, destructive query has violated nothing. This is the hardest problem in agent security and nothing in this platform addresses it.
  • Data leakage within permitted paths. An allowlisted endpoint can receive anything the agent can read. Egress control constrains destinations, not payloads.
  • Insider threats and application vulnerabilities. Out of scope, and worth saying explicitly because “agent safety platform” sounds broader than it is.

The clean summary: this architecture is strong on reachability and silent on intent. It makes it much harder for an agent to touch something it was never permitted to touch. It does nothing about an agent doing the wrong thing with what it was permitted to touch.

The Portability Question

The two halves port very differently, which is the most practically useful thing to understand about this announcement.

OpenShell is Apache 2.0 and runs on Docker, Podman, Kubernetes and VM isolation, across local machines, on-prem, hybrid, cloud and air-gapped environments. Nothing in its documented design requires NVIDIA accelerators for the runtime boundary itself, and NVIDIA states the platform is compatible with other hardware even though it is optimised for its own. Policies are YAML, version-controllable, and enforced by software you can read.

Sentry does not port. It is a reference design for specific NVIDIA silicon, and the capability it provides — enforcement that survives host compromise — has no drop-in equivalent on a generic server.

That asymmetry has a practical consequence for heterogeneous estates. A security control you can apply to 60% of your fleet is not a security baseline; it is a tier. Organisations running mixed infrastructure will need to decide whether to hold the whole estate to what OpenShell alone can enforce, or accept that some workloads have a stronger boundary than others and document which is which.

The nearest non-NVIDIA equivalents are partial: smart NICs with programmable enforcement, service-mesh egress control, network-level policy at the hypervisor or switch. Each achieves some of the out-of-band property. None is packaged as an agent-aware watchdog today.

The Vendor-Control Question

When one vendor supplies the accelerators, the networking, the DPUs, the agent runtime, the enforcement layer and the monitoring, the security architecture and the procurement relationship become the same relationship. This deserves examination without accusation.

The case for it is real. Vertical integration means the layers are designed to work together, the enforcement point is closer to the hardware than any third party can reach, and the trust chain is shorter. Security benefits from fewer seams.

The costs are equally real:

  • Infrastructure dependency. The strongest control in the stack is available only on one vendor’s silicon.
  • Asymmetric portability. The open-source half travels; the hardware half does not.
  • Refresh coupling. Security capability becomes tied to hardware refresh cycles rather than software releases.
  • Concentration. One vendor’s firmware becomes a component of the security boundary for a large share of the industry’s agent workloads.

The mitigating detail is that NVIDIA has open-sourced the layer that does most of the day-to-day work. OpenShell without Sentry is a substantially complete runtime boundary, and it is the part most organisations will actually deploy. Whether that remains true as the platform develops is the thing to watch.

What Enterprises Should Actually Evaluate

Ten questions. They apply to this platform and to every competing approach, which is rather the point.

  1. Where is the enforcement point, and what trust domain is it in? Same process, same host, or separate hardware?
  2. Can the agent modify its own controls? Directly, or by influencing whoever can?
  3. What happens on policy violation — block, alert, or log? All three are defensible; only one of them is prevention.
  4. Is egress independently enforced, or does it depend on the agent using the SDK you gave it?
  5. Are credentials brokered outside the agent environment, or injected into it?
  6. Can the control plane be compromised by the workload it governs?
  7. Is there hardware dependency, and what does the control degrade to without it?
  8. Is every decision observable and auditable, including denials?
  9. Can policies be reviewed, versioned and diffed before they reach production?
  10. What happens during infrastructure failure — does enforcement fail open or closed?

Question 10 is the one that separates a security control from a monitoring feature, and it is rarely answered in an announcement.

What We Know and What We Don’t Yet Know

EstablishedNot established
OpenShell is available now under Apache 2.0Any independent security assessment of it
Its components and their documented functionsHow the Policy Prover’s formal verification works in detail
Sentry runs out-of-band on BlueField-4Whether BlueField-3 is technically excluded or merely unsupported
NVIDIA claims millisecond quarantineAny third-party measurement of that latency
100+ organisations are involvedWhat most of them have actually deployed
The platform is a reference designSentry’s price, general availability, or deployment footprint

For a platform announced days ago, that balance is normal rather than damning. The reason to write it down is that the gap will close unevenly — OpenShell can be evaluated by anyone this week, and Sentry cannot be evaluated by anyone outside NVIDIA’s hardware base for some time.

FAQ

What is the NVIDIA Open Agent Safety Platform?

An open reference system design announced on September 28, 2026, combining the OpenShell secure runtime with NVIDIA Sentry, an in-silicon enforcement layer running on BlueField-4 DPUs. NVIDIA says more than 100 organisations are working with the technologies in it.

What is OpenShell?

An open-source Apache 2.0 runtime that governs what an agent can access. It runs each agent in an isolated sandbox with no direct network access, filters system calls in the kernel, brokers credentials through a supervisor outside the sandbox, enforces egress policy per request, and verifies policy changes with a formal Policy Prover.

What is NVIDIA Sentry?

An out-of-band watchdog design that runs on BlueField-4 DPUs, monitoring agent activity from outside the host and enforcing policy independently. NVIDIA says it can quarantine an agent within milliseconds; no independent testing of that figure has been published.

What is the difference between OpenShell and Sentry?

OpenShell enforces policy while the agent runs, in software, on the host. Sentry watches from separate silicon on the network path, so it continues to operate even if the host is compromised. OpenShell is available now; Sentry depends on BlueField-4 hardware and has no disclosed general availability date.

Does Sentry require BlueField-4?

NVIDIA’s materials specify BlueField-4 for Sentry. NVIDIA has not published a technical explanation of whether BlueField-3 is architecturally excluded or simply unsupported, so treat the requirement as stated rather than explained.

Does OpenShell require BlueField-4?

No. OpenShell runs on Docker, Podman, Kubernetes and VM isolation across local, on-premises, hybrid, cloud and air-gapped environments. BlueField-4 is needed only to add Sentry.

Does OpenShell prevent prompt injection?

No. It constrains what an agent can reach, not what it can be persuaded to attempt. If an injected instruction leads to an action the policy permits, the action proceeds. Egress policy limits the damage; it does not detect the manipulation.

Can this protect agents running outside NVIDIA infrastructure?

Partly. OpenShell’s runtime boundary is portable and hardware-agnostic in its documented design. The out-of-band enforcement Sentry provides has no equivalent on generic servers, so heterogeneous estates end up with two tiers of containment strength.

Does hardware-level containment replace application-level security?

No. It addresses reachability — what an agent can touch. Application security, data governance, least-privilege credentials and human review still govern whether a permitted action is a good one.


Keep reading

NVIDIA Open Agent Safety Platform

NVIDIA Open Agent Safety Platform: What It Covers

Every agent security architecture has to answer one question, and most of them answer it badly: who enforces the boundary when the thing being constrained …

Read more

AI accelerator depreciation

AI Accelerator Depreciation: The Assumption Nobody Audits

An A100 with 80GB of HBM2e still computes exactly as well in 2026 as it did when it was installed. Nothing has degraded. It passes …

Read more

Accelerator Utilisation

Accelerator Utilisation: The Number That Decides Your Bill

An accelerator bills the same whether it is generating tokens or waiting for a database query to return. The hardware cost is fixed the moment …

Read more

Isaac ROS 5.0

Isaac ROS 5.0 and the Robotics Agent Boundary Problem

One of the agent skills NVIDIA published alongside Isaac ROS 5.0 is called submit-and-monitor-mission. Its documented sample prompt reads: submit a route mission to carter01 …

Read more