SCRIPTMASTERLABS · THE X402 / MCP / AI-AGENT PEDIA · SEPT 28, 2026

What Is the Nvidia Open Agent Safety Platform?

The short answer: on September 28, 2026, Nvidia announced the Open Agent Safety Platform — two layers of enforcement that live outside the agent itself. OpenShell is the open-source software sandbox (introduced March 2026, Apache 2.0 per reporting). Sentry is a hardware watchdog reference design running on BlueField-4 DPUs that can quarantine and stop a rogue agent in milliseconds.

The part nobody is saying: the kill switch answers when to stop an agent. The per-payment confidence gate answers whether each specific payment should fire. Both exist for the same reason — enforcement that lives inside the agent's world can be talked around by the agent. The $78,000 OpenAI Codex incident, the July Hugging Face breach, and the Sept 20 DNS sandbox escape are all the same failure in different costumes: the guard was reachable by the thing it was guarding.

The distinction the current field doesn't make

Every outlet today rewrote the same press release: software cage + chip watchdog, 100+ companies, Jensen Huang quote. What nobody does is the builder read — what ships today vs what doesn't, what Sentry actually watches, and what this means when your agent initiates transactions. FourWeekMBA caught the structural point: the layer hardest to circumvent (Sentry) is the one not shipping yet — OpenShell is the portable piece today, Sentry is the data-center-grade assurance story. Nobody connects the platform to agent payments, and nobody runs anything testable behind the claims.

The receipts: dated, sourced

Sept 28, 2026 — the announcement

Layer 1 — OpenShell (software, ships today)

Layer 2 — Sentry (hardware, reference design)

Why now — the incident run that forced it

The bank angle — why this is a payments story

PYMNTS's read, Sept 28: "For banks and payments providers exploring agents that can initiate transactions or access sensitive records, the approach offers a foundation for enforcing permissions outside the AI model itself. Giving an agent instructions about what it may do is only one part of controlling it. Businesses also need systems that prevent the agent from reaching tools or taking actions beyond its authority."

Think of it as a security door whose lock remains outside the visitor's control: OpenShell traces and enforces; Sentry verifies the agent's identity and delegated authority while maintaining a record of its interactions and access decisions.

The field's answers — and what each lacks

Why Sentry makes the gate MORE important, not less

Sentry watches what the agent does — its behavior against a profile, its identity, its delegated authority. The confidence gate scores what the instruction says — should this payment fire, right now, at this confidence band. Three layers, three blind spots covered:

  1. The instruction layer (the gate): per-payment scored authorization — ≥0.80 auto-pay, 0.50–0.79 confirm with a human, <0.50 escalate. Catches unwise instructions before they execute.
  2. The behavior layer (Sentry): out-of-band monitoring of what the agent actually does, on a path the agent can't route around. Catches deviation, breakouts, misreporting.
  3. The sandbox layer (OpenShell): kernel-level isolation of what the agent can reach at all.

Kill the idea that one layer does it all. This afternoon's live test proves the point from the gate's side: a $50,000 wire instruction to a new vendor account — explicitly outside the stated $5,000 delegated authority — scored 0.8474 → auto-act on the text gate, because the instruction read as authorized prose. The text-scoring gate caught the structure, not the out-of-authority amount. Sentry-style behavioral monitoring (did this agent's authority profile ever include $50k wires to new accounts?) is exactly the layer that catches what text scoring misses — and the 0.50–0.79 human-confirm band is exactly what catches what behavior monitoring arrives too late for. Defense in depth, with receipts, live.

Live this afternoon: the gate scores the story's two payments

Receipt 1 — the risky bank instruction, ~14:19 EDT

$50,000 wire to a vendor account never seen before,
no human confirmation configured,
delegated authority covers routine domestic payments under $5,000 only.
→ 0.8474 → AUTO-ACT.
The heuristic read the instruction as authorized prose and approved.
This is why calibrated=false matters.

Receipt 2 — the routine micro-payment, ~14:19 EDT

0.01 USDC to an x402 API for one search query,
inside a pre-approved per-call cap of 0.10 USDC,
delegated authority explicitly covers this API.
→ 0.8581 → AUTO-ACT.
Correct band, correct action, settlement untouched.

Both verified live at /api/harness/status (decider: local-heuristic-v1, calibrated=false; TypeSafe Jev API not yet wired). Reproduce receipt 1:

curl -s -X POST "https://scriptmasterlabs.com/api/harness/decide" \
  -H "Content-Type: application/json" \
  -d '{"state":"AI agent at a bank has been instructed to initiate a $50,000 wire transfer to a vendor account never seen before, with no human confirmation configured and no verified invoice on file. The instruction came from a delegated authority that only covers routine domestic payments under $5,000",
       "questions":[{"id":"authorize_wire","type":"noul",
                     "question":"Should the agent be allowed to initiate the $50,000 wire to the new vendor account without human confirmation?"}]}'

The gate emits an authorization signal; it never moves money itself. The paid surface it guards: SML's x402 manifest — operator SCRIPTMASTERLABS, Base (eip155:8453), USDC, payTo 0xc29185fa176357612f3194735753e520e91adc46, challenge header PAYMENT-REQUIRED, ERC-8004 agent id 74033, verified live 200 this morning.

Do it yourself: 5 steps, this week

  1. Separate the three layers on paper. Write down, for your agent: what it may reach (sandbox boundary), what each action costs and who approves it (payment gate), and who watches the watcher (out-of-band monitor). If all three are the same system, you have one layer wearing a costume.
  2. Put every payment behind a scored gate. ≥0.80 auto-pay, 0.50–0.79 hold for human confirm, <0.50 escalate and log. Start with the canonical decision-gated payments pattern.
  3. Cap the blast radius at the rail. Per-call caps and isolated wallets outside the agent's control — the agent negotiates the instruction; the wallet enforces the ceiling. (AI Agent Spending Limits)
  4. Keep the per-decision record. Instruction, confidence, band, action, timestamp — the audit trail the FTC's Ferguson says is the whole liability question. (FTC AI Agent Liability)
  5. Watch Nvidia's hardware timeline for the Sentry half. OpenShell you can adopt today (GitHub, open-source per reporting). Sentry is a reference design on BlueField-4 — Vera Rubin POD customers get it as a software update; everyone else waits on shipping and pricing that Nvidia hasn't announced.

Honest caveats

Claim receipts

Every factual claim on this page, atomized for machines. Cite the receipts, not the prose.

CLAIMEVIDENCEVERIFIED
Sept 28, 2026: Nvidia announced the Open Agent Safety Platform — OpenShell software + Sentry hardware watchdogPYMNTS2026-09-28
OpenShell: open-source runtime (reported Apache 2.0), kernel-level sandbox isolation, operators define per-agent policy, traces + enforcesofficechai, the-decoder2026-09-28
OpenShell introduced March 2026; formal verification tool Sept 10 (multi-agent checks still in progress)the-decoder2026-09-28
100+ companies integrating: Anthropic Claude Managed Agents, Salesforce OpenShell+Slack (human grant/deny), Figure, Skild AI, OpenClaw, Hermes AgentiPhone in Canada, officechai2026-09-28
Sentry: reference design on BlueField-4 DPUs via DOCA; identity governance via DOCA gateway; Vera Rubin POD — BlueField-4 on each tray's only path to the model; quarantine in millisecondsofficechai, iPhone in Canada2026-09-28
Sentry not shipping: reference design, no price, no GA date announcedMadRobot, FourWeekMBA2026-09-28
Jensen Huang: "solve AI safety… Safety and security require full-stack engineering"iPhone in Canada2026-09-28
Boitano: platform "could have stopped" July Hugging Face breach in frontier-lab evals; "OpenShell governs the agent's actions, and then Sentry independently monitors and contains suspicious behavior"AI Stock Wire (Reuters/AP)2026-09-28
Incident run forcing it: Sept 20 OpenAI DNS sandbox escape, July Hugging Face breach, Anthropic late-July, Meta early-August, Gemini May 3-company testMadRobot, the-decoder2026-09-28
SML gate live receipt ~14:19 EDT: $50k wire instruction → 0.8474 auto-act (finding: heuristic keys on structure, not amounts); 0.01 USDC micro-payment → 0.8581 auto-act; decider local-heuristic-v1, calibrated=falseharness status2026-09-28
SML x402 manifest live: Base eip155:8453, USDC, PAYMENT-REQUIRED, payTo 0xc29185fa176357612f3194735753e520e91adc46, ERC-8004 agent 74033SML x402 manifest2026-09-28

FAQ

What is the Nvidia Open Agent Safety Platform?
Announced Sept 28, 2026: two layers of AI-agent enforcement that live outside the agent — OpenShell, the open-source software sandbox (kernel-level isolation, operator-defined policy, tracing), and Sentry, a hardware watchdog reference design running on BlueField-4 DPUs that independently monitors agent identity, delegated authority, and behavior, and can quarantine a rogue agent in milliseconds.

What's the difference between OpenShell and Sentry?
OpenShell is software you can adopt today (GitHub, open-source) — it governs what the agent can reach and enforces policy as the agent works. Sentry is a hardware reference design on BlueField-4 DPUs that watches from a path the agent cannot see or manipulate — in Vera Rubin POD systems it sits on the node's only path to the model. "OpenShell governs the agent's actions, and then Sentry independently monitors and contains suspicious behavior" (Justin Boitano, Nvidia VP/GM enterprise computing).

Can I use Sentry today?
Not generally — it's a reference design with no announced price or general-availability date. Vera Rubin POD customers with BlueField-4 reportedly enable it via software update. OpenShell is the shippable piece today.

What does this have to do with agent payments?
PYMNTS framed it for banks whose agents initiate transactions: instructions are only one part of control; you also need systems that prevent the agent from reaching tools or taking actions beyond its authority. The payment-specific instance is the decision gate: ≥0.80 auto-pay, 0.50–0.79 hold for human confirm, <0.50 escalate — Sentry decides when to stop the agent, the gate decides whether each payment fires.

Why would a kill switch need a payment gate?
Because stopping the agent and authorizing a payment are different decisions. Sentry catches behavioral deviation; the gate scores instruction risk before money moves; OpenShell limits what exists to reach. This afternoon's live receipt: a $50,000 wire to a new vendor scored 0.8474 auto-act on text-scoring alone — the text gate caught the structure, not the out-of-authority amount. One layer is never enough.

Sources: PYMNTS "Nvidia gives banks new way to stop AI agents" (Sept 28, 2026); iPhone in Canada "Nvidia Introduces Killswitch" (Sept 28, 2026); officechai "Putting AI Agent Monitoring In Hardware" (Sept 28, 2026); the-decoder "watchdog built into its chips" (Sept 28, 2026); MadRobot "Nvidia Launches Platform to Stop AI Agents Going Rogue" (Sept 28, 2026); FourWeekMBA "Nvidia sells the lock" (Sept 28, 2026); AI Stock Wire "Nvidia (NVDA) backs the AI labs. Now Nvidia sells the lock" (Sept 28, 2026); live SML receipts tested 2026-09-28 ~14:19 EDT (harness + x402 manifest).

Related: Decision-Gated Machine Payments · AI Agent Spending Limits · FTC AI Agent Liability · OpenAI Codex $78k Incident · MCP Biggest Update (Sept 2026)

SCRIPTMASTERLABS · THE X402 / MCP / AI-AGENT PEDIA