Blog · 6 October 2026 · Agent infrastructure

Give agents room to work.
Keep a boundary you can check.

Captain Code now includes an experimental OpenShell execution path. It connects scoped sandbox work to verification and inspectable run records. Our next step is stronger evidence for agentic compliance and replayability.

A coding agent needs to read files, run tests and call a model. Giving it those tools also creates a boundary problem. Which files can it change? Which destinations can it contact? What evidence remains when it finishes?

A prompt that says "stay inside these limits" cannot enforce those limits. We integrated NVIDIA OpenShell into Captain Code to put runtime controls around a defined task and retain evidence of the result.

The integration is in the latest Captain Code release. It is experimental, requires a prepared runtime, and is selected explicitly. It does not put every Captain Code session inside a sandbox.

What the OpenShell initiative provides

OpenShell is NVIDIA's open-source runtime for AI agents, licensed under Apache 2.0. It places agents in isolated sandboxes and enforces policies for filesystem and network access. Its provider mechanism supplies credentials to approved endpoints without putting those provider keys in the worker environment.

Captain Code supplies the task and orchestration. OpenShell supplies the execution boundary. The combination lets us ask a more useful question than "did the agent finish?": "what was it allowed to do, and which checks passed before we accepted its output?"

What you can do in Captain Code

You select the repository revision, the existing files a worker may edit, and the command that verifies its work. Captain snapshots that revision. The worker edits and tests inside the sandbox. Captain checks the exported patch and verifies the combined result in a fresh sandbox before returning it.

The result is an exported patch, not an automatic edit to your working tree. You can inspect the result and its evidence before applying it. Passing your chosen tests is evidence about those tests, not proof that the change is correct in every case.

The released path also supports sandbox-only teams and staged workflows. Mixed host and sandbox stages are refused. A failed sandbox worker does not silently fall back to an unrestricted host worker.

After the pilot setup in the latest release (examples/openshell-pilot/README.md), an explicit task looks like this:

# Set the prepared runtime, trusted pilot path and qualified profile first.
# These are example paths and tests for your own repository.
export CAPTAIN_OPENSHELL_ALLOWED='src/parser.py'
export CAPTAIN_OPENSHELL_VERIFY='["python3","-m","unittest"]'
captain with openshell "Fix the parser regression within the allowed file"

This is not a complete bootstrap command. The prepared image must include the test runtime and dependencies. Profile qualification must pass before selection. Our tested MicroVM recovery path still requires a local driver fix; upstream PR 3940 is closed and unmerged as of 6 October 2026. Follow the versioned setup guide and its limits.

From sandbox controls to agentic compliance

Regulated and mission-critical teams need evidence that an action stayed within an approved mandate. A sandbox can limit access. An audit record can explain what the controller observed. They serve different purposes.

Captain's OpenShell records include the base revision, file list, patch digest, verification command and run report. A task identifier connects those artifacts to the attempt and handoff. Failed or interrupted runs retain evidence without being presented as a successful verified export.

That is a useful foundation for agentic compliance. It is not a compliance certification. A local record and its hash are not, on their own, a tamper-resistant proof: a party able to replace both can conceal a change. Logs can also contain sensitive information and need access and retention controls.

Our intended next layer is to bind the approved policy, input commitments, execution evidence and result into an independently checkable receipt. Our work on zero-knowledge state proofs aims to let a verifier check a defined predicate without receiving the private inputs. This proof-backed workflow integration is planned; it is not an OpenShell feature or a shipped guarantee of Captain Code.

A proof would still depend on its statement and input provenance. It would not prove that an off-system fact is true, that an omitted event never happened, or that a model's advice is correct.

Determinism makes replay a testable goal

A saved checkpoint supports recovery. It does not guarantee that rerunning an agent produces the same tool calls or output.

Replay needs a controlled starting state, tool versions, model serving configuration and external responses. It also needs an explicit comparison rule. OpenShell helps control the execution environment; it does not make a remote model deterministic.

Captain Code's separate /deterministic mode uses measurements from our Agentic Determinism Index to prefer measured serving configurations. Its documentation in the latest release (docs/ADI.md) explains an important limit: when no registered configuration is green, the turn can proceed without that filter and reports why. A green result is a measurement, not a guarantee for the next request or a complete agent workflow. Selecting OpenShell does not establish ADI eligibility.

We want to turn replay into a check teams can run, with mismatches reported as evidence. We do not claim identical end-to-end replay today.

How we intend to contribute

We intend to contribute reproducible boundary tests, small failure cases and runtime fixes through the upstream review process. Our existing pilot publishes checks for denied file and network access, cancellation, recovery and exact patch export. An unmerged patch remains our responsibility to maintain.

We also want to share evidence formats that connect a task, policy, snapshot, test result and output. The goal is to make audit and replay tools work together, while keeping private payloads out of shared evidence where possible. These are proposed contribution areas, not an upstream roadmap commitment.

Thank you to NVIDIA and the OpenShell maintainers and contributors. A shared runtime gives application teams a place to test boundaries, report failures and improve the controls together.

Explore Captain Code, read the OpenShell documentation, or inspect the run-record format in the latest release (docs/RUN_RECORDS.md).