Claude Code harness
Run the agent loop inside the session sandbox in Orca Agent Engine, and the capabilities that trades away.
The claude_code harness runs the agent loop inside the session sandbox instead of in the harness server. The sandbox image ships the loop as an HTTP server, and the harness server becomes a bridge: it opens a session, forwards your events in, and streams the sandbox's events back out.
Reach for it when the loop itself needs to be inside the isolation boundary - long autonomous work on a repository, or a compliance requirement that the process touching the code runs in a container you control.
This harness runs only in colocated mode, which is not fully implemented in this release. For a harness that runs today, use the Claude Agent SDK harness, which runs separate, the default.
Get your registry endpoint
Registry endpoint
Examples on this page target your registry endpoint - the deployment host root, with no path
suffix. For CLI, set ORCA_REGISTRY_URL and exactly one of ORCA_ACCESS_TOKEN (Bearer) or
ORCA_API_KEY (x-api-key). For TypeScript SDK, set ORCA_BASE_URL / ORCA_API_KEY (Bearer).
To find the endpoint, see Connect to the registry.
Select this harness
ork agent create \
--name "repo-worker" \
--model claude-sonnet-4-6 \
--metadata harness=claude_codeThis harness runs only in colocated mode. Pairing it with mode: "separate" returns 400.
Despite the name, this is the same agent engine as the Claude Agent SDK harness - the difference is where the process runs, not which model loop drives it. Choose between them on isolation and capability, not on expected model behavior.
What it gives up
Every limitation below is silent. The registry accepts the configuration, returns 200, and echoes it back; the capability simply is not there at run time. Read this list before you switch an existing agent over.
Tools with an always_ask policy are removed from the session, not prompted for. There is no channel from inside the sandbox to your application for a confirmation, so the tool is dropped. always_ask is also the default for MCP toolsets, which means an MCP toolset with no explicit policy contributes no tools at all. Set always_allow on anything this agent needs.
MCP servers are not wired. An agent's mcp_servers and mcp_toolset entries validate and store normally, and the model never sees those tools. Sessions emit no agent.mcp_tool_use or agent.mcp_tool_result. If the agent needs MCP tools, use the Claude Agent SDK harness.
The list and delete built-in tools are unavailable. agent_toolset expands to six tools here - bash, read, write, edit, glob, and grep. The agent can still list and delete through bash.
Delegates lose the same tools. A coordinator's delegates run here, but each delegate's allowlist passes through both filters above: list and delete are dropped, and so is anything whose resolved policy is not always_allow. The default harness applies neither filter to a delegate, so a roster that works there can reach the sandbox with fewer tools here.
Also absent:
| Capability | Behavior here |
|---|---|
| Token-level streaming | Not emitted. agent.message arrives per turn, not as deltas. |
user.interrupt | Rejected. The session emits session.error. The type reads unknown_error on the default wire dialect; only orca-beta clients see the underlying unapplied_event. |
user.tool_confirmation | Rejected, for the reason above. |
user.define_outcome | Rejected. No outcome-evaluation spans are produced. |
user.tool_result | Rejected at the API with 400. |
agent.thread_context_compacted | Not emitted, so context compaction is not visible. |
session.status_rescheduled | Not emitted. Provider retries are not surfaced separately. |
| Guardrails | Not fully supported yet in colocated mode. See Govern agents. |
Skills, custom tools, agent.thinking, and the full file, memory store, and GitHub repository mount lifecycle all work exactly as they do on the default harness. Delegation works too, with the tool reduction noted above.
Environment requirements
The loop listens on a port inside the sandbox, so the sandbox runtime has to be able to expose one.
This harness requires a runtime that publishes a sandbox endpoint. Starting a session on a runtime without one fails with colocated requires a runtime that exposes the harness port. The local runtime used by the evaluation stack does not qualify - see Self-hosted install.
The sandbox image
The harness ships as a container image, and the environment decides which one runs. Leave image unset and the platform uses the operator-selected image. The Agent Engine v0.5.1 release chart pins the digest below; the same multi-platform image is publicly available on Docker Hub:
| Value | |
|---|---|
| Public image | docker.io/streamnative/sandbox-harness-claude-code@sha256:80cfee4761746bc167b5073e6d7236baddea6756cf31a00bbeb5b600e31cdd24 |
| Chart default | ghcr.io/orca-ae/sandbox-harness-claude-code@sha256:80cfee4761746bc167b5073e6d7236baddea6756cf31a00bbeb5b600e31cdd24 |
| Port | 4096 |
The Docker Hub tag is docker.io/streamnative/sandbox-harness-claude-code:0.5.1. A self-hosted operator using that mirror must also authorize docker.io/streamnative in OpenSandbox's trusted repository prefixes. The image is chosen for you. claude_code runs colocated, and the environment's image field
is a top-level field the registry accepts but the sandbox refuses.
Setting image on an environment used by this harness makes every session on it fail to start.
Sandbox setup rejects a user-supplied image in colocated mode - the same check that rejects
config.packages - and the session emits session.error with type setup_failed, staying idle and resumable rather than running. There is no way to pin or replace the
harness image from the API; the operator-owned catalog image is what runs. Ask your operator to
change it.
The contract that image satisfies is worth knowing even though you cannot swap it: it serves the
HTTP protocol above on port 4096 and runs as the node user, which is why the platform applies
that ownership to uploaded files - the harness reads them without a recursive chown.
Model access and credentials
Provider credentials never enter the sandbox. The platform mints a short-lived, session-scoped token and injects it into the sandbox alongside a gateway URL, so the loop reaches models through Orca AI Gateway rather than holding a provider key. A leaked sandbox yields a token that is scoped to one session and expires.
The in-sandbox harness always governs model traffic through the gateway. The separate harness can
also use the gateway through its deployment-wide default or a per-session override. See
Use the gateway with Agent Engine.
Write isolation and outputs
The loop runs under a kernel namespace inside the image that restricts writes to the paths the session declares. The check runs before the agent starts and fails closed - if isolation cannot be established, the session does not run.
Session outputs are captured to a directory inside the sandbox rather than mounted from object storage. They are collected when the session ends, so they survive the sandbox but are not visible externally while it runs.
Session continuity
State lives in two places. The loop keeps its own conversation state inside the sandbox and resumes natively across turns, which is fast. If the sandbox is replaced - a cold start, a rebuild - that state is gone, and the platform replays the session's transcript from Orca's store to rebuild it. You do not choose between them; the fallback is automatic.