Orca Agent Engine
Build, run, and govern AI agents in Orca Agent Engine on any harness, model, and sandbox - self-hosted or on StreamNative Cloud.
Orca Agent Engine is a managed runtime for production AI agents. You define agents as versioned registry resources, run them in isolated sandboxes, route model and tool calls through a governed gateway, and keep every event, file, memory, and credential inside your own Workspace. Its core API shares 72 operations with the Claude Managed Agents API, so a client written for those operations can call Orca after you change its base URL. Capabilities beyond that contract are extension groups, and which ones exist depends on the deployment.
Why a runtime for agents
Your organization already knows how to run services in production: identity, isolation, change control, audit. Agents break those assumptions. They hold credentials, execute code they were not shipped with, and call external systems autonomously - usually from a framework process with none of the controls your platform team expects.
Agent Engine closes that gap:
- Agents are resources, not processes. An agent is a versioned registry record - model, system prompt, tools, MCP servers, skills - that you review, pin, and roll back like any other config.
- Every run is a session. Sessions produce an ordered, replayable event stream you can tail live, store, and audit.
- Credentials never reach the agent. Vaults hold the real secrets; the gateway injects short-lived tokens per call and writes an attributable audit trail.
- Code runs in sandboxes. Each session executes inside an isolated environment with declared packages and network policy.
Architecture
- Registry - the control plane. Its core API manages agents, sessions, environments, skills, files, vaults, memory stores, and Git credentials on any deployment; a deployment may add extension groups on top. See Registry.
- Harness server - hosts the agent loop on the harness the agent declares. See Runtime.
- Sandbox - isolated execution for each session's code and tool calls, built from an environment template.
- AI Gateway - vault-aware egress for model providers and MCP servers. Real credentials stay in vaults; calls carry short-lived tokens and land in the audit log. The gateway is a separate Orca product that also runs standalone; Agent Engine routes MCP tool calls through it and can use it as the deployment-wide default for model calls. See Use with Agent Engine for configuration.
- Transcript stream - every session event is appended to a durable stream you can tail, replay, and monitor.
Open at every layer
- Any harness. Each agent declares its harness - Claude Agent SDK, Claude Code, Codex SDK, or Pi SDK - and where its loop runs: in the harness server, or inside the session sandbox. See Choose a harness.
- Any model. The agent names its model in a
modelfield - swap the id, or set a speed tier and reasoning effort, without touching agent code. See Model access. - Any sandbox. Environments declare packages and networking for pluggable sandbox runtimes. See Environments.
Compatibility with Claude Managed Agents
The registry accepts the anthropic-beta: managed-agents-2026-04-01 header and mirrors the Claude Managed Agents resource shapes, error envelope, and pagination, so Anthropic SDKs and tools can call the shared operations by pointing their base URL at your registry endpoint:
curl "$ORCA_REGISTRY_URL/v1/agents" \
-H "Authorization: Bearer $ORCA_ACCESS_TOKEN"Orca also ships its own surfaces for the same API: the ork CLI, the TypeScript SDK, the Python SDK, the Go SDK, and the rendered API reference.
Resources you manage
| Resource | What it is |
|---|---|
| Agent | The versioned definition: model, system prompt, tools, MCP servers, skills. |
| Environment | A reusable sandbox template - packages and network policy - shared across sessions. |
| Session | A running instance of an agent inside an environment. |
| Events & streaming | The events you send to a session and the stream it emits back. |
| Skill | A filesystem-shaped capability bundle agents load on demand. |
| File | An uploaded asset sessions can mount as a resource. |
| Memory store | Durable agent memory with versioning and redaction. |
| Vault | Credential storage the gateway uses to authenticate MCP calls. |
| Trigger | Starts sessions on a cron schedule in every deployment. StreamNative Cloud extends this to Pulsar and Kafka events. |
| Harness | Selects the agent loop: Claude Agent SDK, Claude Code, Codex SDK, or Pi SDK, subject to the deployed runtime version. |
| Guardrail | Applies request, tool, and budget rules according to its scope and the harness's enforcement support. |
| Model pricing | Provides the prices used by cost guardrails and usage accounting. |
Deployment
There is one registry API and two ways to run it. Self-host the open-source engine (Apache-2.0) in your own cloud, or use a managed Workspace on StreamNative Cloud. Guides on this site are deployment-neutral: they target the registry endpoint you configure. See Self-hosted deployment.
Next steps
Quickstart
Define an agent, start a session, and stream events end-to-end.
Self-hosted deployment
Run the open-source engine with the Helm chart or the local stack.
CLI
Manage registry resources from the terminal with ork.
API reference
Browse every registry operation, generated from the OpenAPI spec.
AI Gateway
Route, govern, and observe model and MCP tool traffic - with or without Agent Engine.