Model access
Select the model an agent runs on in Orca Agent Engine, and understand what decides the endpoint that model resolves to.
An agent names the model it runs on through the model field. This page covers that field - the shapes it accepts, what each part controls, and what decides the endpoint a model id ultimately resolves to.
Two different credential stores are easy to confuse. Vaults hold credentials a session uses to call MCP servers, on any deployment. Model-endpoint credentials are separate and are configured by the deployment.
Get your registry endpoint
Registry endpoint
Examples on this page target your registry endpoint - the deployment host root, with no path
suffix. For CLI, set ORCA_REGISTRY_URL and exactly one of ORCA_ACCESS_TOKEN (Bearer) or
ORCA_API_KEY (x-api-key). For TypeScript SDK, set ORCA_BASE_URL / ORCA_API_KEY (Bearer).
To find the endpoint, see Connect to the registry.
Select a model for an agent
The model field on an agent commonly uses three shapes. All three normalize to the same stored value, so pick whichever is most convenient:
| Shape | Example | When to use |
|---|---|---|
| Bare string | "claude-sonnet-4-6" | You only need to name the model. |
{ id, speed } | { "id": "claude-opus-5", "speed": "fast" } | You want a speed tier alongside the model ID. |
{ id, speed, effort } | { "id": "claude-sonnet-4-6", "effort": "high" } | You want to set reasoning effort. |
The object shapes accept these fields:
| Field | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Model identifier the harness recognizes. This is what selects the model. |
speed | enum | No | standard or fast. Responses report standard when you omit it. fast is accepted only for claude-opus-5 and claude-opus-4-8; any other model id returns 400. |
effort | enum | No | Reasoning effort. Accepts either a bare string or { "type": "high" }, but supported levels depend on id. |
provider | string | No | Provider identity. Omit it to use the selected harness's default (anthropic for Claude, openai for Codex SDK and Pi SDK). Pi SDK uses it with the model ID to select a supported native protocol. It appears in orca-beta responses. |
id, speed, and effort are part of the Claude Managed Agents model object, so a client written against that API sets them the same way here. effort accepts a bare level string or { "type": "high" } on both sides. Omitting it resolves the known model's default at save time. provider is an Orca extension.
Claude reasoning effort by model
This table applies to the Claude harnesses. Codex SDK and Pi SDK use their own pinned model catalogs. Deployments that expose runtime.runorca.ai/v1 list them through GET /apis/runtime.runorca.ai/v1/harnesses. An explicit effort must be supported by the selected harness and model. For Claude, a dated snapshot suffix, such as
-20251101, has the same capability as its base model.
| Model | Supported levels | Default |
|---|---|---|
claude-opus-5, claude-opus-4-8, claude-opus-4-7, claude-sonnet-5, claude-fable-5 | low, medium, high, xhigh, max | high |
claude-opus-4-6, claude-sonnet-4-6 | low, medium, high, max | high |
claude-opus-4-5 | low, medium, high | high |
For another Claude model ID, omit effort. An explicit effort for an unsupported model or level returns 400 before the session starts.
The harness validates the provider and model pair, but the deployment still controls credentials and model grants. Even where the harness catalog endpoint is exposed, a listed model may be unavailable until an operator configures those resources. See Codex SDK and Pi SDK for their model and Gateway requirements.
Set the model when you create the agent:
ork agent create \
--name "support-triage" \
--system "You triage support tickets." \
--model-json '{"id":"claude-opus-5","speed":"fast","effort":"high"}'The ork CLI takes either --model <id> for the bare-string form or --model-json <json> for an object. The two are mutually exclusive, and ork agent create requires one of them.
Reading the agent back returns the model as an object. Default clients receive id, speed, and - when you set it - a nested effort:
{
"model": {
"id": "claude-opus-5",
"speed": "fast",
"effort": { "type": "high" }
}
}Clients that send the orca-beta header receive the first-party shape instead, which reports the stored legacy provider value and omits speed and effort:
{
"model": { "provider": "anthropic", "id": "claude-sonnet-4-6" }
}Where a model resolves to
The harness validates the provider and model ID you set. The deployment decides which endpoint and credentials that selection reaches.
- Self-hosted. The model endpoint and its credential are configured on the harness at deploy time. There is no provider API to query.
- StreamNative Cloud. A Workspace declares its LLM endpoints as providers, a read-only Workspace-level resource in the Cloud extensions group.
Permissions
The model and its controls are part of the agent resource. Treat every Workspace API key as full access to its Workspace's resources, and separate access with Workspaces. See Control registry access.