Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Tools

Declare the tools an agent can call in Orca Agent Engine - the built-in sandbox toolset, tools from MCP servers, and custom tools your own application executes.

Tools are what an agent can do beyond producing text. You declare them in the agent's tools array, alongside the model and system prompt, and the agent carries them into every session created from that version. Three kinds are available: a built-in toolset that operates on the session sandbox, toolsets exposed by MCP servers, and custom tools that your own application executes and returns results for.

Get your registry endpoint

Registry endpoint

Examples on this page target your registry endpoint - the deployment host root, with no path suffix. For CLI, set ORCA_REGISTRY_URL and exactly one of ORCA_ACCESS_TOKEN (Bearer) or ORCA_API_KEY (x-api-key). For TypeScript SDK, set ORCA_BASE_URL / ORCA_API_KEY (Bearer). To find the endpoint, see Connect to the registry.

Tool types

Every entry in tools carries a type. The registry accepts four values and rejects anything else:

typeWhat it declaresRequired fields
agent_toolsetThe built-in sandbox toolset.-
agent_toolset_20260401Dated alias for agent_toolset, accepted for Claude Managed Agents API compatibility.-
mcp_toolsetThe tools exposed by one declared MCP server.mcp_server_name
customA tool your application executes.name, description, input_schema

agent_toolset_20260401 and agent_toolset are the same toolset. The registry stores the canonical form and echoes the dated form back, so a request written against the Claude Managed Agents API works unchanged.

An agent can declare at most 128 tools.

The built-in toolset

Adding agent_toolset gives the agent tools that operate on its session sandbox:

ToolWhat it does
bashRun a shell command in the sandbox.
readRead a file from the sandbox filesystem.
writeWrite a file to the sandbox filesystem.
editReplace a string in a file.
globMatch files by glob pattern.
grepSearch file contents by regular expression.
listList a directory.
deleteDelete a file.

All eight are available whenever the toolset is declared, unless you turn one off with enabled - see Configure the toolset. On the claude_code harness the toolset expands to six: list and delete have no equivalent there and are silently absent.

Agent Engine sessions have no general-purpose web access. The registry still accepts web_fetch and web_search - both as configs[].name values and as reserved custom-tool names - for Claude Managed Agents API compatibility, but neither tool reaches the model. They are filtered out before the session starts, and the equivalent harness tools are explicitly disallowed. To give an agent network access, declare an MCP server that provides it.

Declare the toolset when you create the agent:

ork agent create \
  --name "repo-assistant" \
  --model claude-sonnet-4-6 \
  --system "You help engineers navigate and edit this repository." \
  --tool-json '{"type":"agent_toolset"}'

The response echoes the toolset with the defaults the registry filled in:

{
  "id": "agt_01H8...",
  "type": "agent",
  "name": "repo-assistant",
  "tools": [
    {
      "type": "agent_toolset_20260401",
      "default_config": {
        "enabled": true,
        "permission_policy": { "type": "always_allow" }
      },
      "configs": []
    }
  ],
  "version": 1
}

Configure the toolset

default_config sets the baseline for every tool in the set, and configs overrides it per tool. configs accepts either an array of entries carrying a name, or an object keyed by tool name.

The setting worth reaching for is permission_policy, which decides whether a tool runs automatically or waits for your approval. This example allows the toolset by default but requires confirmation before any shell command:

ork agent create \
  --name "repo-assistant" \
  --model claude-sonnet-4-6 \
  --tool-json '{
    "type": "agent_toolset",
    "default_config": { "permission_policy": { "type": "always_allow" } },
    "configs": [ { "name": "bash", "permission_policy": { "type": "always_ask" } } ]
  }'

See Permission policies for what each policy does and how to answer a confirmation request.

Which names configs accepts depends on which form you use, and the two do not agree.

In the array form, a name must be one the registry recognizes as built-in - bash, edit, read, write, glob, grep, web_fetch, or web_search - and anything else returns 400. In the object form that validation does not run. The runtime recognizes list and delete through that form, so {"type": "agent_toolset", "configs": {"list": {"enabled": true}}} works while the array spelling is rejected. Any other unknown object key is stored but ignored at run time.

Prefer the array form for the names it accepts. Use the object form only for list or delete, which the runtime supports but the array validator omits. Do not use object keys to invent tool names: a typo stores cleanly and governs nothing.

The other setting is enabled, which decides whether a tool is handed to the agent at all. It defaults to true at both levels, and a configs entry inherits whatever default_config resolved to. A tool whose resolved enabled is false is left out of the tools the session exposes, and denied if it is called anyway. So default_config with "enabled": false, plus one configs entry per tool you want back, is how you declare the toolset and hand over only part of it. list and delete can be named this way only through the object form, since the array form's allow-list rejects them. If the agent also declares skills, keep read enabled - a skill-bearing agent whose read is disabled fails at session start rather than at create time.

Tools from MCP servers

On the default claude_agent_sdk harness in separate mode, an mcp_toolset entry exposes the tools of one MCP server declared in the agent's mcp_servers array. The two are checked against each other in both directions: mcp_server_name must name a declared server, and every declared server must be referenced by an mcp_toolset, or the request returns 400.

{
  "mcp_servers": [
    { "name": "tickets", "type": "url", "url": "https://mcp.example.com/tickets" }
  ],
  "tools": [
    { "type": "agent_toolset" },
    { "type": "mcp_toolset", "mcp_server_name": "tickets" }
  ]
}

On that harness, MCP toolsets take the same default_config and configs fields, with configs[].name set to a tool name the MCP server reports. See MCP servers for declaring servers and authenticating to them.

Custom tools

A custom tool is one your application executes. You declare the contract - what the tool is called, what it does, and what input it takes - and the agent decides when to call it. The engine never runs the tool itself: it pauses the session, hands you the call, and waits for a result.

ork agent create \
  --name "weather-agent" \
  --model claude-sonnet-4-6 \
  --tool-json '{"type":"agent_toolset"}' \
  --tool-json '{
    "type": "custom",
    "name": "get_weather",
    "description": "Get the current weather for a city. Returns temperature in Celsius and a one-word conditions summary. Use this whenever the user asks about current conditions; it has no forecast data.",
    "input_schema": {
      "type": "object",
      "properties": { "location": { "type": "string", "description": "City name" } },
      "required": ["location"]
    }
  }'

input_schema must be a JSON Schema object with type: "object". description is required and capped at 4096 characters. A custom tool cannot take the name of a built-in tool: bash, read, write, edit, list, delete, glob, and grep are reserved, along with the two compatibility names noted above.

Answer a custom tool call

When the agent calls a custom tool, the session emits an agent.custom_tool_use event carrying the tool's name and input, then waits. Run the tool and send back a user.custom_tool_result carrying the event's id as custom_tool_use_id.

This happens at run time, so it needs a session. sessionId and $SESSION_ID below are the id from creating a session against this agent; getWeather is your own function.

const stream = await orca.sessions.events.stream(sessionId);

for await (const event of stream) {
  if (
    event.type !== 'agent.custom_tool_use' ||
    !('name' in event) ||
    event['name'] !== 'get_weather' ||
    !('input' in event)
  ) continue;

  const input = event['input'];
  if (
    !input ||
    typeof input !== 'object' ||
    Array.isArray(input) ||
    !('location' in input) ||
    typeof input['location'] !== 'string'
  ) continue;

  const result = await getWeather(input['location']);

  await orca.sessions.events.send(sessionId, {
    events: [
      {
        type: 'user.custom_tool_result',
        // The streamed event carries the id, so nothing has to be extracted.
        custom_tool_use_id: event.id,
        content: [{ type: 'text', text: JSON.stringify(result) }],
      },
    ],
  });
}

content is an array of content blocks - text, image, document, or search_result - not a bare string. Set is_error: true when the tool failed, so the agent can react rather than treating the error text as data.

The TypeScript tab reads event.id off the streamed event directly. The shell tab cannot: it has to pull the ID out of the event stream first, which is why $CUSTOM_TOOL_USE_ID is supplied rather than derived here. See Events and streaming for reading the stream from a shell.

Your application decides whether to run a custom tool. Declared custom tools bypass both exact-name and mcp__orca__* permission policies, so the custom handler can emit its own agent.custom_tool_use request and wait for user.custom_tool_result. Apply any approval or authorization check in your application before you run the tool.

See Events and streaming for the full event stream.

Write good tool descriptions

Tool selection is driven by the description, so it is the highest-leverage thing you control:

  • Be specific about when the tool applies, and when it does not. Three or four sentences is a reasonable target. Say what the tool returns, what each parameter means, and what it cannot do - a tool that returns current conditions should say it has no forecast data, or the agent will call it for forecasts.
  • Consolidate related operations. One manage_pr tool with an action parameter beats separate create_pr, review_pr, and merge_pr tools. Fewer, broader tools reduce selection ambiguity.
  • Namespace names by resource when tools span several systems - db_query, storage_read. Ambiguity grows with the size of the tool surface.
  • Return only what the agent needs next. Stable identifiers and the fields that inform the next step. Bloated responses spend context and bury the signal.

Permissions

Tools are part of the agent resource. Treat every Workspace API key as full access to its Workspace's resources, and separate access with Workspaces. See Control registry access.

What's next

On this page