Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Builtin guardrails

The catalog of builtin guardrail types in Orca Agent Engine - what each type checks, its parameters, and the phases and verdicts it supports.

A builtin guardrail is a rule type from the Agent Engine registry's catalog. You select it by name in rule.builtin and configure it with rule.params. The catalog covers tool access, shell commands, integrations, personal data, session behavior, and spend. For a check no builtin covers, write a custom guardrail.

This request body creates a guardrail that keeps any agent that lists it read-only:

{
  "name": "read-only-agent",
  "scope": "explicit",
  "rule": { "kind": "builtin", "builtin": "read_only_os" }
}

Guardrails shows how to create and attach a rule. These rules apply to every builtin:

  • Parameters are checked when you write the rule. An unknown parameter, a wrong type, or a value out of range returns 400. The registry stores params as written and applies the defaults below when the rule runs.
  • phases defaults to the type's phases. You can narrow it to a subset but not add a phase.
  • Only tool_call can ask. The session pauses for a user.tool_confirmation event. A type that asks at another phase denies instead, except the budget thresholds, which wait for the next tool call.
  • Tool names are matched exactly and case-sensitively. A * matches any run of characters, and mcp__<server>__* matches every tool on one MCP server. In separate mode, the built-in tools are named mcp__orca__<tool>, such as mcp__orca__bash.

Catalog

ork guardrails list-types returns the same 22 types with their parameter schemas.

BuiltinChecksPhasesVerdictsState
require_approval_for_toolsNamed tools need approvaltool_callask-
block_toolsNamed tools are deniedtool_calldeny-
ask_on_os_toolsFilesystem and shell tools need approvaltool_callask-
read_only_osFilesystem writes and shell are deniedtool_calldeny-
block_skillsNamed skills cannot loadtool_calldeny-
headless_subagent_purpose_guardSubagent dispatches declare a purposetool_calldeny-
blast_radiusDestructive and risky shell commandstool_callask, deny-
block_working_dir_changesShell commands that change directorytool_callask, deny-
worktree_guardShell writes outside a roottool_calldeny-
github_policyRepository reads and writestool_callask, deny-
gdrive_policyDrive access and confidential filestool_call, tool_resultask, denySession
gmail_policyMail capabilitiestool_calldeny-
gcalendar_policyCalendar capabilitiestool_calldeny-
deny_pii_in_llm_requestPersonal data in user messagesrequest, llm_requestdeny-
max_tool_calls_per_sessionTool calls per sessiontool_calldenySession
spawn_boundsSubagent dispatches per turntool_calldenyTurn
detect_loopRepeated identical tool callstool_callask, denySession
detect_thrashingConsecutive failing tool resultstool_resultdenySession
cost_budgetSession spendrequest, tool_callask, denySession
user_daily_cost_budgetDaily spend per user or API keyrequest, tool_callask, denyPer principal per day
subagent_cost_budgetSpend of one subagentrequest, tool_callask, denySession
token_budgetSession tokensrequest, tool_callask, denySession

A stateful type keeps counters or history, which not every runtime can enforce. Check Enforcement before you apply one widely. AI Gateway runs only block_tools and the four budget types; see Policy guardrails.

Tool access

require_approval_for_tools

Asks before the named tools run. It is the check an always_ask permission policy makes, set at a level the agent's author cannot relax.

ParameterTypeDefaultDescription
toolsstring[]RequiredTool names or patterns. At least one.
{ "kind": "builtin", "builtin": "require_approval_for_tools", "params": { "tools": ["mcp__github__*"] } }

block_tools

Denies the named tools. The agent receives the reason in place of the tool result.

ParameterTypeDefaultDescription
toolsstring[]RequiredTool names or patterns. At least one.
reasonstringTool <name> is blocked.Message returned to the agent.
{ "kind": "builtin", "builtin": "block_tools", "params": { "tools": ["mcp__orca__bash"], "reason": "Shell access is disabled for this agent." } }

ask_on_os_tools

Asks before any filesystem or shell tool runs, whatever its arguments. It covers the mcp__orca__ tools read, glob, grep, list, write, edit, delete, and bash. It has no parameters.

read_only_os

Denies every tool that can change the filesystem: mcp__orca__write, mcp__orca__edit, mcp__orca__delete, and mcp__orca__bash. Reads and searches stay available. The shell counts as a write whatever the command, because proving a command read-only would require parsing it.

ParameterTypeDefaultDescription
reasonstring<tool> can modify the filesystem; this agent is read-only.Message returned to the agent.

block_skills

Prevents the named skills from loading. Patterns match case-insensitively against the skill's name and against its bare name after the last : or /, so deploy-* also blocks tools:deploy-prod. Agent Engine applies the rule when it copies skills into the sandbox, so a blocked skill is never available to the agent. On a self_hosted environment, the session-runner refuses the rule; see Enforcement.

ParameterTypeDefaultDescription
blockedstring[]RequiredSkill names or patterns, at least one.
{ "kind": "builtin", "builtin": "block_skills", "params": { "blocked": ["deploy-*"] } }

headless_subagent_purpose_guard

Requires each subagent dispatch through the Agent tool to declare a purpose from an allowed list, and denies dispatches that do not.

ParameterTypeDefaultDescription
allowed_purposesstring[]["implement", "review", "explore", "search"]Purposes a dispatch may declare. [] turns the check off.
deny_reasonstringSubagent dispatch must declare a purpose: ...Message returned to the agent.

No Agent Engine runtime adds a purpose field to a subagent dispatch. With the default list, this rule denies every subagent dispatch it evaluates.

Shell and working directory

These three types read the command argument of mcp__orca__bash calls. They split chained and piped commands and look through wrappers such as sudo, nested shells such as bash -c, and command substitutions. They do not expand variables, so a path that contains one is treated as unknown. A command nested too deeply to parse counts as unsafe.

blast_radius

Classifies each command as catastrophic, risky, or safe.

  • Catastrophic commands are always denied: recursive rm, find with -delete, writes to a device, mkfs, fork bombs, and a download piped or substituted into a shell.
  • Risky commands get risky_action: git push when gate_pushes is on, git rebase, git filter-branch, git filter-repo, git commit --amend, git reflog expire, git reset --hard, recursive chmod, chown, or chgrp, and chmod 777 or a+rwx.
ParameterTypeDefaultDescription
gate_pushesbooleantrueTreat every git push as risky.
risky_actionask | denyaskVerdict for a risky command.
deny_reasonstringDescribes the commandMessage returned when a command is denied.

block_working_dir_changes

Blocks commands that change the working directory, which would otherwise let a later command escape a path-based rule: cd, chdir, pushd, popd, a global git -C <dir>, and git worktree add, move, or remove.

ParameterTypeDefaultDescription
block_cdbooleantrueBlock directory changes, including git -C.
block_worktreebooleantrueBlock worktree creation, moves, and removal.
allowed_dirsstring[][]Directories a change may target despite the blocks.
actiondeny | askdenyVerdict for a blocked change.

worktree_guard

Denies shell commands that write outside a root directory, or whose target cannot be resolved. It reads shell commands only: mcp__orca__write and mcp__orca__edit are not checked, so pair it with read_only_os or block_tools if the agent has them.

ParameterTypeDefaultDescription
allowed_rootstring.worktreesRoot that shell writes must stay within.
deny_reasonstringDescribes the writeMessage returned when a write is denied.

Integrations

These types recognize a remote MCP tool by its server's name or by words in the tool's own name, so a tool on an unrelated server can match too. A matching tool with an unfamiliar verb is treated as a write, so a new tool is gated rather than allowed.

github_policy

Restricts repository access through Git-family MCP tools and through git and gh shell commands. A tool counts as Git-family when its server's name starts with git or contains github, gitlab, or gitea, or when its own name contains words such as github, repo, or gist. A call that names no identifiable repository, such as a push to a remote named origin, asks.

ParameterTypeDefaultDescription
read_allbooleantrueAllow reads from any repository. When false, only read_repos can be read.
read_reposstring[][]Repositories readable when read_all is false.
write_reposstring[][]Repositories that may be written to.
write_branchesstring[][]Branch names or patterns that may be written to. When set, a write that names no branch asks.

With the default parameters, github_policy denies every repository write: an empty write_repos means no repository is writable. List the repositories the agent may change.

{ "kind": "builtin", "builtin": "github_policy", "params": { "write_repos": ["acme/docs"], "write_branches": ["docs/*"] } }

gdrive_policy

Restricts Google Drive, Docs, Sheets, and Slides access through servers whose names contain drive, docs, sheet, or slide. It also contains confidential material: after the session reads a file in confidential_files, a write to a file outside that list is denied, or asks with write_down_action: "ask", because it could copy the confidential content somewhere less protected. At tool_result, it records the files the session creates, and the agent may keep writing to them.

ParameterTypeDefaultDescription
read_allbooleantrueAllow reads of any file. When false, only read_files can be read.
read_filesstring[][]Files readable when read_all is false.
allow_createbooleanfalseAllow creating files.
write_filesstring[][]Existing files that may be modified.
comment_filesstring[][]Files that may be commented on without being modified.
confidential_filesstring[][]Files whose contents may not be written elsewhere once read.
write_down_actiondeny | askdenyVerdict for a write after a confidential read.

gmail_policy

Turns mail capabilities on or off for servers whose names contain mail.

ParameterTypeDefaultDescription
allow_readbooleantrueAllow reading and searching.
allow_sendbooleanfalseAllow sending mail.
allow_draftsbooleantrueAllow creating and editing drafts.
allow_modifybooleanfalseAllow modifying existing mail, such as labeling or deleting.

gcalendar_policy

Turns calendar capabilities on or off for servers whose names contain calendar or gcal.

ParameterTypeDefaultDescription
allow_readbooleantrueAllow reading and searching events.
allow_create_eventsbooleanfalseAllow creating events.
allow_modify_eventsbooleanfalseAllow modifying or deleting events.

Personal data

deny_pii_in_llm_request

Denies a user message that appears to contain personal data, before it reaches the model, with a message such as Message appears to contain credit card and was blocked. The patterns match the shape of a value and do not validate it, so a string shaped like a Social Security number matches whether or not it is one.

ParameterTypeDefaultDescription
pii_typesstring[]["ssn", "credit_card", "email", "phone"]Categories to scan for. credit_card matches 13 to 19 digits.

Its default phases include llm_request, which no Agent Engine runtime evaluates. Set phases explicitly: claude_agent_sdk enforces only the request phase and emits a warning, while the Codex SDK and Pi SDK harnesses and self_hosted environments refuse a session with an llm_request rule. A rule already stored with the default phases rejects every update, including {"enabled": false}, until the update also sets phases: ["request"].

{
  "name": "deny-personal-data",
  "phases": ["request"],
  "rule": { "kind": "builtin", "builtin": "deny_pii_in_llm_request", "params": { "pii_types": ["ssn", "credit_card"] } }
}

Session behavior

max_tool_calls_per_session

Counts the session's allowed tool calls and denies once the count reaches limit. Denied calls do not count.

ParameterTypeDefaultDescription
limitinteger100Maximum tool calls in the session.

spawn_bounds

Caps how many subagents one turn may dispatch. The count resets with each user message.

ParameterTypeDefaultDescription
max_dispatches_per_turninteger5Maximum dispatches in one turn.
dispatch_toolsstring[]["Agent"]Tools that count as a dispatch.

detect_loop

Fires when the same tool is called with identical arguments threshold times within the last window calls. When the agent is allowed to continue, the history resets.

ParameterTypeDefaultDescription
thresholdinteger3Identical calls that count as a loop. At least 2.
windowinteger10Recent tool calls to consider.
actionask | denyaskVerdict when a loop is detected.

detect_thrashing

Denies at tool_result after consecutive_threshold tool results in a row look like failures: content with a non-empty error field, or text containing words such as error, failed, or not found. The runtime's own error flag on a result is not considered. How a runtime applies a tool_result deny differs; see Enforcement.

ParameterTypeDefaultDescription
consecutive_thresholdinteger3Consecutive failures that count as thrashing.
windowinteger10Recent tool results to consider.

Cost and usage budgets

The cost budgets use spend that the registry prices from session usage with model pricing. The runtime reports usage after each model call, so a single call can cross a cap before the next request or tool_call check stops the session.

cost_budget

Caps what one session spends, in USD. Set at least one of max_cost_usd and ask_thresholds_usd.

ParameterTypeDefaultDescription
max_cost_usdnumberNoneHard cap, greater than 0. At or above it, the rule denies.
ask_thresholds_usdnumber[][]Soft thresholds. Each asks once, the first time spend crosses it.
expensive_modelsstring[][]Limits the cap to models whose ID contains one of these strings, compared without case.
on_unpricedask | deny | allowaskVerdict when usage cannot be priced.
{ "kind": "builtin", "builtin": "cost_budget", "params": { "max_cost_usd": 10, "ask_thresholds_usd": [2, 5] } }

user_daily_cost_budget

Caps a principal's spend per UTC day across sessions. The principal is the authenticated user, or the API key when there is no user. Spend is counted per Workspace, even when the rule has organization scope. The rule must have workspace or organization scope, and max_cost_usd is required. It takes the same parameters as cost_budget.

subagent_cost_budget

Caps spend attributed to one dispatched subagent within a session. It applies only to actions a subagent takes, and dispatches of the same subagent share one counter. It takes the same parameters as cost_budget.

token_budget

Caps the tokens a session uses. It needs no price data, so it also works for unpriced models.

ParameterTypeDefaultDescription
max_total_tokensintegerRequiredHard cap. At or above it, the rule denies.
ask_thresholdsinteger[][]Soft thresholds in tokens. Each asks once.

How budgets are evaluated

A cost budget makes three checks, in order:

  1. Unpriced usage. If the session used tokens that could not be priced, on_unpriced decides. allow continues without enforcing the budget, and deny denies. ask, the default, asks at the next tool call and runs without budget enforcement once approved. A request has no approval exchange: the first unpriced request is allowed, and a later one without an approval is denied.
  2. The hard cap. At or above max_cost_usd, the rule denies with This session has reached its $<cap> budget (spent $<spent>). Switch to a less expensive model to continue. The daily and subagent budgets start with Today's spend and This subagent. With expensive_models, it denies only while the session uses a matching model, so work can continue on a cheaper one. A model the rule cannot identify counts as matching.
  3. Soft thresholds. Past a threshold, the rule asks once at the next tool call. An approval is recorded; a refusal means the next call asks again. Thresholds never fire at request.

token_budget applies the same cap and threshold steps to tokens.

The Codex SDK and Pi SDK harnesses enforce budgets only at request. They refuse to start a session with ask_thresholds, ask_thresholds_usd, subagent_cost_budget, or a budget that keeps its default tool_call phase, and they treat on_unpriced: "ask" as deny. Set phases: ["request"] on budgets that apply to those agents.

What's next

On this page