Builtin guardrails
The catalog of builtin guardrail types in Orca Agent Engine - what each type checks, its parameters, and the phases and verdicts it supports.
A builtin guardrail is a rule type from the Agent Engine registry's catalog. You select it by name
in rule.builtin and configure it with rule.params. The catalog covers tool access, shell
commands, integrations, personal data, session behavior, and spend. For a check no builtin covers,
write a custom guardrail.
This request body creates a guardrail that keeps any agent that lists it read-only:
{
"name": "read-only-agent",
"scope": "explicit",
"rule": { "kind": "builtin", "builtin": "read_only_os" }
}Guardrails shows how to create and attach a rule. These rules apply to every builtin:
- Parameters are checked when you write the rule. An unknown parameter, a wrong type, or a value
out of range returns
400. The registry storesparamsas written and applies the defaults below when the rule runs. phasesdefaults to the type's phases. You can narrow it to a subset but not add a phase.- Only
tool_callcan ask. The session pauses for auser.tool_confirmationevent. A type that asks at another phase denies instead, except the budget thresholds, which wait for the next tool call. - Tool names are matched exactly and case-sensitively. A
*matches any run of characters, andmcp__<server>__*matches every tool on one MCP server. Inseparatemode, the built-in tools are namedmcp__orca__<tool>, such asmcp__orca__bash.
Catalog
ork guardrails list-types returns the same 22 types with their parameter schemas.
| Builtin | Checks | Phases | Verdicts | State |
|---|---|---|---|---|
require_approval_for_tools | Named tools need approval | tool_call | ask | - |
block_tools | Named tools are denied | tool_call | deny | - |
ask_on_os_tools | Filesystem and shell tools need approval | tool_call | ask | - |
read_only_os | Filesystem writes and shell are denied | tool_call | deny | - |
block_skills | Named skills cannot load | tool_call | deny | - |
headless_subagent_purpose_guard | Subagent dispatches declare a purpose | tool_call | deny | - |
blast_radius | Destructive and risky shell commands | tool_call | ask, deny | - |
block_working_dir_changes | Shell commands that change directory | tool_call | ask, deny | - |
worktree_guard | Shell writes outside a root | tool_call | deny | - |
github_policy | Repository reads and writes | tool_call | ask, deny | - |
gdrive_policy | Drive access and confidential files | tool_call, tool_result | ask, deny | Session |
gmail_policy | Mail capabilities | tool_call | deny | - |
gcalendar_policy | Calendar capabilities | tool_call | deny | - |
deny_pii_in_llm_request | Personal data in user messages | request, llm_request | deny | - |
max_tool_calls_per_session | Tool calls per session | tool_call | deny | Session |
spawn_bounds | Subagent dispatches per turn | tool_call | deny | Turn |
detect_loop | Repeated identical tool calls | tool_call | ask, deny | Session |
detect_thrashing | Consecutive failing tool results | tool_result | deny | Session |
cost_budget | Session spend | request, tool_call | ask, deny | Session |
user_daily_cost_budget | Daily spend per user or API key | request, tool_call | ask, deny | Per principal per day |
subagent_cost_budget | Spend of one subagent | request, tool_call | ask, deny | Session |
token_budget | Session tokens | request, tool_call | ask, deny | Session |
A stateful type keeps counters or history, which not every runtime can enforce. Check
Enforcement before you apply one widely. AI Gateway runs only
block_tools and the four budget types; see
Policy guardrails.
Tool access
require_approval_for_tools
Asks before the named tools run. It is the check an
always_ask permission policy makes, set at a level the
agent's author cannot relax.
| Parameter | Type | Default | Description |
|---|---|---|---|
tools | string[] | Required | Tool names or patterns. At least one. |
{ "kind": "builtin", "builtin": "require_approval_for_tools", "params": { "tools": ["mcp__github__*"] } }block_tools
Denies the named tools. The agent receives the reason in place of the tool result.
| Parameter | Type | Default | Description |
|---|---|---|---|
tools | string[] | Required | Tool names or patterns. At least one. |
reason | string | Tool <name> is blocked. | Message returned to the agent. |
{ "kind": "builtin", "builtin": "block_tools", "params": { "tools": ["mcp__orca__bash"], "reason": "Shell access is disabled for this agent." } }ask_on_os_tools
Asks before any filesystem or shell tool runs, whatever its arguments. It covers the mcp__orca__
tools read, glob, grep, list, write, edit, delete, and bash. It has no parameters.
read_only_os
Denies every tool that can change the filesystem: mcp__orca__write, mcp__orca__edit,
mcp__orca__delete, and mcp__orca__bash. Reads and searches stay available. The shell counts as a write whatever the command, because proving a
command read-only would require parsing it.
| Parameter | Type | Default | Description |
|---|---|---|---|
reason | string | <tool> can modify the filesystem; this agent is read-only. | Message returned to the agent. |
block_skills
Prevents the named skills from loading. Patterns match
case-insensitively against the skill's name and against its bare name after the last : or /, so
deploy-* also blocks tools:deploy-prod. Agent Engine applies the rule when it copies skills into
the sandbox, so a blocked skill is never available to the agent. On a self_hosted environment, the
session-runner refuses the rule; see Enforcement.
| Parameter | Type | Default | Description |
|---|---|---|---|
blocked | string[] | Required | Skill names or patterns, at least one. |
{ "kind": "builtin", "builtin": "block_skills", "params": { "blocked": ["deploy-*"] } }headless_subagent_purpose_guard
Requires each subagent dispatch through the Agent tool to declare a purpose from an allowed list,
and denies dispatches that do not.
| Parameter | Type | Default | Description |
|---|---|---|---|
allowed_purposes | string[] | ["implement", "review", "explore", "search"] | Purposes a dispatch may declare. [] turns the check off. |
deny_reason | string | Subagent dispatch must declare a purpose: ... | Message returned to the agent. |
No Agent Engine runtime adds a purpose field to a subagent dispatch. With the default list, this
rule denies every subagent dispatch it evaluates.
Shell and working directory
These three types read the command argument of mcp__orca__bash calls. They split
chained and piped commands and look through wrappers such as sudo, nested shells such as
bash -c, and command substitutions. They do not expand variables, so a path that contains one is
treated as unknown. A command nested too deeply to parse counts as unsafe.
blast_radius
Classifies each command as catastrophic, risky, or safe.
- Catastrophic commands are always denied: recursive
rm,findwith-delete, writes to a device,mkfs, fork bombs, and a download piped or substituted into a shell. - Risky commands get
risky_action:git pushwhengate_pushesis on,git rebase,git filter-branch,git filter-repo,git commit --amend,git reflog expire,git reset --hard, recursivechmod,chown, orchgrp, andchmod 777ora+rwx.
| Parameter | Type | Default | Description |
|---|---|---|---|
gate_pushes | boolean | true | Treat every git push as risky. |
risky_action | ask | deny | ask | Verdict for a risky command. |
deny_reason | string | Describes the command | Message returned when a command is denied. |
block_working_dir_changes
Blocks commands that change the working directory, which would otherwise let a later command escape
a path-based rule: cd, chdir, pushd, popd, a global git -C <dir>, and
git worktree add, move, or remove.
| Parameter | Type | Default | Description |
|---|---|---|---|
block_cd | boolean | true | Block directory changes, including git -C. |
block_worktree | boolean | true | Block worktree creation, moves, and removal. |
allowed_dirs | string[] | [] | Directories a change may target despite the blocks. |
action | deny | ask | deny | Verdict for a blocked change. |
worktree_guard
Denies shell commands that write outside a root directory, or whose target cannot be resolved. It
reads shell commands only: mcp__orca__write and mcp__orca__edit are not checked, so pair it with
read_only_os or block_tools if the agent has them.
| Parameter | Type | Default | Description |
|---|---|---|---|
allowed_root | string | .worktrees | Root that shell writes must stay within. |
deny_reason | string | Describes the write | Message returned when a write is denied. |
Integrations
These types recognize a remote MCP tool by its server's name or by words in the tool's own name, so a tool on an unrelated server can match too. A matching tool with an unfamiliar verb is treated as a write, so a new tool is gated rather than allowed.
github_policy
Restricts repository access through Git-family MCP tools and through git and gh shell commands.
A tool counts as Git-family when its server's name starts with git or contains github, gitlab,
or gitea, or when its own name contains words such as github, repo, or gist. A call that names no
identifiable repository, such as a push to a remote named origin, asks.
| Parameter | Type | Default | Description |
|---|---|---|---|
read_all | boolean | true | Allow reads from any repository. When false, only read_repos can be read. |
read_repos | string[] | [] | Repositories readable when read_all is false. |
write_repos | string[] | [] | Repositories that may be written to. |
write_branches | string[] | [] | Branch names or patterns that may be written to. When set, a write that names no branch asks. |
With the default parameters, github_policy denies every repository write: an empty write_repos
means no repository is writable. List the repositories the agent may change.
{ "kind": "builtin", "builtin": "github_policy", "params": { "write_repos": ["acme/docs"], "write_branches": ["docs/*"] } }gdrive_policy
Restricts Google Drive, Docs, Sheets, and Slides access through servers whose names contain drive,
docs, sheet, or slide. It also contains confidential material: after the session reads a file
in confidential_files, a write to a file outside that list is denied, or asks with
write_down_action: "ask", because it could copy the confidential content somewhere less
protected. At tool_result, it records the files the session creates, and the agent may keep
writing to them.
| Parameter | Type | Default | Description |
|---|---|---|---|
read_all | boolean | true | Allow reads of any file. When false, only read_files can be read. |
read_files | string[] | [] | Files readable when read_all is false. |
allow_create | boolean | false | Allow creating files. |
write_files | string[] | [] | Existing files that may be modified. |
comment_files | string[] | [] | Files that may be commented on without being modified. |
confidential_files | string[] | [] | Files whose contents may not be written elsewhere once read. |
write_down_action | deny | ask | deny | Verdict for a write after a confidential read. |
gmail_policy
Turns mail capabilities on or off for servers whose names contain mail.
| Parameter | Type | Default | Description |
|---|---|---|---|
allow_read | boolean | true | Allow reading and searching. |
allow_send | boolean | false | Allow sending mail. |
allow_drafts | boolean | true | Allow creating and editing drafts. |
allow_modify | boolean | false | Allow modifying existing mail, such as labeling or deleting. |
gcalendar_policy
Turns calendar capabilities on or off for servers whose names contain calendar or gcal.
| Parameter | Type | Default | Description |
|---|---|---|---|
allow_read | boolean | true | Allow reading and searching events. |
allow_create_events | boolean | false | Allow creating events. |
allow_modify_events | boolean | false | Allow modifying or deleting events. |
Personal data
deny_pii_in_llm_request
Denies a user message that appears to contain personal data, before it reaches the model, with a
message such as Message appears to contain credit card and was blocked. The patterns match the shape of a value and do
not validate it, so a string shaped like a Social Security number matches whether or not it is one.
| Parameter | Type | Default | Description |
|---|---|---|---|
pii_types | string[] | ["ssn", "credit_card", "email", "phone"] | Categories to scan for. credit_card matches 13 to 19 digits. |
Its default phases include llm_request, which no Agent Engine runtime evaluates. Set phases
explicitly: claude_agent_sdk enforces only the request phase and emits a warning, while the Codex
SDK and Pi SDK harnesses and self_hosted environments refuse a session with an llm_request rule.
A rule already stored with the
default phases rejects every update, including {"enabled": false}, until the update also sets
phases: ["request"].
{
"name": "deny-personal-data",
"phases": ["request"],
"rule": { "kind": "builtin", "builtin": "deny_pii_in_llm_request", "params": { "pii_types": ["ssn", "credit_card"] } }
}Session behavior
max_tool_calls_per_session
Counts the session's allowed tool calls and denies once the count reaches limit. Denied calls do
not count.
| Parameter | Type | Default | Description |
|---|---|---|---|
limit | integer | 100 | Maximum tool calls in the session. |
spawn_bounds
Caps how many subagents one turn may dispatch. The count resets with each user message.
| Parameter | Type | Default | Description |
|---|---|---|---|
max_dispatches_per_turn | integer | 5 | Maximum dispatches in one turn. |
dispatch_tools | string[] | ["Agent"] | Tools that count as a dispatch. |
detect_loop
Fires when the same tool is called with identical arguments threshold times within the last
window calls. When the agent is allowed to continue, the history resets.
| Parameter | Type | Default | Description |
|---|---|---|---|
threshold | integer | 3 | Identical calls that count as a loop. At least 2. |
window | integer | 10 | Recent tool calls to consider. |
action | ask | deny | ask | Verdict when a loop is detected. |
detect_thrashing
Denies at tool_result after consecutive_threshold tool results in a row look like failures:
content with a non-empty error field, or text containing words such as error, failed, or
not found. The runtime's own error flag on a result is not considered. How a runtime applies a
tool_result deny differs; see
Enforcement.
| Parameter | Type | Default | Description |
|---|---|---|---|
consecutive_threshold | integer | 3 | Consecutive failures that count as thrashing. |
window | integer | 10 | Recent tool results to consider. |
Cost and usage budgets
The cost budgets use spend that the registry prices from session usage with
model pricing. The runtime reports usage after each model call,
so a single call can cross a cap before the next request or tool_call check stops the session.
cost_budget
Caps what one session spends, in USD. Set at least one of max_cost_usd and ask_thresholds_usd.
| Parameter | Type | Default | Description |
|---|---|---|---|
max_cost_usd | number | None | Hard cap, greater than 0. At or above it, the rule denies. |
ask_thresholds_usd | number[] | [] | Soft thresholds. Each asks once, the first time spend crosses it. |
expensive_models | string[] | [] | Limits the cap to models whose ID contains one of these strings, compared without case. |
on_unpriced | ask | deny | allow | ask | Verdict when usage cannot be priced. |
{ "kind": "builtin", "builtin": "cost_budget", "params": { "max_cost_usd": 10, "ask_thresholds_usd": [2, 5] } }user_daily_cost_budget
Caps a principal's spend per UTC day across sessions. The principal is the authenticated user, or
the API key when there is no user. Spend is counted per Workspace, even when the rule has
organization scope. The rule must have workspace or organization scope, and max_cost_usd is
required. It takes the same parameters as cost_budget.
subagent_cost_budget
Caps spend attributed to one dispatched subagent within a session. It applies only to actions a
subagent takes, and dispatches of the same subagent share one counter. It takes the same parameters
as cost_budget.
token_budget
Caps the tokens a session uses. It needs no price data, so it also works for unpriced models.
| Parameter | Type | Default | Description |
|---|---|---|---|
max_total_tokens | integer | Required | Hard cap. At or above it, the rule denies. |
ask_thresholds | integer[] | [] | Soft thresholds in tokens. Each asks once. |
How budgets are evaluated
A cost budget makes three checks, in order:
- Unpriced usage. If the session used tokens that could not be priced,
on_unpriceddecides.allowcontinues without enforcing the budget, anddenydenies.ask, the default, asks at the next tool call and runs without budget enforcement once approved. Arequesthas no approval exchange: the first unpriced request is allowed, and a later one without an approval is denied. - The hard cap. At or above
max_cost_usd, the rule denies withThis session has reached its $<cap> budget (spent $<spent>). Switch to a less expensive model to continue.The daily and subagent budgets start withToday's spendandThis subagent. Withexpensive_models, it denies only while the session uses a matching model, so work can continue on a cheaper one. A model the rule cannot identify counts as matching. - Soft thresholds. Past a threshold, the rule asks once at the next tool call. An approval is
recorded; a refusal means the next call asks again. Thresholds never fire at
request.
token_budget applies the same cap and threshold steps to tokens.
The Codex SDK and Pi SDK harnesses enforce budgets only at request. They refuse to start a
session with ask_thresholds, ask_thresholds_usd, subagent_cost_budget, or a budget that keeps
its default tool_call phase, and they treat on_unpriced: "ask" as deny. Set
phases: ["request"] on budgets that apply to those agents.
What's next
Guardrails
Create, attach, update, and retire guardrails in Orca Agent Engine at session, agent, Workspace, or organization scope.
Custom guardrails
Write guardrail rules as CEL expressions in Orca Agent Engine - the rule shape, the event fields an expression can read, write-time validation, and examples.