Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Guardrail enforcement

Which runtimes in Orca Agent Engine enforce each guardrail phase, what a session reports when a rule fires, and how AI Gateway applies the same rules.

The Agent Engine registry stores guardrails and composes the rules that apply to a session. The runtime that runs the session enforces them, and runtimes differ in which phases they can evaluate and whether they can keep state. This page covers harnesses in separate mode.

colocated mode - the claude_code harness, and Codex SDK or Pi SDK agents with mode: colocated - does not fully support guardrails yet.

Enforcement by runtime

The runtime depends on the agent's harness and mode and on the target of the session's environment:

Runtimerequesttool_calltool_resultStateful rulesSubagent rules
claude_agent_sdk on a cloud environmentYesYes, including askRecorded output onlyYesYes
codex_sdk or pi_sdk in separate modeYesStateless rules, including askStateless rulesrequest onlyRefused
claude_agent_sdk on a self_hosted environmentStateless rulesRefusedRefusedRefusedRefused
  • claude_agent_sdk evaluates every phase and pauses for a user.tool_confirmation event when a rule asks. A tool_result deny replaces the output in the recorded agent.tool_result event and marks it is_error; the model's own conversation keeps the original output.
  • codex_sdk and pi_sdk run in separate mode on cloud environments and refuse a session whose rules they cannot enforce. A tool_result deny replaces the output the model receives. Budget rules have extra limits; see Cost and usage budgets.
  • On a self_hosted environment, a session runs in a session-runner, which enforces stateless request rules and refuses every other rule, including block_skills.

On cloud environments, Agent Engine applies block_skills when it copies skills into the sandbox, so a blocked skill is never staged.

When a runtime cannot enforce a rule

A rule that a runtime cannot fully enforce has one of four outcomes:

  • The session does not start. The session emits session.error and goes idle. The message names the rule and the reason, for example invalid runtime binding: guardrail codex_sdk/separate cannot enforce guardrail: tool-call-cap (<guardrail-id>): stateful phase tool_call is not supported.
  • The turn does not run, and nothing reports it. The Codex SDK and Pi SDK harnesses refuse soft budget thresholds and subagent_cost_budget when they start, and a session-runner on a self_hosted environment refuses unsupported rules when a turn begins. In both cases no session event reports the refusal; only the service logs record it.
  • The session runs with a warning. claude_agent_sdk emits session.warning with warning.type guardrail_not_enforced, the guardrail's ID and name, and a message when a rule declares a phase it does not evaluate, such as the llm_request phase of deny_pii_in_llm_request. It enforces the rule's other phases.
  • The rule is not enforced, and nothing reports it. This happens for an explicit rule that no agent or session lists, a disabled rule, and a reference to an archived or deleted rule.

A workspace or organization rule reaches every session in its scope, including sessions whose runtime refuses it. One Workspace-wide stateful tool_call rule, such as max_tool_calls_per_session or a budget left on its default phases, stops every Codex SDK and Pi SDK session in that Workspace from starting, and any tool_call rule stops every turn on a self_hosted environment. Test a rule at explicit scope, and narrow its phases to what the runtimes in its scope enforce before you widen it.

What a session shows when a guardrail fires

OutcomeEvents
A request is deniedsession.error, then session.status_idle with stop_reason.type end_turn. The message is the rule's reason, or a default: The request was denied by a managed-agent guardrail. on claude_agent_sdk and Request denied by policy. on Codex SDK and Pi SDK. On a self_hosted environment, the session-runner records agent.error with harness turn failed: <reason> instead, where the default reason is This message was blocked by a guardrail., and only orca-beta clients see it.
A tool call is deniedThe agent.tool_use or agent.mcp_tool_use event, then its result event with is_error: true. The content carries the rule's reason, which the agent also receives.
A tool call asksThe tool use event, then session.status_idle with stop_reason.type requires_action. Answer with a user.tool_confirmation event, as for a permission policy. The event does not include the guardrail's reason.
A tool result is deniedThe result event carries is_error: true and Tool output suppressed by policy, followed by the reason on claude_agent_sdk.
Usage was not acknowledgedsession.error, then session.status_idle with stop_reason.type retries_exhausted. Stateful rules stay locked until the registry records the session's usage.

By default, the registry reports guardrail errors in the Claude Managed Agents error format, with error.type unknown_error and the reason in error.message:

{
  "id": "evt_01H8...",
  "type": "session.error",
  "processed_at": "2026-09-29T17:04:11.512Z",
  "error": {
    "type": "unknown_error",
    "message": "Message appears to contain email and was blocked.",
    "retry_status": { "type": "exhausted" }
  }
}

List or stream events with the orca-beta header to see the Orca error types instead: policy_denied with a reasons array for a denied request, setup_failed for a refused session, and guardrail_usage_unavailable for unacknowledged usage. session.warning events appear only with orca-beta. See Events and streaming.

Enforcement in AI Gateway

When a deployment routes a session's traffic through AI Gateway with registry policy turned on, the Gateway loads the same composed rules and evaluates them again on the traffic it proxies: before a model request, after a model response, and before and after an MCP tool call. It runs five of the builtins, block_tools and the four budget types, and treats budget request phases as model requests. It does not run expression rules or the other builtins.

With the Gateway's default on_unsupported_guardrail: deny, a rule the Gateway cannot run denies every event at its phase when that phase is tool_call or llm_request. Any other builtin or expression rule with a tool_call phase makes the Gateway deny every MCP tool call, and deny_pii_in_llm_request with its default llm_request phase makes it deny every model request. Unsupported tool_result rules pass, and rules with only a request phase never reach the Gateway, except budgets. skip_with_alert skips unsupported rules without a log entry and leaves them to the harness.

The Gateway resolves an ask as a deny, because it has no approval exchange.

Audit

The registry writes organization guardrail changes to its admin audit log, which no API reads back. Workspace guardrail changes are not audited, and Agent Engine keeps no separate log of verdicts: a session's transcript is the record of what its guardrails decided. AI Gateway records its own decisions in audit and usage records.

What's next

On this page