Payload guardrails
Inspect, redact, and block request and response content in Orca AI Gateway.
A payload guardrail inspects a request or response body and returns one of three decisions:
allow it, sanitize it and forward the rewritten body, or block it. Payload guardrails are configured
under plugins.guardrails[] and run on every data-plane route, including MCP. They are separate from
policy guardrails, which apply Agent Engine rules to model
choice, tool calls, and spend.
Where payload guardrails run
| Phase | Runs | Kinds |
|---|---|---|
input | On the request body, after authentication, authorization, route selection, and rate-limit admission, and before the request is dispatched | pii_regex, llm_evaluator, ext_proc |
output | On a buffered response body, after the provider or MCP server responds | llm_evaluator, ext_proc |
Only pii_regex and ext_proc rewrite a body; an llm_evaluator sanitize verdict does not. Because
the route is chosen first, a rewrite cannot change where a request goes.
For pii_regex and llm_evaluator, the gateway reads every string under messages[].content,
system, input, prompt, and params, which holds an MCP call's tool name and arguments. If
those fields hold no text, it reads every string in the body. A body whose Content-Type is not
JSON has no text for them to read: pii_regex passes it, while llm_evaluator still calls the
evaluator, as it does for requests with no body such as GET /v1/models. ext_proc receives the raw
body. A sanitize rewrites each match in every string of the JSON body, not only in the fields it
read.
Output guardrails inspect buffered response bodies. Send a request without streaming when its
response must pass an output guardrail, and set each guardrail's phase to input or output.
While any payload guardrail is configured, /v1/responses, /v1/llm/responses,
/v1/llm/v1/responses, and /v1/proxy/* reject every request with 400 invalid_request, because
their bodies cannot be inspected safely. Route OpenAI Responses and native provider traffic through
a gateway without payload guardrails if you need both.
Decisions and failure modes
A block decision stops the request. With the default failure_mode: deny, the gateway returns
400 with the guardrail's reason and code:
{
"error": {
"code": "guardrail_blocked_input",
"message": "input guardrail blocked: PII detected: email",
"reason": "PII detected: email",
"guardrail_code": "pii_blocked"
}
}An output block uses guardrail_blocked_output. guardrail_code is pii_blocked,
evaluator_blocked, the code an ext_proc sidecar returns, or guardrail_backend_error when the
guardrail's backend failed. An llm_evaluator reply that the gateway cannot parse returns
400 invalid_request.
failure_mode decides what a block or an error means for each guardrail:
failure_mode | Block | Error |
|---|---|---|
deny (default) | The request fails with 400. | The request fails with 400. |
allow | The request continues. | The request continues. |
log_only | Same as allow. | Same as allow. |
allow and log_only behave the same way; they differ only in the failure_mode attribute on
the audit event. A sequential guardrail logs a warning when it lets a block or an error through; a
parallel one does not. Use deny for compliance work, because the other modes let a block through.
Evaluation order
Guardrails in a phase run one at a time, in the order you declare them, so each one sees the
previous sanitizer's rewrite. After them, every llm_evaluator with parallel: true runs
concurrently against a copy of that rewritten body. The parallel group does not run alongside the
upstream call. Only llm_evaluator can run in parallel, and it never rewrites the body, so parallel
evaluation cannot change the payload.
pii_regex
A regex detector and redactor for input bodies. It has no phase setting.
plugins:
guardrails:
- name: pii
kind: pii_regex
action: sanitize
patterns:
- label: customer_id
regex: 'CUST-\d{8}'
extend_default: true| Field | Default | Notes |
|---|---|---|
action | sanitize | sanitize redacts matches, block refuses the request with code pii_blocked, and log allows the request and logs the label of each match, not the matched text. |
patterns | [] | Custom { label, regex } entries. When omitted, the built-in set is used. An invalid regex fails at boot. |
extend_default | false | Appends custom patterns to the built-in set instead of replacing it. |
The built-in set is email, phone, ssn, credit_card, and iban. credit_card matches 13 to 19
digits without a checksum. Redaction replaces each match in place: contact me at alice@x.com
becomes contact me at [REDACTED:email].
Treat audit sinks that receive pii_regex sanitize events as holding personal data, and restrict
access to them. action: block records only the matched category.
llm_evaluator
Sends the payload to a model with a rubric you supply, for judgments a regex cannot express, such as jailbreak detection or topic restriction.
plugins:
guardrails:
- name: jailbreak
kind: llm_evaluator
phase: input
evaluator_route: prompt-evaluator
evaluator_url: "http://evaluator:8080/v1/chat/completions"
parallel: false
template: |
Classify this payload. Reply with JSON:
{"verdict":"allow|sanitize|block","reason":"..."}
PAYLOAD: {payload}| Field | Default | Notes |
|---|---|---|
phase | input | input or output. |
evaluator_url | - | Required. An absolute OpenAI chat-completions endpoint. |
evaluator_route | - | Required. A label that appears only in debug logs. |
template | - | Required. {payload} is replaced with the extracted text, joined by newlines. |
action | sanitize | log turns a sanitize verdict into an allow. It does not change a block verdict. |
parallel | false | Runs with other parallel guardrails; see Evaluation order. |
The gateway posts {"model": "evaluator", "messages": [{"role": "user", "content": <prompt>}]} with no
authorization header, and waits up to five seconds. The first choice's message content must be a
JSON object, or a string holding one, of the form
{ "verdict": "allow" | "sanitize" | "block", "reason": "..." }. A block verdict blocks with code
evaluator_blocked whatever action says, subject to failure_mode. A sanitize verdict records the
reason in the audit patch but does not change the payload.
Each evaluated request makes an HTTP call to the evaluator, with a fixed five-second timeout. Keep the evaluator on a private network that only the gateway can reach.
ext_proc
For heavier checks, such as an in-house classifier, run a sidecar and reference it over gRPC:
plugins:
guardrails:
- name: ml-classifier
kind: ext_proc
endpoint: "http://ml-sidecar:9000"
timeout_ms: 2000
phase: inputendpoint is required. timeout_ms defaults to 5000 and applies to both connecting and each
request; phase defaults to input and also accepts output. The gateway connects lazily, so it
starts even when the sidecar is down. The sidecar's reply maps to a decision: allow passes,
sanitize replaces the body with rewritten_body when that is not empty, block uses the reply's
reason and code, and a non-empty error counts as a guardrail error. The same endpoint and
timeout_ms fields are used when mounting ext_proc as an authorizer or audit sink.
The gateway also defines a WebAssembly plugin interface for guardrails, but the shipped binary
does not start the WASM host, so WASM guardrails cannot be installed. Use the built-in kinds or
ext_proc.
Only pii_regex, llm_evaluator, and ext_proc exist. Any other kind fails at boot. Unknown
fields next to kind are ignored, so check spelling against the tables above.
Audit and metrics
With an audit sink configured, each decision emits an audit
event whose action is policy.guardrail.input.<outcome> or policy.guardrail.output.<outcome>,
where the outcome is allow, sanitize, block, or error. The event carries the guardrail's name
and failure_mode, plus the patch, the reason and code, or the error. Audit events do not
include the request or response body. The only payload guardrail metric is the stage duration,
orca_gw_pipeline_stage_duration_seconds with stage="guardrails_input" or
stage="guardrails_output".