Use the gateway with Agent Engine
Route model and MCP traffic from Orca Agent Engine through Orca AI Gateway, including the deployment-wide model egress default.
The two Orca products are independent - each runs without the other - but they are designed to compose. When both are deployed, Agent Engine sends its agents' MCP tool calls through the gateway. You can also make the gateway the deployment-wide default for model calls. Credentials, authorization, guardrails, usage metering, and audit then apply at the shared egress point.
What flows through the gateway
MCP tool calls: yes, and this is the production path. At session start, the Agent Engine runtime
rewrites every entry in an agent's mcp_servers list to the /v1/mcp endpoint derived from
AI_GATEWAY_URL. The rewritten entry keeps the logical server name in X-Orca-Backend, adds
X-Orca-Session-Id, a short-lived session JWT, and an optional X-Orca-Credential-Id. The gateway
validates the token, resolves the logical name and bound credential through the registry, then
forwards the JSON-RPC envelope upstream. The agent never sees the real credential or URL.
Model calls: always in Cloud colocated mode, optional in separate mode. Colocated harnesses
use LLM_GATEWAY_URL and a session-scoped JWT. For a separate harness, Agent Engine first checks
the session's metadata.orca_llm_egress value, then the deployment's LLM_EGRESS_DEFAULT, and
finally falls back to direct. Set the deployment default to gateway to route model calls through
the gateway without changing each session. A colocated session always uses the gateway, even if its
metadata asks for direct egress. The gateway accepts the Anthropic
compatibility path /v1/llm/v1/messages and the native OpenAI Responses paths
/v1/llm/responses and /v1/llm/v1/responses. Pi SDK uses provider-native proxy routes.
colocated mode is not fully implemented in this release. Use separate, the default.
Gateway model egress for a separate harness requires a registry connection and LLM_GATEWAY_URL.
Configuring LLM_EGRESS_DEFAULT=gateway without the URL fails Helm validation; an individual
session requesting gateway egress cannot start without the required runtime settings. A session
requesting direct egress bypasses the gateway even when the deployment default is gateway.
Configure the gateway for MCP egress
An Agent Engine deployment needs the gateway configured with three things: a validator that trusts
the registry's JWTs, an HTTP vault resolver pointed at the registry's internal resolve endpoint, and
a wildcard mcp destination resolved through the registry.
identity:
scope_dims: [workspace_id, session_id, mcp_server_names, credential_ids]
validators:
- name: orca-registry
kind: jwt
issuers: ["orca-registry"]
audiences: ["ai-gateway"]
static_public_key_pem: /etc/orca-gateway/registry-pubkey.pem
scope_from:
jwt_claims:
workspace_id: workspace_id
session_id: session_id
mcp_server_names: mcp_server_names
credential_ids: credential_ids
vaults:
- name: registry-vaults
resolver: http
url_template: "http://orca-registry:8081/internal/v1/workspaces/{scope.workspace_id}/sessions/{scope.session_id}/vault-credentials/{credential_id}/resolve"
bearer_token_file: /var/run/secrets/orca/registry/token
timeout_ms: 5000
destinations:
"*":
kind: mcp
destination_resolver:
kind: http
url_template: "http://orca-registry:8081/internal/v1/workspaces/{scope.workspace_id}/sessions/{scope.session_id}/mcp-destination/resolve"
bearer_token_file: /var/run/secrets/orca/registry/token
timeout_ms: 5000
credentials:
vault: registry-vaults
plugins:
authorizers:
- name: session-mcp-allowlist
kind: yaml_acl
required: true
failure_mode: deny
rules:
- effect: allow
action: mcp
resource:
kind: mcp_server
header: X-Orca-Backend
conditions:
- header_in_scope_list:
header: X-Orca-Backend
scope: mcp_server_names
- header_in_scope_list:
header: X-Orca-Credential-Id
scope: credential_ids
if_missing: allow
audit_sinks:
- name: audit
kind: stdoutThe authorizer is the enforcement point for the session JWT's MCP grants. The two conditions bind the requested backend and credential headers to the lists the registry signed into the session token. Keep this authorizer in place, and add a rule that denies anonymous principals.
Registry-minted Agent Engine session tokens use aud=ai-gateway. During a migration, add a legacy
audience only while tokens minted for that audience remain in circulation.
Add Registry policy, pricing, and usage
The MCP egress configuration above is sufficient to route tools. You can independently add three Agent Engine registry integrations: session policy bundles, the Workspace model-price catalog, and model-usage delivery.
First, map the session claims that the policy source and usage sink consume:
identity:
scope_dims:
- org_id
- workspace_id
- session_id
- agent_id
- guardrail_ids
- runtime_config_revision
- mcp_server_names
- credential_ids
validators:
- name: orca-registry
kind: jwt
issuers: ["orca-registry"]
audiences: ["ai-gateway"]
static_public_key_pem: /etc/orca-gateway/registry-pubkey.pem
scope_from:
jwt_claims:
org_id: org_id
workspace_id: workspace_id
session_id: session_id
agent_id: agent_id
guardrail_ids: guardrail_ids
runtime_config_revision: runtime_config_revision
mcp_server_names: mcp_server_names
credential_ids: credential_idsThen add the integrations you use. The pricing source calls the registry's public Workspace API and requires exactly one Workspace credential. The policy source and usage sink call internal routes and use the registry workload token.
plugins:
cost_model:
name: registry-prices
kind: registry
base_url: http://orca-registry:8080
api_key_file: /var/run/secrets/orca/workspace-api-key
refresh_interval_secs: 300
usage_sinks:
- name: registry-session-usage
kind: registry
base_url: http://orca-registry:8081
bearer_token_file: /var/run/secrets/orca/registry/token
spool:
directory: /var/lib/orca-gateway/registry-usage
guardrail_source:
name: registry-policy
kind: registry
base_url: http://orca-registry:8081
bearer_token_file: /var/run/secrets/orca/registry/token
on_backend_error: last_known
guardrail_state:
name: registry-policy-state
kind: redis
url: redis://redis:6379/1
guardrails_policy:
mode: enforce
phases: [llm_request, llm_response, tool_call, tool_result]Replace api_key_file with oidc_token_file when the Workspace uses OIDC. Registry policy bundles
load lazily on the first authenticated request for a session. The Registry usage sink sends model
usage only; keep Kafka, Postgres, or another sink when you also need MCP usage records. See the
configuration reference for failure modes
and standalone file-backed policy.
Point Agent Engine at the gateway
The harness server reads the MCP gateway endpoint from AI_GATEWAY_URL. You can supply the host
root or the full endpoint; it appends /v1/mcp only when the value does not already end with that
path. The deployed chart uses the explicit endpoint:
AI_GATEWAY_URL="http://orca-gateway:8080/v1/mcp"On Kubernetes, set the equivalent Helm value on the Agent Engine chart:
harness:
aiGatewayUrl: "http://orca-gateway.orca-system.svc.cluster.local:8080/v1/mcp"Roll the harness, and tool calls begin flowing through the gateway. Rolling the value back restores the previous path; the gateway holds no state that makes the change one-way.
Route model calls through the gateway by default
Self-hosted deployment
Set the model gateway endpoint and deployment default in the Agent Engine Helm values:
registry:
aiGatewayLlmUrl: "http://orca-gateway.orca-system.svc.cluster.local:8080/v1/llm"
harness:
aiGatewayUrl: "http://orca-gateway.orca-system.svc.cluster.local:8080/v1/mcp"
llmEgressDefault: gatewayregistry.aiGatewayLlmUrl configures both the registry and harness model gateway URL.
harness.llmEgressDefault: gateway renders LLM_EGRESS_DEFAULT=gateway. Use the Kubernetes
Service name and namespace from your deployment. Keep the model compatibility base ending in
/v1/llm; the Gateway exposes compatible Anthropic Messages and OpenAI Responses paths below that
base. Pi SDK requires its provider-native proxy routes in addition to this endpoint.
Download the Agent Engine v0.5.1 chart as described in Deploy on Kubernetes, then upgrade the release and check the rendered configuration. Replace the release and namespace names with the values from your installation:
helm upgrade <release-name> ./orca-managed-agents-0.5.1.tgz \
--namespace <agent-engine-namespace> \
--reuse-values \
--set registry.aiGatewayLlmUrl=http://orca-gateway.orca-system.svc.cluster.local:8080/v1/llm \
--set harness.aiGatewayUrl=http://orca-gateway.orca-system.svc.cluster.local:8080/v1/mcp \
--set harness.llmEgressDefault=gateway
kubectl -n <agent-engine-namespace> get configmap \
-l app.kubernetes.io/instance=<release-name> \
-o json | jq -r \
'.items[] | select(.data.LLM_EGRESS_DEFAULT != null) | .data.LLM_EGRESS_DEFAULT, .data.LLM_GATEWAY_URL'
kubectl -n <agent-engine-namespace> get configmap \
-l app.kubernetes.io/instance=<release-name> \
-o json | jq -r \
'.items[] | select(.data.AI_GATEWAY_LLM_URL != null) | .data.AI_GATEWAY_LLM_URL'The first command prints gateway and the /v1/llm endpoint. The second command prints the same
model endpoint.
Override one session
For a separate harness, a session can override the deployment default with
metadata.orca_llm_egress. Use direct to bypass the gateway for one session or gateway to route
one session through it:
{
"metadata": {
"orca_llm_egress": "direct"
}
}Omit this metadata to use the deployment default. A Cloud colocated harness always routes model
calls through the gateway, regardless of this setting. To verify the deployment-wide default,
create a session without orca_llm_egress and confirm its session ID appears in the Gateway usage
or audit output. The Guardrail validation demo
uses this method.
Migrating from the earlier mcp-gateway service
Deployments that ran the earlier mcp-gateway service have two migration aids. The gateway serves
POST /mcp as a deprecated alias for /v1/mcp, and it accepts the old configuration verbatim under
a legacy_mcp_gateway: key, expanding it into the modern shape at load time. Both are deprecated.
See Migrate for the rewrite.