Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Use the gateway with Agent Engine

Route model and MCP traffic from Orca Agent Engine through Orca AI Gateway, including the deployment-wide model egress default.

The two Orca products are independent - each runs without the other - but they are designed to compose. When both are deployed, Agent Engine sends its agents' MCP tool calls through the gateway. You can also make the gateway the deployment-wide default for model calls. Credentials, authorization, guardrails, usage metering, and audit then apply at the shared egress point.

What flows through the gateway

MCP tool calls: yes, and this is the production path. At session start, the Agent Engine runtime rewrites every entry in an agent's mcp_servers list to the /v1/mcp endpoint derived from AI_GATEWAY_URL. The rewritten entry keeps the logical server name in X-Orca-Backend, adds X-Orca-Session-Id, a short-lived session JWT, and an optional X-Orca-Credential-Id. The gateway validates the token, resolves the logical name and bound credential through the registry, then forwards the JSON-RPC envelope upstream. The agent never sees the real credential or URL.

Model calls: always in Cloud colocated mode, optional in separate mode. Colocated harnesses use LLM_GATEWAY_URL and a session-scoped JWT. For a separate harness, Agent Engine first checks the session's metadata.orca_llm_egress value, then the deployment's LLM_EGRESS_DEFAULT, and finally falls back to direct. Set the deployment default to gateway to route model calls through the gateway without changing each session. A colocated session always uses the gateway, even if its metadata asks for direct egress. The gateway accepts the Anthropic compatibility path /v1/llm/v1/messages and the native OpenAI Responses paths /v1/llm/responses and /v1/llm/v1/responses. Pi SDK uses provider-native proxy routes.

colocated mode is not fully implemented in this release. Use separate, the default.

Gateway model egress for a separate harness requires a registry connection and LLM_GATEWAY_URL. Configuring LLM_EGRESS_DEFAULT=gateway without the URL fails Helm validation; an individual session requesting gateway egress cannot start without the required runtime settings. A session requesting direct egress bypasses the gateway even when the deployment default is gateway.

Configure the gateway for MCP egress

An Agent Engine deployment needs the gateway configured with three things: a validator that trusts the registry's JWTs, an HTTP vault resolver pointed at the registry's internal resolve endpoint, and a wildcard mcp destination resolved through the registry.

gateway-config.yaml
identity:
  scope_dims: [workspace_id, session_id, mcp_server_names, credential_ids]
  validators:
    - name: orca-registry
      kind: jwt
      issuers: ["orca-registry"]
      audiences: ["ai-gateway"]
      static_public_key_pem: /etc/orca-gateway/registry-pubkey.pem
      scope_from:
        jwt_claims:
          workspace_id: workspace_id
          session_id: session_id
          mcp_server_names: mcp_server_names
          credential_ids: credential_ids

vaults:
  - name: registry-vaults
    resolver: http
    url_template: "http://orca-registry:8081/internal/v1/workspaces/{scope.workspace_id}/sessions/{scope.session_id}/vault-credentials/{credential_id}/resolve"
    bearer_token_file: /var/run/secrets/orca/registry/token
    timeout_ms: 5000

destinations:
  "*":
    kind: mcp
    destination_resolver:
      kind: http
      url_template: "http://orca-registry:8081/internal/v1/workspaces/{scope.workspace_id}/sessions/{scope.session_id}/mcp-destination/resolve"
      bearer_token_file: /var/run/secrets/orca/registry/token
      timeout_ms: 5000
    credentials:
      vault: registry-vaults

plugins:
  authorizers:
    - name: session-mcp-allowlist
      kind: yaml_acl
      required: true
      failure_mode: deny
      rules:
        - effect: allow
          action: mcp
          resource:
            kind: mcp_server
            header: X-Orca-Backend
          conditions:
            - header_in_scope_list:
                header: X-Orca-Backend
                scope: mcp_server_names
            - header_in_scope_list:
                header: X-Orca-Credential-Id
                scope: credential_ids
                if_missing: allow
  audit_sinks:
    - name: audit
      kind: stdout

The authorizer is the enforcement point for the session JWT's MCP grants. The two conditions bind the requested backend and credential headers to the lists the registry signed into the session token. Keep this authorizer in place, and add a rule that denies anonymous principals.

Registry-minted Agent Engine session tokens use aud=ai-gateway. During a migration, add a legacy audience only while tokens minted for that audience remain in circulation.

Add Registry policy, pricing, and usage

The MCP egress configuration above is sufficient to route tools. You can independently add three Agent Engine registry integrations: session policy bundles, the Workspace model-price catalog, and model-usage delivery.

First, map the session claims that the policy source and usage sink consume:

identity:
  scope_dims:
    - org_id
    - workspace_id
    - session_id
    - agent_id
    - guardrail_ids
    - runtime_config_revision
    - mcp_server_names
    - credential_ids
  validators:
    - name: orca-registry
      kind: jwt
      issuers: ["orca-registry"]
      audiences: ["ai-gateway"]
      static_public_key_pem: /etc/orca-gateway/registry-pubkey.pem
      scope_from:
        jwt_claims:
          org_id: org_id
          workspace_id: workspace_id
          session_id: session_id
          agent_id: agent_id
          guardrail_ids: guardrail_ids
          runtime_config_revision: runtime_config_revision
          mcp_server_names: mcp_server_names
          credential_ids: credential_ids

Then add the integrations you use. The pricing source calls the registry's public Workspace API and requires exactly one Workspace credential. The policy source and usage sink call internal routes and use the registry workload token.

plugins:
  cost_model:
    name: registry-prices
    kind: registry
    base_url: http://orca-registry:8080
    api_key_file: /var/run/secrets/orca/workspace-api-key
    refresh_interval_secs: 300
  usage_sinks:
    - name: registry-session-usage
      kind: registry
      base_url: http://orca-registry:8081
      bearer_token_file: /var/run/secrets/orca/registry/token
      spool:
        directory: /var/lib/orca-gateway/registry-usage

guardrail_source:
  name: registry-policy
  kind: registry
  base_url: http://orca-registry:8081
  bearer_token_file: /var/run/secrets/orca/registry/token
  on_backend_error: last_known

guardrail_state:
  name: registry-policy-state
  kind: redis
  url: redis://redis:6379/1

guardrails_policy:
  mode: enforce
  phases: [llm_request, llm_response, tool_call, tool_result]

Replace api_key_file with oidc_token_file when the Workspace uses OIDC. Registry policy bundles load lazily on the first authenticated request for a session. The Registry usage sink sends model usage only; keep Kafka, Postgres, or another sink when you also need MCP usage records. See the configuration reference for failure modes and standalone file-backed policy.

Point Agent Engine at the gateway

The harness server reads the MCP gateway endpoint from AI_GATEWAY_URL. You can supply the host root or the full endpoint; it appends /v1/mcp only when the value does not already end with that path. The deployed chart uses the explicit endpoint:

AI_GATEWAY_URL="http://orca-gateway:8080/v1/mcp"

On Kubernetes, set the equivalent Helm value on the Agent Engine chart:

values.yaml
harness:
  aiGatewayUrl: "http://orca-gateway.orca-system.svc.cluster.local:8080/v1/mcp"

Roll the harness, and tool calls begin flowing through the gateway. Rolling the value back restores the previous path; the gateway holds no state that makes the change one-way.

Route model calls through the gateway by default

Self-hosted deployment

Set the model gateway endpoint and deployment default in the Agent Engine Helm values:

values.yaml
registry:
  aiGatewayLlmUrl: "http://orca-gateway.orca-system.svc.cluster.local:8080/v1/llm"
harness:
  aiGatewayUrl: "http://orca-gateway.orca-system.svc.cluster.local:8080/v1/mcp"
  llmEgressDefault: gateway

registry.aiGatewayLlmUrl configures both the registry and harness model gateway URL. harness.llmEgressDefault: gateway renders LLM_EGRESS_DEFAULT=gateway. Use the Kubernetes Service name and namespace from your deployment. Keep the model compatibility base ending in /v1/llm; the Gateway exposes compatible Anthropic Messages and OpenAI Responses paths below that base. Pi SDK requires its provider-native proxy routes in addition to this endpoint.

Download the Agent Engine v0.5.1 chart as described in Deploy on Kubernetes, then upgrade the release and check the rendered configuration. Replace the release and namespace names with the values from your installation:

Terminal
helm upgrade <release-name> ./orca-managed-agents-0.5.1.tgz \
  --namespace <agent-engine-namespace> \
  --reuse-values \
  --set registry.aiGatewayLlmUrl=http://orca-gateway.orca-system.svc.cluster.local:8080/v1/llm \
  --set harness.aiGatewayUrl=http://orca-gateway.orca-system.svc.cluster.local:8080/v1/mcp \
  --set harness.llmEgressDefault=gateway

kubectl -n <agent-engine-namespace> get configmap \
  -l app.kubernetes.io/instance=<release-name> \
  -o json | jq -r \
  '.items[] | select(.data.LLM_EGRESS_DEFAULT != null) | .data.LLM_EGRESS_DEFAULT, .data.LLM_GATEWAY_URL'

kubectl -n <agent-engine-namespace> get configmap \
  -l app.kubernetes.io/instance=<release-name> \
  -o json | jq -r \
  '.items[] | select(.data.AI_GATEWAY_LLM_URL != null) | .data.AI_GATEWAY_LLM_URL'

The first command prints gateway and the /v1/llm endpoint. The second command prints the same model endpoint.

Override one session

For a separate harness, a session can override the deployment default with metadata.orca_llm_egress. Use direct to bypass the gateway for one session or gateway to route one session through it:

{
  "metadata": {
    "orca_llm_egress": "direct"
  }
}

Omit this metadata to use the deployment default. A Cloud colocated harness always routes model calls through the gateway, regardless of this setting. To verify the deployment-wide default, create a session without orca_llm_egress and confirm its session ID appears in the Gateway usage or audit output. The Guardrail validation demo uses this method.

Migrating from the earlier mcp-gateway service

Deployments that ran the earlier mcp-gateway service have two migration aids. The gateway serves POST /mcp as a deprecated alias for /v1/mcp, and it accepts the old configuration verbatim under a legacy_mcp_gateway: key, expanding it into the modern shape at load time. Both are deprecated. See Migrate for the rewrite.

On this page