Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Routes

Match requests to destinations in Orca AI Gateway, and control fallback, load balancing, retries, timeouts, and request shaping.

A route binds a set of requests to a strategy for picking a destination. The gateway evaluates strategy groups in this order: conditional, loadbalance, fallback, then direct selection. Within a group, matching routes retain configuration order. If no route yields a destination, model traffic uses the lexicographically first compatible candidate as route default; 404 no_route means there is no candidate.

Match

routes:
  - name: prod-chat
    match:
      path: "/v1/chat/completions"
      model: "gpt-*"
      header:
        x-team: payments
      scope:
        env: prod
    strategy: { }
FieldDefaultNotes
pathmatches anyExact path match. Path globs are not implemented.
modelmatches anyExact match, *, prefix (gpt-*), suffix (*-sonnet), or contains (*-mini-*).
header{}All entries must match; header names are case-insensitive.
scope{}Each dim: value pair must match the principal's resolved scope.

Every dimension named in match.scope must be declared in identity.scope_dims, and orca-gateway check fails if it is not.

Strategy modes

fallback

Try each target in order. Within a target, retry up to the retry budget; when that is exhausted, walk to the next target. A route without retry gives each target one attempt. A non-retryable failure still advances a fallback chain; retry triggers control repeated attempts on one target. When every target fails, the final provider failure is mapped to a gateway error response.

strategy:
  mode: fallback
  timeout_ms: 30000
  retry:
    triggers: [5xx, timeout]
    max_attempts: 3
    backoff_ms: 100
  targets:
    - destination: openai-primary
    - destination: azure-gpt4
    - destination: bedrock-claude

loadbalance

Pick one target by weighted random selection. Either every target has a weight that sums to 100, or no target has a weight and the split is equal. A zero-weight target is excluded from the pool.

strategy:
  mode: loadbalance
  targets:
    - destination: openai-primary
      weight: 70
    - destination: azure-gpt4
      weight: 30
  retry:
    triggers: [5xx]
    max_attempts: 2

With retry triggers set, the retry engine retries the destination selected by the load balancer. It does not draw another pool member for that request. Use a fallback route when a failed destination must move traffic to another upstream.

conditional

Pick a target from a predicate over headers and scope. The first matching condition wins. If no condition matches, the gateway continues with the next strategy group and eventually the model default candidate when one exists.

strategy:
  mode: conditional
  conditions:
    - if: { header: x-team, equals: gold }
      target: openai-primary
    - if: { header: x-team, equals: silver }
      target: azure-gpt4
    - if: { scope: { env: dev } }
      target: openai-cheap

Each predicate is a conjunction of header == value and scope.dim == value checks. For richer logic, use an authorizer and let OPA decide.

Route composition

The runtime evaluates conditional, load-balancing, fallback, and direct route policies in that priority order. This is not recursive routing: a conditional target must name a destination, not another route. Put fallback or load-balancing behavior directly on the route that needs it.

Retries

TriggerMeaning
http_5xx or 5xxUpstream HTTP 5xx response, including network failures.
http_429 or 429Upstream rate-limit response.
timeout_ms or timeoutPer-attempt timer expired.

There is no default retry trigger list: omitting strategy.retry means one attempt per target. The retry engine does not retry a route marked idempotency_safe: false; Retry-After from an upstream rate-limit response overrides configured backoff for that retry.

Idempotency

idempotency_safe defaults to true. Set it to false on any route whose upstream mutates state, and the retry engine propagates the first attempt's outcome verbatim instead of retrying or falling back.

Set idempotency_safe: false on MCP routes that carry side-effectful tools. Chat completions, embeddings, and Anthropic messages are idempotent at the application level - nothing is persisted upstream - so the default is correct for them.

Timeouts

strategy.timeout_ms bounds a single upstream attempt; when it fires, the attempt is cancelled and the retry engine decides what to do next. The config model also accepts server.request_timeout_ms, but the current CLI server does not enforce it as an end-to-end deadline.

Always set a per-attempt strategy.timeout_ms in production. If you leave it unset, no configured server-wide deadline bounds the attempt, retry backoff, or fallback walk.

Request shaping

Each target can rewrite chat, Messages, and embeddings requests before they leave the gateway:

targets:
  - destination: openai-primary
    default_params:
      temperature: 0.7
    override_params:
      max_tokens: 1024
    drop_params:
      - logit_bias
      - logprobs

default_params fills in fields the client omitted, override_params replaces whatever the client sent, and drop_params removes fields entirely. They apply in that order, so a key in both override_params and drop_params ends up dropped.

Input guardrails inspect the inbound request before target shaping. Shaping runs after route selection, immediately before the provider adapter receives the request.

Worked example: cheapest first

routes:
  - name: cost-first
    match:
      path: "/v1/chat/completions"
    strategy:
      mode: fallback
      timeout_ms: 20000
      retry:
        triggers: [5xx, 429, timeout]
        max_attempts: 2
      targets:
        - destination: bedrock-haiku
        - destination: openai-4o-mini
        - destination: openai-primary

Traffic lands on the cheapest destination and escalates only when one fails or rate-limits, so cost optimization degrades into availability rather than trading one for the other.

On this page