Routes
Match requests to destinations in Orca AI Gateway, and control fallback, load balancing, retries, timeouts, and request shaping.
A route binds a set of requests to a strategy for picking a destination. The gateway evaluates
strategy groups in this order: conditional, loadbalance, fallback, then direct selection.
Within a group, matching routes retain configuration order. If no route yields a destination, model
traffic uses the lexicographically first compatible candidate as route default; 404 no_route
means there is no candidate.
Match
routes:
- name: prod-chat
match:
path: "/v1/chat/completions"
model: "gpt-*"
header:
x-team: payments
scope:
env: prod
strategy: { }| Field | Default | Notes |
|---|---|---|
path | matches any | Exact path match. Path globs are not implemented. |
model | matches any | Exact match, *, prefix (gpt-*), suffix (*-sonnet), or contains (*-mini-*). |
header | {} | All entries must match; header names are case-insensitive. |
scope | {} | Each dim: value pair must match the principal's resolved scope. |
Every dimension named in match.scope must be declared in identity.scope_dims, and
orca-gateway check fails if it is not.
Strategy modes
fallback
Try each target in order. Within a target, retry up to the retry budget; when that is exhausted,
walk to the next target. A route without retry gives each target one attempt. A non-retryable
failure still advances a fallback chain; retry triggers control repeated attempts on one target.
When every target fails, the final provider failure is mapped to a gateway error response.
strategy:
mode: fallback
timeout_ms: 30000
retry:
triggers: [5xx, timeout]
max_attempts: 3
backoff_ms: 100
targets:
- destination: openai-primary
- destination: azure-gpt4
- destination: bedrock-claudeloadbalance
Pick one target by weighted random selection. Either every target has a weight that sums to 100,
or no target has a weight and the split is equal. A zero-weight target is excluded from the pool.
strategy:
mode: loadbalance
targets:
- destination: openai-primary
weight: 70
- destination: azure-gpt4
weight: 30
retry:
triggers: [5xx]
max_attempts: 2With retry triggers set, the retry engine retries the destination selected by the load balancer. It
does not draw another pool member for that request. Use a fallback route when a failed destination
must move traffic to another upstream.
conditional
Pick a target from a predicate over headers and scope. The first matching condition wins. If no condition matches, the gateway continues with the next strategy group and eventually the model default candidate when one exists.
strategy:
mode: conditional
conditions:
- if: { header: x-team, equals: gold }
target: openai-primary
- if: { header: x-team, equals: silver }
target: azure-gpt4
- if: { scope: { env: dev } }
target: openai-cheapEach predicate is a conjunction of header == value and scope.dim == value checks. For richer
logic, use an authorizer and let OPA decide.
Route composition
The runtime evaluates conditional, load-balancing, fallback, and direct route policies in that
priority order. This is not recursive routing: a conditional target must name a destination, not
another route. Put fallback or load-balancing behavior directly on the route that needs it.
Retries
| Trigger | Meaning |
|---|---|
http_5xx or 5xx | Upstream HTTP 5xx response, including network failures. |
http_429 or 429 | Upstream rate-limit response. |
timeout_ms or timeout | Per-attempt timer expired. |
There is no default retry trigger list: omitting strategy.retry means one attempt per target. The
retry engine does not retry a route marked idempotency_safe: false; Retry-After from an upstream
rate-limit response overrides configured backoff for that retry.
Idempotency
idempotency_safe defaults to true. Set it to false on any route whose upstream mutates state,
and the retry engine propagates the first attempt's outcome verbatim instead of retrying or
falling back.
Set idempotency_safe: false on MCP routes that carry side-effectful tools. Chat completions,
embeddings, and Anthropic messages are idempotent at the application level - nothing is persisted
upstream - so the default is correct for them.
Timeouts
strategy.timeout_ms bounds a single upstream attempt; when it fires, the attempt is cancelled and
the retry engine decides what to do next. The config model also accepts
server.request_timeout_ms, but the current CLI server does not enforce it as an end-to-end
deadline.
Always set a per-attempt strategy.timeout_ms in production. If you leave it unset, no configured
server-wide deadline bounds the attempt, retry backoff, or fallback walk.
Request shaping
Each target can rewrite chat, Messages, and embeddings requests before they leave the gateway:
targets:
- destination: openai-primary
default_params:
temperature: 0.7
override_params:
max_tokens: 1024
drop_params:
- logit_bias
- logprobsdefault_params fills in fields the client omitted, override_params replaces whatever the client
sent, and drop_params removes fields entirely. They apply in that order, so a key in both
override_params and drop_params ends up dropped.
Input guardrails inspect the inbound request before target shaping. Shaping runs after route selection, immediately before the provider adapter receives the request.
Worked example: cheapest first
routes:
- name: cost-first
match:
path: "/v1/chat/completions"
strategy:
mode: fallback
timeout_ms: 20000
retry:
triggers: [5xx, 429, timeout]
max_attempts: 2
targets:
- destination: bedrock-haiku
- destination: openai-4o-mini
- destination: openai-primaryTraffic lands on the cheapest destination and escalates only when one fails or rate-limits, so cost optimization degrades into availability rather than trading one for the other.