Rate limits
Cap request and token rates per verified scope in Orca AI Gateway, per replica or shared through Redis.
A rate limit caps requests or tokens for the scope and principal that match an entry. Configure both the top-level limits and the limiter plugin that enforces them.
plugins:
rate_limiter:
name: local
kind: token_bucket_local
required: true
failure_mode: deny
rate_limits:
- match:
scope: { env: prod }
qpm: 1000
tpm: 200000
qph: 50000
tph: 1000000
- match:
scope: { env: dev }
qpm: 100Set request and token windows
| Field | Window |
|---|---|
qpm | Requests per rolling minute. |
tpm | Input plus output tokens per rolling minute. |
qph | Requests per rolling hour. |
tph | Input plus output tokens per rolling hour. |
The limiter compares match.scope with verified principal scope. It chooses the matching entry with
the most dimensions; configuration order breaks a tie. A request rejected at this stage returns
429 before any provider call.
budget_usd_per_day remains functional for compatibility but emits a startup deprecation warning.
For new daily budgets, use a user_daily_cost_budget rule in
Agent Engine. The
policy guardrail guide explains how Gateway consumes a
Registry policy bundle.
Set bucket identity
By default, a bucket key contains every verified scope.<dim> attribute and the authenticated
principal id. Restrict the scope dimensions with spend.key_dims:
spend:
key_dims: [org_id, workspace_id]
trusted_header_dims: []
headers:
enabled: truetrusted_header_dims explicitly permits a dimension from X-Orca-Scope-* only when no verified
value exists. Keep it empty unless a trusted proxy stamps the header. A verified value always wins.
When headers.enabled is true, successful responses include the tightest request-window values in
X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset.
Choose local or shared enforcement
token_bucket_local keeps rolling windows in process memory. It is suitable for one replica and
for development.
Each token_bucket_local replica enforces a separate limit. Three replicas can admit roughly
three times the configured traffic. Use redis_window for one shared admission ceiling.
Configure Redis for shared windows:
plugins:
rate_limiter:
name: shared
kind: redis_window
required: true
failure_mode: deny
url: redis://redis.svc:6379/0
reservation_ttl_secs: 900
on_backend_error: denyon_backend_error accepts deny, allow, or fallback_local. fallback_local continues with
per-replica enforcement until Redis recovers. Plugin fields named burst and refill_per_sec are
accepted by older examples but are not read by the server; limits come from rate_limits.
Coordinate provider retries
Gateway rate limiting happens after route selection and before provider dispatch. A provider-side
429 happens later and can trigger a route fallback. If callers should receive the upstream
backpressure, omit 429 from the route's retry triggers. See
Routes.
Apply concurrency backpressure elsewhere
server.max_concurrent_requests is parsed but not enforced by the current CLI server. Enforce a
process-wide concurrency ceiling at an ingress or proxy.
Dollar budgets share the limiter's reservation lifecycle but require model prices and additional admission settings. Configure them in Spend controls.