Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Rate limits

Cap request and token rates per verified scope in Orca AI Gateway, per replica or shared through Redis.

A rate limit caps requests or tokens for the scope and principal that match an entry. Configure both the top-level limits and the limiter plugin that enforces them.

gateway.yaml
plugins:
  rate_limiter:
    name: local
    kind: token_bucket_local
    required: true
    failure_mode: deny

rate_limits:
  - match:
      scope: { env: prod }
    qpm: 1000
    tpm: 200000
    qph: 50000
    tph: 1000000
  - match:
      scope: { env: dev }
    qpm: 100

Set request and token windows

FieldWindow
qpmRequests per rolling minute.
tpmInput plus output tokens per rolling minute.
qphRequests per rolling hour.
tphInput plus output tokens per rolling hour.

The limiter compares match.scope with verified principal scope. It chooses the matching entry with the most dimensions; configuration order breaks a tie. A request rejected at this stage returns 429 before any provider call.

budget_usd_per_day remains functional for compatibility but emits a startup deprecation warning. For new daily budgets, use a user_daily_cost_budget rule in Agent Engine. The policy guardrail guide explains how Gateway consumes a Registry policy bundle.

Set bucket identity

By default, a bucket key contains every verified scope.<dim> attribute and the authenticated principal id. Restrict the scope dimensions with spend.key_dims:

gateway.yaml
spend:
  key_dims: [org_id, workspace_id]
  trusted_header_dims: []
  headers:
    enabled: true

trusted_header_dims explicitly permits a dimension from X-Orca-Scope-* only when no verified value exists. Keep it empty unless a trusted proxy stamps the header. A verified value always wins.

When headers.enabled is true, successful responses include the tightest request-window values in X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset.

Choose local or shared enforcement

token_bucket_local keeps rolling windows in process memory. It is suitable for one replica and for development.

Each token_bucket_local replica enforces a separate limit. Three replicas can admit roughly three times the configured traffic. Use redis_window for one shared admission ceiling.

Configure Redis for shared windows:

gateway.yaml
plugins:
  rate_limiter:
    name: shared
    kind: redis_window
    required: true
    failure_mode: deny
    url: redis://redis.svc:6379/0
    reservation_ttl_secs: 900
    on_backend_error: deny

on_backend_error accepts deny, allow, or fallback_local. fallback_local continues with per-replica enforcement until Redis recovers. Plugin fields named burst and refill_per_sec are accepted by older examples but are not read by the server; limits come from rate_limits.

Coordinate provider retries

Gateway rate limiting happens after route selection and before provider dispatch. A provider-side 429 happens later and can trigger a route fallback. If callers should receive the upstream backpressure, omit 429 from the route's retry triggers. See Routes.

Apply concurrency backpressure elsewhere

server.max_concurrent_requests is parsed but not enforced by the current CLI server. Enforce a process-wide concurrency ceiling at an ingress or proxy.

Dollar budgets share the limiter's reservation lifecycle but require model prices and additional admission settings. Configure them in Spend controls.

On this page