Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Spend controls

Price model calls and enforce calendar spend budgets in Orca AI Gateway.

A spend control prices a model request, reserves its worst-case cost before dispatch, and reconciles the reservation with provider-reported usage. A cost model supplies prices; a rate limiter supplies the budget ledger.

Choose a price source

plugins.cost_model.kindSource
seedThe price catalog vendored into the gateway release.
static_tableInline entries or a local YAML file.
http_refreshA remote catalog refreshed on an interval.
registryThe Agent Engine registry Workspace pricing API.
layeredSeveral of these sources combined into one resolver.

For explicit operator prices, start with a static table:

gateway.yaml
plugins:
  cost_model:
    name: operator-prices
    kind: static_table
    models:
      "openai/gpt-4o":
        input_per_1m: 2.50
        output_per_1m: 10.00
      "anthropic/claude-*":
        input_per_1m: 3.00
        output_per_1m: 15.00
        cache_read_per_1m: 0.30
        cache_write_per_1m: 3.75

Keys use provider/model glob patterns. Input and output rates are USD per million tokens and must be positive. Optional cache rates can be zero. Set file instead of models to load the same models: shape from a local YAML file.

A price directly on a destination overrides the cost model for that destination:

gateway.yaml
destinations:
  openai-contract-price:
    kind: openai
    credentials: { vault: openai-key }
    pricing:
      input_per_1m: 2.00
      output_per_1m: 8.00

Destinations without an override continue to use the configured cost model.

Combine catalogs

Use a layered model to combine the vendored catalog, a registry catalog, and operator overrides:

gateway.yaml
plugins:
  cost_model:
    name: prices
    kind: layered
    sources:
      - kind: seed
      - kind: registry
        base_url: http://orca-registry:8080
        api_key_file: /var/run/secrets/orca/workspace-api-key
        refresh_interval_secs: 300
      - kind: static_table
        models:
          "openai/gpt-4o":
            input_per_1m: 2.00
            output_per_1m: 8.00

The registry base_url is the public Workspace API service root. Supply exactly one api_key_file or oidc_token_file; the registry internal-service bearer token cannot authenticate this API. Registry and HTTP refreshers keep their last successful snapshot when a later refresh fails.

Configure admission

gateway.yaml
spend:
  key_dims: [org_id, workspace_id]
  budget_timezone: UTC
  admission:
    default_max_output_tokens: 4096
    unpriced_model: deny
  headers:
    enabled: true

Before dispatch, the gateway estimates input tokens and uses the request's maximum output tokens. When the request omits that maximum, it uses default_max_output_tokens. A fallback route is priced at the most expensive destination it may attempt instead of only its first target.

If a configured cost model cannot price any candidate destination, unpriced_model decides what happens:

ValueBehavior
denyReturn 403 model_unpriced. This is the default.
estimateReserve estimate_usd_per_request, which is required with this value.
allow_untrackedAdmit without a dollar reservation and mark usage as unpriced.

This policy runs whenever a cost model is configured, even without a rate limiter. With no cost model configured, the gateway treats unpriced usage as a supported observation-only setup rather than applying unpriced_model.

Enforce a calendar budget

Add a limiter and a budget to the matching scope:

gateway.yaml
plugins:
  rate_limiter:
    name: shared-spend
    kind: redis_window
    required: true
    failure_mode: deny
    url: redis://redis:6379/0
    reservation_ttl_secs: 900
    on_backend_error: deny

rate_limits:
  - match:
      scope: { workspace_id: ws_prod }
    qpm: 1000
    tpm: 200000
    budget_usd_per_month: 1000.00

spend:
  key_dims: [workspace_id]
  budget_timezone: America/Los_Angeles
  admission:
    default_max_output_tokens: 4096
    unpriced_model: deny
  headers:
    enabled: true

The gateway returns 429 budget_exhausted when the reservation would cross the calendar-month cap. After the provider response, it settles the reservation with actual usage. A failure before observed usage releases the reservation.

budget_usd_per_day remains functional but emits a startup deprecation warning. Use the user_daily_cost_budget policy guardrail for new daily budgets. The Policy guardrails page covers that stateful rule.

Use redis_window when replicas must share a budget. token_bucket_local gives every replica a separate ledger and can multiply the effective cap.

Observe headroom and usage

With spend.headers.enabled, budgeted responses include X-Orca-Budget-Remaining-Usd and X-Orca-Budget-Window when pricing is available. Usage sinks receive the settled token counts and cost. You can also query price resolution and usage aggregation through the Admin API.

On this page