Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Pricing

Configure model prices for usage and spend controls in Orca AI Gateway.

Pricing converts observed model tokens into USD for Gateway usage records and dollar budgets. Set a pricing override on a destination, configure plugins.cost_model, or use both. The destination override takes precedence for that destination. Request and token limits work without a price; dollar accounting needs one.

Set a destination price

Add rates in USD per million tokens to a model destination:

destinations:
  openai-primary:
    kind: openai
    credentials:
      vault: openai-key
    pricing:
      input_per_1m: 2.5
      output_per_1m: 10
      cache_read_per_1m: 1.25
FieldMeaning
input_per_1m, output_per_1mInput and output token rates.
cache_read_per_1m, cache_write_per_1m, cache_write_1h_per_1mOptional cache-token rates.
reasoning_output_per_1mOptional reasoning-output rate.
per_request_usdOptional fixed amount per request.

This price applies to traffic sent to openai-primary. See Destinations for model maps and provider setup.

Use Agent Engine prices

To use the effective prices from an Agent Engine Workspace, configure the registry cost model:

plugins:
  cost_model:
    name: registry-prices
    kind: registry
    base_url: http://orca-registry:8080
    api_key_file: /var/run/secrets/orca/workspace-api-key
    refresh_interval_secs: 300

base_url is the public Registry listener's host root. Supply exactly one Workspace credential: api_key_file for a Workspace API key or oidc_token_file for a Workspace OIDC token. The Registry internal-service token cannot read this public pricing route. Registry resolves seed, upstream, and organization overrides into effective model prices before the Gateway consumes them. The Gateway refreshes its price snapshot on the configured interval; changing a registry price is not an immediate Gateway config change.

plugins.cost_model also supports seed, static_table, http_refresh, and layered. Use layered.sources[] when several catalogs need an explicit order. A destination pricing override still takes precedence over the configured cost model for that destination.

Handle unpriced models

When a configured cost model cannot resolve a selected model, spend.admission.unpriced_model controls admission:

ValueBehavior
denyReject the request. This is the default.
estimateReserve estimate_usd_per_request as a flat amount.
allow_untrackedAdmit without dollar accounting for that model.
spend:
  admission:
    unpriced_model: deny

An unpriced request is not a measured zero-cost request. Usage records include cost and pricing provenance when available. For calendar spend caps, see Rate limits and budgets. For per-session or per-principal limits in Agent Engine, see cost guardrails.

What's next

On this page