Pricing
Configure model prices for usage and spend controls in Orca AI Gateway.
Pricing converts observed model tokens into USD for Gateway usage records and dollar budgets.
Set a pricing override on a destination, configure plugins.cost_model, or use both. The
destination override takes precedence for that destination. Request and token limits work without
a price; dollar accounting needs one.
Set a destination price
Add rates in USD per million tokens to a model destination:
destinations:
openai-primary:
kind: openai
credentials:
vault: openai-key
pricing:
input_per_1m: 2.5
output_per_1m: 10
cache_read_per_1m: 1.25| Field | Meaning |
|---|---|
input_per_1m, output_per_1m | Input and output token rates. |
cache_read_per_1m, cache_write_per_1m, cache_write_1h_per_1m | Optional cache-token rates. |
reasoning_output_per_1m | Optional reasoning-output rate. |
per_request_usd | Optional fixed amount per request. |
This price applies to traffic sent to openai-primary. See
Destinations for model maps and provider setup.
Use Agent Engine prices
To use the effective prices from an Agent Engine Workspace, configure the registry cost model:
plugins:
cost_model:
name: registry-prices
kind: registry
base_url: http://orca-registry:8080
api_key_file: /var/run/secrets/orca/workspace-api-key
refresh_interval_secs: 300base_url is the public Registry listener's host root. Supply exactly one Workspace credential:
api_key_file for a Workspace API key or oidc_token_file for a Workspace OIDC token. The
Registry internal-service token cannot read this public pricing route. Registry resolves seed,
upstream, and organization overrides into effective model prices
before the Gateway consumes them. The Gateway refreshes its price snapshot on the configured
interval; changing a registry price is not an immediate Gateway config change.
plugins.cost_model also supports seed, static_table, http_refresh, and layered. Use
layered.sources[] when several catalogs need an explicit order. A destination pricing override
still takes precedence over the configured cost model for that destination.
Handle unpriced models
When a configured cost model cannot resolve a selected model,
spend.admission.unpriced_model controls admission:
| Value | Behavior |
|---|---|
deny | Reject the request. This is the default. |
estimate | Reserve estimate_usd_per_request as a flat amount. |
allow_untracked | Admit without dollar accounting for that model. |
spend:
admission:
unpriced_model: denyAn unpriced request is not a measured zero-cost request. Usage records include cost and pricing provenance when available. For calendar spend caps, see Rate limits and budgets. For per-session or per-principal limits in Agent Engine, see cost guardrails.