Spend controls
Price model calls and enforce calendar spend budgets in Orca AI Gateway.
A spend control prices a model request, reserves its worst-case cost before dispatch, and reconciles the reservation with provider-reported usage. A cost model supplies prices; a rate limiter supplies the budget ledger.
Choose a price source
plugins.cost_model.kind | Source |
|---|---|
seed | The price catalog vendored into the gateway release. |
static_table | Inline entries or a local YAML file. |
http_refresh | A remote catalog refreshed on an interval. |
registry | The Agent Engine registry Workspace pricing API. |
layered | Several of these sources combined into one resolver. |
For explicit operator prices, start with a static table:
plugins:
cost_model:
name: operator-prices
kind: static_table
models:
"openai/gpt-4o":
input_per_1m: 2.50
output_per_1m: 10.00
"anthropic/claude-*":
input_per_1m: 3.00
output_per_1m: 15.00
cache_read_per_1m: 0.30
cache_write_per_1m: 3.75Keys use provider/model glob patterns. Input and output rates are USD per million tokens and must
be positive. Optional cache rates can be zero. Set file instead of models to load the same
models: shape from a local YAML file.
A price directly on a destination overrides the cost model for that destination:
destinations:
openai-contract-price:
kind: openai
credentials: { vault: openai-key }
pricing:
input_per_1m: 2.00
output_per_1m: 8.00Destinations without an override continue to use the configured cost model.
Combine catalogs
Use a layered model to combine the vendored catalog, a registry catalog, and operator overrides:
plugins:
cost_model:
name: prices
kind: layered
sources:
- kind: seed
- kind: registry
base_url: http://orca-registry:8080
api_key_file: /var/run/secrets/orca/workspace-api-key
refresh_interval_secs: 300
- kind: static_table
models:
"openai/gpt-4o":
input_per_1m: 2.00
output_per_1m: 8.00The registry base_url is the public Workspace API service root. Supply exactly one
api_key_file or oidc_token_file; the registry internal-service bearer token cannot authenticate
this API. Registry and HTTP refreshers keep their last successful snapshot when a later refresh
fails.
Configure admission
spend:
key_dims: [org_id, workspace_id]
budget_timezone: UTC
admission:
default_max_output_tokens: 4096
unpriced_model: deny
headers:
enabled: trueBefore dispatch, the gateway estimates input tokens and uses the request's maximum output tokens.
When the request omits that maximum, it uses default_max_output_tokens. A fallback route is priced
at the most expensive destination it may attempt instead of only its first target.
If a configured cost model cannot price any candidate destination, unpriced_model decides what
happens:
| Value | Behavior |
|---|---|
deny | Return 403 model_unpriced. This is the default. |
estimate | Reserve estimate_usd_per_request, which is required with this value. |
allow_untracked | Admit without a dollar reservation and mark usage as unpriced. |
This policy runs whenever a cost model is configured, even without a rate limiter. With no cost
model configured, the gateway treats unpriced usage as a supported observation-only setup rather
than applying unpriced_model.
Enforce a calendar budget
Add a limiter and a budget to the matching scope:
plugins:
rate_limiter:
name: shared-spend
kind: redis_window
required: true
failure_mode: deny
url: redis://redis:6379/0
reservation_ttl_secs: 900
on_backend_error: deny
rate_limits:
- match:
scope: { workspace_id: ws_prod }
qpm: 1000
tpm: 200000
budget_usd_per_month: 1000.00
spend:
key_dims: [workspace_id]
budget_timezone: America/Los_Angeles
admission:
default_max_output_tokens: 4096
unpriced_model: deny
headers:
enabled: trueThe gateway returns 429 budget_exhausted when the reservation would cross the calendar-month cap.
After the provider response, it settles the reservation with actual usage. A failure before observed
usage releases the reservation.
budget_usd_per_day remains functional but emits a startup deprecation warning. Use the
user_daily_cost_budget policy guardrail for new daily budgets. The
Policy guardrails page covers that stateful rule.
Use redis_window when replicas must share a budget. token_bucket_local gives every replica a
separate ledger and can multiply the effective cap.
Observe headroom and usage
With spend.headers.enabled, budgeted responses include X-Orca-Budget-Remaining-Usd and
X-Orca-Budget-Window when pricing is available. Usage sinks receive the settled token counts and
cost. You can also query price resolution and usage aggregation through the
Admin API.