Quickstart
Write a minimal config for Orca AI Gateway, validate it, and route your first model call through it.
This walkthrough takes you from nothing to a running gateway that proxies chat completions to OpenAI, with the API key resolved from a vault instead of being held by the caller.
Prerequisites
- Docker and the public
docker.io/streamnative/orca-ai-gateway:0.4.3image. See Install for the check and run commands. - An API key for at least one model provider.
1. Write a config
The gateway reads one YAML (or JSON) document. This is the smallest config that does something useful - one destination, one vault, one route:
server:
listen: 0.0.0.0:8080
admin_listen: 0.0.0.0:9099
identity:
scope_dims: [workspace_id]
validators:
- name: platform-jwt
kind: jwt
issuers: ["https://issuer.example.com"]
audiences: ["orca-gateway"]
jwks_url: "https://issuer.example.com/.well-known/jwks.json"
scope_from:
jwt_claims:
ws: workspace_id
destinations:
openai-primary:
kind: openai
base_url: "https://api.openai.com"
credentials:
vault: openai_key
pricing:
input_per_1m: 2.5
output_per_1m: 10.0
routes:
- name: chat
match:
path: "/v1/chat/completions"
strategy:
mode: fallback
targets:
- destination: openai-primary
timeout_ms: 30000
vaults:
- name: openai_key
resolver: env
env_var: OPENAI_API_KEY
scheme: bearerTop-level parsing uses deny_unknown_fields, so a misspelled key fails at boot rather than being
silently ignored.
scope_dims declares your tenancy dimensions. Every scope key you later reference in a route
match or a rate limit must appear here. Single-tenant deployments can leave the list empty.
2. Validate it
docker run --rm \
-v "$PWD/config.yaml:/etc/orca-gateway/config.yaml:ro" \
docker.io/streamnative/orca-ai-gateway:0.4.3 \
check /etc/orca-gateway/config.yamlcheck goes beyond schema validation. It confirms that every route target resolves to a declared
destination, that plugin instance names are unique within their trait, that every scope key used in
a match is declared in scope_dims, and that each strategy mode has the shape it requires. Errors
name the offending YAML path, such as routes[2].strategy.targets[0].destination, so you can go
straight to the line. The command exits non-zero when validation fails.
3. Run it
The env_var setting makes this vault read OPENAI_API_KEY. Without env_var, an env vault
derives its name as ORCA_VAULT__<vault_id>, or
ORCA_VAULT__<workspace_id>__<vault_id> when scope carries workspace_id. See
Vaults.
export OPENAI_API_KEY="sk-..."
docker run --rm \
-p 127.0.0.1:8080:8080 -p 127.0.0.1:9099:9099 \
-v "$PWD/config.yaml:/etc/orca-gateway/config.yaml:ro" \
-e OPENAI_API_KEY \
docker.io/streamnative/orca-ai-gateway:0.4.3The gateway binds two listeners inside the container: the data plane on :8080 and the admin plane
on :9099. Docker publishes both on the host's loopback address. Keep the admin plane there unless a
network policy fronts it.
Confirm it is healthy:
curl -fsS http://localhost:9099/healthz
curl -fsS http://localhost:9099/readyzhealthz answers as soon as the process is up. readyz becomes healthy after AppState
construction.
4. Route a call through it
Point an OpenAI-compatible client at the gateway instead of at the provider. The provider key never leaves the vault. This local example sends no caller token. Before you expose the gateway to untrusted clients, configure an authorizer that denies anonymous principals.
curl -fsS -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Say hello in five words."}]
}'Streaming works the same way - set "stream": true and read the Server-Sent Events response.
5. Watch what it did
The admin plane exposes Prometheus metrics:
curl -fsS http://localhost:9099/metrics | grep orca_gatewayTo get per-call detail, add a usage sink and a trace exporter. The stdout usage sink needs no
external infrastructure:
plugins:
usage_sinks:
- name: local
kind: stdout
required: false
failure_mode: log_onlySee Observe for OpenTelemetry spans, usage sinks, and audit sinks.
Where to go next
Model providers
Add providers, choose credential schemes, and map model ids.
Routes
Load balance, fall back, and retry across destinations.
Rate limits
Cap requests and tokens per verified scope.
Spend controls
Price model calls and enforce calendar budgets.
Payload guardrails
Redact sensitive data and evaluate prompts before they reach a provider.
Policy guardrails
Enforce model, tool, token, and spend policy.