Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Quickstart

Write a minimal config for Orca AI Gateway, validate it, and route your first model call through it.

This walkthrough takes you from nothing to a running gateway that proxies chat completions to OpenAI, with the API key resolved from a vault instead of being held by the caller.

Prerequisites

  • Docker and the public docker.io/streamnative/orca-ai-gateway:0.4.3 image. See Install for the check and run commands.
  • An API key for at least one model provider.

1. Write a config

The gateway reads one YAML (or JSON) document. This is the smallest config that does something useful - one destination, one vault, one route:

config.yaml
server:
  listen: 0.0.0.0:8080
  admin_listen: 0.0.0.0:9099

identity:
  scope_dims: [workspace_id]
  validators:
    - name: platform-jwt
      kind: jwt
      issuers: ["https://issuer.example.com"]
      audiences: ["orca-gateway"]
      jwks_url: "https://issuer.example.com/.well-known/jwks.json"
      scope_from:
        jwt_claims:
          ws: workspace_id

destinations:
  openai-primary:
    kind: openai
    base_url: "https://api.openai.com"
    credentials:
      vault: openai_key
    pricing:
      input_per_1m: 2.5
      output_per_1m: 10.0

routes:
  - name: chat
    match:
      path: "/v1/chat/completions"
    strategy:
      mode: fallback
      targets:
        - destination: openai-primary
      timeout_ms: 30000

vaults:
  - name: openai_key
    resolver: env
    env_var: OPENAI_API_KEY
    scheme: bearer

Top-level parsing uses deny_unknown_fields, so a misspelled key fails at boot rather than being silently ignored.

scope_dims declares your tenancy dimensions. Every scope key you later reference in a route match or a rate limit must appear here. Single-tenant deployments can leave the list empty.

2. Validate it

docker run --rm \
  -v "$PWD/config.yaml:/etc/orca-gateway/config.yaml:ro" \
  docker.io/streamnative/orca-ai-gateway:0.4.3 \
  check /etc/orca-gateway/config.yaml

check goes beyond schema validation. It confirms that every route target resolves to a declared destination, that plugin instance names are unique within their trait, that every scope key used in a match is declared in scope_dims, and that each strategy mode has the shape it requires. Errors name the offending YAML path, such as routes[2].strategy.targets[0].destination, so you can go straight to the line. The command exits non-zero when validation fails.

3. Run it

The env_var setting makes this vault read OPENAI_API_KEY. Without env_var, an env vault derives its name as ORCA_VAULT__<vault_id>, or ORCA_VAULT__<workspace_id>__<vault_id> when scope carries workspace_id. See Vaults.

export OPENAI_API_KEY="sk-..."
docker run --rm \
  -p 127.0.0.1:8080:8080 -p 127.0.0.1:9099:9099 \
  -v "$PWD/config.yaml:/etc/orca-gateway/config.yaml:ro" \
  -e OPENAI_API_KEY \
  docker.io/streamnative/orca-ai-gateway:0.4.3

The gateway binds two listeners inside the container: the data plane on :8080 and the admin plane on :9099. Docker publishes both on the host's loopback address. Keep the admin plane there unless a network policy fronts it.

Confirm it is healthy:

curl -fsS http://localhost:9099/healthz
curl -fsS http://localhost:9099/readyz

healthz answers as soon as the process is up. readyz becomes healthy after AppState construction.

4. Route a call through it

Point an OpenAI-compatible client at the gateway instead of at the provider. The provider key never leaves the vault. This local example sends no caller token. Before you expose the gateway to untrusted clients, configure an authorizer that denies anonymous principals.

curl -fsS -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Say hello in five words."}]
  }'

Streaming works the same way - set "stream": true and read the Server-Sent Events response.

5. Watch what it did

The admin plane exposes Prometheus metrics:

curl -fsS http://localhost:9099/metrics | grep orca_gateway

To get per-call detail, add a usage sink and a trace exporter. The stdout usage sink needs no external infrastructure:

config.yaml
plugins:
  usage_sinks:
    - name: local
      kind: stdout
      required: false
      failure_mode: log_only

See Observe for OpenTelemetry spans, usage sinks, and audit sinks.

Where to go next

On this page