Deploy with Helm
Install Orca AI Gateway on Kubernetes, deliver its configuration, and wire probes and metrics.
The public 0.4.3 OCI chart installs the gateway as a Deployment with a production posture by
default. See Install for the basic command; this page covers what
to change and why.
Delivering configuration
config:
source: file
file:
content: |
server:
listen: 0.0.0.0:8080
admin_listen: 0.0.0.0:9099
destinations: {}
routes: []The chart mounts the config document at /etc/orca-gateway/config.yaml. Supply it inline with
config.file.content, from a local file with --set-file config.file.content=gateway.yaml, or from
a ConfigMap or Secret you manage with config.existingConfigMap or config.existingSecret, so config
changes are not chart releases. Prefer config.existingSecret when the document embeds credentials.
The gateway reads its config once at startup, so roll the pods after a change.
Use config.source: file with the 0.4.3 chart. Its postgres value passes
--config postgres-env:<var>, which the gateway rejects at startup. postgres.enabled, with its
default postgres.runMigrations: true, adds an init container that runs
orca-gateway admin migrate, a subcommand the 0.4.3 binary does not have. To run the
Postgres control plane, start the gateway with
orca-gateway run --config postgres://... outside the chart.
Secrets
vaultSecrets maps Kubernetes Secret keys to the environment variables your vaults resolve from.
Each entry names existingSecret, existingSecretKey, and envVar. An env resolver with
env_var reads that fixed variable. Without env_var, it derives ORCA_VAULT__<vault_id>, or
ORCA_VAULT__<workspace_id>__<vault_id> when a scope is present. See
Vaults.
The chart can annotate its service account, but the current Bedrock and Vertex adapters do not use ambient AWS or Google credential chains. They still require credentials from the destination's configured vault. Service-account annotations alone do not authenticate provider calls.
Set annotations when another container in the pod, such as a sidecar, needs the cloud identity:
serviceAccount:
create: true
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/orca-gateway-prod
# or, on GKE:
# iam.gke.io/gcp-service-account: orca-gateway@my-proj.iam.gserviceaccount.comExposure
The ClusterIP Service exposes the data plane on 8080 and the admin plane on 9099.
Do not route ingress to the admin port. It serves GET /admin/v1/config, which returns the full
active configuration. Expose 8080 and leave 9099 reachable only from inside the cluster, with
networkPolicy.enabled: true once you have audited your traffic.
Before you enable ingress or otherwise expose the data plane to untrusted clients, configure an auth validator for your callers' credentials and an authorizer that denies anonymous principals in the gateway config.
Ingress is off by default. Enable it for the data plane only:
ingress:
enabled: true
hosts:
- host: gateway.example.com
paths:
- path: /
pathType: Prefix
tls:
- secretName: gateway-tls
hosts:
- gateway.example.comTerminate TLS at the ingress or a sidecar. server.tls is parsed and live-validated, but the
current CLI binds a plain TCP listener.
Probes
Wire readinessProbe to /readyz and livenessProbe to /healthz. Invalid config and plugin
initialization errors stop the process before it serves; readiness becomes true after AppState
construction and does not track per-plugin health afterward.
Metrics
The chart sets Prometheus scrape annotations on pods by default. With Prometheus Operator, prefer a ServiceMonitor:
metrics:
serviceMonitor:
enabled: trueUpgrading
Write the same gateway config document to config.yaml for validation, then roll:
docker run --rm -v "$PWD/config.yaml:/etc/orca-gateway/config.yaml:ro" \
docker.io/streamnative/orca-ai-gateway:0.4.3 \
check /etc/orca-gateway/config.yaml
helm upgrade orca-ai-gateway oci://ghcr.io/orca-ae/charts/orca-ai-gateway \
--version 0.4.3 --namespace orca-system -f values.yaml \
--set image.repository=docker.io/streamnative/orca-ai-gateway \
--set image.tag=0.4.3The Deployment rolls with maxUnavailable: 0, so existing pods keep serving until new pods pass
readiness, and a new config that cannot initialize stalls the rollout. The PodDisruptionBudget
separately keeps 75% of the fleet up during voluntary disruptions such as node drains. To roll back,
run helm rollback; the gateway holds no state that makes a downgrade one-way.