Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

High availability

Run Orca AI Gateway as a multi-replica fleet, and understand what fails when a dependency goes down.

Any gateway replica can serve any request, and each loads the same config independently. Optional features still keep local state: token_bucket_local, in-memory policy state, sink queues, and usage spools. Use shared backends where enforcement or durability must span replicas.

Topology

                  Ingress / load balancer
                            │
        ┌───────────────────┼───────────────────┐
        ▼                   ▼                   ▼
   gateway replica     gateway replica     gateway replica
        │                   │                   │
        └───────────────────┼───────────────────┘
                            ▼
              Postgres        Redis        Kafka
          (control/usage) (limits/policy)  (sinks)

The gateway does not manage any of those dependencies. Run them as operator-managed or cloud-managed clusters.

Chart defaults

The Helm chart defaults to a production HA posture:

SettingDefaultNotes
replicaCount3Spread across zones with preferred pod anti-affinity.
pdb.enabledtrueKeeps 75% of the fleet available during voluntary disruption.
autoscaling.enabledfalseTurn on once a metrics pipeline exists.
networkPolicy.enabledfalseTurn on after auditing your ingress.
metrics.serviceMonitor.enabledfalseTurn on with Prometheus Operator.

For development, drop to replicaCount: 1 and pdb.enabled: false. For multi-region, deploy the chart once per region rather than stretching one deployment.

What a replica loses when it dies

A lost replica drops in-flight requests and its in-memory queues, limiter buckets, and policy state. A usage spool on a persistent volume can replay after replacement; an ephemeral spool disappears with the pod. Audit sinks do not use the usage spool, so failed audit deliveries are not replayed after a restart. Kafka, Postgres, and Redis keep their remote state independently of the replica.

Failure modes

Dependency downEffectWhat to do
Postgres control planeRunning replicas keep serving the last config; new replicas cannot bootFix before scaling or rolling. Traffic is unaffected in the meantime.
Redisredis_window and Redis policy state follow their configured failure behaviorChoose deny, allow, or fallback_local for redis_window based on whether availability or strict shared admission matters more.
Kafka usage sinkUsage records queue, spool when configured, then drop once both capacities fillPut the usage spool on persistent storage and alert on sink drops.
Kafka audit sinkAudit delivery runs on the caller's path with no local durable spoolRestore Kafka promptly and alert on delivery failures; failed audit records are not replayed after restart.
A providerFallback walks to the next destinationConfigure a multi-destination strategy. See Routes.

token_bucket_local is per replica, so a three-replica fleet can admit roughly three times each configured limit. Use redis_window for one atomic shared ceiling. Its fallback_local backend error mode deliberately returns to per-replica limits during a Redis outage.

Rolling safely

Plugin initialization is synchronous. A rollout carrying a bad config crashes new processes before they serve; /readyz becomes healthy after AppState construction. To use that behavior safely:

  • Wire readinessProbe to /readyz, not /healthz.
  • Inspect restart and rollout status as well as readiness. Initialization errors fail boot regardless of the plugin's required value.

Validate the config in CI with orca-gateway check so a rollout is not the first thing to notice a typo.

Sizing

CPU scales roughly linearly with request rate. Memory is dominated by sink queues and policy state. Start from the chart's defaults, watch the per-stage duration metrics, and scale on the stage that saturates - which is usually the upstream call, not the gateway.

On this page