High availability
Run Orca AI Gateway as a multi-replica fleet, and understand what fails when a dependency goes down.
Any gateway replica can serve any request, and each loads the same config independently. Optional
features still keep local state: token_bucket_local, in-memory policy state, sink queues, and usage
spools. Use shared backends where enforcement or durability must span replicas.
Topology
Ingress / load balancer
│
┌───────────────────┼───────────────────┐
▼ ▼ ▼
gateway replica gateway replica gateway replica
│ │ │
└───────────────────┼───────────────────┘
▼
Postgres Redis Kafka
(control/usage) (limits/policy) (sinks)The gateway does not manage any of those dependencies. Run them as operator-managed or cloud-managed clusters.
Chart defaults
The Helm chart defaults to a production HA posture:
| Setting | Default | Notes |
|---|---|---|
replicaCount | 3 | Spread across zones with preferred pod anti-affinity. |
pdb.enabled | true | Keeps 75% of the fleet available during voluntary disruption. |
autoscaling.enabled | false | Turn on once a metrics pipeline exists. |
networkPolicy.enabled | false | Turn on after auditing your ingress. |
metrics.serviceMonitor.enabled | false | Turn on with Prometheus Operator. |
For development, drop to replicaCount: 1 and pdb.enabled: false. For multi-region, deploy the
chart once per region rather than stretching one deployment.
What a replica loses when it dies
A lost replica drops in-flight requests and its in-memory queues, limiter buckets, and policy state. A usage spool on a persistent volume can replay after replacement; an ephemeral spool disappears with the pod. Audit sinks do not use the usage spool, so failed audit deliveries are not replayed after a restart. Kafka, Postgres, and Redis keep their remote state independently of the replica.
Failure modes
| Dependency down | Effect | What to do |
|---|---|---|
| Postgres control plane | Running replicas keep serving the last config; new replicas cannot boot | Fix before scaling or rolling. Traffic is unaffected in the meantime. |
| Redis | redis_window and Redis policy state follow their configured failure behavior | Choose deny, allow, or fallback_local for redis_window based on whether availability or strict shared admission matters more. |
| Kafka usage sink | Usage records queue, spool when configured, then drop once both capacities fill | Put the usage spool on persistent storage and alert on sink drops. |
| Kafka audit sink | Audit delivery runs on the caller's path with no local durable spool | Restore Kafka promptly and alert on delivery failures; failed audit records are not replayed after restart. |
| A provider | Fallback walks to the next destination | Configure a multi-destination strategy. See Routes. |
token_bucket_local is per replica, so a three-replica fleet can admit roughly three times each
configured limit. Use redis_window for one atomic shared ceiling. Its fallback_local backend
error mode deliberately returns to per-replica limits during a Redis outage.
Rolling safely
Plugin initialization is synchronous. A rollout carrying a bad config crashes new processes before
they serve; /readyz becomes healthy after AppState construction. To use that behavior safely:
- Wire
readinessProbeto/readyz, not/healthz. - Inspect restart and rollout status as well as readiness. Initialization errors fail boot
regardless of the plugin's
requiredvalue.
Validate the config in CI with orca-gateway check so a rollout is not the first thing to notice a
typo.
Sizing
CPU scales roughly linearly with request rate. Memory is dominated by sink queues and policy state. Start from the chart's defaults, watch the per-stage duration metrics, and scale on the stage that saturates - which is usually the upstream call, not the gateway.