Orca AI Gateway
Route, govern, and observe model and MCP tool traffic with Orca AI Gateway - credentials, guardrails, budgets, and telemetry.
Orca AI Gateway is a gateway, written in Rust and distributed as a Docker image, that sits in front of the model providers and Model Context Protocol (MCP) servers your applications call. It resolves credentials so callers never hold them, routes across providers with fallback, enforces guardrails and budgets, and emits a consistent telemetry and audit record for every call it handles.
It runs standalone. Nothing on this page requires Orca Agent Engine - point any OpenAI- or Anthropic-compatible client at the gateway and it works. When you do run both, Agent Engine routes its agents' MCP tool calls through the gateway automatically. See Use with Agent Engine.
Why put a gateway in front
Without a gateway, every application that calls a model carries a provider API key, picks a provider, and produces whatever telemetry its SDK happens to emit. Rotating a key means redeploying everything that holds it, switching providers means a code change in each caller, and answering "what did we spend, and on what" means correlating logs from services that never agreed on a format.
The gateway makes those cross-cutting concerns configuration instead of code:
- Credentials stay out of callers. A destination names a vault; the gateway resolves the credential at call time and injects it upstream. Callers authenticate to the gateway, not to OpenAI.
- Provider choice becomes a route. Weighted load balancing, sequential fallback, and conditional routing are config, so failing over from one provider to another does not touch application code.
- One telemetry shape. Spans follow the OpenTelemetry GenAI semantic conventions regardless of which provider served the request.
- One enforcement point. Identity, authorization, rate limits, budgets, and input guardrails run before traffic reaches an upstream. Output guardrails inspect non-streaming responses after provider dispatch.
What it handles
| Traffic | Endpoints | Status |
|---|---|---|
| Model calls | /v1/chat/completions, /v1/embeddings, /v1/messages, /v1/responses | Implemented |
| MCP tool calls | POST / GET /v1/mcp, GET /v1/mcp/sse | Implemented |
| Native API-key proxy | /v1/proxy/{provider}/* | Implemented for explicitly configured native_api_key destinations and allowed paths |
| Model discovery | /v1/models | Stub - returns an empty list |
Six provider adapters are implemented: openai, anthropic, azure_openai, openai_compatible,
bedrock, and vertex. Streaming is supported on all model endpoints.
Agent-session endpoints (/v1/agent/sessions) return 501. To govern agent traffic, run
agents on Agent Engine, which routes their MCP tool calls through the gateway.
How a request flows
- An auth validator turns the caller's credential into a principal carrying a scope - a
set of deployer-chosen dimensions such as
workspace_idandenv. No part of the gateway hard-codes what a tenant is. - Authorizers decide whether that principal may make this call.
- Rate limits apply per scope: request and token windows, plus calendar spend budgets.
- Guardrails inspect the request, and later the response, and can redact or deny.
- A route matches the request and its strategy picks a destination, with retry and fallback.
- The vault resolves that destination's credential, and the provider adapter translates the call to the provider's native shape.
- Trace exporters, usage sinks, and audit sinks record what happened.