Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Model providers

Configure model-provider destinations and credentials in Orca AI Gateway.

A model provider destination translates the gateway's OpenAI-compatible request into the provider's protocol. Configure a destination, a compatible vault credential, and a route that selects it.

Choose an adapter

kindRequired provider fieldsCredential schemeEmbeddings
openainonebearerYes
openai_compatibleSet base_url to the compatible servicebearerYes
anthropicnoneapi_keyNo
azure_openaibase_url; optional deploymentapi_key or bearerYes
bedrockoptional region, default us-east-1aws-sig-v4No
vertexproject and location; optional publishergcp-service-accountNo

The OpenAI and Anthropic adapters use their public API roots by default. Azure OpenAI defaults api_version to 2024-10-21. Bedrock derives its runtime endpoint from region, and Vertex AI derives its endpoint from location.

Configure OpenAI or a compatible endpoint

The vault's scheme must match the destination adapter:

gateway.yaml
destinations:
  openai-primary:
    kind: openai
    credentials:
      vault: openai-key

  local-openai:
    kind: openai_compatible
    base_url: http://model-server:8000
    credentials:
      vault: local-key

vaults:
  - name: openai-key
    resolver: env
    env_var: OPENAI_API_KEY
    scheme: bearer
  - name: local-key
    resolver: env
    env_var: LOCAL_MODEL_TOKEN
    scheme: bearer

The gateway appends the provider path to base_url. Supply the service root, not /v1/chat/completions or /v1/embeddings.

Configure Anthropic

gateway.yaml
destinations:
  anthropic-primary:
    kind: anthropic
    credentials:
      vault: anthropic-key

vaults:
  - name: anthropic-key
    resolver: env
    env_var: ANTHROPIC_API_KEY
    scheme: api_key

Clients can call either the OpenAI-compatible /v1/chat/completions endpoint or the Anthropic compatibility endpoint at /v1/llm/v1/messages. The adapter does not implement embeddings.

Configure Azure OpenAI

gateway.yaml
destinations:
  azure-prod:
    kind: azure_openai
    base_url: https://example.openai.azure.com
    deployment: gpt-4o-prod
    api_version: "2024-10-21"
    credentials:
      vault: azure-key

vaults:
  - name: azure-key
    resolver: env
    env_var: AZURE_OPENAI_API_KEY
    scheme: api_key

When deployment is absent, the adapter uses the mapped request model as the deployment name. Use scheme: bearer instead when the secret is an Azure bearer token.

Configure Bedrock

gateway.yaml
destinations:
  bedrock-claude:
    kind: bedrock
    region: us-west-2
    credentials:
      vault: aws-creds
    model_map:
      claude-sonnet: anthropic.claude-3-sonnet-20240229-v1:0

vaults:
  - name: aws-creds
    resolver: env
    env_var: BEDROCK_CREDENTIALS
    scheme: aws-sig-v4

Set BEDROCK_CREDENTIALS to access_key:secret_key or access_key:secret_key:session_token.

Configure Vertex AI

gateway.yaml
destinations:
  vertex-gemini:
    kind: vertex
    project: my-gcp-project
    location: us-central1
    publisher: google
    credentials:
      vault: gcp-service-account

vaults:
  - name: gcp-service-account
    resolver: env
    env_var: GCP_SERVICE_ACCOUNT_JSON
    scheme: gcp-service-account

Set GCP_SERVICE_ACCOUNT_JSON to the complete service-account JSON document.

Bedrock and Vertex AI require credentials from the configured vault. The adapters do not use the AWS default credential chain, IRSA, Google Application Default Credentials, or GCP Workload Identity.

Map a stable model name

model_map rewrites the model before the provider request. This lets a route keep one public model name while changing its provider-native target:

gateway.yaml
destinations:
  azure-prod:
    kind: azure_openai
    base_url: https://example.openai.azure.com
    credentials: { vault: azure-key }
    model_map:
      support-model: gpt-4o-prod

A request for support-model is sent to Azure as gpt-4o-prod. Pricing lookups and usage records retain the destination and provider context needed for attribution.

Verify the configuration

Validate and start the gateway
orca-gateway check gateway.yaml
orca-gateway run --config gateway.yaml

Then call the data plane:

Send a chat request
curl http://localhost:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"support-model","messages":[{"role":"user","content":"Hello"}]}'

On this page