Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Metrics and traces

Prometheus metrics and OpenTelemetry GenAI spans emitted by Orca AI Gateway.

Prometheus metrics

The admin plane serves metrics at /metrics with no configuration required. In Kubernetes, the Helm chart sets scrape annotations by default and can create a ServiceMonitor:

values.yaml
metrics:
  serviceMonitor:
    enabled: true

Metrics cover request rates and latencies, per-stage durations across the pipeline, retry and fallback counts, sink drops, and spend-control behavior. Because they are per replica, aggregate across the fleet before alerting on them. Cache stages currently record timing only; response cache lookup and storage are not wired.

OpenTelemetry traces

Spans follow the OpenTelemetry GenAI semantic conventions directly - there is no translation shim, so any GenAI-aware backend understands them without custom mapping.

Configuration

plugins:
  trace_exporters:
    - name: otlp
      kind: otlp
      endpoint: "http://otel-collector.svc:4317"
      transport: grpc
      service_name: orca-gateway

transport accepts grpc, http_protobuf, or http_json. Multiple exporters can run at once. The current server does not read per-exporter sampling, timeout, queue, or batch-size fields.

Span structure

The chat, embeddings, and Messages handlers emit one completed model span per request. The current export path does not emit a gateway root, stage child spans, per-attempt spans, or MCP spans. Retry and fallback details remain available through metrics and logs.

Attributes

Standard GenAI attributes on the gen_ai.client span:

AttributeValue
gen_ai.systemConfigured provider kind, such as openai, anthropic, bedrock, vertex, or azure_openai
gen_ai.request.modelThe model id the client asked for
gen_ai.response.modelThe model the provider actually served, after model_map
gen_ai.usage.input_tokens, gen_ai.usage.output_tokensToken counts

The span also carries two agent-trace attributes:

AttributeValue
contributor.typeai for current gateway-emitted spans
model_idProvider and model slug, such as openai/gpt-5

Use usage records for route, destination, principal, and verified-scope attribution.

Trace context

The gateway parses the trace-id segment of an inbound W3C traceparent and reuses that trace id for its model span. It does not preserve the inbound parent span id or tracestate, and the exported model span has no parent. A request without a valid traceparent gets a new trace id.

On this page