Metrics and traces
Prometheus metrics and OpenTelemetry GenAI spans emitted by Orca AI Gateway.
Prometheus metrics
The admin plane serves metrics at /metrics with no configuration required. In Kubernetes, the Helm
chart sets scrape annotations by default and can create a ServiceMonitor:
metrics:
serviceMonitor:
enabled: trueMetrics cover request rates and latencies, per-stage durations across the pipeline, retry and fallback counts, sink drops, and spend-control behavior. Because they are per replica, aggregate across the fleet before alerting on them. Cache stages currently record timing only; response cache lookup and storage are not wired.
OpenTelemetry traces
Spans follow the OpenTelemetry GenAI semantic conventions directly - there is no translation shim, so any GenAI-aware backend understands them without custom mapping.
Configuration
plugins:
trace_exporters:
- name: otlp
kind: otlp
endpoint: "http://otel-collector.svc:4317"
transport: grpc
service_name: orca-gatewaytransport accepts grpc, http_protobuf, or http_json. Multiple exporters can run at once.
The current server does not read per-exporter sampling, timeout, queue, or batch-size fields.
Span structure
The chat, embeddings, and Messages handlers emit one completed model span per request. The current export path does not emit a gateway root, stage child spans, per-attempt spans, or MCP spans. Retry and fallback details remain available through metrics and logs.
Attributes
Standard GenAI attributes on the gen_ai.client span:
| Attribute | Value |
|---|---|
gen_ai.system | Configured provider kind, such as openai, anthropic, bedrock, vertex, or azure_openai |
gen_ai.request.model | The model id the client asked for |
gen_ai.response.model | The model the provider actually served, after model_map |
gen_ai.usage.input_tokens, gen_ai.usage.output_tokens | Token counts |
The span also carries two agent-trace attributes:
| Attribute | Value |
|---|---|
contributor.type | ai for current gateway-emitted spans |
model_id | Provider and model slug, such as openai/gpt-5 |
Use usage records for route, destination, principal, and verified-scope attribution.
Trace context
The gateway parses the trace-id segment of an inbound W3C traceparent and reuses that trace id for
its model span. It does not preserve the inbound parent span id or tracestate, and the exported
model span has no parent. A request without a valid traceparent gets a new trace id.