Cloud and BYOC for Orca Agent Engine are in Private Preview — request an invite
Docs

Monitoring and observability

Where to find health, throughput, usage, and logs across agents, functions, and connectors in Orca Agent Engine.

Agent Engine observability is split across sessions, runtime services, and StreamNative Cloud integration workloads. This page is the operator's index: it tells you which signals live where and points to the detailed monitoring guide for each.

The two layers to watch

Every agent runs at two layers, and each emits its own signals:

  • The agent layer - the session: conversation events, status transitions, token usage, and agent-side errors. Observed through the registry.
  • The runtime layer - the registry, harness, and sandbox services that execute the session. Observe them through your deployment's service logs and metrics.

A healthy session on an unhealthy function (or the reverse) is a real state, so watch both layers together.

Monitoring guides

Signal map

You want to checkWhere it livesGuide
What an agent did, step by stepSession events list and SSE streamSessions and agents
Session status, timing, and token usagestatus, stats, and usage fields on the SessionSessions and agents
Whether runtime services are healthyRegistry, harness, and sandbox service health and logsSessions and agents
Function status, throughput, latency, restarts, and exceptions for Pulsar or KafkaFunction status and statistics APIsFunctions
Whether a connector is running, and its lagSource, sink, or Kafka Connect status plus broker metricsSources and Kafka Connect
Registry request activityStreamNative Cloud Console - Workspace activity logSessions and agents

Alerting

Agent Engine does not ship session-level alerting today. To build alerts:

  • Tail the events stream and pattern-match session.status_rescheduled, session.error, and session.deleted into your alerting pipeline. session.status_terminated is reserved and never emitted, so nothing keyed to it will fire.
  • Scrape session usage fields on a schedule and alert on cost or token-burn anomalies.
  • For runtime-service alerts, use the registry, harness, and sandbox metrics exposed by your deployment. For Cloud integration workloads, alert on function or connector status and error counters.

Access for operators

Operators need a Workspace credential to read sessions and Cloud workload status. Treat every Workspace API key as full access to its Workspace's resources, even when an operator only reads status, and separate access with Workspaces. See Control registry access.

What's next

On this page