Best practices
Guidance for building against Orca Agent Engine - credentials, session lifecycle, idempotency, pagination, streaming, retries, and environment separation.
This page collects practices for building reliably against Orca Agent Engine, grounded in how the TypeScript SDK and the ork CLI actually behave. The examples use the SDK, but the principles apply equally to direct Registry API calls.
Handle credentials and tokens safely
Keep tokens out of source and out of browsers. Both tools read credentials from the environment by default - the SDK from ORCA_API_KEY and ORCA_BASE_URL, the CLI from ORCA_ACCESS_TOKEN and ORCA_REGISTRY_URL. Set them from your secret manager or CI secrets; never commit them. The SDK refuses to run in a browser unless you set dangerouslyAllowBrowser: true, because doing so leaks your token to end users - keep the client server-side.
The SDK and CLI use different environment variable names for the same two values. The SDK reads ORCA_API_KEY / ORCA_BASE_URL; the CLI reads ORCA_ACCESS_TOKEN / ORCA_REGISTRY_URL. Point both base URLs at the same per-Workspace registry endpoint and use the same service-account token. If credentials seem to work in one tool but not the other, check which variable names each one reads.
Rotate tokens with an async provider. Rather than constructing a client with a static string that expires, pass a function for apiKey. The SDK awaits it on every request, so a refreshed token takes effect immediately with no reconstruction:
const orca = new Orca({
apiKey: async () => getTokenFromSecretManager(), // must return a non-empty string
baseURL: process.env.ORCA_BASE_URL,
});Store external-service credentials in vaults, not in agent definitions. When an agent needs to authenticate to a third-party service, put the secret in a vault credential and reference the vault from the session, rather than embedding secrets in the agent's configuration or system prompt. Use orca.vaults.credentials.validate(vaultId, credentialId) to confirm an MCP OAuth credential is configured correctly before a session relies on it.
Manage the session lifecycle deliberately
Archive instead of delete unless you mean it. Sessions, environments, vaults, and memory stores expose both archive and delete. Agents expose only archive - there is no orca.agents.delete and no CLI delete command for them, so archiving is the only way to retire one. Archiving is reversible and keeps the object and its history available (list accepts include_archived to surface archived items); deleting is permanent. Prefer archive for anything you might want to inspect or audit later, and reserve delete for true cleanup:
await orca.sessions.archive(session.id); // reversible, retains history
await orca.sessions.delete(session.id); // permanentDrive a session to a terminal state, then stop reading. A session moves through statuses as it runs. When streaming events, treat session.status_idle as the signal that the agent has finished its turn, and break out of the loop - otherwise the stream stays open waiting for more events:
for await (const event of await orca.sessions.events.stream(session.id)) {
if (event.type === 'session.status_idle') break;
}Use the session handle for single-session work. When you operate on one session repeatedly, orca.session(sessionId) removes the repeated ID argument and keeps call sites short and less error-prone. See Session handle.
Make writes idempotent
Use optimistic concurrency on agent updates. orca.agents.update takes an optional version, and the concurrency check only happens when you send it - omit it and a blind write wins silently. Always send it. Read the agent (or keep the version from the last write), pass that version, and handle a ConflictError (409) by re-reading and retrying - this prevents two writers from silently clobbering each other:
import { ConflictError } from '@runorca/orca-sdk';
async function setDescription(agentId: string, description: string) {
for (let attempt = 0; attempt < 3; attempt++) {
const agent = await orca.agents.retrieve(agentId);
try {
return await orca.agents.update(agentId, { version: agent.version, description });
} catch (err) {
if (err instanceof ConflictError) continue; // someone else updated; re-read and retry
throw err;
}
}
throw new Error('Exhausted retries updating agent');
}Treat skill and version uploads as additive. Skills are versioned. To change a skill, upload a new version with orca.skills.versions.create rather than trying to mutate one in place. Agents that pin a specific version keep using it; agents that omit a version pick up the new one on the next session start.
Make retried creates safe. Because the client retries transient failures automatically (see below), design create flows so a retry can't produce duplicate side effects you care about - for example, key off a natural identifier and reconcile, or check for an existing object before creating.
Page through lists, never assume one page
Iterate, do not read one page. Every list method is paginated, and reading only .data from the first page silently misses everything past the page limit. The returned cursor is async-iterable and fetches each next page for you:
for await (const agent of orca.agents.list()) {
// visits every agent across every page
}When you genuinely want one page - to render a UI table, say - read .data and check .next_page to decide whether to offer another:
const page = await orca.agents.list({ limit: 50 });
render(page.data);
if (page.next_page) showLoadMoreButton();Filter server-side, with a filter the resource actually has. Filtering in the registry is faster than fetching everything and filtering in memory, but each resource accepts its own set and an unrecognized query parameter is ignored rather than rejected - you get a 200 and a full, unfiltered page. Over HTTP, sessions accept agent_id, agent_version, statuses, memory_store_id, deployment_id, include_archived, order, and created_at[gte] / created_at[lte]; agents accept include_archived and the same created_at bounds, but not agent_id. The SDK types a narrower set - SessionListParams is agent_id, limit, page, and include_archived - so statuses and the rest are a compile error there and need a direct request. Check the operation in the API reference before relying on a filter.
Stream resiliently
Consume a stream once, and split with tee() if you need two readers. A Stream can be iterated only once. If two parts of your code need the same events, call .tee() to get two independent streams rather than iterating the same one twice.
Release streams you stop reading. If you break out of a stream loop early, the underlying connection should be released. Calling stream.controller.abort() cancels the request explicitly - do this when a user navigates away or a deadline passes, so connections don't leak:
const stream = await orca.sessions.events.stream(session.id);
const deadline = Date.now() + 30_000;
for await (const event of stream) {
if (event.type === 'session.status_idle') break;
if (Date.now() > deadline) {
stream.controller.abort();
break;
}
}Reconnect by replaying from persisted events. If a long-lived stream drops, the durable record of the session is its event history. Reconnect by re-opening the stream, and reconcile against orca.sessions.events.list(sessionId) to recover any events you may have missed while disconnected - the persisted list is the source of truth, the stream is the live view.
Handle errors and retries intentionally
Catch specific error subclasses. Every SDK error extends OrcaError; HTTP errors extend APIError with .status, .headers, and .error. Branch on the specific subclass so each failure gets the right response - a NotFoundError is a permanent condition to handle, a RateLimitError is transient:
import { APIError, NotFoundError, RateLimitError } from '@runorca/orca-sdk';
try {
await orca.sessions.retrieve(id);
} catch (err) {
if (err instanceof NotFoundError) return null; // gone - don't retry
if (err instanceof RateLimitError) return scheduleRetry(err.headers); // back off
if (err instanceof APIError) logApiFailure(err.status, err.message);
throw err;
}Lean on built-in retries, and bound them per request. The client already retries transient failures (network errors, 408, 409, 429, 5xx) with exponential backoff and jitter - you usually don't need your own retry loop on top. Tune maxRetries and timeout on the client, and override per request for operations with different latency or idempotency profiles:
// A cheap read you want to fail fast, with no retries
await orca.agents.list({}, { maxRetries: 0, timeout: 5_000 });Set a realistic timeout for streaming work. The default per-request timeout is 10 minutes. Long agent runs are expected, so don't lower the timeout so far that legitimate sessions are cut off; instead enforce your own deadline in the stream loop and abort() when it passes.
Separate environments cleanly
Use distinct Workspaces or environments per stage. A Workspace scopes registry resources and has its own registry endpoint, on StreamNative Cloud and on a self-hosted deployment alike. Keep development, staging, and production in separate Workspaces (each with its own baseURL and token), or at minimum separate Agent Engine environments per stage, so a session never runs against the wrong configuration.
Make the target explicit in configuration, not in code. Select the stage by setting ORCA_BASE_URL / ORCA_API_KEY (SDK) or ORCA_REGISTRY_URL / ORCA_ACCESS_TOKEN (CLI) from your deployment configuration. Avoid hardcoding endpoints or branching on stage in application code - a single misconfigured variable should be the only thing that ever points the client at the wrong place, and it should be visible in your environment, not buried in source.
Scope tokens to the Workspace they serve. Use a separate service-account token per Workspace with only the permissions it needs, so a leaked or misused development token can't reach production data.