Files
orca/cloud/apps/relay-ops
Jinwoo Hong 07dad6739a refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run (#24443)
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run

A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed,
single-use, five-minute-fresh evidence. Each apply wave now samples fleet health
itself right before isolation, with the monitor's evaluator, thresholds, and
tolerances, for a window sized to the cell's host count (3/5/8 min), plus three
lookback rules: no cell container exit in 10 min, no minute over 500 director
503s in 10 min, and director concurrency p99 within the monitor bar over 4 min.

Removes the monitor-run inputs, the gate's consume/authorize steps, the
break-glass override, and the same-cap-only authorization shapes in
relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are
unchanged.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): bound the pre-drain sample overrun and keep the drain token fresh

Review follow-ups: alternating tolerated readings could hold the sample open
until its step timeout, so cap the overrun at three samples past the window;
record why a read failed; mint a fresh admin ID token for the drain after the
sample; raise the job timeout to 90 min so a long sample cannot cancel the
job past the failsafe.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule

The exit rule counted every relay container exit fleet-wide, so a cell that
crashes every few hours (c25, 12 a week) blocked the very roll that fixes it,
and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part
in. Exits are now grouped by instance, each instance is named by its own newest
runtime-metrics log line, and only exits on general or migration-only cells
other than the target count. An exit no configured cell can be named for trips
the rule; a failed lookup is a failed read. relay-observability.tf joins the
evidence-code set because the rule depends on its filter.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 15:36:17 -04:00
..

Orca Relay Operations

A private, aggregate dashboard for the Orca Relay control and data planes. It reads local gcloud and gh credentials on the server; credentials and per-user Relay state never enter the browser. One cached gcloud auth print-access-token refresh feeds concurrent read-only Google APIs so the collector does not stampede the local credential store.

Run locally

Prerequisites:

  • Node 24 and pnpm 10
  • gcloud authenticated for onorca-cloud and onorca-cloud-staging
  • gh authenticated with read access to stablyai/orca-cloud

From the repository root:

pnpm install
pnpm ops:relay

Open http://127.0.0.1:2455. The server binds only to loopback and refreshes aggregate data every minute. Production and staging are read-only by default.

The cost panel is a labeled planning estimate. There is currently no Cloud Billing export in either project, so the dashboard cannot claim exact billed spend. GCP Billing remains authoritative.

Share through Tailscale

Keep the dashboard bound to loopback and let Tailscale provide identity, TLS, and tailnet ACL enforcement:

tailscale serve --bg http://127.0.0.1:2455
tailscale serve status

Share the HTTPS URL printed by tailscale serve status with the team. Limit access to the intended operator group in the tailnet ACL. Do not use a public funnel. Stop sharing with:

tailscale serve reset

For a persistent host, run pnpm --filter @orca-cloud/relay-ops build and supervise pnpm --filter @orca-cloud/relay-ops start with the host's normal process manager. The process needs the same non-interactive gcloud and gh identities.

Optional staging controls

Controls are intentionally local-only and disabled unless explicitly enabled:

RELAY_OPS_ENABLE_STAGING_CONTROLS=1 pnpm ops:relay

Even in this mode the service never changes GCP directly. It dispatches .github/workflows/power-relay-staging.yml, preserves the workflow's typed WAKE_STAGING / SLEEP_STAGING confirmation, and always wakes only configured-admission cells. Requests require the loopback origin and a per-process CSRF token, so controls stay unavailable through the Tailscale view.

Data and security boundaries

  • Browser payloads contain aggregate Monitoring points, resource health, immutable image digests, alert-policy metadata, and workflow metadata.
  • Account IDs, host IDs, device IDs, pairing state, assignment rows, bearer tokens, service-account tokens, startup scripts, secret values, and individual Relay-admin state are excluded.
  • Sleeping staging is inventory-only. Viewing it does not probe or cold-start Cloud Run services and cannot resize empty MIGs.
  • Partial GCP or GitHub failures degrade the affected panel and produce a sanitized warning.
  • Missing cell inventory renders as Unknown, never Sleeping. After one successful read, transient credential or collector failures retain the last good snapshot and mark it stale.
  • Responses use no-store, a restrictive CSP, frame denial, and no-referrer headers.

Verification

pnpm --filter @orca-cloud/relay-ops test
pnpm --filter @orca-cloud/relay-ops typecheck
pnpm --filter @orca-cloud/relay-ops build