Incident 2026-09-04 ~01:05Z: after a background/foreground cycle the phone's
relay dial timed out inside the cell's acceptClient DB phase while the fleet
was in a cell-inventory lock storm (55P03 retries ~7.5k/h vs a ~1k/h floor).
Relay (cloud/apps/relay)
- acceptClient checks socket.readyState after each serialized Postgres call
and abandons the accept once the phone has hung up, releasing the capacity
reservation, failing the credential reservation, and releasing the activity
lease it just acquired instead of leaking it to expiry cleanup and then
throwing host_data_reservation_already_bound at bind.
- New structured event orca_relay_client_accept_abandoned {stage, elapsedMs}
and runtime-metric fields clientAcceptsAbandonedByStageDelta /
clientAcceptAbandonedMsMax so the "phone gave up behind the lock" rate is
quantifiable per cell.
- Control lease grants are jittered: 55 min minus [0, 10 min). Every host
that (re)connected in the same minute rebound as one cohort every ~54 min
(c27 autoheal recreate at 23:23Z re-homed ~420 controls; ~1.1k 1006 +
~1k 4408 "control rebound" closes landed in a 3 s window at 00:50:14Z),
and each rebind is an activateControl transaction on the inventory lock.
No wire change: leaseExpiresAt was always a server-chosen absolute time.
Desktop (src/main/runtime/relay)
- Control rotation rebinds 1-6 min early instead of 1-2 min, so a re-homed
cohort spreads across cycles rather than pinning one phase for the life of
the process.
Phone (mobile/src/transport)
- openAuthenticatedDirectEndpoint treats 'reconnecting' as a failed probe. On
a dead LAN the foreground direct dial dies with an instant 1006 and the
direct client enters its own 500/1000/2000 ms backoff; the probe used to
wait out its full 12 s bound holding the supervisor mutex, so relay recovery
queued behind three doomed redials. The stage-aware bound from #18518 is
unaffected.
Not done here: rolling the 23 GCE cells onto the post-#18521 image (500 ms
lock_timeout) is a deploy owned by cloud-deploy-relay-production-same-cap.
Orca Relay
The relay that connects the Orca mobile app to a desktop host. Phones and desktops never talk to each other directly: each opens an outbound WebSocket to a relay cell, the relay pairs the two sessions, and it splices frames between them. A director assigns hosts to cells and coordinates migrations; cells carry the user connections.
This directory is an independent pnpm workspace inside the Orca monorepo. Run
its commands from cloud/, not the repository root. The source is covered by
the repository's root MIT license.
Packages
packages/relay-contract: the wire contract shared by the relay, the desktop app, and the mobile app (frame shapes, close codes, admission budgets, splice state machine).apps/relay: the relay server. The same image runs as a director or a cell depending onORCA_RELAY_ROLE.apps/relay-fence-broker: a private, IAM-only service that owns the durable mutation lease, the Terraform checkout, and the narrow Compute mutation used when a registered target is superseded. The workflow that calls it holds read and invoke rights only, never those mutation permissions.apps/relay-ops: the relay operations console and the incident monitor behindpnpm ops:relay,pnpm incident:relay, andpnpm incident:relay-preflight.
Infrastructure and operations
infra/terraform: the relay Terraform root. It owns the cells, the director, the shared Cloud SQL instance, DNS, observability, and every GitHub Workload Identity provider the relay workflows authenticate through.backend/holds the per-environment backend configuration andenvironments/the tfvars. Drive it throughpnpm infra:init,pnpm infra:plan, andpnpm infra:apply.dev/scripts: the deploy, capacity, admission, rehome, monitoring, and load scripts the workflows call, plus the contract tests that pin each workflow and Terraform surface. Run them withpnpm test.dev/contractsanddev/fixtures: the checked-in data those contract tests read, including the Terraform root partition.docs/: the relay runbooks, capacity-testing guide, incident-monitor reference, and the workflow variable reference indocs/relay-workflows.md.
Workflows
The 24 .github/workflows/cloud-*.yml workflows are the relay's deploy and
operate surface: publish and deploy the director, roll GCE cell capacity,
operate Asia admission and regional rehoming, prove staging capacity, monitor
production, and power staging up and down. .github/actions/cloud-sql-rollout-lease
is the compare-and-swap lease that serializes every rollout against the shared
Cloud SQL instance.
Every one of them is inert. Each top-level job is gated on
vars.ORCA_CLOUD_OPERATIONS_ENABLED == 'true', a repository variable that is
unset here, so the two scheduled triggers and every manual dispatch skip
without running a step. Only the repository owner, holding the GCP identities
these workflows authenticate as, can turn them on.
Cloud Verify is not gated. It builds, typechecks, lints, tests, secret-scans,
and validates the relay Terraform on every change under cloud/, and it runs
on fork pull requests, so it configures no backend and holds no credential.
What is not here
The terraform-foundation and terraform-apps roots and the API and auth
services live in the private stablyai/orca-cloud repository. Scripts and
tests that spanned both trees were narrowed to the relay side rather than
carrying a dangling reference.
Local development
cd cloud
pnpm install
pnpm build
pnpm test
pnpm test runs the SQLite-backed suites. Tests that need PostgreSQL run only
when ORCA_RELAY_TEST_POSTGRES_URL points at a disposable PostgreSQL 16 or 17
database, for example:
docker run --rm -d --name orca-relay-pg -e POSTGRES_HOST_AUTH_METHOD=trust \
-e POSTGRES_DB=orca_relay_test -p 55440:5432 postgres:16-alpine
ORCA_RELAY_TEST_POSTGRES_URL=postgres://postgres@127.0.0.1:55440/orca_relay_test \
pnpm --filter @orca-cloud/relay test
docker rm -f orca-relay-pg
Configuration is read from environment variables validated in
apps/relay/src/config.ts. ORCA_RELAY_ASSIGNMENT_SIGNING_KEY (at least 32
bytes) is the only required value; everything else has a local default.