* fix(relay): stop taking the fleet-wide cell inventory lock on per-connection paths activateControl, acquireActivity, changeActivity and removeSupersededSameCellControls each adjust exactly one cell's reservation, yet took SELECT * FROM relay_cells FOR UPDATE, so every desktop rebind and phone reconnect in the fleet queued behind every other one and behind placement. They now use the single-row atomic update (or lock only their own cell row), leaving the inventory lock to placement and sweeps. Fleet-wide 55P03 retries ran p50 430 / p99 1320 per five minutes on 2026-09-03, every cell pinned sqlLatencyMsMax at the lock timeout, and the old cell image crashed on the resulting pool timeouts ~every 15 minutes. A real-Postgres test holds another cell's row and asserts a rebind proceeds; re-adding the inventory lock fails it. * fix(relay): lock the touched cell rows in order on cross-cell activity moves Review found that acquireActivity's existing-lease branch could lock the old lease's cell row (via removeActivityLease) before the new cell's row, which cycles with placement's ascending inventory lock; reproduced on real Postgres as paired 55P03 retries. lockCellRows now takes the one or two rows a per-connection path touches in cell_id order with the 500 ms request bound, and the census fails on any inline relay_cells FOR UPDATE outside the named lock helpers. A three-cell Postgres test moves an activity from the highest cell to a lower one while the target row is held and asserts the mover holds nothing else; five revert-mutants (inventory lock on each path, dropped ordering, dropped ORDER BY) fail it. * test(relay): make the inline relay_cells lock census scan whole statements Review showed two evasions: a FOR UPDATE inside query() and a queryLocked whose FROM relay_cells sat past a fixed line window. The guard now matches every query()/queryLocked() template statement in full; both evasions fail it. Also clears relay_cell_connection_snapshots in the connection- headroom Postgres suite so an aborted run does not poison the next.
Orca Relay
The relay that connects the Orca mobile app to a desktop host. Phones and desktops never talk to each other directly: each opens an outbound WebSocket to a relay cell, the relay pairs the two sessions, and it splices frames between them. A director assigns hosts to cells and coordinates migrations; cells carry the user connections.
This directory is an independent pnpm workspace inside the Orca monorepo. Run
its commands from cloud/, not the repository root. The source is covered by
the repository's root MIT license.
Packages
packages/relay-contract: the wire contract shared by the relay, the desktop app, and the mobile app (frame shapes, close codes, admission budgets, splice state machine).apps/relay: the relay server. The same image runs as a director or a cell depending onORCA_RELAY_ROLE.apps/relay-fence-broker: a private, IAM-only service that owns the durable mutation lease, the Terraform checkout, and the narrow Compute mutation used when a registered target is superseded. The workflow that calls it holds read and invoke rights only, never those mutation permissions.apps/relay-ops: the relay operations console and the incident monitor behindpnpm ops:relay,pnpm incident:relay, andpnpm incident:relay-preflight.
Infrastructure and operations
infra/terraform: the relay Terraform root. It owns the cells, the director, the shared Cloud SQL instance, DNS, observability, and every GitHub Workload Identity provider the relay workflows authenticate through.backend/holds the per-environment backend configuration andenvironments/the tfvars. Drive it throughpnpm infra:init,pnpm infra:plan, andpnpm infra:apply.dev/scripts: the deploy, capacity, admission, rehome, monitoring, and load scripts the workflows call, plus the contract tests that pin each workflow and Terraform surface. Run them withpnpm test.dev/contractsanddev/fixtures: the checked-in data those contract tests read, including the Terraform root partition.docs/: the relay runbooks, capacity-testing guide, incident-monitor reference, and the workflow variable reference indocs/relay-workflows.md.
Workflows
The 24 .github/workflows/cloud-*.yml workflows are the relay's deploy and
operate surface: publish and deploy the director, roll GCE cell capacity,
operate Asia admission and regional rehoming, prove staging capacity, monitor
production, and power staging up and down. .github/actions/cloud-sql-rollout-lease
is the compare-and-swap lease that serializes every rollout against the shared
Cloud SQL instance.
Every one of them is inert. Each top-level job is gated on
vars.ORCA_CLOUD_OPERATIONS_ENABLED == 'true', a repository variable that is
unset here, so the two scheduled triggers and every manual dispatch skip
without running a step. Only the repository owner, holding the GCP identities
these workflows authenticate as, can turn them on.
Cloud Verify is not gated. It builds, typechecks, lints, tests, secret-scans,
and validates the relay Terraform on every change under cloud/, and it runs
on fork pull requests, so it configures no backend and holds no credential.
What is not here
The terraform-foundation and terraform-apps roots and the API and auth
services live in the private stablyai/orca-cloud repository. Scripts and
tests that spanned both trees were narrowed to the relay side rather than
carrying a dangling reference.
Local development
cd cloud
pnpm install
pnpm build
pnpm test
pnpm test runs the SQLite-backed suites. Tests that need PostgreSQL run only
when ORCA_RELAY_TEST_POSTGRES_URL points at a disposable PostgreSQL 16 or 17
database, for example:
docker run --rm -d --name orca-relay-pg -e POSTGRES_HOST_AUTH_METHOD=trust \
-e POSTGRES_DB=orca_relay_test -p 55440:5432 postgres:16-alpine
ORCA_RELAY_TEST_POSTGRES_URL=postgres://postgres@127.0.0.1:55440/orca_relay_test \
pnpm --filter @orca-cloud/relay test
docker rm -f orca-relay-pg
Configuration is read from environment variables validated in
apps/relay/src/config.ts. ORCA_RELAY_ASSIGNMENT_SIGNING_KEY (at least 32
bytes) is the only required value; everything else has a local default.