Files
orca/cloud
Jinwoo Hong ed462b2caa fix(relay): re-place hosts off a draining cell without locking its row (#24216)
* test(relay): reproduce drain-release contention against director placement

Adds a Postgres harness that drives releases from an isolated cell over a
171 ms per-statement pool while five directors re-place reconnecting hosts
through the sticky lane. At 18 releases/s placements fall from 20/s to about
5/s and every active director backend is blocked on relay_cells.

Moves the per-statement delay pool into a shared test fixture so the
rehome target-row test and this harness use one implementation.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): re-place hosts off a draining cell without locking its row

Sticky re-placement of a host whose cell is isolated for a roll locked
every relay_cells row. During an Asia drain the source row is held by the
cell's own releases for a round trip each, so the placement waited on it,
the single sticky slot backed up, and /v1/assign returned 503 fleet-wide.

The isolation decision now comes from an unlocked read, and the placement
locks only the same-region general rows it can move to. It no longer writes
the source row: the host's source leases stay, and each one's own release
or expiry takes its units back off the source. The assignment keeps its
activity counters and adds one control instead of resetting them. With no
same-region headroom the path falls back to the all-rows lock, as before.

Dormant hosts hold no units, so their placement also locks only the
general rows and skips the zero write to their old cell.

The drain harness now asserts the after picture: 20 placements/s at Asia
latency with no director lock waits, against 9.2/s and 4.8/s before.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): take a moved host's units off the cell that holds them

After a narrowed re-placement a host keeps leases on its old cell while
its assignment row names the new one. Two paths charged the row's whole
counted total to the row's cell: aggregate expiry, and the lease deletion
in dead-cell and stranded re-placement. Both over-charged the new cell and
left the old cell's units stranded.

Aggregate expiry now skips hosts that still hold any lease; the lease
sweep takes each lease's units off its own cell. Placement frees each
deleted lease's units on that lease's cell, charges the old cell only for
units no lease backs, and sets the counters from the leases it keeps plus
the new control. The narrowed path runs only when the counters already
match the leases, so it never needs to write the old cell's row.

The all-rows re-placement off an isolated cell follows the same rule, so
it no longer decrements the source at placement either.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): try other regions before the all-rows lock when re-placing off a roll

With every same-region neighbour at its connection cap, the narrowed path
found no target and fell back to the all-rows lock behind the busy source
row, which is the drain brownout again. It now tries a second tier, general
cells in every other region, in its own transaction over one ordered
lockCellRows, still never the source row. Only when no general cell in any
region has room does it fall back to the all-rows path, which keeps the pin.

This changes the policy from #21911, which refused to re-place an isolated
host across a region. The unit tests that encoded that rule now assert the
tier order instead.

The drain harness gains a US cell and an arm with every Asia neighbour
capped: 200 of 200 dials placed cross-region at 20/s with no lock waits,
against 0 placed and 144 sticky rejections on the previous head. Its pass
bars are now the rejection share and the lock-waiting share, not the
placement rate a slow runner's pacing can move.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): take the narrowed path for hosts whose counters sit below their leases

Main's old placement reset a moved host's counters while keeping its source
leases, and those leases' releases floored the counters at zero. Such hosts
hold fewer counted units than lease units, and requiring equality sent them
down the all-rows path behind the busy source row.

The narrowed path now requires only that the host holds no units no lease
backs, the one case that needs a write to the old row. Its placement
already rebuilds the counters from the kept leases plus the new control,
so a drifted host is healed by its next re-placement.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): judge the local drain arm on completion and lock waits, not rate

The local-latency arm asserted at least 18 placements/s at a 20/s dial rate,
which a slow runner's pacing alone can miss. It now asserts what the Asia
arm does: no dial failures, sticky rejections under 10% of dials, every
other dial placed, and director lock-waiting under half a backend.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-30 17:00:03 -04:00
..

Orca Relay

The relay that connects the Orca mobile app to a desktop host. Phones and desktops never talk to each other directly: each opens an outbound WebSocket to a relay cell, the relay pairs the two sessions, and it splices frames between them. A director assigns hosts to cells and coordinates migrations; cells carry the user connections.

This directory is an independent pnpm workspace inside the Orca monorepo. Run its commands from cloud/, not the repository root. The source is covered by the repository's root MIT license.

Packages

  • packages/relay-contract: the wire contract shared by the relay, the desktop app, and the mobile app (frame shapes, close codes, admission budgets, splice state machine).
  • apps/relay: the relay server. The same image runs as a director or a cell depending on ORCA_RELAY_ROLE.
  • apps/relay-fence-broker: a private, IAM-only service that owns the durable mutation lease, the Terraform checkout, and the narrow Compute mutation used when a registered target is superseded. The workflow that calls it holds read and invoke rights only, never those mutation permissions.
  • apps/relay-ops: the relay operations console and the incident monitor behind pnpm ops:relay, pnpm incident:relay, and pnpm incident:relay-preflight.
  • apps/push and packages/push-contract: the mobile push gateway that holds the APNs key and sends to phones through APNs and FCM, and its wire contract. It is deployed and operated from here but is not part of the relay data path; see docs/push-gateway.md.

Mobile push gateway

apps/push is a separate Cloud Run service from the relay. Phones never hold an Orca credential for it: the desktop host authenticates with the same X25519 key it uses for the relay, answering an encrypted challenge to mint a 24 hour session, then registers each paired phone's native push token and asks the gateway to push. The gateway queues each event as its own notification, enforces per-host quotas and request limits, and retires a registration as soon as Apple or Google reports the token unregistered. Provider push is the only ordinary mobile OS-banner path. The notification socket is retained only for live dismissal and reconnect tray reconciliation; it never creates or recovers banners. Desktop notification categories remain authoritative. Each delivery is persisted as one notification event. Before deploying an incompatible queue format, stop all older push gateway revisions and clear only unpublished push delivery fixtures; no queue preservation or migration is required. FCM notification messages are inherently collapsible while offline and have a small concurrent collapse-key budget, so every pending alert is not guaranteed.

Storage follows the relay pattern: PostgreSQL in production, SQLite for tests and local development. Configure it with ORCA_PUSH_PUBLIC_URL, ORCA_PUSH_FCM_PROJECT_ID, ORCA_PUSH_DATABASE_URL, the three APNs variables (ORCA_PUSH_APNS_KEY, ORCA_PUSH_APNS_KEY_ID, ORCA_PUSH_APPLE_TEAM_ID, all three or none), and optionally ORCA_PUSH_APNS_TOPIC. The FCM credential comes from the runtime service account, so no key material is configured for Android. See push gateway operations for deployment and recovery.

Logging is aggregate counters only. Tokens, notification titles, notification bodies, and full host fingerprints never reach a log line.

Infrastructure and operations

  • infra/terraform: the relay Terraform root. It owns the cells, the director, the shared Cloud SQL instance, DNS, observability, and every GitHub Workload Identity provider the relay workflows authenticate through. backend/ holds the per-environment backend configuration and environments/ the tfvars. Drive it through pnpm infra:init, pnpm infra:plan, and pnpm infra:apply.
  • dev/scripts: the deploy, capacity, admission, rehome, monitoring, and load scripts the workflows call, plus the contract tests that pin each workflow and Terraform surface. Run them with pnpm test.
  • dev/contracts and dev/fixtures: the checked-in data those contract tests read, including the Terraform root partition.
  • docs/: the relay runbooks, capacity-testing guide, incident-monitor reference, the workflow variable reference in docs/relay-workflows.md, and the push gateway runbook in docs/push-gateway.md.

Workflows

The 25 .github/workflows/cloud-*.yml workflows are the deploy and operate surface: publish and deploy the director, roll GCE cell capacity, operate Asia admission and regional rehoming, prove staging capacity, monitor production, power staging up and down, and deploy the mobile push gateway. .github/actions/cloud-sql-rollout-lease is the compare-and-swap lease that serializes rollouts against the shared Cloud SQL instance. Push reuses that action with its own lease object and deployment concurrency group.

Every one of them is inert. Each top-level job is gated on vars.ORCA_CLOUD_OPERATIONS_ENABLED == 'true', a repository variable that is unset here, so the two scheduled triggers and every manual dispatch skip without running a step. Only the repository owner, holding the GCP identities these workflows authenticate as, can turn them on.

Cloud Verify is not gated. It builds, typechecks, lints, tests, secret-scans, and validates the relay Terraform on every change under cloud/, and it runs on fork pull requests, so it configures no backend and holds no credential.

What is not here

The terraform-foundation and terraform-apps roots and the API and auth services live in the private stablyai/orca-cloud repository. Scripts and tests that spanned both trees were narrowed to the relay side rather than carrying a dangling reference.

Local development

cd cloud
pnpm install
pnpm build
pnpm test

pnpm test runs the SQLite-backed suites. Tests that need PostgreSQL run only when ORCA_RELAY_TEST_POSTGRES_URL points at a disposable PostgreSQL 16 or 17 database, for example:

docker run --rm -d --name orca-relay-pg -e POSTGRES_HOST_AUTH_METHOD=trust \
  -e POSTGRES_DB=orca_relay_test -p 55440:5432 postgres:16-alpine
ORCA_RELAY_TEST_POSTGRES_URL=postgres://postgres@127.0.0.1:55440/orca_relay_test \
  pnpm --filter @orca-cloud/relay test
docker rm -f orca-relay-pg

Configuration is read from environment variables validated in apps/relay/src/config.ts. ORCA_RELAY_ASSIGNMENT_SIGNING_KEY (at least 32 bytes) is the only required value; everything else has a local default.