Files
orca/cloud
Jinwoo Hong 61b09b7a02 fix(relay): abandon dead client accepts, jitter and lengthen the control lease, fail direct probes fast (#18959)
* fix(relay): abandon a client accept once the phone hangs up; jitter the control lease

The accept runs several serialized Postgres calls behind the contended
cell-inventory lock, and phones bound their dial. Finishing that work for a
phone that had already left acquired (and leaked for 90s) an activity lease and
then failed at bind with host_data_reservation_already_bound. Check the client
socket between the DB steps and unwind what was taken, reporting the stage on
orca_relay_client_accept_abandoned.

Jitter the control lease grant so a cohort that reconnected in the same minute
(a cell recreate dumps hundreds at once) walks apart instead of rebinding
together every cycle.

On the phone, treat a probe session that enters 'reconnecting' as a failed
probe: it is the direct client's own backoff after a dead-LAN 1006, and waiting
it out held the supervisor's operation mutex for the full 12s bound.

* perf(relay): lengthen the control lease to 6h

The lease bounds how long a host lingers on a cell after a missed drain, and
rebinding it is the only passive rebalancing we have, so it stays finite. 6h
keeps both properties while cutting control-activation traffic on the contended
cell-inventory lock ~6x. The relay JWT (5 min, refreshed by the desktop) and the
75s silence watchdog are enforced separately, so the longer grant authorizes
nothing extra. The jitter widens with it, to +/-30 min.

* fix(relay): let one flap recover the direct probe; correct the leak window

'reconnecting' is published on any socket close, so rejecting on it outright
turned a single access-point flap into a booked direct failure and a 60s
cooldown. Give the first 'reconnecting' a 2s grace in which a 'connected'
transition still resolves; a dead LAN still fails in ~2s rather than holding the
supervisor's operation mutex for the 12s bound.

The abandoned accept held its activity lease for the 10s attach deadline, not
90s -- the attach timer is armed before bind throws and already unwinds it.

Also cover the assignment-stage check that guards reserveCredential, and drop a
spread assertion the two exact-value assertions above already imply.

* fix(relay): extend the probe grace once on a handshake; pin the lease band top

The redial fires at 500ms but 'connected' waits on the Noise handshake and a
capability RPC, so one 2s window is too tight for real work. A 'handshaking'
transition is evidence the peer answered, so extend the grace once; a stalled
handshake still fails at ~3.5s, far inside the 12s bound.

The longest-lease case only had an upper bound, which a jitter clamped to one
side would satisfy. Pin it to the exact top of the band instead, and assert the
assignment resolve ran so the third-guard test cannot pass vacuously.
2026-09-05 20:47:27 -04:00
..

Orca Relay

The relay that connects the Orca mobile app to a desktop host. Phones and desktops never talk to each other directly: each opens an outbound WebSocket to a relay cell, the relay pairs the two sessions, and it splices frames between them. A director assigns hosts to cells and coordinates migrations; cells carry the user connections.

This directory is an independent pnpm workspace inside the Orca monorepo. Run its commands from cloud/, not the repository root. The source is covered by the repository's root MIT license.

Packages

  • packages/relay-contract: the wire contract shared by the relay, the desktop app, and the mobile app (frame shapes, close codes, admission budgets, splice state machine).
  • apps/relay: the relay server. The same image runs as a director or a cell depending on ORCA_RELAY_ROLE.
  • apps/relay-fence-broker: a private, IAM-only service that owns the durable mutation lease, the Terraform checkout, and the narrow Compute mutation used when a registered target is superseded. The workflow that calls it holds read and invoke rights only, never those mutation permissions.
  • apps/relay-ops: the relay operations console and the incident monitor behind pnpm ops:relay, pnpm incident:relay, and pnpm incident:relay-preflight.

Infrastructure and operations

  • infra/terraform: the relay Terraform root. It owns the cells, the director, the shared Cloud SQL instance, DNS, observability, and every GitHub Workload Identity provider the relay workflows authenticate through. backend/ holds the per-environment backend configuration and environments/ the tfvars. Drive it through pnpm infra:init, pnpm infra:plan, and pnpm infra:apply.
  • dev/scripts: the deploy, capacity, admission, rehome, monitoring, and load scripts the workflows call, plus the contract tests that pin each workflow and Terraform surface. Run them with pnpm test.
  • dev/contracts and dev/fixtures: the checked-in data those contract tests read, including the Terraform root partition.
  • docs/: the relay runbooks, capacity-testing guide, incident-monitor reference, and the workflow variable reference in docs/relay-workflows.md.

Workflows

The 24 .github/workflows/cloud-*.yml workflows are the relay's deploy and operate surface: publish and deploy the director, roll GCE cell capacity, operate Asia admission and regional rehoming, prove staging capacity, monitor production, and power staging up and down. .github/actions/cloud-sql-rollout-lease is the compare-and-swap lease that serializes every rollout against the shared Cloud SQL instance.

Every one of them is inert. Each top-level job is gated on vars.ORCA_CLOUD_OPERATIONS_ENABLED == 'true', a repository variable that is unset here, so the two scheduled triggers and every manual dispatch skip without running a step. Only the repository owner, holding the GCP identities these workflows authenticate as, can turn them on.

Cloud Verify is not gated. It builds, typechecks, lints, tests, secret-scans, and validates the relay Terraform on every change under cloud/, and it runs on fork pull requests, so it configures no backend and holds no credential.

What is not here

The terraform-foundation and terraform-apps roots and the API and auth services live in the private stablyai/orca-cloud repository. Scripts and tests that spanned both trees were narrowed to the relay side rather than carrying a dangling reference.

Local development

cd cloud
pnpm install
pnpm build
pnpm test

pnpm test runs the SQLite-backed suites. Tests that need PostgreSQL run only when ORCA_RELAY_TEST_POSTGRES_URL points at a disposable PostgreSQL 16 or 17 database, for example:

docker run --rm -d --name orca-relay-pg -e POSTGRES_HOST_AUTH_METHOD=trust \
  -e POSTGRES_DB=orca_relay_test -p 55440:5432 postgres:16-alpine
ORCA_RELAY_TEST_POSTGRES_URL=postgres://postgres@127.0.0.1:55440/orca_relay_test \
  pnpm --filter @orca-cloud/relay test
docker rm -f orca-relay-pg

Configuration is read from environment variables validated in apps/relay/src/config.ts. ORCA_RELAY_ASSIGNMENT_SIGNING_KEY (at least 32 bytes) is the only required value; everything else has a local default.