Files
orca/cloud
Jinwoo Hong c8a5580659 fix(relay): re-place hosts off a cell isolated for a roll (#21911)
* fix(relay): re-place hosts off a cell isolated for a roll

A roll isolates a cell by moving it out of the 'general' admission class; the
cell then refuses every attach with 4503. The director never noticed, because
the only liveness test it applies to a host's current cell reads
`relay_cell_runtime.ready` and the heartbeat, and an isolated cell keeps
heartbeating ready=1 for the whole drain. So every host on that cell was handed
its own dead cell, closed, and handed it back — 500-1,900 hosts looping for
13-16 minutes per cell roll, at ~6 dials each per minute, with no neighbour
absorbing anything.

The sticky lane now treats a live incumbent whose admission is 'migration-only'
— the state a roll's isolate step writes — the same way it treats a dead one:
it returns null, which means "fall through to placement". The placement lane
had the identical hole eleven lines further down, so it takes the same
predicate; without that second swap the sticky change is inert, because
placement would hand the pin straight back (a draining cell has more headroom
than anyone). An isolated incumbent skips the dead-cell fence branch: that
branch exists to prove an unreachable cell stopped serving a host, and this one
is reachable and enforces the epoch itself.

'existing-only' is deliberately untouched — those cells serve the hosts they
already hold, and only `assignmentStrandedOnUnservedCell` may release that pin.
A host with an open `relay_assignment_migrations` row keeps its pin too, so
this stays disjoint from the migration machinery.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): gate re-placement on a roll-isolation marker, not on admission

Review of the first commit found the predicate wrong. `migration-only` is an
admission class, not a drain signal: an Asia `--mode rollback`, an evacuation or
forward-recovery target awaiting a separate promote dispatch, a failed same-cap
wave's re-isolate, an abandoned migration retired on its target and a rehome
settlement all park loaded cells there durably, with no migration lease and no
open migration row. All five were indistinguishable from a roll's isolate, so
the first commit would have converted `operate-relay-asia-admission --mode
rollback` from a reversible admission flip into a mass move of ~4,000 hosts —
and, because `leastLoadedCell` treated region as a preference, into us-central1.

The signal is now an explicit stamp. `relay_cell_admission` gains a nullable
`roll_isolated_at`, added through the shared schema runner's catalog pre-check
so a migrated database takes no relation lock on boot and an un-migrated one
gets a catalog-only rewrite. The same-cap isolate step is its only writer, via a
new optional `rollIsolatedCells` on the selector apply; the same UPDATE that
writes the state clears the stamp whenever a cell leaves 'migration-only', so a
restore cannot leave one behind and a failed wave's re-isolate keeps the one it
has. Every other admission writer omits the field, so its cells stay unmarked
and their hosts stay pinned. Old directors ignore the field; old callers never
send it.

Region is now a constraint rather than a preference on this path only: a
re-placement must find a general, live cell with connection headroom in the
host's own region, or the pin is kept and one
`orca_relay_sticky_replacement_deferred` event is logged. Cross-region spill is
no longer reachable here.

The fence bypass is narrowed to a live incumbent. It was always a no-op for the
intended case, and for a stamped cell that stops heartbeating while still
holding sockets it reopened split-brain; that cell now takes the dead-cell path
unchanged.

Also: the hot-path admission reader no longer throws on an unrecognised state —
it sits on every sticky dial and the rule it feeds is "move the host", so an
unreadable row has to mean "don't". And the sticky lane reads the admission row
once for both the stranded rule and the stamp instead of twice.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): emit the re-placement events after the transaction commits

CodeRabbit on assignment-store.ts:1075. Both events were written where they are
decided, which is inside assignOnce's transaction. A reservation or lease write
failing after that point rolls the placement back, but a line already on stdout
cannot be rolled back with it — so the canary this PR asks an operator to read
would count re-placements that never happened, and a Postgres transaction retry
could leave a stale line behind as well.

The transaction now returns its events alongside the RelayAssignment and the
caller flushes them once it has resolved. Returning them rather than setting a
variable in the enclosing scope is what makes the retry case safe too: only the
attempt that committed can carry its events out. assign()'s signature is
unchanged; the extra shape lives entirely inside assignOnce.

orca_relay_sticky_replacement_deferred was moved the same way. It cost one more
push into the array that already existed, and it is decided inside the same
transaction, so leaving it behind would have been the odd case rather than the
cheap one.

The new test injects a failure on the first write after the decision, asserts no
event is emitted, and asserts the assignment is still on its original cell —
without that second assertion the absence would only prove the emit was early,
not that it would have been wrong. A control dial with nothing injected emits
exactly one event, so the case cannot pass on a broken harness. With the emit
put back inside the transaction, it fails.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): expire the roll stamp, correct the wire note, assert the stamp landed

Delta review findings B, D and E. A (the deferral path's cost) is deliberately
not implemented; it is now written up under Follow-ups in the PR body as
required before any Asia roll, because it cannot fire in a US canary.

B, which also closes C: the stamp was written, carried and never compared to
anything. A roll isolates and restores one cell inside ~15 minutes, so a stamp
older than two hours is not a roll in progress. It is a failed wave whose
failsafe re-isolated a possibly healthy cell and is waiting on an operator — the
postmortem in this tree records gaps of hours — or an orphan left by a director
rollback whose restore wrote 'general' without the clause that clears the stamp,
which the selector's 'keep' branch would then preserve until some later park
reactivated it. Both want the same answer and it is the pre-existing one: keep
the pin. One comparison against a value already on the row.

The bound takes the caller's `now` rather than reading the clock again, so one
assign reasons about one instant; the stamp's age is now a thing that decides
whether a host moves, and two clock reads could disagree across it.

D: the comment beside the new request field claimed an updated caller reaching
an older director "is simply ignored". The schema is .strict(), so it is a 400.
That fails closed — the isolate aborts before MUTATION_STARTED is set and
nothing is written — but it is a deploy ordering constraint, and it was
undocumented. The comment now says so and the PR body's rollout notes carry it.

E: nothing read the `rollIsolated` the script already prints, so an older script
against a newer director would silently produce today's behaviour and the canary
would read as "the fix did nothing" with no way to tell that from a wrong
premise. Both isolate steps now assert it, beside the generation they already
parse.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 04:18:57 -04:00
..

Orca Relay

The relay that connects the Orca mobile app to a desktop host. Phones and desktops never talk to each other directly: each opens an outbound WebSocket to a relay cell, the relay pairs the two sessions, and it splices frames between them. A director assigns hosts to cells and coordinates migrations; cells carry the user connections.

This directory is an independent pnpm workspace inside the Orca monorepo. Run its commands from cloud/, not the repository root. The source is covered by the repository's root MIT license.

Packages

  • packages/relay-contract: the wire contract shared by the relay, the desktop app, and the mobile app (frame shapes, close codes, admission budgets, splice state machine).
  • apps/relay: the relay server. The same image runs as a director or a cell depending on ORCA_RELAY_ROLE.
  • apps/relay-fence-broker: a private, IAM-only service that owns the durable mutation lease, the Terraform checkout, and the narrow Compute mutation used when a registered target is superseded. The workflow that calls it holds read and invoke rights only, never those mutation permissions.
  • apps/relay-ops: the relay operations console and the incident monitor behind pnpm ops:relay, pnpm incident:relay, and pnpm incident:relay-preflight.
  • apps/push and packages/push-contract: the mobile push gateway that holds the APNs key and sends to phones through APNs and FCM, and its wire contract. It is deployed and operated from here but is not part of the relay data path; see docs/push-gateway.md.

Mobile push gateway

apps/push is a separate Cloud Run service from the relay. Phones never hold an Orca credential for it: the desktop host authenticates with the same X25519 key it uses for the relay, answering an encrypted challenge to mint a 24 hour session, then registers each paired phone's native push token and asks the gateway to push. The gateway queues each event as its own notification, enforces per-host quotas and request limits, and retires a registration as soon as Apple or Google reports the token unregistered. Provider push is the only ordinary mobile OS-banner path. The notification socket is retained only for live dismissal and reconnect tray reconciliation; it never creates or recovers banners. Desktop notification categories remain authoritative. Each delivery is persisted as one notification event. Before deploying an incompatible queue format, stop all older push gateway revisions and clear only unpublished push delivery fixtures; no queue preservation or migration is required. FCM notification messages are inherently collapsible while offline and have a small concurrent collapse-key budget, so every pending alert is not guaranteed.

Storage follows the relay pattern: PostgreSQL in production, SQLite for tests and local development. Configure it with ORCA_PUSH_PUBLIC_URL, ORCA_PUSH_FCM_PROJECT_ID, ORCA_PUSH_DATABASE_URL, the three APNs variables (ORCA_PUSH_APNS_KEY, ORCA_PUSH_APNS_KEY_ID, ORCA_PUSH_APPLE_TEAM_ID, all three or none), and optionally ORCA_PUSH_APNS_TOPIC. The FCM credential comes from the runtime service account, so no key material is configured for Android. See push gateway operations for deployment and recovery.

Logging is aggregate counters only. Tokens, notification titles, notification bodies, and full host fingerprints never reach a log line.

Infrastructure and operations

  • infra/terraform: the relay Terraform root. It owns the cells, the director, the shared Cloud SQL instance, DNS, observability, and every GitHub Workload Identity provider the relay workflows authenticate through. backend/ holds the per-environment backend configuration and environments/ the tfvars. Drive it through pnpm infra:init, pnpm infra:plan, and pnpm infra:apply.
  • dev/scripts: the deploy, capacity, admission, rehome, monitoring, and load scripts the workflows call, plus the contract tests that pin each workflow and Terraform surface. Run them with pnpm test.
  • dev/contracts and dev/fixtures: the checked-in data those contract tests read, including the Terraform root partition.
  • docs/: the relay runbooks, capacity-testing guide, incident-monitor reference, the workflow variable reference in docs/relay-workflows.md, and the push gateway runbook in docs/push-gateway.md.

Workflows

The 25 .github/workflows/cloud-*.yml workflows are the deploy and operate surface: publish and deploy the director, roll GCE cell capacity, operate Asia admission and regional rehoming, prove staging capacity, monitor production, power staging up and down, and deploy the mobile push gateway. .github/actions/cloud-sql-rollout-lease is the compare-and-swap lease that serializes rollouts against the shared Cloud SQL instance. Push reuses that action with its own lease object and deployment concurrency group.

Every one of them is inert. Each top-level job is gated on vars.ORCA_CLOUD_OPERATIONS_ENABLED == 'true', a repository variable that is unset here, so the two scheduled triggers and every manual dispatch skip without running a step. Only the repository owner, holding the GCP identities these workflows authenticate as, can turn them on.

Cloud Verify is not gated. It builds, typechecks, lints, tests, secret-scans, and validates the relay Terraform on every change under cloud/, and it runs on fork pull requests, so it configures no backend and holds no credential.

What is not here

The terraform-foundation and terraform-apps roots and the API and auth services live in the private stablyai/orca-cloud repository. Scripts and tests that spanned both trees were narrowed to the relay side rather than carrying a dangling reference.

Local development

cd cloud
pnpm install
pnpm build
pnpm test

pnpm test runs the SQLite-backed suites. Tests that need PostgreSQL run only when ORCA_RELAY_TEST_POSTGRES_URL points at a disposable PostgreSQL 16 or 17 database, for example:

docker run --rm -d --name orca-relay-pg -e POSTGRES_HOST_AUTH_METHOD=trust \
  -e POSTGRES_DB=orca_relay_test -p 55440:5432 postgres:16-alpine
ORCA_RELAY_TEST_POSTGRES_URL=postgres://postgres@127.0.0.1:55440/orca_relay_test \
  pnpm --filter @orca-cloud/relay test
docker rm -f orca-relay-pg

Configuration is read from environment variables validated in apps/relay/src/config.ts. ORCA_RELAY_ASSIGNMENT_SIGNING_KEY (at least 32 bytes) is the only required value; everything else has a local default.