* infra(relay): raise asia-east2 cell pools to 16 and retire four idle cells The three asia-east2 cells sit 176 ms from the Cloud SQL instance in us-central1. Server-side statement time there is 0.2 ms, so a pool slot is held by the round trip, not by the query. At a pool of 10 they measured 94-156 waiters and 2 s waits, and client accepts ran a ~4 s p95 against 222-646 ms in us-central1. Raising those three pools to 16 is the agreed first step; every other cell stays at 10. c4 and c5 join the committed fence set. Both are existing-only capacity the admission selector can never place on again, they carried ~1 connection each on 40-day-old images, and each still holds 10 Postgres connections. The fence set is the prerequisite the fence-source workflow confirms before it drains and attests a cell; it is not itself the resize. c17 and c18 are not fenced here. They are migration-only, and the runbook requires retire-migration-cell to move a migration-only cell to existing-only through a generation-bound selector CAS before it can be fenced. Terraform cannot express that step. The Cloud SQL consumer contract carried two stale numbers: auth at 2 instances when production has run a cap of 20 since 2026-09-04, and a 400-connection ceiling when the live instance reports 500. Both are corrected, and the budget now asserts its headroom in two named gates instead of one aggregate boolean. Those gates fail: auth alone accounts for 200 configured connections and a 215-connection rollout overlap, so the operating maximum is 713 against a usable ceiling of 490. Nothing here caused that, and no pool was lowered to hide it. * infra(relay): move the Cloud SQL contract correction out of this branch The contract correction (auth at its real 20-instance cap, the measured 500-connection ceiling) makes the budget gate fail for reasons that have nothing to do with asia pools or fenced cells, and it held this branch red. It moves to its own branch where the failure is the subject. production-cloud-sql-app-consumers.json returns to main unchanged. The budget test keeps main's single gate and only repins the cell figure that this branch genuinely moves: 230 -> 228, being +18 for three asia pools at 16 and -20 for fencing c4 and c5. Against main's 400-connection model that leaves an operating maximum of 383 under a usable ceiling of 390. * infra(relay): move the c4/c5 fence entries out of this branch Terraform now sets a cell's MIG target size directly from relay_gce_fenced_cells (relay-gce-cells.tf); the lifecycle ignore that used to protect operational target_size drift is gone. So a fence entry sitting on main ahead of its fence-source run is a standing instruction that any apply reaching that cell may execute without the documented drain and attestation. Keeping the entry in the same merge as an unrelated pool change widens that blast radius for no reason. The two entries move to their own branch, to be merged immediately before fence-source runs for c4 and then c5. This branch keeps the multi-line reflow of the list, which makes that later diff two added lines instead of a rewritten one. The cell figure in the budget test follows: 230 + 18 for the three asia-east2 pools at 16, with no fenced-cell subtraction. That is 403 operating against a usable ceiling of 390, so the headroom gate now fails by 13. It fails against a ceiling of 400 that is itself wrong; the instance reports 500. See the PR body. * infra(cloud-sql): record the measured 500-connection ceiling The budget's usable ceiling came from maxConnections: 400, described as the tier default. It is a tier default, since no max_connections flag is set, but the instance does not report 400. SHOW max_connections on it returns 500, measured 2026-09-16. On main the model sat at 385 against a usable ceiling of 390, five connections of margin, so raising the three asia-east2 pools by 18 failed the gate by 13 against a ceiling that was never checked. Against the measured one it is 403 against 490, clearing by 87. Only the ceiling and its source note change here. auth stays recorded at 2 instances, which is also wrong; PR #21165 corrects it, and with the true auth figure the budget is over by 225 for reasons that have nothing to do with these pools. * test(cloud): state the cell pool arithmetic literally in the budget pin comment
Orca Relay
The relay that connects the Orca mobile app to a desktop host. Phones and desktops never talk to each other directly: each opens an outbound WebSocket to a relay cell, the relay pairs the two sessions, and it splices frames between them. A director assigns hosts to cells and coordinates migrations; cells carry the user connections.
This directory is an independent pnpm workspace inside the Orca monorepo. Run
its commands from cloud/, not the repository root. The source is covered by
the repository's root MIT license.
Packages
packages/relay-contract: the wire contract shared by the relay, the desktop app, and the mobile app (frame shapes, close codes, admission budgets, splice state machine).apps/relay: the relay server. The same image runs as a director or a cell depending onORCA_RELAY_ROLE.apps/relay-fence-broker: a private, IAM-only service that owns the durable mutation lease, the Terraform checkout, and the narrow Compute mutation used when a registered target is superseded. The workflow that calls it holds read and invoke rights only, never those mutation permissions.apps/relay-ops: the relay operations console and the incident monitor behindpnpm ops:relay,pnpm incident:relay, andpnpm incident:relay-preflight.apps/pushandpackages/push-contract: the mobile push gateway that holds the APNs key and sends to phones through APNs and FCM, and its wire contract. It is deployed and operated from here but is not part of the relay data path; see docs/push-gateway.md.
Mobile push gateway
apps/push is a separate Cloud Run service from the relay. Phones never hold an
Orca credential for it: the desktop host authenticates with the same X25519
key it uses for the relay, answering an encrypted challenge to mint a 24 hour
session, then registers each paired phone's native push token and asks the
gateway to push. The gateway queues each event as its own notification,
enforces per-host quotas and request limits, and retires a
registration as soon as Apple or Google reports the token unregistered.
Provider push is the only ordinary mobile OS-banner path. The notification
socket is retained only for live dismissal and reconnect tray reconciliation;
it never creates or recovers banners. Desktop notification categories remain
authoritative.
Each delivery is persisted as one notification event. Before deploying an
incompatible queue format, stop all older push gateway revisions and clear only
unpublished push delivery fixtures; no queue preservation or migration is required.
FCM notification messages are inherently collapsible while offline and have a
small concurrent collapse-key budget, so every pending alert is not guaranteed.
Storage follows the relay pattern: PostgreSQL in production, SQLite for tests
and local development. Configure it with ORCA_PUSH_PUBLIC_URL, ORCA_PUSH_FCM_PROJECT_ID,
ORCA_PUSH_DATABASE_URL, the three APNs variables (ORCA_PUSH_APNS_KEY,
ORCA_PUSH_APNS_KEY_ID, ORCA_PUSH_APPLE_TEAM_ID, all three or none), and
optionally ORCA_PUSH_APNS_TOPIC. The FCM credential comes from
the runtime service account, so no key material is configured for Android. See
push gateway operations for deployment and recovery.
Logging is aggregate counters only. Tokens, notification titles, notification bodies, and full host fingerprints never reach a log line.
Infrastructure and operations
infra/terraform: the relay Terraform root. It owns the cells, the director, the shared Cloud SQL instance, DNS, observability, and every GitHub Workload Identity provider the relay workflows authenticate through.backend/holds the per-environment backend configuration andenvironments/the tfvars. Drive it throughpnpm infra:init,pnpm infra:plan, andpnpm infra:apply.dev/scripts: the deploy, capacity, admission, rehome, monitoring, and load scripts the workflows call, plus the contract tests that pin each workflow and Terraform surface. Run them withpnpm test.dev/contractsanddev/fixtures: the checked-in data those contract tests read, including the Terraform root partition.docs/: the relay runbooks, capacity-testing guide, incident-monitor reference, the workflow variable reference indocs/relay-workflows.md, and the push gateway runbook indocs/push-gateway.md.
Workflows
The 25 .github/workflows/cloud-*.yml workflows are the deploy and operate
surface: publish and deploy the director, roll GCE cell capacity, operate Asia
admission and regional rehoming, prove staging capacity, monitor production,
power staging up and down, and deploy the mobile push gateway.
.github/actions/cloud-sql-rollout-lease is the compare-and-swap lease that
serializes rollouts against the shared Cloud SQL instance. Push reuses that
action with its own lease object and deployment concurrency group.
Every one of them is inert. Each top-level job is gated on
vars.ORCA_CLOUD_OPERATIONS_ENABLED == 'true', a repository variable that is
unset here, so the two scheduled triggers and every manual dispatch skip
without running a step. Only the repository owner, holding the GCP identities
these workflows authenticate as, can turn them on.
Cloud Verify is not gated. It builds, typechecks, lints, tests, secret-scans,
and validates the relay Terraform on every change under cloud/, and it runs
on fork pull requests, so it configures no backend and holds no credential.
What is not here
The terraform-foundation and terraform-apps roots and the API and auth
services live in the private stablyai/orca-cloud repository. Scripts and
tests that spanned both trees were narrowed to the relay side rather than
carrying a dangling reference.
Local development
cd cloud
pnpm install
pnpm build
pnpm test
pnpm test runs the SQLite-backed suites. Tests that need PostgreSQL run only
when ORCA_RELAY_TEST_POSTGRES_URL points at a disposable PostgreSQL 16 or 17
database, for example:
docker run --rm -d --name orca-relay-pg -e POSTGRES_HOST_AUTH_METHOD=trust \
-e POSTGRES_DB=orca_relay_test -p 55440:5432 postgres:16-alpine
ORCA_RELAY_TEST_POSTGRES_URL=postgres://postgres@127.0.0.1:55440/orca_relay_test \
pnpm --filter @orca-cloud/relay test
docker rm -f orca-relay-pg
Configuration is read from environment variables validated in
apps/relay/src/config.ts. ORCA_RELAY_ASSIGNMENT_SIGNING_KEY (at least 32
bytes) is the only required value; everything else has a local default.