Commit Graph
24 Commits
Author SHA1 Message Date
Jinwoo-H f8dd8bd2ac docs(cloud): c7 up on target image 2026-09-04 02:25:15 -04:00
Jinwoo-H 60f60edf0b docs(cloud): attribute the 06:20Z 503 burst 2026-09-04 02:23:01 -04:00
Jinwoo-H f61a879c17 docs(cloud): record c27/c29 crash storm during c7 apply 2026-09-04 02:19:12 -04:00
Jinwoo-H 169d6ff816 docs(cloud): record c7 drain recovery numbers 2026-09-04 02:16:45 -04:00
Jinwoo-H 903da34b95 docs(cloud): record c7 drain effect on the director 2026-09-04 02:12:14 -04:00
Jinwoo-H 87047a2927 docs(cloud): record dry-run #6 pass and c7 canary dispatch 2026-09-04 02:08:02 -04:00
Jinwoo-H 4d80d24108 docs(cloud): record dry-run #6 dispatch and PR states 2026-09-04 01:47:45 -04:00
Jinwoo-H 5a96670d7f docs(cloud): record Asia-cell autoheal recreate amplifier 2026-09-04 01:43:09 -04:00
Jinwoo-H 262a18867d docs(cloud): record dry-run #5 freeze on c27 autoheal recreate 2026-09-04 01:42:28 -04:00
Jinwoo-H bad866ac8a docs(cloud): record dry-run #4 freeze (c27 crash) and dry-run #5 2026-09-04 01:40:33 -04:00
Jinwoo-H d85fe9e414 docs(cloud): record canary blast radius and dry-run #4 2026-09-04 01:27:30 -04:00
Jinwoo-H 7a8d1fa948 docs(cloud): record monitor dry-run dispatch inputs 2026-09-04 01:17:26 -04:00
Jinwoo-H 16c6aef916 docs(cloud): record retries-bar decision and PR #18580 basis 2026-09-04 01:15:27 -04:00
Jinwoo-H 8ebff89106 docs(cloud): move relay reconnect findings under cloud/docs
The root directory guard blocks new root-level files.
2026-09-04 01:13:22 -04:00
Jinwoo-H aa3a7ef147 fix(relay): make the control lease jitter symmetric
Pullfrog: shortening-only jitter raised the mean rebind rate ~15%. Grant
55 min +/- 5 min instead; same cohort spread, unchanged steady-state load.
Auth expiry (5 min token) and the 75 s silence watchdog are enforced
separately, so a grant up to 60 min risks nothing.
2026-09-03 23:58:17 -04:00
Jinwoo-H e5ccd4a1a4 fix(relay): check client abandonment between each accept lookup
CodeRabbit: a phone that hung up while resolveResume was in flight still
paid for resolveInviteForMove and assignments.resolve. Check after each
awaited lookup; regression test asserts neither later lookup runs.
2026-09-03 23:33:08 -04:00
Jinwoo-H fdf6fb2f27 fix(relay): stop dead accept work, spread control rotations, fail direct probes fast
Incident 2026-09-04 ~01:05Z: after a background/foreground cycle the phone's
relay dial timed out inside the cell's acceptClient DB phase while the fleet
was in a cell-inventory lock storm (55P03 retries ~7.5k/h vs a ~1k/h floor).

Relay (cloud/apps/relay)
- acceptClient checks socket.readyState after each serialized Postgres call
  and abandons the accept once the phone has hung up, releasing the capacity
  reservation, failing the credential reservation, and releasing the activity
  lease it just acquired instead of leaking it to expiry cleanup and then
  throwing host_data_reservation_already_bound at bind.
- New structured event orca_relay_client_accept_abandoned {stage, elapsedMs}
  and runtime-metric fields clientAcceptsAbandonedByStageDelta /
  clientAcceptAbandonedMsMax so the "phone gave up behind the lock" rate is
  quantifiable per cell.
- Control lease grants are jittered: 55 min minus [0, 10 min). Every host
  that (re)connected in the same minute rebound as one cohort every ~54 min
  (c27 autoheal recreate at 23:23Z re-homed ~420 controls; ~1.1k 1006 +
  ~1k 4408 "control rebound" closes landed in a 3 s window at 00:50:14Z),
  and each rebind is an activateControl transaction on the inventory lock.
  No wire change: leaseExpiresAt was always a server-chosen absolute time.

Desktop (src/main/runtime/relay)
- Control rotation rebinds 1-6 min early instead of 1-2 min, so a re-homed
  cohort spreads across cycles rather than pinning one phase for the life of
  the process.

Phone (mobile/src/transport)
- openAuthenticatedDirectEndpoint treats 'reconnecting' as a failed probe. On
  a dead LAN the foreground direct dial dies with an instant 1006 and the
  direct client enters its own 500/1000/2000 ms backoff; the probe used to
  wait out its full 12 s bound holding the supervisor mutex, so relay recovery
  queued behind three doomed redials. The stage-aware bound from #18518 is
  unaffected.

Not done here: rolling the 23 GCE cells onto the post-#18521 image (500 ms
lock_timeout) is a deploy owned by cloud-deploy-relay-production-same-cap.
2026-09-03 23:07:52 -04:00
Jinwoo Hong 7d27c841b4 fix(cloud): run the rehome control job under pipefail (#18537)
The five `node ... | tee` steps in cloud-operate-relay-production-rehome-job.yml
reported tee's exit code, so a thrown inspect or apply passed green. The Aug 28
21:25Z and Aug 29 inspects and today's first inspect all printed
"director returned an invalid regional rehome control" (the durable control had
moved to generation 12 when the Aug 28 rehome aborted) and still succeeded.
`shell: bash` adds `-o pipefail`. A test pins the default and the tee count.
2026-09-03 18:38:33 -04:00
Jinwoo Hong f974e98162 test(cloud): derive both reachability directions for the relay inventory census (#18524)
Mirrors stablyai/orca-cloud#472 (c82f98f), byte-identical under cloud/.
2026-09-03 17:43:40 -04:00
Jinwoo Hong d66386bc82 fix(cloud): bound and yield the relay's global cell-inventory lock (#18521)
Mirrors stablyai/orca-cloud#471 (squash c3354e8), byte-identical under cloud/.

The relay's cell-inventory lock (SELECT ... FROM relay_cells FOR UPDATE over
all 23 rows) is one global critical section shared by the assignment hot path
and every director sweep; with the pool's 1s lock_timeout a blocked waiter held
a pooled client for a full second, producing ~690 55P03 retries per 5 minutes
in production. Request paths now bound the wait at 500ms with a SET LOCAL that
is restored to the pool default before the next statement; director-only sweeps
take the lock NOWAIT and skip the tick; sweep timers are jittered; hold time is
exported as additive runtime-metrics fields so the bound can be tuned.
2026-09-03 17:27:57 -04:00
Jinwoo Hong 0746d82c01 chore(cloud): close the Workload Identity cutover onto stablyai/orca (#18509)
Mirrors stablyai/orca-cloud#470. The private relay workflows are retired, so
the dual accept has one live arm left. Add `github_workflow_file_prefix` for
the primary repository's workflow filenames, point `github_repo`/
`github_repo_id` at `stablyai/orca` (`1183888342`), and empty
`github_accepted_repositories` in both environments. Every relay provider goes
back to a single arm naming `cloud-` prefixed workflow refs.

`cloud/infra/terraform` stays byte-identical to the private branch. The two
identity tests diverge here as they already did, so they take the same change
rather than the same bytes: both now render the trusted ref head from the
Terraform variable instead of this checkout's own workflow filenames, which is
what lets the length pin be the same 791 characters in either repository.
2026-09-03 15:57:00 -04:00
Jinwoo Hong fbea749d07 chore(cloud): pin staging relay c3 to the director's image (#18508)
* chore(cloud): pin staging relay c3 to the director's image

Mirrors stablyai/orca-cloud#468. c3 stayed on sha-c91439af after the
director and c4 moved to sha-e3e92d95, so the staging capacity proof's
compatible-director-image check has failed since 2026-08-14.

* test(cloud): scope the launch-image pin to staging C4 now that C3 shares the digest

* test(cloud): keep the public workflow assertions; scope only the launch-image pin to C4
2026-09-03 15:36:41 -04:00
Jinwoo Hong 67e22345da fix(cloud): stop passing manage_artifact_dns to the relay root (#18442)
The relay root does not declare it (it belongs to the private apps root), and
Terraform rejects an undeclared -var, so the first public Deploy Relay Staging
run failed at the C4 image bind.
2026-09-03 07:18:29 -04:00
Jinwoo Hong 3eec77c11a chore(cloud): add the relay fence broker, ops console, Terraform root, scripts, and 24 cloud-* workflows (#18413)
Phase 6 of the relay split: the relay's deploy/operate surface moves under cloud/ with 24 cloud-* workflows gated on ORCA_CLOUD_OPERATIONS_ENABLED, the Cloud SQL rollout lease action, the relay Terraform root (dual-accept identities for both repositories), scripts, docs, CODEOWNERS, and a terraform validate job in Cloud Verify.
2026-09-03 06:55:14 -04:00