Commit Graph
57 Commits
Author SHA1 Message Date
Jinwoo-H 0fdbfe6990 docs(cloud): dry-run #11 froze; #12 armed with longer window 2026-09-04 06:09:04 -04:00
Jinwoo-H 7605acdda5 docs(cloud): dry-run #11 dispatched into a crash 2026-09-04 06:08:26 -04:00
Jinwoo-H f18962efd7 docs(cloud): cascade gap math; dry-run #11 armed 2026-09-04 06:02:22 -04:00
Jinwoo-H e37089e400 docs(cloud): stall census 2026-09-04 06:01:49 -04:00
Jinwoo-H c623f121b2 docs(cloud): checkpoint-phase check negative 2026-09-04 06:00:59 -04:00
Jinwoo-H 153340917e docs(cloud): dry-run #10 froze on a 4 s Postgres connect stall 2026-09-04 05:58:27 -04:00
Jinwoo-H 65c2fe98da docs(cloud): deploy-vs-crash-rate check 2026-09-04 05:41:46 -04:00
Jinwoo-H 5632975cd9 docs(cloud): 09:39 cascade; per-instance crash ranking 2026-09-04 05:40:07 -04:00
Jinwoo-H c8b4cc76fc docs(cloud): dry-run #9 froze; #10 armed; gate decision put to owner 2026-09-04 05:35:41 -04:00
Jinwoo-H 97898f50d4 docs(cloud): 09:34 cascade 2026-09-04 05:35:01 -04:00
Jinwoo-H 580630c015 docs(cloud): dry-run #9 dispatched; 09:31 cascade 2026-09-04 05:32:10 -04:00
Jinwoo-H 2c9399fb61 docs(cloud): dry-run #8 froze on the monitor's own 500; #9 armed 2026-09-04 05:27:02 -04:00
Jinwoo-H 8724fb6040 docs(cloud): dry-run #8 dispatched with auto-canary on success 2026-09-04 05:09:41 -04:00
Jinwoo-H 48286f8314 docs(cloud): dry-run #7 froze on aged director 500s 2026-09-04 05:06:17 -04:00
Jinwoo-H b869093b32 docs(cloud): verify passed; dry-run #7 running; selector gen 112 2026-09-04 05:05:01 -04:00
Jinwoo-H d52ae4bc89 docs(cloud): director 500 shape; verify + dry-run #7 dispatched 2026-09-04 05:04:41 -04:00
Jinwoo-H ad043fd246 docs(cloud): fourth cascade 09:00Z, five cells 2026-09-04 05:01:58 -04:00
Jinwoo-H 017b3c0781 docs(cloud): Finding 9, director lock retries down ~10x on 519f4914 2026-09-04 04:56:21 -04:00
Jinwoo-H d907776811 docs(cloud): director rollback command 2026-09-04 04:48:02 -04:00
Jinwoo-H 63ceeef632 docs(cloud): director on 519f4914; Finding 8 ten-cell cascade 2026-09-04 04:47:10 -04:00
Jinwoo-H 5a33e5ac9e docs(cloud): director deploy dispatched; record required predecessor input 2026-09-04 04:38:23 -04:00
Jinwoo-H 817b0a85c6 docs(cloud): image 519f4914 published; director deploy dispatched 2026-09-04 04:37:13 -04:00
Jinwoo-H e4b54d6abb docs(cloud): #18606 merged, image publish dispatched 2026-09-04 04:34:59 -04:00
Jinwoo-H cc36ced9b3 docs(cloud): post-merge dispatch plan for the lock fix 2026-09-04 04:33:56 -04:00
Jinwoo-H 5de21d4fe3 docs(cloud): lock PR re-verification outcome 2026-09-04 04:32:22 -04:00
Jinwoo-H efaefdc594 docs(cloud): PR #18606 opened; c7 two-hour evidence 2026-09-04 04:26:28 -04:00
Jinwoo-H ffb833a2ae docs(cloud): lock-removal PR review outcome 2026-09-04 04:25:16 -04:00
Jinwoo-H 9f162c151d docs(cloud): correct the rollout-speed plan after reading the cell job 2026-09-04 04:11:45 -04:00
Jinwoo-H 600b95921b docs(cloud): faster same-cap rollout design from measured limits 2026-09-04 04:09:55 -04:00
Jinwoo-H 21f69ded3b docs(cloud): lock-removal PR status 2026-09-04 04:06:24 -04:00
Jinwoo-H 6fffb33a4c docs(cloud): record the agreed lock-first plan 2026-09-04 03:59:51 -04:00
Jinwoo-H d0ff6f6141 docs(cloud): c7 post-restore observations 2026-09-04 02:27:24 -04:00
Jinwoo-H c4af2a21cf docs(cloud): c7 canary succeeded 2026-09-04 02:26:23 -04:00
Jinwoo-H f8dd8bd2ac docs(cloud): c7 up on target image 2026-09-04 02:25:15 -04:00
Jinwoo-H 60f60edf0b docs(cloud): attribute the 06:20Z 503 burst 2026-09-04 02:23:01 -04:00
Jinwoo-H f61a879c17 docs(cloud): record c27/c29 crash storm during c7 apply 2026-09-04 02:19:12 -04:00
Jinwoo-H 169d6ff816 docs(cloud): record c7 drain recovery numbers 2026-09-04 02:16:45 -04:00
Jinwoo-H 903da34b95 docs(cloud): record c7 drain effect on the director 2026-09-04 02:12:14 -04:00
Jinwoo-H 87047a2927 docs(cloud): record dry-run #6 pass and c7 canary dispatch 2026-09-04 02:08:02 -04:00
Jinwoo-H 4d80d24108 docs(cloud): record dry-run #6 dispatch and PR states 2026-09-04 01:47:45 -04:00
Jinwoo-H 5a96670d7f docs(cloud): record Asia-cell autoheal recreate amplifier 2026-09-04 01:43:09 -04:00
Jinwoo-H 262a18867d docs(cloud): record dry-run #5 freeze on c27 autoheal recreate 2026-09-04 01:42:28 -04:00
Jinwoo-H bad866ac8a docs(cloud): record dry-run #4 freeze (c27 crash) and dry-run #5 2026-09-04 01:40:33 -04:00
Jinwoo-H d85fe9e414 docs(cloud): record canary blast radius and dry-run #4 2026-09-04 01:27:30 -04:00
Jinwoo-H 7a8d1fa948 docs(cloud): record monitor dry-run dispatch inputs 2026-09-04 01:17:26 -04:00
Jinwoo-H 16c6aef916 docs(cloud): record retries-bar decision and PR #18580 basis 2026-09-04 01:15:27 -04:00
Jinwoo-H 8ebff89106 docs(cloud): move relay reconnect findings under cloud/docs
The root directory guard blocks new root-level files.
2026-09-04 01:13:22 -04:00
Jinwoo-H aa3a7ef147 fix(relay): make the control lease jitter symmetric
Pullfrog: shortening-only jitter raised the mean rebind rate ~15%. Grant
55 min +/- 5 min instead; same cohort spread, unchanged steady-state load.
Auth expiry (5 min token) and the 75 s silence watchdog are enforced
separately, so a grant up to 60 min risks nothing.
2026-09-03 23:58:17 -04:00
Jinwoo-H e5ccd4a1a4 fix(relay): check client abandonment between each accept lookup
CodeRabbit: a phone that hung up while resolveResume was in flight still
paid for resolveInviteForMove and assignments.resolve. Check after each
awaited lookup; regression test asserts neither later lookup runs.
2026-09-03 23:33:08 -04:00
Jinwoo-H fdf6fb2f27 fix(relay): stop dead accept work, spread control rotations, fail direct probes fast
Incident 2026-09-04 ~01:05Z: after a background/foreground cycle the phone's
relay dial timed out inside the cell's acceptClient DB phase while the fleet
was in a cell-inventory lock storm (55P03 retries ~7.5k/h vs a ~1k/h floor).

Relay (cloud/apps/relay)
- acceptClient checks socket.readyState after each serialized Postgres call
  and abandons the accept once the phone has hung up, releasing the capacity
  reservation, failing the credential reservation, and releasing the activity
  lease it just acquired instead of leaking it to expiry cleanup and then
  throwing host_data_reservation_already_bound at bind.
- New structured event orca_relay_client_accept_abandoned {stage, elapsedMs}
  and runtime-metric fields clientAcceptsAbandonedByStageDelta /
  clientAcceptAbandonedMsMax so the "phone gave up behind the lock" rate is
  quantifiable per cell.
- Control lease grants are jittered: 55 min minus [0, 10 min). Every host
  that (re)connected in the same minute rebound as one cohort every ~54 min
  (c27 autoheal recreate at 23:23Z re-homed ~420 controls; ~1.1k 1006 +
  ~1k 4408 "control rebound" closes landed in a 3 s window at 00:50:14Z),
  and each rebind is an activateControl transaction on the inventory lock.
  No wire change: leaseExpiresAt was always a server-chosen absolute time.

Desktop (src/main/runtime/relay)
- Control rotation rebinds 1-6 min early instead of 1-2 min, so a re-homed
  cohort spreads across cycles rather than pinning one phase for the life of
  the process.

Phone (mobile/src/transport)
- openAuthenticatedDirectEndpoint treats 'reconnecting' as a failed probe. On
  a dead LAN the foreground direct dial dies with an instant 1006 and the
  direct client enters its own 500/1000/2000 ms backoff; the probe used to
  wait out its full 12 s bound holding the supervisor mutex, so relay recovery
  queued behind three doomed redials. The stage-aware bound from #18518 is
  unaffected.

Not done here: rolling the 23 GCE cells onto the post-#18521 image (500 ms
lock_timeout) is a deploy owned by cloud-deploy-relay-production-same-cap.
2026-09-03 23:07:52 -04:00