Files
orca/mobile/src
Jinwoo-H fdf6fb2f27 fix(relay): stop dead accept work, spread control rotations, fail direct probes fast
Incident 2026-09-04 ~01:05Z: after a background/foreground cycle the phone's
relay dial timed out inside the cell's acceptClient DB phase while the fleet
was in a cell-inventory lock storm (55P03 retries ~7.5k/h vs a ~1k/h floor).

Relay (cloud/apps/relay)
- acceptClient checks socket.readyState after each serialized Postgres call
  and abandons the accept once the phone has hung up, releasing the capacity
  reservation, failing the credential reservation, and releasing the activity
  lease it just acquired instead of leaking it to expiry cleanup and then
  throwing host_data_reservation_already_bound at bind.
- New structured event orca_relay_client_accept_abandoned {stage, elapsedMs}
  and runtime-metric fields clientAcceptsAbandonedByStageDelta /
  clientAcceptAbandonedMsMax so the "phone gave up behind the lock" rate is
  quantifiable per cell.
- Control lease grants are jittered: 55 min minus [0, 10 min). Every host
  that (re)connected in the same minute rebound as one cohort every ~54 min
  (c27 autoheal recreate at 23:23Z re-homed ~420 controls; ~1.1k 1006 +
  ~1k 4408 "control rebound" closes landed in a 3 s window at 00:50:14Z),
  and each rebind is an activateControl transaction on the inventory lock.
  No wire change: leaseExpiresAt was always a server-chosen absolute time.

Desktop (src/main/runtime/relay)
- Control rotation rebinds 1-6 min early instead of 1-2 min, so a re-homed
  cohort spreads across cycles rather than pinning one phase for the life of
  the process.

Phone (mobile/src/transport)
- openAuthenticatedDirectEndpoint treats 'reconnecting' as a failed probe. On
  a dead LAN the foreground direct dial dies with an instant 1006 and the
  direct client enters its own 500/1000/2000 ms backoff; the probe used to
  wait out its full 12 s bound holding the supervisor mutex, so relay recovery
  queued behind three doomed redials. The stage-aware bound from #18518 is
  unaffected.

Not done here: rolling the 23 GCE cells onto the post-#18521 image (500 ms
lock_timeout) is a deploy owned by cloud-deploy-relay-production-same-cap.
2026-09-03 23:07:52 -04:00
..