Files
orca/cloud/apps
Jinwoo Hong b82307937a fix(relay): admit drained hosts through their own lane and stagger their return (#24446)
* fix(relay): admit drained hosts through their own lane and stagger their return

A same-cap roll's drain sends every host on the isolated cell back to the
director at once. Those hosts reconnect through the sticky lane (one slot per
director), and each one's re-placement holds that slot for most of a second
behind the region-wide inventory lock, so ordinary reconnects time out behind
them and the drained hosts retry every 2 s: ~30k 503s per drain.

The reconnect verification read now also says whether the host's home cell is
isolated for a roll right now (same predicate re-placement uses). Those hosts
release the sticky slot after the read and take a separate drain-return lane
(1 per director, matching the store's per-director placement serialization).
When that lane is full the host gets a Retry-After that reserves the next
free service slot, paced by the measured re-placement time and capped at 300 s.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): keep drain returns inside the placement pool budget and their own slot

Review follow-ups for the drain-return lane:
- The lane now borrows placement permits (never placement's last, never ahead
  of a queued placement), so placement + sticky still bounds the database pool.
- A host's own early retry (row-busy redial, duplicate dial) gets the 2 s lane
  interval instead of a fresh slot behind the cohort, and a host that returns
  early to the same director keeps its reserved slot.
- Classification also excludes an open migration row whose lease counter
  lapsed, matching the re-placement rule.
- The load test now runs five directors behind random routing with a shared
  inventory lock.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 15:34:23 -04:00
..