mirror of
https://github.com/stablyai/orca.git
synced 2026-10-03 08:02:12 +00:00
* fix(relay): admit drained hosts through their own lane and stagger their return A same-cap roll's drain sends every host on the isolated cell back to the director at once. Those hosts reconnect through the sticky lane (one slot per director), and each one's re-placement holds that slot for most of a second behind the region-wide inventory lock, so ordinary reconnects time out behind them and the drained hosts retry every 2 s: ~30k 503s per drain. The reconnect verification read now also says whether the host's home cell is isolated for a roll right now (same predicate re-placement uses). Those hosts release the sticky slot after the read and take a separate drain-return lane (1 per director, matching the store's per-director placement serialization). When that lane is full the host gets a Retry-After that reserves the next free service slot, paced by the measured re-placement time and capped at 300 s. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): keep drain returns inside the placement pool budget and their own slot Review follow-ups for the drain-return lane: - The lane now borrows placement permits (never placement's last, never ahead of a queued placement), so placement + sticky still bounds the database pool. - A host's own early retry (row-busy redial, duplicate dial) gets the 2 s lane interval instead of a fresh slot behind the cohort, and a host that returns early to the same director keeps its reserved slot. - Classification also excludes an open migration row whose lease counter lapsed, matching the re-placement rule. - The load test now runs five directors behind random routing with a shared inventory lock. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010