Files
orca/cloud/apps/relay
Jinwoo Hong a6e6de93c4 fix(relay): keep failed rehome polls out of the durable failure budget (#19915)
* fix(relay): keep failed rehome polls out of the durable failure budget

The regional rehome worker polls claimRegionalRehome about once a second.
Any error thrown before an attempt was claimed - in practice a director pool
timeout on the pre-claim control read, 52-74 a day against a pool of 3 - was
charged to relay_region_rehome_worker_state.consecutive_failures, which
durably disables the control at three. That counter only ever resets on a
drain receipt, so while the control is disabled it never resets: production
sits at 1068 and still climbing. Enabling the control leaves the stale
counter in place, so the next pool timeout latches it straight back off.
That is what ended the 2026-08-28 enable after ten minutes.

- A poll that never claimed an attempt drained nothing, so it no longer feeds
  the dispatch-failure budget and logs .._poll_failed instead of
  .._dispatch_failed. recordRegionalRehomeWorkerFailure had no other caller
  and is removed.
- Enabling the control clears consecutive_failures and paused_until, so a
  budget spent under a previous enable cannot kill a fresh one. The dispatch
  interval in next_dispatch_at is deliberately left alone.
- The budget's auto-disable now emits
  orca_relay_regional_rehome_failure_budget_disabled, matching the existing
  .._safety_disabled precedent. It wrote no event before, which is why this
  went unnoticed for two weeks.

No change to region selection, the candidate query, or host eligibility.

* fix(relay): serialize rehome failure accounting with control updates
2026-09-10 16:04:25 -04:00
..