Files
orca/.github/workflows
Jinwoo Hong a07f4b2cbd feat(relay): slower same-cap drain paces (15 and 20 min) for every cell class (#26350)
* feat(relay): slower same-cap drain paces (900000, 1200000) for every cell class

c28's 1,901-host drain at the 300 s pace (~6.3 hosts/s, 2026-10-07) held
database lock time over its bar for about 10 minutes. Add 15- and 20-minute
pace windows to the same-cap closed set, allowed on every cell class.

- The cell's /v1/admin/drain cap rises from 5 to 20 minutes.
- The drain script fails instead of falling back to an unpaced drain when a
  cell rejects a window above 300000, since an older image answers both with
  the same 400.
- The fast-pace US-only rule and the canary PASS rule apply only below the
  default; a canary still authorizes its own pace or slower.
- The restart-safe timeout adds the window's excess over 300000 (quiet starts
  after the last host leaves), unchanged at 300000/60000/30000; the job
  timeout rises from 90 to 120 minutes.

* fix(relay): refuse a slow drain pace before isolating a cell that cannot accept it

The cell advertises its drain pace cap on /health; the same-cap job reads it
before the isolate, so a pace above an old image's 300000 cap stops with the
cell untouched instead of isolated. Also refreshes two stale timeout comments
and pins the drain step under its one-hour ID token and the job's 70-min
remainder.
2026-10-07 19:07:02 -04:00
..