* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run
A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed,
single-use, five-minute-fresh evidence. Each apply wave now samples fleet health
itself right before isolation, with the monitor's evaluator, thresholds, and
tolerances, for a window sized to the cell's host count (3/5/8 min), plus three
lookback rules: no cell container exit in 10 min, no minute over 500 director
503s in 10 min, and director concurrency p99 within the monitor bar over 4 min.
Removes the monitor-run inputs, the gate's consume/authorize steps, the
break-glass override, and the same-cap-only authorization shapes in
relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are
unchanged.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): bound the pre-drain sample overrun and keep the drain token fresh
Review follow-ups: alternating tolerated readings could hold the sample open
until its step timeout, so cap the overrun at three samples past the window;
record why a read failed; mint a fresh admin ID token for the drain after the
sample; raise the job timeout to 90 min so a long sample cannot cancel the
job past the failsafe.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule
The exit rule counted every relay container exit fleet-wide, so a cell that
crashes every few hours (c25, 12 a week) blocked the very roll that fixes it,
and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part
in. Exits are now grouped by instance, each instance is named by its own newest
runtime-metrics log line, and only exits on general or migration-only cells
other than the target count. An exit no configured cell can be named for trips
the rule; a failed lookup is a failed read. relay-observability.tf joins the
evidence-code set because the rule depends on its filter.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): let a same-cap drain finish when only unplaceable hosts remain
The c28 canary on 2026-10-01 drained the cell to zero live connections, but four
hosts with no free slot anywhere kept redialling and held director leases on it,
so the restart-safe wait timed out and left the cell isolated and empty.
The drain wait now also passes once the runtime has carried nothing for a
sustained quiet window while a small, capped number of leases remain, and logs
the escape. Apply modes also refuse a cell whose hosts exceed 80% of the free
slots on the other general cells, so a wave cannot strand hosts in the first place.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): define restart-safe by the cell runtime, not director leases
Replaces the opt-in stranded-host escape with a corrected definition. A
restart is safe when the cell runtime carries nothing live and no migration
is open, sustained for the drain pace window. Director activity leases lag
hosts that already left or cannot be placed, so they are reported in a
progress line and the verified result instead of blocking the restart.
The same-cap drain passes its existing pace window. The headroom script is
added to the trusted evidence code paths with the other production scripts.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): print stranded director counts on every restart-safe sample
Each restart-safe poll now prints its sample count and the director's
restart-blocking leases, request units, reserved remainder, and migrations
under `stranded`; the verified line carries the same object. Open migrations
still block because each is pinned to the cell incarnation a restart replaces.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): require the pace window for every live restart-safe wait
Pre-auth and total connections no longer reset the restart-safe window:
on drained c28 they flickered with unauthenticated redials in a third of
samples, which a restart does not lose. They stay in the progress output.
Every live restart-safe call must now pass --pace-window-ms. The capacity
job and staging proof drain unpaced, so they pass the production 300000 ms
window, and the calls that relied on the 180000 ms default get 480000 ms.
Headroom free slots now follow the director's placement rule: the admission
pause minus the larger of observed and enforced units, minus outstanding
control reservations.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010