* fix(relay): commit the cell counter in one round trip; cells boot without the DB
Step 2 (option B) cell image:
- One-round-trip counter commit at acquireActivity, releaseActivity and
activateControl: the final counter UPDATE and COMMIT go as one simple-query
message. Server errors mean COMMIT never ran (retry as today; 22012 = no row,
rolled back and disambiguated outside the transaction); a lost connection is
never retried.
- Cells skip the schema apply and region backfill, so they listen while the
database is down and turn ready on their first successful query.
- G13: rehome target connection headroom folded into the existing NOWAIT
UPDATE, excluding the host's own reservation by key.
- fixLevel on every runtime metrics line, plus declared (not applied) cell
fix-level metrics and alert.
- Per-desktop drain disconnect-gap measurement from existing log lines.
- Census test that fails on floating database promises; fixes two shutdown
sites. Lock-wait sample keeps the combined role.
* fix(relay): make the outdated-image alert creatable: one PromQL condition, 1 h lookback, fixed floor
A PromQL condition must be the only condition in its policy, and alerts on
log-based metrics may look back at most 25 h. Replace the 6-day/7-day design
with relay_cell_min_fix_level (tfvars, raised by a targeted apply after each
wave) and one query: a serving cell below the floor or reporting no level,
sustained 6 h. Drops the separate without-level metric.
* fix(relay): review fixes: gap-script ordering, wider promise census, fused-path guard, row-busy as scheduled
- Drain gap script: sort closes by time (gcloud exports newest first) and
refuse an invalid drain start.
- Census: any floating promise in relay src, including callback-discarded
and never-read ones, with a reviewed never-rejects list.
- Test the fused counter commit through the store the server builds, so a
wrapper that stops forwarding commitWithFinal fails CI.
- Same-cap shadow gate: a row-busy refusal (the host's own release still
holds its row) is a scheduled 503, like an own early retry. No client
change.
* test(relay): judge drain redials by no host refused twice, not a refusal count
The row-busy count tracks how many releases are still in flight at the
dial (80 of 180 every run at 1 s, against a bar of 90). What matters is
that the release has finished by the next dial: assert no host is refused
twice, keep the time-to-placed p95 bound.
* fix(relay): cap row-busy as scheduled at the drain-return admissions; bound the gap script's window
Shadow gate: a row-busy refusal of a drained host follows its drain-return
lane admission, so per minute only that many (plus a rounding margin of 2)
are scheduled; the rest stay non-drain, so row contention the drain does
not explain still fails the budget. Gap script: --drain-ended-at excludes
the new container's closes after the roll; later grants still close a gap.