Commit Graph
3 Commits
Author SHA1 Message Date
Jinwoo Hong f1ed355d06 feat(relay): drain pace window as a reviewed same-cap input, with drain-aware 503 gates (#25639)
* feat(relay): drain pace window as a reviewed same-cap input, with drain-aware 503 gates

The same-cap roll drained every cell over a fixed 300 s window, so a US roll
re-placed hosts at ~2/s and spent ~10 minutes draining and waiting for quiet.
The window is now a dispatch input from a closed set (300000, 60000, 30000),
defaulting to today's 300000.

- Below the default is refused for anything but US general cells; Asia drains
  are bound by their targets' accept rate, and migration-only cells hold no hosts.
- A non-default window must be named in the confirmation, the canary authority
  records it (v2), and a batch may run at its canary's window or slower only.
- Each cell job re-checks the window, scales the restart-safe timeout with it
  (15-min lease + window, unchanged at the default), and records what the cell
  applied and when it settled.
- The report-only shadow gate takes the director's drain-return deferrals out
  of the 503 count (window and baselines), adds the rung's 5-min non-drain 503
  budget and a Retry-After check, and reports the measured re-placement rate.
- relay-workflows.md documents the pace ladder and what each rung records.

* fix(relay): judge paced drains on counted 503s against the pre-drain minutes, and seal the canary's pace verdict

Review of #25639 replayed the shadow gate: it read 10-01 c29 as unverified (the
log read stopped at 20k entries), false-blocked 10-02 c22, and was blind on four
cells whose 24 h/48 h baseline held an incident.

- Director 503s now come from Cloud Run's request_count, aligned per minute by
  Cloud Monitoring, so volume cannot truncate the count.
- The background is the median of the 10 same-day minutes before the drain;
  the 24 h/48 h baselines are gone.
- Scheduled 503s come out: drain-return deferrals and sticky/placement answers
  to a host's own early retry (host-rate-limited, host-in-flight), each split
  across the minutes its 30 s sample covers.
- The rung budget counts only sustained breaches: two straight minutes over
  max(1.5x, +20) warn, over max(2x, +40) would-block.
- The report carries a paceVerdict over the three pace checks. seal_canary
  downloads the canary cell's report and seals that verdict; a batch below the
  default pace needs PASS, from a report on the same cell that drained at that
  pace.
- Docs: the step-down rule reads paceVerdict, 30 s waits for the lane service
  time (#25645), and the staging step is dropped since staging drains unpaced.

Replayed read-only: 10-01 c29 would-block (9 minutes over 41.5/min); all nine
10-02 cells and 10-01 c25 paceVerdict PASS.

* fix(relay): a partial count already past a block line blocks, in the shadow gate's Cloud SQL and pool checks

A truncated FATAL count is a floor, and one runtime sample over the SQL-failure
line is a fact, so neither waits for a complete read. The waiter-run rule still
needs a complete run, since holes can join two runs into one.

* fix(relay): a canary pace PASS needs a real cohort and whole telemetry; one median-based 503 check

From the final review of #25639:
- canaryPaceVerdict seals PASS only from a report that drained at least 400
  hosts (about half a 10-02 US cell), so a near-empty canary cannot authorize
  a fast batch.
- A director-metrics sub-window with fewer samples than one instance emits
  is unverified, so an empty or short Logging answer is never a calm drain.
- seal_canary names the shadow artifact, report path and cell from the gate's
  normalized cell list, as cell_1 uploads it.
- director503 folds into nonDrain503Budget as a single-minute spike rule,
  max(10x median, 200), dropping the pre-drain peak statistic. All 30 replayed
  windows keep their verdicts.
2026-10-05 17:31:29 -04:00
Jinwoo Hong e476193bf5 chore(relay): bound the shadow health gate and apply a pending backend update on resume (#21865)
* fix(relay): bound the same-cap shadow gate and apply a resumed backend update

Two findings both adversarial reviews of tonight's merged set agree on.

The report-only shadow health gate (#21849) had `continue-on-error: true` but
no step timeout. That bounds the step's contribution to the job outcome, not
its clock. Its reads are serialised, and a failure that answers nothing slowly
— an expired credential, a project-wide Logging 429 storm — makes every read
cost its full 3 x 60 s retry budget, so the cost scales with the roll window:
roughly 8S + 2 reads for S ten-minute sub-windows. A 40-minute window is about
34 reads, or 108 minutes, against the job's `timeout-minutes: 75`. A cancelled
job cannot be absorbed by continue-on-error, fires the failure-gated cleanup
isolation on an already-restored cell, and stops the strict next-cell chain.

Give the step `timeout-minutes: 5` and the artifact upload `timeout-minutes: 2`.
A timed-out step is a failed step, which continue-on-error covers, so the job
stays green. Inside the script, stop reading after an overall four-minute
deadline and report the remaining checks unverified, so the normal outcome is a
written verdict rather than a killed process; the step timeout is then only for
a hung process. The census test pins both timeouts and that the deadline leaves
the step time to write its verdict.

The resume branch (#21860) accepted `changes == 0` with a non-empty
`backendUpdate` as complete and applied nothing, so a resumed cell silently
kept the 300-second drain and no request logging behind a green resume. That
shape means the template and MIG are converged and only this cell's reviewed
backend update is left, so apply the saved resume plan — the validator has
already bounded it to this cell's backend and neither attribute restarts an
instance — then continue as converged. Template-and-MIG drift still applies
nothing, which is what a resume means, and a stranded cell's explicit MIG
replace is unchanged.

Claude-Session: relay-same-cap-gate-timeout-and-resume

* fix(relay): raise the shadow gate bounds clear of a healthy gate's read time

A healthy gate is already minutes of serial reads on the 2-vcpu runner, so a
four-minute deadline would report unverified tails on ordinary days and stop
the shadow roll measuring the comparison it exists for. Raise both together:
the step to eight minutes and the script's own deadline to seven, keeping the
census pin that the deadline leaves the step room to write its verdict. The
job budget is unaffected: a ~14-minute cell plus eight is well inside 75.

Claude-Session: relay-same-cap-gate-timeout-and-resume
2026-09-20 21:35:09 -04:00
Jinwoo Hong 5b8ac36f41 chore(relay): add a report-only post-wave health gate to the same-cap cell job (#21849)
* feat(relay): report a post-wave health verdict on each same-cap cell, without gating on it

After a same-cap cell finishes rolling, an operator reads five things by hand
before dispatching the next cell: director 503s against the same clock hour a day
and two days earlier, whether the cell's new container announced its listener and
has stayed up, the cell's own pool pressure, the asia-east2 pool trio, and Cloud
SQL FATALs. This runs those same reads automatically and records PASS / WARN /
WOULD_BLOCK with its numbers, so its calls can be compared with the operator's
over a full roll before it is ever allowed to stop one.

It cannot fail a cell in this change. The script exits 0 on every verdict, and
the step is continue-on-error, so even a crash stays off the job's outcome and the
failure failsafe cannot fire on anything it observes. It also runs after the
restore, so no cell waits on it to go back into admission.

Cloud Logging returns only --limit entries and says nothing when it truncates, so
every count is split into sub-windows of ten minutes and a sub-window that comes
back at the limit is reported unverified rather than as a count. Windows are
always explicitly bounded: --freshness does not bind on these logs.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): bound the shadow gate's cell reads at the apply start and cap every read

Four fixes from review, all in the report-only shadow health gate.

The boot search opened at apply-completed-at, which is stamped after
`terraform apply` and `wait-until --stable`. The new container announces its
listener while the MIG is still converging, so that bound is already past the
announcement it looks for and a healthy roll read as would-block. The job now
stamps apply-started-at immediately before the apply, and the boot search opens
there; apply-completed-at is kept, recorded rather than judged, so an operator
comparing verdicts can see apply time next to boot time.

The crash query started at the newest listener timestamp, which erased any crash
before it. A crash-restart loop ends with an announcement that looks like a clean
boot, so that is exactly the case it hid: against production, the 2026-09-20 c28
crash at 20:18:10 was dropped because the listener landed at 20:18:27. It now
runs from the apply start, still scoped to the instance id the listener
identified, and that crash is counted.

A runtime-metrics read that came back at its 500-entry limit fed judgePool as
though it were a complete sample run. A truncated run has holes and the
consecutive-sample rule reads a hole as a recovery, so it now reports unverified.

gcloud reads had no timeout. continue-on-error bounds the job's outcome but not
its clock, so a stalled read could have spent the rollout's remaining minutes.
Each read now gets 60 s and a timed-out read is just a failed read.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): require each shadow-gate stamp's presence before asserting its order

The ordering assertion used indexOf, which answers -1 for an absent stamp, and
-1 precedes every real offset. Deleting the apply-started-at line left the test
green, so the census could not see the fix it was written to pin.

Each stamp's presence is now asserted first, with a message naming the stamp and
the step, and presence is judged inside the step that owns the stamp rather than
anywhere in the file: a stamp written into a neighbouring step records the wrong
instant but would satisfy a whole-file match.

Control-run against a scratch copy of the job. Deleting drain-started-at,
apply-started-at, or apply-completed-at each reds with its own message, and
moving apply-started-at after terraform apply reds on the ordering assertion, so
presence and order both fail independently.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-20 19:10:35 -04:00