* perf(relay): cut the cell LB drain to 60s and widen the same-cap batch to ten cells
Two independent sources of relay roll wall clock, neither of which protects a
host:
1. `connection_draining_timeout_sec` on the per-cell backend services was 300s.
The same-cap job drains every host off the cell to a restart-safe condition
before Terraform runs, so the LB drain only ever covers a host still
mid-handshake. Measured 2026-09-16 over ten same-cap cell jobs, it sat as
~5m55s of dead time between `Apply complete` and the old VM powering off,
inside an 8.5-minute `wait-until --stable` step. Now 60s, and pinned in the
topology `check` block beside the other fixed-one invariants.
2. The same-cap wave capped a batch at four cells, so a 22-cell roll needed six
batches, six single-use monitor gates, and a human handoff per batch. The
wave workflow now declares cell_1..cell_10 with the identical serial shape
and chaining, and the validator accepts two to ten.
The shared wave-index rule (`relay-monitor-evidence.mjs` and the relay-ops
preflight CLI) widens from 0-3 to 0-9 so the later cells can present the same
evidence; each job workflow keeps its own narrower range, so the capacity wave
stays at four. Cells remain strictly serial, one at a time behind the rollout
lease, each with its own live preflight.
Claude-Session: https://claude.ai/session/relay-roll-drain-timeout-and-batch-cap
* fix(relay): align the Asia topology plan validator with the 60s cell drain
`validate-relay-asia-topology-plan.mjs` rejected any Asia backend whose
`connection_draining_timeout_sec` was not 300, and
`cloud-deploy-relay-asia-topology.yml` targets
`google_compute_backend_service.relay_gce_cell["<cell>"]` per cell. With the
Terraform local at 60 that workflow would have failed its own plan review.
The validator's two restated topology values are now named exports, and a new
census test reads `relay-gce-cells.tf` and equates three statements of each:
the `relay_gce_topology` local, the topology `check` assert that pins it, and
the validator constant. Terraform cannot export a local to JS, so reading the
source is the only way to stop them drifting; the test was confirmed to fail
when the local alone is moved back to 300.
Repo-wide grep finds no other pin of the drain value.
Claude-Session: https://claude.ai/session/relay-roll-drain-timeout-and-batch-cap