mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 00:02:29 +00:00
017ad743fa52087d5662836abe070fb02609ba64
68
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
54ded3bc18 |
fix(relay): commit the cell counter in one round trip; cells boot without the database (#25765)
* fix(relay): commit the cell counter in one round trip; cells boot without the DB Step 2 (option B) cell image: - One-round-trip counter commit at acquireActivity, releaseActivity and activateControl: the final counter UPDATE and COMMIT go as one simple-query message. Server errors mean COMMIT never ran (retry as today; 22012 = no row, rolled back and disambiguated outside the transaction); a lost connection is never retried. - Cells skip the schema apply and region backfill, so they listen while the database is down and turn ready on their first successful query. - G13: rehome target connection headroom folded into the existing NOWAIT UPDATE, excluding the host's own reservation by key. - fixLevel on every runtime metrics line, plus declared (not applied) cell fix-level metrics and alert. - Per-desktop drain disconnect-gap measurement from existing log lines. - Census test that fails on floating database promises; fixes two shutdown sites. Lock-wait sample keeps the combined role. * fix(relay): make the outdated-image alert creatable: one PromQL condition, 1 h lookback, fixed floor A PromQL condition must be the only condition in its policy, and alerts on log-based metrics may look back at most 25 h. Replace the 6-day/7-day design with relay_cell_min_fix_level (tfvars, raised by a targeted apply after each wave) and one query: a serving cell below the floor or reporting no level, sustained 6 h. Drops the separate without-level metric. * fix(relay): review fixes: gap-script ordering, wider promise census, fused-path guard, row-busy as scheduled - Drain gap script: sort closes by time (gcloud exports newest first) and refuse an invalid drain start. - Census: any floating promise in relay src, including callback-discarded and never-read ones, with a reviewed never-rejects list. - Test the fused counter commit through the store the server builds, so a wrapper that stops forwarding commitWithFinal fails CI. - Same-cap shadow gate: a row-busy refusal (the host's own release still holds its row) is a scheduled 503, like an own early retry. No client change. * test(relay): judge drain redials by no host refused twice, not a refusal count The row-busy count tracks how many releases are still in flight at the dial (80 of 180 every run at 1 s, against a bar of 90). What matters is that the release has finished by the next dial: assert no host is refused twice, keep the time-to-placed p95 bound. * fix(relay): cap row-busy as scheduled at the drain-return admissions; bound the gap script's window Shadow gate: a row-busy refusal of a drained host follows its drain-return lane admission, so per minute only that many (plus a rounding margin of 2) are scheduled; the rest stay non-drain, so row contention the drain does not explain still fails the budget. Gap script: --drain-ended-at excludes the new container's closes after the roll; later grants still close a gap. |
||
|
|
c7d74b1160 |
feat(relay): give Asia cell c34 a promotion wave so it can become a general cell (#25757)
* feat(relay): give Asia cell c34 a promotion wave so it can become a general cell c34 launched on 2026-10-05 as a migration-only spare with no promotion path. This adds it to the Asia admission promotion waves, the workflow's promote and canary cases, and the canary evidence map, so the reviewed Asia workflow can promote it with the same five-minute canary c30 and c31 ran. The same-cap migration-only list is deliberately unchanged: a same-cap job reads a cell's class from that list, and c34 must be rolled to the director's image as a migration-only cell before promotion can run. The list moves after promotion, in its own change. Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041 * docs(relay): scope the c34 same-cap pause to the window after promotion Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041 * docs(relay): rewrap the c34 paragraph Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041 |
||
|
|
0e9e273b91 |
fix(relay): deploy driver builds the checked commit, prompts on their own line, summarises progress (#25755)
* fix(relay): publish the reviewed commit before main can move, and quiet the deploy driver The publish workflow builds main's head at dispatch. The driver checked main at preflight but dispatched the publish about four minutes later, after the inspects and the typed phrase, so a busy main stopped the first real deploy. It now dispatches the publish seconds after the check, before anything else, and a build of a moved main stops with the --commit/--publish-run command that deploys it once reviewed. Typed prompts end in a newline, and runs are summarised (status changes plus every 5 min) instead of streaming gh run watch. * fix(relay): a main-moved stop prints only the command that reuses the build The generic re-run line named the reviewed commit without --publish-run, which would only build the moved main again. |
||
|
|
ab91559fb4 |
feat(relay): one-command director deploy driver (#25642)
* feat(relay): add an operator-local driver for director deploys One command runs the audited director deploy: preflight, pause rehome if enabled, publish, deploy, optional cell configure, a digest-bound inspect, the monitor dry-run, and re-enable. It only dispatches the existing workflows, reads the published digest from the registry and the run log, reads the monitor verdict from its sealed state, and enables with the digests gcloud reports after the deploy. It stops at the first failure, records state, and resumes from it. * fix(relay): report and own the rehome pause; operator types every phrase - Read the control back from pause and enable runs whatever their conclusion, and report PAUSED or UNCONFIRMED loudly. - Resume re-enables only the pause this driver recorded (generation and run). - Ctrl-C and SIGTERM print the same state and resume report. - The operator types every workflow confirmation. A 5-minute soak gated on director 5xx runs before configure. - A dry run keeps no state file. The quiet check pages through all runs. Run IDs come only from the printed URL. One step table drives execute, dry run and resume. The monitor verdict reuses verify-authority. * fix(relay): read back only the driver's own rehome run; anchor the soak at the traffic switch - The pause and enable read-back accepts only the control line its own step's mode prints, at the generation its own dispatch expected. A run adopted after a crash is settled even when green. - The soak window opens a minute before the deploy run completed and is read a minute after it ends, for log ingestion lag. * fix(relay): say what typing ENABLE_REGIONAL_REHOMING commits to * fix(relay): the ENABLE prompt also names the 150 s evidence budget * refactor(relay): derive every deploy decision from live state; no resume machinery The driver keeps no state between runs. Each run reads the serving director, its configured cells and the rehome control, and skips what is already done. - The only rehome fact it owns is the run that paused rehome. A re-run names it (--pause-run), and the driver checks it against that run's log and the live generation. - A recover-enable line counts as the driver's own pause only with recovered: true. A director safety pause is never adopted (F3). - Publish runs before the pause. A fresh run that finds rehome paused stops unless given --pause-run or --rehome-disabled (F2). - Interrupts report a pause or enable still in flight as REHOME IS CHANGING (F1). PAUSE UNCONFIRMED and ENABLE UNCONFIRMED are distinct. - A tripped soak is judged again on fresh traffic (F5). SIGHUP is handled, and a pending signal stops the driver before its next dispatch. - Every stop prints the single command that finishes the deploy. Removes the state file, --resume, step statuses, monitor adoption and interrupted-dispatch adoption. * fix(relay): never report done over an unexplained pause; prove pause ownership by actor - G1: always read rehome; no early DONE. - G2: --pause-run must be a rehome-control run by the same user. - G3: a failed enable run is never an enable. - C1: a pause or enable is reported as changing from the moment it is dispatched. - C2: --leave-rehome-paused (was --rehome-disabled) refuses an enabled switch. Every printed command parses. - G4: the enable-in-flight report prints both finishing commands. - G5: the driver's own runs never block the quiet-lane check. * test(relay): port the round-3 probes: hard kill mid-pause, unexplained disable after a failed enable Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041 |
||
|
|
f1ed355d06 |
feat(relay): drain pace window as a reviewed same-cap input, with drain-aware 503 gates (#25639)
* feat(relay): drain pace window as a reviewed same-cap input, with drain-aware 503 gates The same-cap roll drained every cell over a fixed 300 s window, so a US roll re-placed hosts at ~2/s and spent ~10 minutes draining and waiting for quiet. The window is now a dispatch input from a closed set (300000, 60000, 30000), defaulting to today's 300000. - Below the default is refused for anything but US general cells; Asia drains are bound by their targets' accept rate, and migration-only cells hold no hosts. - A non-default window must be named in the confirmation, the canary authority records it (v2), and a batch may run at its canary's window or slower only. - Each cell job re-checks the window, scales the restart-safe timeout with it (15-min lease + window, unchanged at the default), and records what the cell applied and when it settled. - The report-only shadow gate takes the director's drain-return deferrals out of the 503 count (window and baselines), adds the rung's 5-min non-drain 503 budget and a Retry-After check, and reports the measured re-placement rate. - relay-workflows.md documents the pace ladder and what each rung records. * fix(relay): judge paced drains on counted 503s against the pre-drain minutes, and seal the canary's pace verdict Review of #25639 replayed the shadow gate: it read 10-01 c29 as unverified (the log read stopped at 20k entries), false-blocked 10-02 c22, and was blind on four cells whose 24 h/48 h baseline held an incident. - Director 503s now come from Cloud Run's request_count, aligned per minute by Cloud Monitoring, so volume cannot truncate the count. - The background is the median of the 10 same-day minutes before the drain; the 24 h/48 h baselines are gone. - Scheduled 503s come out: drain-return deferrals and sticky/placement answers to a host's own early retry (host-rate-limited, host-in-flight), each split across the minutes its 30 s sample covers. - The rung budget counts only sustained breaches: two straight minutes over max(1.5x, +20) warn, over max(2x, +40) would-block. - The report carries a paceVerdict over the three pace checks. seal_canary downloads the canary cell's report and seals that verdict; a batch below the default pace needs PASS, from a report on the same cell that drained at that pace. - Docs: the step-down rule reads paceVerdict, 30 s waits for the lane service time (#25645), and the staging step is dropped since staging drains unpaced. Replayed read-only: 10-01 c29 would-block (9 minutes over 41.5/min); all nine 10-02 cells and 10-01 c25 paceVerdict PASS. * fix(relay): a partial count already past a block line blocks, in the shadow gate's Cloud SQL and pool checks A truncated FATAL count is a floor, and one runtime sample over the SQL-failure line is a fact, so neither waits for a complete read. The waiter-run rule still needs a complete run, since holes can join two runs into one. * fix(relay): a canary pace PASS needs a real cohort and whole telemetry; one median-based 503 check From the final review of #25639: - canaryPaceVerdict seals PASS only from a report that drained at least 400 hosts (about half a 10-02 US cell), so a near-empty canary cannot authorize a fast batch. - A director-metrics sub-window with fewer samples than one instance emits is unverified, so an empty or short Logging answer is never a calm drain. - seal_canary names the shadow artifact, report path and cell from the gate's normalized cell list, as cell_1 uploads it. - director503 folds into nonDrain503Budget as a single-minute spike rule, max(10x median, 200), dropping the pre-drain peak statistic. All 30 replayed windows keep their verdicts. |
||
|
|
b94cfa7fd0 |
feat(relay): declare Asia spare cell c34 as migration-only (#25336)
Adds a sixth asia-east2 cell at the C31 shape (cap 3000, 6000 request units, pool 16, e2-standard-4) in asia-east2-c, pinned to the f30b5cb1 cell image. It gets its own topology and registration wave but no promotion wave, so the admission script and workflow refuse to promote it; it stays a migration-only landing zone and out of the fleet pool list. Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041 |
||
|
|
aac1c8f9c2 |
chore(relay): move US cells c32 and c33 to the general same-cap list after promotion (#25367)
Both were promoted to general on 2026-10-01 (selector 344 shows them general), but the same-cap job still classed them migration-only. Rollback mode on either would isolate (general -> migration-only) before its no-op check failed, demoting a live cell. They stay out of the fleet pool list (pool 10, not 16). Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041 |
||
|
|
07dad6739a |
refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run (#24443)
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed, single-use, five-minute-fresh evidence. Each apply wave now samples fleet health itself right before isolation, with the monitor's evaluator, thresholds, and tolerances, for a window sized to the cell's host count (3/5/8 min), plus three lookback rules: no cell container exit in 10 min, no minute over 500 director 503s in 10 min, and director concurrency p99 within the monitor bar over 4 min. Removes the monitor-run inputs, the gate's consume/authorize steps, the break-glass override, and the same-cap-only authorization shapes in relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are unchanged. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): bound the pre-drain sample overrun and keep the drain token fresh Review follow-ups: alternating tolerated readings could hold the sample open until its step timeout, so cap the overrun at three samples past the window; record why a read failed; mint a fresh admin ID token for the drain after the sample; raise the job timeout to 90 min so a long sample cannot cancel the job past the failsafe. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule The exit rule counted every relay container exit fleet-wide, so a cell that crashes every few hours (c25, 12 a week) blocked the very roll that fixes it, and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part in. Exits are now grouped by instance, each instance is named by its own newest runtime-metrics log line, and only exits on general or migration-only cells other than the target count. An exit no configured cell can be named for trips the rule; a failed lookup is a failed read. relay-observability.tf joins the evidence-code set because the rule depends on its filter. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
051ca4d34e |
feat(relay): declare US cells c32 and c33 at the 3,000-host shape (#24444)
* feat(relay): declare US cells c32 and c33 at the 3,000-host shape Declares two us-central1 cells at the Asia shape (cap 3000, 6000 request units, e2-standard-4) with the US default pool of 10, and generalises the Asia topology and admission ladder to derive each wave's region from its reviewed zone, leaving every Asia wave's behaviour unchanged. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * docs(relay): note the US canary tie-break and leave the fleet pool list to promotion Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): plan C32 and C33 as one topology wave The live-image overlay refuses a declared non-target cell with no template, so a lone C32 plan would fail on C33. Registration and promotion stay one cell at a time. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
744e7722c2 |
fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG (#24373)
* fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG The stranded-rollback recovery ran a gcloud rolling action, which renames the MIG version outside Terraform. Every later plan for that cell then reverted the label, and the capacity-plan validator refused the revert as an unreviewed MIG change, so the cell could be neither rolled nor rolled back. The validator now accepts a MIG field moving back to what relay-gce-cells.tf declares (version name and update policy), in every mode, and a test pins those values to the Terraform file. The stranded branch recreates the cell's single instance with recreate-instances, which leaves the MIG untouched. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): let a label-only MIG plan through and recreate on it in a stranded rollback A stranded rollback whose template is already in place plans only the version name revert. The validator still required the MIG template to move, so that plan was refused, and the recreate gate (changes == 0) would have skipped a plan of one change and left the drain flag set. Require the template move only when no declared field reconciles, and recreate whenever the template was not replaced (changes < 2). Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
c422936a71 |
fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup (#24349)
* fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup The same-cap gate now verifies the dry-run on its own clock and records the authorisation instant in the single-use consumed marker; each cell job checks the evidence age at that instant and bounds its own start after it. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): refuse a re-run same-cap gate before it consumes evidence; tighten order tests Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
5f728f7862 |
fix(relay): let a draining cell pass restart-safe through refused redials (#24347)
* fix(relay): let a draining cell pass restart-safe through refused redials A draining cell refuses every control and host proof, so once no session, splice, or queued byte remains, an in-flight or reserved connection unit can only belong to a dial the cell is about to refuse. The restart-safe wait no longer resets its pace-window streak on those units, and its progress line now prints every counter the gate reads. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * test(relay): pin fail-closed parsing of handshake counters on a draining cell Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
9ed7b39b3c |
fix(relay): run the same-cap headroom gate in the modes the job actually receives (#24343)
The parent workflow collapses canary-apply and batch-apply into the job mode apply, so the headroom step's canary-apply/batch-apply condition never held and the gate was skipped on every real roll. Run it wherever the drain runs (apply, rollback before its restart) and in read-only verify; skip only a resumed rollback, which drains nothing. A new workflow-shape test fails on any job step comparing against a mode the parent cannot pass, and on a drain that can run without the headroom check. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
6b8ca27d85 |
chore(relay): move Asia cell c31 to the general same-cap lists after promotion (#24318)
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
94f6c53387 |
fix(relay): gate the Asia canary on Asia-targeted region fallbacks only (#24320)
* fix(relay): gate the Asia canary on region fallbacks against a pre-canary baseline Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): gate the Asia canary on region fallbacks against a pre-canary baseline Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
3e5c8d9f8e |
feat(relay): declare Asia cell c31 at the c30 shape (#24310)
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
7c119465b0 |
fix(relay): define restart-safe by the cell runtime, and refuse waves without headroom (#24259)
* fix(relay): let a same-cap drain finish when only unplaceable hosts remain The c28 canary on 2026-10-01 drained the cell to zero live connections, but four hosts with no free slot anywhere kept redialling and held director leases on it, so the restart-safe wait timed out and left the cell isolated and empty. The drain wait now also passes once the runtime has carried nothing for a sustained quiet window while a small, capped number of leases remain, and logs the escape. Apply modes also refuse a cell whose hosts exceed 80% of the free slots on the other general cells, so a wave cannot strand hosts in the first place. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): define restart-safe by the cell runtime, not director leases Replaces the opt-in stranded-host escape with a corrected definition. A restart is safe when the cell runtime carries nothing live and no migration is open, sustained for the drain pace window. Director activity leases lag hosts that already left or cannot be placed, so they are reported in a progress line and the verified result instead of blocking the restart. The same-cap drain passes its existing pace window. The headroom script is added to the trusted evidence code paths with the other production scripts. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): print stranded director counts on every restart-safe sample Each restart-safe poll now prints its sample count and the director's restart-blocking leases, request units, reserved remainder, and migrations under `stranded`; the verified line carries the same object. Open migrations still block because each is pinned to the cell incarnation a restart replaces. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): require the pace window for every live restart-safe wait Pre-auth and total connections no longer reset the restart-safe window: on drained c28 they flickered with unauthenticated redials in a third of samples, which a restart does not lose. They stay in the progress output. Every live restart-safe call must now pass --pace-window-ms. The capacity job and staging proof drain unpaced, so they pass the production 300000 ms window, and the calls that relied on the 180000 ms default get 480000 ms. Headroom free slots now follow the director's placement rule: the admission pause minus the larger of observed and enforced units, minus outstanding control reservations. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
31012aeb09 |
test: remove assertion-free probes, copied inventories and export-shape checks (#23816)
Second audit wave, targeting three more junk patterns: - assertion-free cases that run code and assert nothing, so they pass no matter what the code does; - inventory literals re-typed from a production declaration, where the only way the assertion can fail is someone editing one of the two copies; - export key-set and export-shape loops (`typeof x === 'function'` over every export) that restate what TypeScript already enforces. Yield is much smaller than wave 1 on purpose: the assertion-free scanner has a high false-positive rate, because many flagged blocks assert through a shared helper or their oracle is "this must not throw". Those were kept. `mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is de-exported — after the inventory comparison went away, nothing outside the module read it. |
||
|
|
6e1b7e7fa3 |
test: remove junk tests that assert source text instead of behavior (#23815)
Deletes 101 test files and trims 112 more, all matching documented junk patterns: exact source/import/string greps, copied inventories and export lists, duplicate invocations of a contract another test already owns, typeof-shape checks TypeScript already enforces, and self-comparisons. The largest group read a production `.ts` file and asserted on its text — for example a TaskPage test that required the source to contain `selectedRepos.find((r) => r.id === newIssueRepoId) ?? selectedRepos[0] ?? null`. Any behavior-preserving rename broke it; no behavior change ever did. Production-side follow-through: exports that only these tests imported are de-exported or deleted, stale comments pointing at removed censuses are dropped, and the reliability-gate registry, `cloud/package.json` test lists, and orphaned source-reading helpers are updated so nothing references a deleted file. Two files kept their real coverage and lost only the census scaffolding: `agent-status-producer-census.test.ts` now drives all five producers end to end instead of grepping the source tree, and `config-toml-trust-stale-writes` replaces an export-list parity check. |
||
|
|
818237b9e3 | fix: stop retired relay load connections from rescheduling refreshes (#23071) | ||
|
|
128e97ffca |
fix(relay): drain a same-cap cell over 5 minutes, not 2 (#22584)
* fix(relay): drain a same-cap cell over 5 minutes, not 2 The 2026-09-23 c27 roll drained 2,145 hosts over the 2-minute window, about 18 re-dials/s, while the director re-places roughly 8/s through its single-slot sticky lane. The overflow queued behind slow re-placements and timed out, so /v1/assign returned 503 fleet-wide for about 5 minutes. 5 minutes is the cell's maximum pace window and keeps the remaining 2,650-host cells near lane capacity. The transition wait already outlasts a 5-minute window (17 min). Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * test(relay): pin the same-cap drain window contract at 5 minutes Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): keep the drain wait at lease plus the 5-minute window The transition wait after a drain was set to the 15-minute migration lease plus the pace window. Widening the window to 5 minutes without moving the wait left 12 minutes for a migration that can hold for 15, so a late-window migration would time the wave out into rollback. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
1116c54230 |
fix(relay): accept production cells past c29 in the regional rehome operator (#22518)
The selector membership check capped cell ids at c29, so enable failed closed with "selector membership is invalid" once c30 went general. Accept c1-c99 with no leading zero. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
a2a78ab335 |
feat(relay): alert on relay cell table lock convoys (#22446)
* feat(relay): alert on relay cell table lock convoys Adds a log-based metric and alert for cell-inventory lock holds of at least 1,000 ms, and a Cloud SQL log metric and alert for relay-only lock timeout cancels at 20 or more per minute. NOWAIT refusals are excluded: background sweeps produce about 160 per minute even with rehome paused. Replayed over 2026-09-20 14:00 to 2026-09-22 15:00 UTC: the hold filter matches all 93 asia-east2 rehome holds plus 9 director holds, and every one of the 88 cancel burst minutes overlaps an asia-east2 hold. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): page only on cell lock holds; director holds stay visible Director holds of 1-2.5 s recur several times a day with rehoming paused, and pausing rehome does not stop them. The paging hold policy now selects role=cell samples only; a separate policy with no notification channel keeps director holds visible. The burst documentation no longer claims no burst happens while paused, and the runbook points a burst with no cell hold at the director policy. Replayed cell-only: 93 of 93 asia-east2 holds, 0 director holds over 2026-09-20 14:00 to 2026-09-22 15:00 UTC; 0 from then to 2026-09-23 07:30. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
51c3434851 |
chore(relay): treat Asia cell c30 as a general cell now that it is promoted (#22439)
C30 was promoted to general on 2026-09-23 (selector generation 286). The same-cap wave now rolls it as a general cell instead of handing it back isolated, and the shadow gate reads its pool beside C27-C29. Follow-up to #22375. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
9fef7a0f04 |
fix(cloud): gate the Asia canary on its own cell's SQL failures, not the directors' (#22405)
* fix(cloud): gate the Asia canary on its own cell's SQL failures, not the directors' The production canary summed sqlFailuresDelta over every director and the canary cell and required zero. Directors log a steady baseline of relay_cells NOWAIT and lock-timeout refusals unrelated to the canary cell, so a C30 canary failed most attempts on that noise. The canary now requires zero SQL failures from the canary cell's own metrics and records the director sum as directorSqlFailures without gating on it. Directors keep every other rule (unavailable regions, fallbacks, pool waiting, transient waiter and wait-time bounds). Staging keeps the combined zero rule and its evidence shape unchanged. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(cloud): gate the Asia canary's pool bounds on its own cell too Directors also show a steady pool-wait baseline (waiting above zero and waits over 50 ms in about 6 of every 60 minutes), so a five-minute canary still failed about half the time on director pool pressure unrelated to the canary cell. With gateDirectorDatabase off, the production canary now applies databasePoolWaitingMax, databasePoolWaitersMax and databasePoolWaitMsMax to the canary cell's metrics only and records the director values under director-prefixed names. Directors still gate Asia selections, region fallbacks and unavailable regions. The staging path keeps its combined values, key order and validation order. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
483fa0aca2 |
fix(cloud): compare the Asia topology budget gate against the measured 500-connection default (#22386)
The topology workflow's Cloud SQL gate carried a hard-coded 400 for the instance's tier default while the consumer contract records the value measured on the live instance (SHOW max_connections = 500, 2026-09-16, #21163). The gate compares the two and the first production plan run (35815654836) failed silently on that mismatch before Terraform ran. The verified default now lives beside the tier and version it is verified for, as VERIFIED_DEFAULT_MAX_CONNECTIONS, so the contract and the workflow are two independent records of the same measurement and the gate keeps its cross-check. The test pins the new source and forbids a bare literal. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
bf3f95245c |
feat(relay): declare Asia cell c30 at the c27 shape (#22375)
* feat(relay): declare Asia cell c30 at the c27 shape Adds production-gce-c30 in asia-east2-a at the reviewed Asia shape (6,000 request units, 3,000/60 connection limits, 16-connection pool, disabled) and the rehome trust the other Asia cells carry. Every Asia enumeration now knows C30. The topology, admission, and director tools treat it as its own reviewed wave so its plan and registration never touch the live launch cells. C30 promotion requires C27 general and fresh staging evidence. The topology and director validators now pin the committed production pool of 16 instead of the stale 10, which had made the topology workflow reject the committed launch cells. * fix(relay): plan C30 at live images and prove it with its own canary The shared URL map pulls every cell into the C30 topology plan, so the workflow now plans each non-target cell at the image its live template serves, and the validator names any change to a cell outside the wave. C30 promotion runs the same five-minute production canary and automatic rollback C27 used, with the load report proving the canary control was placed on C30, instead of relying on staging evidence. C30 leaves the shadow gate's fleet pool list until it serves, rollback rejects mixed partial sets, and a budget test pins the mixed-Asia-pool refusal. * fix(relay): pin C30 to the production director's live image digest C30 promotion requires the director and C30 to report one digest, so C30 takes the director's sha256:4158d8a2 (read 2026-09-22). C27-C29 keep their committed lines; every Asia check compares only the cells named in a run. * fix(relay): read the committed cell map from a plan, not console terraform console evaluates every output against state, and the Relay deployments output indexes each cell's MIG, so it fails with Invalid index while C30 is declared but not created. Read the map from a no-refresh, unlocked plan over the same targets instead, and refuse empty overlay input. * fix(relay): keep console readers working and C30 migration-only until promotion relay_gce_cell_deployments indexed each cell's MIG, backend, and template, so once C30 is declared but not applied every production terraform console reader printed a warning to stdout and broke its jq parse. Wrap those six lookups in try(..., null). Same-cap listed C30 as general, so a rollback dispatch on a migration-only C30 would restore it with activate and skip its canary. List it with the migration-only cells until the promotion follow-up moves it. |
||
|
|
8568d77b06 |
fix(relay): let a same-cap cell plan delete the deposed template a failed wave left behind (#22170)
A same-cap wave applies with create_before_destroy. When run 35684694704 died after the new template was made, the previous template stayed in Terraform state as a deposed object, so every later plan for that cell carried its delete. The plan validator only tolerated deposed deletes in same-cap-image mode, so the recovery run 35698133226 was refused with "cell plan must change only the exact instance template and MIG" and the cell was stranded. Allow exactly one deposed delete at the cell's own template address in same-cap-cell mode too, mirroring the one-deposed bound the convergence path already applies to non-image modes. Report it as `obsoleteTemplates` rather than in `changes`, the way the cell backend update is already split out, so the job's `changes == 2` resume gate and `changes == 0` stranded roll keep reading the template-and-MIG count. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
bf8d63bc67 |
fix(relay): keep the backend service out of the same-cap wave; the capacity role cannot update it (#22140)
The capacity role the same-cap wave authenticates as, orcaRelayProductionCapacity, has no compute.backendServices.update. Since #21860 added `google_compute_backend_service.relay_gce_cell["${TARGET_CELL_ID}"]` to both of the job's plan invocations, every wave has therefore created the new instance template, modified the MIG, and then failed 403 on the backend, leaving the cell isolated with its trust probe, admission restore, and shadow gate all skipped. Run 35684694704 on production-gce-c7 is the first one that hit it in production. Drop the backend target from both plans and restore the resume gate to exactly `.changes == 2` (the template-and-MIG rollback-image drift) or a converged plan, removing the backend-only resume apply #21865 added on top. A resume applies nothing again, which is what a resume means. The validator keeps its bound on a cell backend update, so it still reports one and refuses anything wider, but a wave plan can no longer contain one. The drain timeout from #21848 and the log_config from #21860 need a root apply by a principal that holds the permission; granting the capacity role that permission is itself a root apply, so it can follow as its own change rather than blocking every wave in the meantime. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
c8a5580659 |
fix(relay): re-place hosts off a cell isolated for a roll (#21911)
* fix(relay): re-place hosts off a cell isolated for a roll A roll isolates a cell by moving it out of the 'general' admission class; the cell then refuses every attach with 4503. The director never noticed, because the only liveness test it applies to a host's current cell reads `relay_cell_runtime.ready` and the heartbeat, and an isolated cell keeps heartbeating ready=1 for the whole drain. So every host on that cell was handed its own dead cell, closed, and handed it back — 500-1,900 hosts looping for 13-16 minutes per cell roll, at ~6 dials each per minute, with no neighbour absorbing anything. The sticky lane now treats a live incumbent whose admission is 'migration-only' — the state a roll's isolate step writes — the same way it treats a dead one: it returns null, which means "fall through to placement". The placement lane had the identical hole eleven lines further down, so it takes the same predicate; without that second swap the sticky change is inert, because placement would hand the pin straight back (a draining cell has more headroom than anyone). An isolated incumbent skips the dead-cell fence branch: that branch exists to prove an unreachable cell stopped serving a host, and this one is reachable and enforces the epoch itself. 'existing-only' is deliberately untouched — those cells serve the hosts they already hold, and only `assignmentStrandedOnUnservedCell` may release that pin. A host with an open `relay_assignment_migrations` row keeps its pin too, so this stays disjoint from the migration machinery. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(relay): gate re-placement on a roll-isolation marker, not on admission Review of the first commit found the predicate wrong. `migration-only` is an admission class, not a drain signal: an Asia `--mode rollback`, an evacuation or forward-recovery target awaiting a separate promote dispatch, a failed same-cap wave's re-isolate, an abandoned migration retired on its target and a rehome settlement all park loaded cells there durably, with no migration lease and no open migration row. All five were indistinguishable from a roll's isolate, so the first commit would have converted `operate-relay-asia-admission --mode rollback` from a reversible admission flip into a mass move of ~4,000 hosts — and, because `leastLoadedCell` treated region as a preference, into us-central1. The signal is now an explicit stamp. `relay_cell_admission` gains a nullable `roll_isolated_at`, added through the shared schema runner's catalog pre-check so a migrated database takes no relation lock on boot and an un-migrated one gets a catalog-only rewrite. The same-cap isolate step is its only writer, via a new optional `rollIsolatedCells` on the selector apply; the same UPDATE that writes the state clears the stamp whenever a cell leaves 'migration-only', so a restore cannot leave one behind and a failed wave's re-isolate keeps the one it has. Every other admission writer omits the field, so its cells stay unmarked and their hosts stay pinned. Old directors ignore the field; old callers never send it. Region is now a constraint rather than a preference on this path only: a re-placement must find a general, live cell with connection headroom in the host's own region, or the pin is kept and one `orca_relay_sticky_replacement_deferred` event is logged. Cross-region spill is no longer reachable here. The fence bypass is narrowed to a live incumbent. It was always a no-op for the intended case, and for a stamped cell that stops heartbeating while still holding sockets it reopened split-brain; that cell now takes the dead-cell path unchanged. Also: the hot-path admission reader no longer throws on an unrecognised state — it sits on every sticky dial and the rule it feeds is "move the host", so an unreadable row has to mean "don't". And the sticky lane reads the admission row once for both the stranded rule and the stamp instead of twice. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(relay): emit the re-placement events after the transaction commits CodeRabbit on assignment-store.ts:1075. Both events were written where they are decided, which is inside assignOnce's transaction. A reservation or lease write failing after that point rolls the placement back, but a line already on stdout cannot be rolled back with it — so the canary this PR asks an operator to read would count re-placements that never happened, and a Postgres transaction retry could leave a stale line behind as well. The transaction now returns its events alongside the RelayAssignment and the caller flushes them once it has resolved. Returning them rather than setting a variable in the enclosing scope is what makes the retry case safe too: only the attempt that committed can carry its events out. assign()'s signature is unchanged; the extra shape lives entirely inside assignOnce. orca_relay_sticky_replacement_deferred was moved the same way. It cost one more push into the array that already existed, and it is decided inside the same transaction, so leaving it behind would have been the odd case rather than the cheap one. The new test injects a failure on the first write after the decision, asserts no event is emitted, and asserts the assignment is still on its original cell — without that second assertion the absence would only prove the emit was early, not that it would have been wrong. A control dial with nothing injected emits exactly one event, so the case cannot pass on a broken harness. With the emit put back inside the transaction, it fails. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(relay): expire the roll stamp, correct the wire note, assert the stamp landed Delta review findings B, D and E. A (the deferral path's cost) is deliberately not implemented; it is now written up under Follow-ups in the PR body as required before any Asia roll, because it cannot fire in a US canary. B, which also closes C: the stamp was written, carried and never compared to anything. A roll isolates and restores one cell inside ~15 minutes, so a stamp older than two hours is not a roll in progress. It is a failed wave whose failsafe re-isolated a possibly healthy cell and is waiting on an operator — the postmortem in this tree records gaps of hours — or an orphan left by a director rollback whose restore wrote 'general' without the clause that clears the stamp, which the selector's 'keep' branch would then preserve until some later park reactivated it. Both want the same answer and it is the pre-existing one: keep the pin. One comparison against a value already on the row. The bound takes the caller's `now` rather than reading the clock again, so one assign reasons about one instant; the stamp's age is now a thing that decides whether a host moves, and two clock reads could disagree across it. D: the comment beside the new request field claimed an updated caller reaching an older director "is simply ignored". The schema is .strict(), so it is a 400. That fails closed — the isolate aborts before MUTATION_STARTED is set and nothing is written — but it is a deploy ordering constraint, and it was undocumented. The comment now says so and the PR body's rollout notes carry it. E: nothing read the `rollIsolated` the script already prints, so an older script against a newer director would silently produce today's behaviour and the canary would read as "the fix did nothing" with no way to tell that from a wrong premise. Both isolate steps now assert it, beside the generation they already parse. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
e476193bf5 |
chore(relay): bound the shadow health gate and apply a pending backend update on resume (#21865)
* fix(relay): bound the same-cap shadow gate and apply a resumed backend update Two findings both adversarial reviews of tonight's merged set agree on. The report-only shadow health gate (#21849) had `continue-on-error: true` but no step timeout. That bounds the step's contribution to the job outcome, not its clock. Its reads are serialised, and a failure that answers nothing slowly — an expired credential, a project-wide Logging 429 storm — makes every read cost its full 3 x 60 s retry budget, so the cost scales with the roll window: roughly 8S + 2 reads for S ten-minute sub-windows. A 40-minute window is about 34 reads, or 108 minutes, against the job's `timeout-minutes: 75`. A cancelled job cannot be absorbed by continue-on-error, fires the failure-gated cleanup isolation on an already-restored cell, and stops the strict next-cell chain. Give the step `timeout-minutes: 5` and the artifact upload `timeout-minutes: 2`. A timed-out step is a failed step, which continue-on-error covers, so the job stays green. Inside the script, stop reading after an overall four-minute deadline and report the remaining checks unverified, so the normal outcome is a written verdict rather than a killed process; the step timeout is then only for a hung process. The census test pins both timeouts and that the deadline leaves the step time to write its verdict. The resume branch (#21860) accepted `changes == 0` with a non-empty `backendUpdate` as complete and applied nothing, so a resumed cell silently kept the 300-second drain and no request logging behind a green resume. That shape means the template and MIG are converged and only this cell's reviewed backend update is left, so apply the saved resume plan — the validator has already bounded it to this cell's backend and neither attribute restarts an instance — then continue as converged. Template-and-MIG drift still applies nothing, which is what a resume means, and a stranded cell's explicit MIG replace is unchanged. Claude-Session: relay-same-cap-gate-timeout-and-resume * fix(relay): raise the shadow gate bounds clear of a healthy gate's read time A healthy gate is already minutes of serial reads on the 2-vcpu runner, so a four-minute deadline would report unverified tails on ordinary days and stop the shadow roll measuring the comparison it exists for. Raise both together: the step to eight minutes and the script's own deadline to seven, keeping the census pin that the deadline leaves the step room to write its verdict. The job budget is unaffected: a ~14-minute cell plus eight is well inside 75. Claude-Session: relay-same-cap-gate-timeout-and-resume |
||
|
|
2524737ef0 |
chore(relay): apply the cell backend drain and request-logging settings inside each same-cap wave (#21860)
* chore(relay): target each cell's backend service from the same-cap job
The 60 s connection drain timeout merged in #21848 has no safe apply path.
A root plan scoped to the backend services alone still pulls every
`google_compute_instance_template.relay_gce_cell` in as a dependency, and
standing image drift turns all 29 into replacements, so applying it would roll
the fleet at once.
Add `google_compute_backend_service.relay_gce_cell["${TARGET_CELL_ID}"]` to
both plan invocations in the per-cell same-cap job, next to the template and
MIG it already targets, and teach the reviewed plan validator to allow exactly
one extra change: an in-place update of that one cell's backend whose only
changed attribute is `connection_draining_timeout_sec`, landing on the
constant `validate-relay-asia-topology-plan.mjs` exports. Any other attribute,
any other resource, or a backend for another cell still fails the validator.
The accepted update is reported as `connectionDrainUpdate` and kept out of
`changes`, so the apply step's stranded branch and the resume step's drift
branch keep reading the template-and-MIG count they were written against; the
resume branch additionally accepts a plan whose only pending change is that
drain update, which restarts nothing.
Claude-Session: relay-same-cap-targets-cell-backend
* fix(relay): also let the same-cap wave apply this cell's LB request logging
A read-only production plan for production-gce-c7 showed the live US cell
backends carry no `log_config` at all, while relay-gce-cells.tf has declared
`log_config { enable = true, sample_rate = var.relay_gce_cell_log_sample_rate }`
on every cell backend since the Terraform root landed in
|
||
|
|
5b8ac36f41 |
chore(relay): add a report-only post-wave health gate to the same-cap cell job (#21849)
* feat(relay): report a post-wave health verdict on each same-cap cell, without gating on it After a same-cap cell finishes rolling, an operator reads five things by hand before dispatching the next cell: director 503s against the same clock hour a day and two days earlier, whether the cell's new container announced its listener and has stayed up, the cell's own pool pressure, the asia-east2 pool trio, and Cloud SQL FATALs. This runs those same reads automatically and records PASS / WARN / WOULD_BLOCK with its numbers, so its calls can be compared with the operator's over a full roll before it is ever allowed to stop one. It cannot fail a cell in this change. The script exits 0 on every verdict, and the step is continue-on-error, so even a crash stays off the job's outcome and the failure failsafe cannot fire on anything it observes. It also runs after the restore, so no cell waits on it to go back into admission. Cloud Logging returns only --limit entries and says nothing when it truncates, so every count is split into sub-windows of ten minutes and a sub-window that comes back at the limit is reported unverified rather than as a count. Windows are always explicitly bounded: --freshness does not bind on these logs. Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): bound the shadow gate's cell reads at the apply start and cap every read Four fixes from review, all in the report-only shadow health gate. The boot search opened at apply-completed-at, which is stamped after `terraform apply` and `wait-until --stable`. The new container announces its listener while the MIG is still converging, so that bound is already past the announcement it looks for and a healthy roll read as would-block. The job now stamps apply-started-at immediately before the apply, and the boot search opens there; apply-completed-at is kept, recorded rather than judged, so an operator comparing verdicts can see apply time next to boot time. The crash query started at the newest listener timestamp, which erased any crash before it. A crash-restart loop ends with an announcement that looks like a clean boot, so that is exactly the case it hid: against production, the 2026-09-20 c28 crash at 20:18:10 was dropped because the listener landed at 20:18:27. It now runs from the apply start, still scoped to the instance id the listener identified, and that crash is counted. A runtime-metrics read that came back at its 500-entry limit fed judgePool as though it were a complete sample run. A truncated run has holes and the consecutive-sample rule reads a hole as a recovery, so it now reports unverified. gcloud reads had no timeout. continue-on-error bounds the job's outcome but not its clock, so a stalled read could have spent the rollout's remaining minutes. Each read now gets 60 s and a timed-out read is just a failed read. Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010 * test(relay): require each shadow-gate stamp's presence before asserting its order The ordering assertion used indexOf, which answers -1 for an absent stamp, and -1 precedes every real offset. Deleting the apply-started-at line left the test green, so the census could not see the fix it was written to pin. Each stamp's presence is now asserted first, with a message naming the stamp and the step, and presence is judged inside the step that owns the stamp rather than anywhere in the file: a stamp written into a neighbouring step records the wrong instant but would satisfy a whole-file match. Control-run against a scratch copy of the job. Deleting drain-started-at, apply-started-at, or apply-completed-at each reds with its own message, and moving apply-started-at after terraform apply reds on the ordering assertion, so presence and order both fail independently. Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
4f839cc8c9 |
chore(relay): cut the cell LB connection drain to 60 s and allow ten-cell same-cap batches (#21848)
* perf(relay): cut the cell LB drain to 60s and widen the same-cap batch to ten cells Two independent sources of relay roll wall clock, neither of which protects a host: 1. `connection_draining_timeout_sec` on the per-cell backend services was 300s. The same-cap job drains every host off the cell to a restart-safe condition before Terraform runs, so the LB drain only ever covers a host still mid-handshake. Measured 2026-09-16 over ten same-cap cell jobs, it sat as ~5m55s of dead time between `Apply complete` and the old VM powering off, inside an 8.5-minute `wait-until --stable` step. Now 60s, and pinned in the topology `check` block beside the other fixed-one invariants. 2. The same-cap wave capped a batch at four cells, so a 22-cell roll needed six batches, six single-use monitor gates, and a human handoff per batch. The wave workflow now declares cell_1..cell_10 with the identical serial shape and chaining, and the validator accepts two to ten. The shared wave-index rule (`relay-monitor-evidence.mjs` and the relay-ops preflight CLI) widens from 0-3 to 0-9 so the later cells can present the same evidence; each job workflow keeps its own narrower range, so the capacity wave stays at four. Cells remain strictly serial, one at a time behind the rollout lease, each with its own live preflight. Claude-Session: https://claude.ai/session/relay-roll-drain-timeout-and-batch-cap * fix(relay): align the Asia topology plan validator with the 60s cell drain `validate-relay-asia-topology-plan.mjs` rejected any Asia backend whose `connection_draining_timeout_sec` was not 300, and `cloud-deploy-relay-asia-topology.yml` targets `google_compute_backend_service.relay_gce_cell["<cell>"]` per cell. With the Terraform local at 60 that workflow would have failed its own plan review. The validator's two restated topology values are now named exports, and a new census test reads `relay-gce-cells.tf` and equates three statements of each: the `relay_gce_topology` local, the topology `check` assert that pins it, and the validator constant. Terraform cannot export a local to JS, so reading the source is the only way to stop them drifting; the test was confirmed to fail when the local alone is moved back to 300. Repo-wide grep finds no other pin of the drain value. Claude-Session: https://claude.ai/session/relay-roll-drain-timeout-and-batch-cap |
||
|
|
68b11282a5 |
fix(relay): let the rehome evidence parser read a line the director grew (#21823)
The enable workflow reads the director's `[orca-relay] regional rehome inventory` line out of Cloud Logging and pins the whole line with one regex. Adding `hostNotArrivedLast24Hours` in #21813 made every healthy line stop matching, so "Read fresh aggregate completion and abort evidence" threw "no aggregate regional rehome inventory evidence" and the fail-closed step disabled the durable switch at control generation 26. The parser now requires the six original fields and tolerates further ones in any order. Extra fields stay fenced by value shape rather than by pinning the whole line: a field must be a bare name and a non-negative integer or `none`, so `hostId=someone` is still not a counter and cannot ride along. An absent count reads as null, not zero, because an older director not reporting leaks is not the same as reporting none. `hostNotArrivedLast24Hours` and `oldestActiveAgeMs` now reach the evidence JSON and the operator step summary. Two guards close the chain, each verified to fail on the regression it exists for: a census in the relay package feeds the real formatter's output to the real parser, and a script-side test pins the parser's output to the fields the workflow summary renders. Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
164c7140fd |
fix(relay): stop reporting an unavailable home cell as exhausted capacity (#21518)
A host whose home cell is not live — readiness false, drained, or inside a boot window — is refused by the committed-fence branch in assignOnce() without any capacity being consulted. It answered relay_capacity_exhausted, so every cell boot and every readiness dip printed capacity rejections at 17% fleet utilisation and sent an investigation after headroom that was never short. The branch now raises RelayHomeCellUnavailableError, which carries the cell id and which of cellIsLive()'s conditions failed (draining / booting / unheard / not_ready). The director logs reason, cause and cell, and returns the new reason in the same retryable 503. Nothing on the wire reads the body: the desktop client discards it unread and branches on status only, and no log-based metric or alert parses the reason. The load harness, the only body-reading consumer, gets its own bucket so a home-cell rejection no longer inflates the capacity count. Hinted grants are now logged on whichever lane served them, so a host that failed sticky verification and was rehomed by placement leaves a record of where it landed. Unhinted placement grants stay silent. |
||
|
|
6c913a917f |
fix(relay-ops): roll a cell a wave stranded after its drain (#21321)
* fix(relay-ops): roll a cell a wave stranded after its drain A wave that stops any time after its drain leaves the cell migration-only and draining on the rollback image, and nothing clears it: the drain flag is a one-way latch on the running process, and the failsafe restarts nothing. Both recovery modes then refuse the cell. Apply wants it general and not draining. Rollback sees the rollback image, reads it as a resume, refuses the draining, and would not have restarted it anyway. The image alone cannot separate a rollback that failed after its template apply from a wave that stopped before one. The restart can: the first left a fresh process, the second did not. Classify on that, so the cell that never restarted takes the rolling path instead of the resuming one. Its template still carries the image it serves, so that is the predecessor its plan is reviewed against, and a template already moved on to the target is refused rather than rolled backwards under a stale review. When the reviewed template is already in place the plan changes nothing, so the MIG is rolled explicitly on the same replacement policy a template change uses; the existing incarnation check is what proves the instance came back. Every other combination of mode, live image, and drain flag keeps the value it had, held by a census that runs the real block over all nine. * fix(relay-ops): pin the replacement method on the explicit MIG roll gcloud persists every rolling-action bound into the group's update policy, and it defaults the replacement method to substitute on a group with no stateful config. Passing surge and unavailable without the method would patch the policy off the declared RECREATE, and the next targeted plan would then carry a MIG change outside version.0.instance_template, which the plan validator refuses. Pass all three so the patch is identical to the declared policy, and read the declared values in the census instead of restating two of them. Dropping the flag, or moving any of the three in Terraform, now fails the census. |
||
|
|
27bddc6198 |
fix(relay-ops): accept a drained predecessor on a cell that holds no hosts (#21315)
c17's canary stopped at the pre-apply predecessor check with `runtime predecessor mismatch fields=draining`. The flag is residue: the previous canary (run 35290908836) drained c17 at 00:26:01, its terraform apply then failed, and the failsafe re-isolates without restarting the VM, so nothing cleared it. The same run had passed this very check a second earlier, which is what proves a parked cell is not draining at rest. Draining means connections are being shed, and a migration-only cell holds none, so the flag is not a precondition there. Accept it on entry for that class only. The replacement VM is still required not to be draining, on every path, and the incarnation check still proves it was replaced. Both predecessor checks now read one decision instead of computing the rule twice, so the assertion and its diagnostic cannot disagree. Every general-cell and rollback path keeps the value it had; a census test runs the real block over all eight mode and class combinations to hold that. |
||
|
|
acedcf2a97 |
fix(relay-ops): pin the capacity identity so a stale same-cap template can roll (#21314)
c17's canary-apply failed closed at plan validation. Its instance template is from 2026-08-07 and predates the ORCA_RELAY_CAPACITY_SERVICE_ACCOUNT line that every cell rolled since already carries, so the plan legitimately added it. The same-cap validator holds the whole startup script identical before and after except the image, and that line is not one it excluded, so the wave stopped with nothing applied. Pin the line for same-cap-cell exactly as bootstrap-cell already does, and exclude it from the before/after comparison. The cell may gain it; the pin is what refuses a roll that drops it or rewrites it to another identity. Both plan validations in the job now pass the capacity identity the job already requires. The same-cap contract is otherwise unchanged: any other stale line still fails closed, and needs a convergence apply before the cell can roll. |
||
|
|
ff8f7085cc |
fix(relay-ops): bind the canary cell's admission class into same-cap batch authority (#21313)
A batch-apply wave verified only that the sealed canary named some approved same-cap cell, so a canary rolled on the migration-only, zero-host, 600-cap c17 or c18 was accepted as authority for a general 1000/3000-cap batch. The verify step now hands the batch's own cells to the check, which requires the sealed cell's entry admission to equal the batch's class. |
||
|
|
399306c171 |
feat(relay-ops): allow the migration-only cells c17 and c18 in same-cap waves (#21307)
c17 and c18 hold no hosts and sit outside general admission, so rolling one displaces nobody. They are the only zero-displacement canary for a new cell image, but the same-cap wave refused them at the dispatch validator and would have promoted them to general at the end if it had not. Add them to the approved list and teach the wave a cell's entry admission class: the precheck demands the class the cell is declared to serve in, the restore hands it back that class, the isolate on an already-isolated cell is asserted to change nothing, and the selector generation advances by 2 for a general cell and by 0 for a migration-only one. One wave may not mix the two, because every cell after the first offsets from a single per-wave delta. Neither cell is a declared regional-rehome source, so its template carries no rehome trust lines. The source-membership guard now fires exactly when a roll expects those lines instead of for every US cell, which is the invariant it was standing in for, and which limits c17 and c18 to rehome protocol 0. |
||
|
|
0e7948fa6d |
feat(relay): pace the drain send during a same-cap cell roll (#21284)
* feat(relay): pace the drain send during a same-cap cell roll A same-cap roll drains a cell with graceMs 0, which sends `drain` to all ~800 controls in one pass. Every desktop re-dials on receipt regardless of graceMs, so the whole cell reconnects inside a second. On 2026-09-16 that stampede hit a Cloud SQL stall: attaches timed out, each leaving 10 minutes of late-arrival debt on connection headroom, and placement answered relay_capacity_exhausted fleet-wide for ~13 minutes. Spreading the sends spreads the re-dials. `HostSessionRegistry.drain` takes an optional pacing window and schedules each session's send evenly across it; admission is fenced for every session up front, and each host keeps its own full grace after its own send. /v1/admin/drain accepts `paceWindowMs` (<= 5 min) and echoes what it applied. The same-cap job asks for 120 s, and the drain-completion wait grew by the same amount. A cell still on an older image rejects the field, so the deploy script falls back to an unpaced drain rather than failing the roll. * fix(relay): scope the drain fence to the hosts already told Review of the paced drain found two problems, both from treating "this cell is draining" as one instant when pacing makes it a window. Timers: the sends queued by a paced drain were neither cleared when a later drain superseded them nor unref'd. A SIGTERM mid-window left up to 800 no-op timers holding the event loop open until systemd escalated to SIGKILL. Drain timers are now tracked, cleared on the next drain, and unref'd, so a retry re-arms a session's teardown instead of stacking a second one. Phones: the client fence read the global draining flag, so every phone was refused for the whole window even though its own host had not been told yet and was still serving. The director keeps pointing phones at this cell until their host moves, so they would have looped for up to two minutes. A session is now fenced when its drain is sent, not when the drain starts, and the client paths key off that. New control connections and re-attaches stay fenced globally: nothing new should land on a cell that is going away. |
||
|
|
09622f0c28 |
feat(relay): add a break-glass override for the same-cap monitor gate (#21270)
* feat(relay): add a break-glass override for the same-cap monitor gate Every mutating same-cap wave consumes a fresh 15-minute aggregate monitor dry-run. When a chronic fault is what the gate freezes on, waiting for a green window means waiting for the condition the wave removes: the gate froze 44 consecutive times on the recurring Cloud SQL stall the rolling image fixes. Add `gate-override-reason` and `gate-override-confirmation` (`SKIP_RELAY_MONITOR_GATE <target-image-digest>`) to the same-cap dispatch. A valid pair skips only the aggregate evidence download, provenance verification, and single-use marker. A partial or mismatched override fails closed before any mutation, in both the caller and the reusable job. Record the actor, reason, and confirmation in the gate run summary and, for a canary, in the sealed artifact. The live per-wave preflight still runs. Give it a `--no-monitor-state` source that takes the expected selector from the dispatch inputs and pins the migration policy to `strict`, rather than synthesising a state file that would claim a dry-run it never ran. Also give `director.instances` the two-consecutive-sample tolerance the cell probes have: Cloud Run replaces an instance in place, so the count leaves the [5, 6] band for one sample roughly twice a day, and a deploy overlap raises it the same way. Min and max share one streak so an alternating count still freezes. * fix(relay): canonicalise the break-glass preflight membership The override path parsed the operator's membership with a bare schema parse, while the live selector read from the director is normalised and the comparison is an ordered `JSON.stringify`. Unsorted dispatch input would therefore read as selector drift on a healthy fleet, and the every-configured-cell-exactly-once check was lost with it. Normalise through the same `normalizeSelectorMembership` call the monitor CLI uses when it seals evidence, against the same durable Terraform cell set. Tests use a collect stub that returns the director's canonical selector rather than echoing the expected one, so the ordering is actually exercised: unsorted input must canonicalise, and a duplicated, missing, or unknown cell must be rejected. |
||
|
|
560c42e1d1 |
fix(cloud): pin the asia cell database pool in the same-cap plan validator (#21171)
* fix(cloud): pin the asia cell database pool in the same-cap plan validator Raising `database_pool_max` from 10 to 16 for production-gce-c27, c28 and c29 made every same-cap roll of those three cells fail closed at plan validation. The cell startup template emits `ORCA_RELAY_DATABASE_POOL_MAX` only for a cell whose region differs from the root region or whose pool is off the default, so the asia cells carry that line while the us-central1 cells do not. The plan validator requires the before and after startup scripts to normalize to the same text, masking only the lines it independently pins to a reviewed value. The pool line was neither masked nor pinned, so the live template's `'10'` and the plan's `'16'` were read as unreviewed drift. The validator gains an optional `--database-pool-max`, accepted in `same-cap-cell` mode alone. When it is supplied the after-script must contain exactly that pool line and the line is masked from the equality check; when it is not supplied the after-script must contain no pool line at all. Masking without the pin would have removed the guard rather than moved it. The same-cap job resolves the expected pool next to the hard cap, cross-checks it against the committed `relay_gce_cells` map (asserting the default 10 for the us-central1 cells), and passes the flag to both validator invocations only for the cells that emit the line. * test(cloud): require the pool pin for a line the live template already carries |
||
|
|
5947d6b269 |
infra(relay): raise asia-east2 cell pools to 16 and record the measured connection ceiling (#21163)
* infra(relay): raise asia-east2 cell pools to 16 and retire four idle cells The three asia-east2 cells sit 176 ms from the Cloud SQL instance in us-central1. Server-side statement time there is 0.2 ms, so a pool slot is held by the round trip, not by the query. At a pool of 10 they measured 94-156 waiters and 2 s waits, and client accepts ran a ~4 s p95 against 222-646 ms in us-central1. Raising those three pools to 16 is the agreed first step; every other cell stays at 10. c4 and c5 join the committed fence set. Both are existing-only capacity the admission selector can never place on again, they carried ~1 connection each on 40-day-old images, and each still holds 10 Postgres connections. The fence set is the prerequisite the fence-source workflow confirms before it drains and attests a cell; it is not itself the resize. c17 and c18 are not fenced here. They are migration-only, and the runbook requires retire-migration-cell to move a migration-only cell to existing-only through a generation-bound selector CAS before it can be fenced. Terraform cannot express that step. The Cloud SQL consumer contract carried two stale numbers: auth at 2 instances when production has run a cap of 20 since 2026-09-04, and a 400-connection ceiling when the live instance reports 500. Both are corrected, and the budget now asserts its headroom in two named gates instead of one aggregate boolean. Those gates fail: auth alone accounts for 200 configured connections and a 215-connection rollout overlap, so the operating maximum is 713 against a usable ceiling of 490. Nothing here caused that, and no pool was lowered to hide it. * infra(relay): move the Cloud SQL contract correction out of this branch The contract correction (auth at its real 20-instance cap, the measured 500-connection ceiling) makes the budget gate fail for reasons that have nothing to do with asia pools or fenced cells, and it held this branch red. It moves to its own branch where the failure is the subject. production-cloud-sql-app-consumers.json returns to main unchanged. The budget test keeps main's single gate and only repins the cell figure that this branch genuinely moves: 230 -> 228, being +18 for three asia pools at 16 and -20 for fencing c4 and c5. Against main's 400-connection model that leaves an operating maximum of 383 under a usable ceiling of 390. * infra(relay): move the c4/c5 fence entries out of this branch Terraform now sets a cell's MIG target size directly from relay_gce_fenced_cells (relay-gce-cells.tf); the lifecycle ignore that used to protect operational target_size drift is gone. So a fence entry sitting on main ahead of its fence-source run is a standing instruction that any apply reaching that cell may execute without the documented drain and attestation. Keeping the entry in the same merge as an unrelated pool change widens that blast radius for no reason. The two entries move to their own branch, to be merged immediately before fence-source runs for c4 and then c5. This branch keeps the multi-line reflow of the list, which makes that later diff two added lines instead of a rewritten one. The cell figure in the budget test follows: 230 + 18 for the three asia-east2 pools at 16, with no fenced-cell subtraction. That is 403 operating against a usable ceiling of 390, so the headroom gate now fails by 13. It fails against a ceiling of 400 that is itself wrong; the instance reports 500. See the PR body. * infra(cloud-sql): record the measured 500-connection ceiling The budget's usable ceiling came from maxConnections: 400, described as the tier default. It is a tier default, since no max_connections flag is set, but the instance does not report 400. SHOW max_connections on it returns 500, measured 2026-09-16. On main the model sat at 385 against a usable ceiling of 390, five connections of margin, so raising the three asia-east2 pools by 18 failed the gate by 13 against a ceiling that was never checked. Against the measured one it is 403 against 490, clearing by 87. Only the ceiling and its source note change here. auth stays recorded at 2 instances, which is also wrong; PR #21165 corrects it, and with the true auth figure the budget is over by 225 for reasons that have nothing to do with these pools. * test(cloud): state the cell pool arithmetic literally in the budget pin comment |
||
|
|
e1bac25041 | fix(cloud): diagnose wrapped relay trust probe failures (#20403) | ||
|
|
9f7fd9a270 | fix(relay): reuse canary across completed rollout batches (#20214) | ||
|
|
113e58f34e |
feat(relay): support protocol 3 in cell rollout gates (#20174)
* feat(relay): support protocol 3 in cell rollout gates * fix(relay): validate and prove protocol-3 cell rollouts * docs(relay): clarify regional capability deployment prerequisite * test(relay): cover protocol-3 plans across rollout cells |
||
|
|
cd9aa43a2c |
feat(relay): correct regional placement only when the source is idle (#20105)
* feat(relay): correct regional placement only at an idle source * test(relay): lock source activity capacity semantics |
||
|
|
eb2f2d52ae |
feat(cloud): native push gateway and dedicated infrastructure (1/3) (#19912)
* refactor(cloud): share PostgreSQL schema startup between services * feat(cloud): add durable native push notification gateway * infra(push): define dedicated gateway resources and operational checks * fix(push): bound cross-host admission and simplify gateway configuration * fix(push): validate deploy configuration and preserve topic-error registrations |