Commit Graph
137 Commits
Author SHA1 Message Date
Jinwoo Hong 744e7722c2 fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG (#24373)
* fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG

The stranded-rollback recovery ran a gcloud rolling action, which renames the
MIG version outside Terraform. Every later plan for that cell then reverted the
label, and the capacity-plan validator refused the revert as an unreviewed MIG
change, so the cell could be neither rolled nor rolled back.

The validator now accepts a MIG field moving back to what relay-gce-cells.tf
declares (version name and update policy), in every mode, and a test pins those
values to the Terraform file. The stranded branch recreates the cell's single
instance with recreate-instances, which leaves the MIG untouched.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): let a label-only MIG plan through and recreate on it in a stranded rollback

A stranded rollback whose template is already in place plans only the version
name revert. The validator still required the MIG template to move, so that
plan was refused, and the recreate gate (changes == 0) would have skipped a
plan of one change and left the drain flag set. Require the template move only
when no declared field reconciles, and recreate whenever the template was not
replaced (changes < 2).

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 08:12:49 -04:00
Jinwoo Hong c422936a71 fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup (#24349)
* fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup

The same-cap gate now verifies the dry-run on its own clock and records the
authorisation instant in the single-use consumed marker; each cell job checks
the evidence age at that instant and bounds its own start after it.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): refuse a re-run same-cap gate before it consumes evidence; tighten order tests

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 06:52:31 -04:00
Jinwoo Hong 5f728f7862 fix(relay): let a draining cell pass restart-safe through refused redials (#24347)
* fix(relay): let a draining cell pass restart-safe through refused redials

A draining cell refuses every control and host proof, so once no session,
splice, or queued byte remains, an in-flight or reserved connection unit can
only belong to a dial the cell is about to refuse. The restart-safe wait no
longer resets its pace-window streak on those units, and its progress line
now prints every counter the gate reads.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): pin fail-closed parsing of handshake counters on a draining cell

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 05:58:43 -04:00
Jinwoo Hong 9ed7b39b3c fix(relay): run the same-cap headroom gate in the modes the job actually receives (#24343)
The parent workflow collapses canary-apply and batch-apply into the job mode
apply, so the headroom step's canary-apply/batch-apply condition never held and
the gate was skipped on every real roll. Run it wherever the drain runs (apply,
rollback before its restart) and in read-only verify; skip only a resumed
rollback, which drains nothing. A new workflow-shape test fails on any job
step comparing against a mode the parent cannot pass, and on a drain that can
run without the headroom check.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 05:18:42 -04:00
Jinwoo Hong 6b8ca27d85 chore(relay): move Asia cell c31 to the general same-cap lists after promotion (#24318)
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 04:07:02 -04:00
Jinwoo Hong 94f6c53387 fix(relay): gate the Asia canary on Asia-targeted region fallbacks only (#24320)
* fix(relay): gate the Asia canary on region fallbacks against a pre-canary baseline

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): gate the Asia canary on region fallbacks against a pre-canary baseline

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 03:51:53 -04:00
Jinwoo Hong 3e5c8d9f8e feat(relay): declare Asia cell c31 at the c30 shape (#24310)
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 03:11:32 -04:00
Jinwoo Hong 7c119465b0 fix(relay): define restart-safe by the cell runtime, and refuse waves without headroom (#24259)
* fix(relay): let a same-cap drain finish when only unplaceable hosts remain

The c28 canary on 2026-10-01 drained the cell to zero live connections, but four
hosts with no free slot anywhere kept redialling and held director leases on it,
so the restart-safe wait timed out and left the cell isolated and empty.

The drain wait now also passes once the runtime has carried nothing for a
sustained quiet window while a small, capped number of leases remain, and logs
the escape. Apply modes also refuse a cell whose hosts exceed 80% of the free
slots on the other general cells, so a wave cannot strand hosts in the first place.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): define restart-safe by the cell runtime, not director leases

Replaces the opt-in stranded-host escape with a corrected definition. A
restart is safe when the cell runtime carries nothing live and no migration
is open, sustained for the drain pace window. Director activity leases lag
hosts that already left or cannot be placed, so they are reported in a
progress line and the verified result instead of blocking the restart.

The same-cap drain passes its existing pace window. The headroom script is
added to the trusted evidence code paths with the other production scripts.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): print stranded director counts on every restart-safe sample

Each restart-safe poll now prints its sample count and the director's
restart-blocking leases, request units, reserved remainder, and migrations
under `stranded`; the verified line carries the same object. Open migrations
still block because each is pinned to the cell incarnation a restart replaces.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): require the pace window for every live restart-safe wait

Pre-auth and total connections no longer reset the restart-safe window:
on drained c28 they flickered with unauthenticated redials in a third of
samples, which a restart does not lose. They stay in the progress output.

Every live restart-safe call must now pass --pace-window-ms. The capacity
job and staging proof drain unpaced, so they pass the production 300000 ms
window, and the calls that relied on the 180000 ms default get 480000 ms.

Headroom free slots now follow the director's placement rule: the admission
pause minus the larger of observed and enforced units, minus outstanding
control reservations.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-30 22:25:36 -04:00
Jinwoo Hong 84f58a1fc7 fix(relay): refuse a redial at once while the host's own release holds its row (#24225)
* fix(relay): refuse a redial at once while the host's own release holds its row

During an Asia drain the host whose socket closes is the one that redials.
Its release on the draining cell locks its assignment row first, then waits
on the cell's busy row for up to the lock timeout. The director's sticky
and placement paths waited on that assignment row inside the single sticky
slot, and the sticky path then locked the busy cell row itself before it
checked isolation. The slot backed up and dials timed out fleet-wide.

Both paths now take the host's assignment row NOWAIT and throw
RelayAssignmentRowBusyError when it is held. /v1/assign answers that with
503, Retry-After 1 and error assignment_row_busy, logged with its own
reason. The sticky path decides isolation before it touches the pinned
cell row. An isolated retry keeps its own tier as its retry scope, so a
busy lock inside it no longer falls to the all-rows path. The local drain
arm drops its zero-release-failures bar, which the Asia arm never had.

The drain harness gains a departing-host arm: each host releases its own
lease, then redials after 150, 400 or 1000 ms on the desktop client's
5-5.5 s pacing. At 400 ms, main rejected 83 of 180 first dials by sticky
wait timeout, placed 11.6/s with 3.1 director backends lock-waiting, and
took 11.2 s at p95 from release to placed. Now: 16 fast refusals, 18/s,
no lock waits, 5.7 s at p95.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): wait briefly for a calm host's row and keep the dead-cell sweep going

The dead-cell sweep treated RelayAssignmentRowBusyError as fatal, so one
busy host ended the sweep for every later host each tick. It now skips
that host and carries on.

The sticky path refused a busy row at once for every host. A calm host
redialling after its own clean close often meets its own short release,
and a refusal costs it the client's 5 s assign gate. When the pinned cell
is general and live, the sticky path now waits up to 1 s for the row
before refusing. A roll-isolated, parked or dead cell still gets the
immediate refusal. Placement keeps NOWAIT, because it holds cell rows
while it would wait. A resume refused for a busy row now carries
Retry-After 1 as well.

The departing-host harness arm now bounds the busy refusals at 20% of
hosts and the p95 at 8 s, and counts unexpected errors apart from
retryable refusals.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): never wait on a host's row while the sticky retry holds its cell row

The inventory-first sticky retry takes the pinned cell row before the
assignment row. With the calm-host bounded wait it could then wait up to
1 s on the assignment row while holding the cell row, the reverse of the
ranked lock order, against this host's own release, which holds its row
and wants the cell's. The bounded wait now applies only when no cell row
is held; the retry stays NOWAIT.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-30 17:58:30 -04:00
Jinwoo Hong ed462b2caa fix(relay): re-place hosts off a draining cell without locking its row (#24216)
* test(relay): reproduce drain-release contention against director placement

Adds a Postgres harness that drives releases from an isolated cell over a
171 ms per-statement pool while five directors re-place reconnecting hosts
through the sticky lane. At 18 releases/s placements fall from 20/s to about
5/s and every active director backend is blocked on relay_cells.

Moves the per-statement delay pool into a shared test fixture so the
rehome target-row test and this harness use one implementation.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): re-place hosts off a draining cell without locking its row

Sticky re-placement of a host whose cell is isolated for a roll locked
every relay_cells row. During an Asia drain the source row is held by the
cell's own releases for a round trip each, so the placement waited on it,
the single sticky slot backed up, and /v1/assign returned 503 fleet-wide.

The isolation decision now comes from an unlocked read, and the placement
locks only the same-region general rows it can move to. It no longer writes
the source row: the host's source leases stay, and each one's own release
or expiry takes its units back off the source. The assignment keeps its
activity counters and adds one control instead of resetting them. With no
same-region headroom the path falls back to the all-rows lock, as before.

Dormant hosts hold no units, so their placement also locks only the
general rows and skips the zero write to their old cell.

The drain harness now asserts the after picture: 20 placements/s at Asia
latency with no director lock waits, against 9.2/s and 4.8/s before.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): take a moved host's units off the cell that holds them

After a narrowed re-placement a host keeps leases on its old cell while
its assignment row names the new one. Two paths charged the row's whole
counted total to the row's cell: aggregate expiry, and the lease deletion
in dead-cell and stranded re-placement. Both over-charged the new cell and
left the old cell's units stranded.

Aggregate expiry now skips hosts that still hold any lease; the lease
sweep takes each lease's units off its own cell. Placement frees each
deleted lease's units on that lease's cell, charges the old cell only for
units no lease backs, and sets the counters from the leases it keeps plus
the new control. The narrowed path runs only when the counters already
match the leases, so it never needs to write the old cell's row.

The all-rows re-placement off an isolated cell follows the same rule, so
it no longer decrements the source at placement either.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): try other regions before the all-rows lock when re-placing off a roll

With every same-region neighbour at its connection cap, the narrowed path
found no target and fell back to the all-rows lock behind the busy source
row, which is the drain brownout again. It now tries a second tier, general
cells in every other region, in its own transaction over one ordered
lockCellRows, still never the source row. Only when no general cell in any
region has room does it fall back to the all-rows path, which keeps the pin.

This changes the policy from #21911, which refused to re-place an isolated
host across a region. The unit tests that encoded that rule now assert the
tier order instead.

The drain harness gains a US cell and an arm with every Asia neighbour
capped: 200 of 200 dials placed cross-region at 20/s with no lock waits,
against 0 placed and 144 sticky rejections on the previous head. Its pass
bars are now the rejection share and the lock-waiting share, not the
placement rate a slow runner's pacing can move.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): take the narrowed path for hosts whose counters sit below their leases

Main's old placement reset a moved host's counters while keeping its source
leases, and those leases' releases floored the counters at zero. Such hosts
hold fewer counted units than lease units, and requiring equality sent them
down the all-rows path behind the busy source row.

The narrowed path now requires only that the host holds no units no lease
backs, the one case that needs a write to the old row. Its placement
already rebuilds the counters from the kept leases plus the new control,
so a drifted host is healed by its next re-placement.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): judge the local drain arm on completion and lock waits, not rate

The local-latency arm asserted at least 18 placements/s at a 20/s dial rate,
which a slow runner's pacing alone can miss. It now asserts what the Asia
arm does: no dial failures, sticky rejections under 10% of dials, every
other dial placed, and director lock-waiting under half a backend.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-30 17:00:03 -04:00
Neil b99462ac1c test: retire mobile, cloud, config and e2e cases their input cannot reach (#24077)
Completes the first pass over every test area in the repository. Sweep over
`mobile/src`, `config/scripts`, `cloud/`, and `tests/` (1,494 files in scope, with
the 24 files under `mobile/src/test-support/rpc-recording/` deliberately excluded).
31 case declarations removed across 17 files, 2 test files deleted, 356 lines gone.

What went, by pattern:

- Cross-boundary replays of a shared helper. A whole mobile file re-ran
  `extractPendingAsk`/`parseAskFromStatus`/`formatAskAnswer`, all owned by
  `src/shared/native-chat-ask.test.ts`, `native-chat-ask-fifo.test.ts` and the
  renderer's interactive-prompt suite — one case title was verbatim identical to the
  owner's, and the owners' inputs are supersets. The mobile file imported the shared
  module directly and exercised no mobile transport, lifecycle or rendering.
- A case whose input cannot reach the behavior its title names: "arms it on Android
  while the drawer is open", where `use-back-claim.ts` has zero
  Platform/OS references, so flipping the mocked OS changes only shadow styles.
- Identity copiers, including one asserting `prSidebarRenderBranch(state) ===
  state.kind` against a production body that is `return state.kind`. The function
  stays; it has three live callers.
- A test of the runtime rather than the product: a case asserting Node's own
  `EventEmitter` crash contract on a bare emitter, with zero production code in the
  path. The guard it documents is exercised behaviourally by the case after it.
- Duplicate invocations, one of them provable rather than eyeballed: with
  `MODULE_SCOPE_ENV_WRITER_PIN = 0`, `files.size <= 0` is strictly implied by the
  sibling's `expect(offenders).toEqual([])`, since a non-empty `offenders` forces
  `files.size >= 1`. The pin's own doc says it may only ever be decreased from 0, so
  it could never become a meaningful bound either. Its policy guidance survives as a
  comment; the file's real ratchet and its regex self-test both stay.
- Expected values produced by the test's own arithmetic, and a p95 case strictly
  implied by a sibling that already pins exact p95 and exact max over a wider range.

One production line goes: the `export` keyword on `assignmentCleanupSteps` in
`cloud/apps/relay/src/assignment-cleanup-steps.ts`. The function itself stays and is
still called internally; only the test-only export was orphaned.

Kept deliberately: everything a gate cites, checked by case title and not only by
file path; a gate-cited case that does not deliver its claim (reported instead — see
below); a cross-version wire cell whose ledger is never invoked, left under the
raised bar for wire coverage; and every limit, bound, quota and provenance guard.

Nothing under `mobile/src/test-support/rpc-recording/` or
`mobile/rpc-foundation/goldens/` was touched — those bytes feed a `recorderSha256`
digest pinning 398 golden recordings.

Verified: `mobile` vitest over the modified mobile files (8 files, 50 cases);
`mobile/scripts/check-tests-typecheck-ratchet.mjs` OK (898 files in program, 125
grandfathered, none @ts-nocheck); relay suite 799 passed; `check-reliability-gates.mjs`
140 gates; both deleted files confirmed absent from the gate manifest,
`cloud/package.json` and `mobile/tests-typecheck-baseline.txt`.

Seven local failures were investigated and none is caused by this change: five
`mobile-web-app-*-render` tests drive `playwright-core` chromium/webkit and need
browsers this machine lacks, `release-checkout.unit.test.ts` needs cross-version git
refs, and `e2e-worker-env-isolation.unit.test.ts` fails identically with its HEAD
content restored — it recurses `tests/e2e` with symlink-following `statSync` and no
depth guard.
2026-09-30 02:05:30 -07:00
Jinwoo Hong e8e09eed46 fix(push): size the claim-attempt budget from the drain count (#24040)
* fix(push): size the claim-attempt budget from the drain count

With twelve drains, up to eleven peers can hold device heads, so a
four-attempt claim budget can run out while claimable rows remain and the
drain exits idle for a tick. Move the drain count into one module and derive
the attempt budget from it.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* style(push): keep the worker and store in repo formatting

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* style(push): drop the stray semicolon in the concurrency constant

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-30 02:59:48 -04:00
Jinwoo Hong 52110982ca fix(push): run twelve delivery drains instead of four (#24038)
Four drains, each holding one provider round trip of ~100 ms plus its
database statements, capped delivery near 30/s. Production inflow reached
35/s on 2026-09-30, so the backlog aged past the five-minute TTL and
notifications expired. Twelve drains lift the ceiling to roughly 90/s; the
pool is now six per instance, so the extra drains queue on connections
instead of starving the request path.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-30 02:43:28 -04:00
Jinwoo Hong f7025d88be fix(push): give the push gateway six database connections per instance (#24026)
Delivery collapsed on 2026-09-29 once send volume doubled: the worker, the
retention pruner and the request path share a two-connection pool, and the
database transaction rate pinned at ~140/s regardless of how many notifications
were delivered. Raise the pool to six so worker and pruner stop serialising on
one connection. The budget precondition stays satisfied (2 x 6 x 3 = 36 <= 64).

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-30 02:16:04 -04:00
Neil 70475e0228 test: stop testing private internals through exports no caller needs (#23829)
Third audit wave. The detector looked for production modules exporting three
or more symbols that no production file imports — only tests do. That shape is
the authoring gate's fourth question failing: a test needing a production seam
no caller needs belongs at the real boundary instead.

Most hits were detector false positives and were left alone; the scanner misses
re-export barrels and dynamic imports, so every module was re-verified with rg
before any edit. Where a private predicate's behavior was already covered
through the module's real entry point, the duplicate cases are gone and the
symbol is module-private again. Where it was NOT covered anywhere else, the test
stays — this audit removes tests, it does not author replacements.

Production code deleted where tests were its only callers: the superseded
`filesystem-directory-listing-limit` module, the unused
`format{Hourly,Daily,Adhoc}Version` helpers and their orphaned prerelease
identifiers, the dead `filterByAutomationListSearch*` family superseded by
`matchAutomationListSearchRowKeys`, and the dead
`getAiVaultResumeWorktreeTargetStatus` copy of the live workspace branch.

Also drops two call-shape source greps in `relay-sweep-schedule.test.ts` that
asserted `index.ts` spells `jitteredSweepIntervalMs(30_000)`; the jitter math
has a behavioral owner at the top of the same file. The structural census that
counts role-gated vs total `setInterval(` calls stays — an ungated sweep runs in
every cell, and nothing else can catch that.
2026-09-29 15:53:18 -07:00
Neil 31012aeb09 test: remove assertion-free probes, copied inventories and export-shape checks (#23816)
Second audit wave, targeting three more junk patterns:

- assertion-free cases that run code and assert nothing, so they pass no
  matter what the code does;
- inventory literals re-typed from a production declaration, where the only
  way the assertion can fail is someone editing one of the two copies;
- export key-set and export-shape loops (`typeof x === 'function'` over every
  export) that restate what TypeScript already enforces.

Yield is much smaller than wave 1 on purpose: the assertion-free scanner has
a high false-positive rate, because many flagged blocks assert through a
shared helper or their oracle is "this must not throw". Those were kept.

`mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is
de-exported — after the inventory comparison went away, nothing outside the
module read it.
2026-09-29 02:21:47 -07:00
Neil 6e1b7e7fa3 test: remove junk tests that assert source text instead of behavior (#23815)
Deletes 101 test files and trims 112 more, all matching documented junk
patterns: exact source/import/string greps, copied inventories and export
lists, duplicate invocations of a contract another test already owns,
typeof-shape checks TypeScript already enforces, and self-comparisons.

The largest group read a production `.ts` file and asserted on its text —
for example a TaskPage test that required the source to contain
`selectedRepos.find((r) => r.id === newIssueRepoId) ?? selectedRepos[0] ?? null`.
Any behavior-preserving rename broke it; no behavior change ever did.

Production-side follow-through: exports that only these tests imported are
de-exported or deleted, stale comments pointing at removed censuses are
dropped, and the reliability-gate registry, `cloud/package.json` test lists,
and orphaned source-reading helpers are updated so nothing references a
deleted file.

Two files kept their real coverage and lost only the census scaffolding:
`agent-status-producer-census.test.ts` now drives all five producers end to
end instead of grepping the source tree, and `config-toml-trust-stale-writes`
replaces an export-list parity check.
2026-09-29 01:21:53 -07:00
NeilandClaude 8da0ca52e7 perf(relay): release rejected first-frame connections [trade-off] (#23011)
* perf(relay): release rejected first-frame connections [trade-off]

* fix(relay): bound the director's rejected and redirected first-frame closes too

The parent PR routed four cell-side first-frame rejections through
closeRelayWebSocket but left two raw socket.close() calls in the same
handler. A real-socket probe shows both still pin a connection unit for
ws's full 30s close timer when the peer ignores the close frame:

- 'invalid invite' is reachable by an unauthenticated peer with a
  well-formed but bogus credential, so the exhaustion the parent PR
  claims to prevent stayed reachable on the director;
- 'connect to assigned cell' is the happy path for every phone's first
  director contact, so it is the highest-volume unbounded close here.

closeWithDrain gets the same treatment; host-session-registry already
closes the identical drain through the helper.

Also records that the bounded close is not a user-facing trade-off: the
close frame is written before the force-close timer can fire and TCP
delivers it ahead of the FIN, so an abandoned peer still reads code and
reason over a graceful close. The new regression asserts that, plus
exactly-once release across concurrent bursts and rejection racing the
peer's own disconnect (the ledger does not clamp at zero, so a double
release would surface as a negative count).

* refactor(relay): drop the unused closeWithDrain helper

`closeWithDrain` has no callers anywhere in the repo, and its `graceMs`
parameter promised a caller-supplied drain window that the body no longer
honours: routing it through `closeRelayWebSocket` force-terminates after 1s
regardless, so a future caller passing `graceMs: 30_000` would have had its
drain silently cut short while the signature still claimed otherwise.

The real drain path is `host-session-registry`, which sends the same
`resolve-director` drain and closes it there. Delete the dead duplicate
rather than bound a helper whose contract says "graceful".

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(relay): call the bounded close's rejection delivery best effort, not guaranteed

The helper claimed the forced terminate costs an abandoned peer nothing it can
observe, and the test presented its fast-peer assertion as proof. Neither holds in
general: `ws` writes the close frame to the socket and `terminate()` destroys that
socket a second later, so under backpressure the frame — and any `relay-moved`
message queued ahead of it — can go unsent even to a peer that never stopped
reading.

Qualifies both comments to describe delivery as best effort and name the
backpressure case. The fast-peer assertion is valid and stays exactly as it was;
only its stated scope narrows, and the test is renamed to say which peer it speaks
for. What the bound actually buys — a stalled peer cannot hold admission — is now
stated on its own rather than resting on a delivery claim.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 00:28:52 -07:00
NeilandClaude da95225ba7 perf(push): keep retention sweeps from overlapping without slowing the drain (#23001)
* perf(push): keep retention sweeps from overlapping

* perf(push): drain a saturated retention sweep instead of idling out the tick

The overlap guard on the shared prune timer removed a side effect the sweeper had
been relying on: overlap was the only thing that let a backlog exceed the
50-batch-per-call cap inside one 60-second tick. With the guard, a sweep that
spent its whole budget went idle for the rest of the interval, so a large backlog
drained far slower exactly when retention matters most.

The timer is now a chained setTimeout rather than an interval. `deleteInBatches`
reports whether it exhausted its batch budget, `prune()` returns
`{ deleted, saturated }`, and a saturated sweep is rescheduled immediately. The
connection gate hands a freed slot to the longest waiter, so one serial sweeper
looping back to back still parks a single statement ahead of a worker claim: claim
latency keeps the value the guard bought while the maximum drain rate returns to
what it was before. The loop is self-limiting and stops once the backlog clears.

A sweep that has not settled a full interval after it started now logs
`orca_push_prune_overdue` with its target. Admission waits have no timeout, so a
lost slot release could previously wedge retention permanently and silently.

Chaining makes the overlap guard structural, so there is no flag to scope. The two
single-DELETE sweeps state through `unbatchedSweep` that they have no batch budget
to exhaust, which keeps the immediate-resume path readable as delivery-only.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-26 20:39:57 -07:00
Neil 71344de12f Scope relay capacity checks to the requested cell (#23036)
* perf: scope relay capacity checks to the requested cell

* fix: restore catalog entries required by the current CI baseline

* fix(i18n): make AI the recipient of diff notes
2026-09-26 14:17:47 -07:00
Neil 668fa71906 Release SQLite transaction queues when startup fails (#22993) 2026-09-26 13:01:16 -07:00
Neil 818237b9e3 fix: stop retired relay load connections from rescheduling refreshes (#23071) 2026-09-26 12:59:47 -07:00
Neil 148296e010 perf(relay): skip abandoned queued control activation (#23029)
* perf(relay): skip abandoned queued control activation

* style: format reliability gate metadata
2026-09-26 12:54:19 -07:00
Neil 1fc0bfe46d perf(relay): share concurrent readiness probes (#23012) 2026-09-26 12:53:09 -07:00
Neil fa018fe520 chore(deps): refresh maintained dependencies (#22964)
* chore(deps): refresh maintained dependencies

* fix(deps): defer upgrades that violate runtime and test contracts

* test: align catalog response assertion and updated formatter
2026-09-25 19:04:12 -07:00
Jinwoo Hong 128e97ffca fix(relay): drain a same-cap cell over 5 minutes, not 2 (#22584)
* fix(relay): drain a same-cap cell over 5 minutes, not 2

The 2026-09-23 c27 roll drained 2,145 hosts over the 2-minute window,
about 18 re-dials/s, while the director re-places roughly 8/s through
its single-slot sticky lane. The overflow queued behind slow
re-placements and timed out, so /v1/assign returned 503 fleet-wide for
about 5 minutes. 5 minutes is the cell's maximum pace window and keeps
the remaining 2,650-host cells near lane capacity. The transition wait
already outlasts a 5-minute window (17 min).

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): pin the same-cap drain window contract at 5 minutes

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): keep the drain wait at lease plus the 5-minute window

The transition wait after a drain was set to the 15-minute migration
lease plus the pace window. Widening the window to 5 minutes without
moving the wait left 12 minutes for a migration that can hold for 15,
so a late-window migration would time the wave out into rollback.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-23 23:38:15 -04:00
Jinwoo Hong 1116c54230 fix(relay): accept production cells past c29 in the regional rehome operator (#22518)
The selector membership check capped cell ids at c29, so enable failed
closed with "selector membership is invalid" once c30 went general.
Accept c1-c99 with no leading zero.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-23 14:37:17 -04:00
Jinwoo Hong ae3d380b23 fix(relay): lock only the target cell row, last and NOWAIT, in the rehome commit (#22449)
* fix(relay): lock only the target cell row, last and NOWAIT, in the rehome commit

The idle-rehome commit runs on the source cell. From an Asia cell each
statement is a cross-region round trip, and the transaction locked every
relay_cells row plus every runtime, capability and safety row before about
twenty more statements, so each Asia-source rehome held the whole fleet's
cell rows for ~3.6 s and every reconnect, renewal and placement queued or
timed out behind it.

The commit now reads the cell inventory and the runtime, capability and
safety tables unlocked, keeps the control and worker rows locked (now
NOWAIT), and takes one cell lock: the target row, in a single statement
that locks it NOWAIT, re-checks enabled, general admission and capacity,
and reserves the units, issued as the last statement before COMMIT. A
target that changed admission, filled up, or is locked by another writer
rolls the whole commit back and answers deferred (candidate-ineligible).

The hold is sampled under a site label, so cellInventoryHoldMsMax still
sees rehome holds and rehomeTargetRowHoldMsMax reports them apart.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): fail closed on the rehome target-row lock clause

The target-row statement now carries FOR UPDATE ... NOWAIT unless the
dialect is explicitly SQLite, so a wrapper that omits the optional dialect
can no longer run the reservation unlocked. Test wrappers and the fault
injection entry forward the dialect they wrap.

The latency test also probes the admission and region tables at every
round trip; only the target's admission row may be locked, and only
before COMMIT. The runbook notes that an Asia-sourced commit holds the
rehome control row for about 6.5 s, so a pause that fails once is retried.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-23 03:45:47 -04:00
Jinwoo Hong a2a78ab335 feat(relay): alert on relay cell table lock convoys (#22446)
* feat(relay): alert on relay cell table lock convoys

Adds a log-based metric and alert for cell-inventory lock holds of at
least 1,000 ms, and a Cloud SQL log metric and alert for relay-only lock
timeout cancels at 20 or more per minute. NOWAIT refusals are excluded:
background sweeps produce about 160 per minute even with rehome paused.

Replayed over 2026-09-20 14:00 to 2026-09-22 15:00 UTC: the hold filter
matches all 93 asia-east2 rehome holds plus 9 director holds, and every
one of the 88 cancel burst minutes overlaps an asia-east2 hold.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): page only on cell lock holds; director holds stay visible

Director holds of 1-2.5 s recur several times a day with rehoming paused,
and pausing rehome does not stop them. The paging hold policy now selects
role=cell samples only; a separate policy with no notification channel
keeps director holds visible. The burst documentation no longer claims no
burst happens while paused, and the runbook points a burst with no cell
hold at the director policy.

Replayed cell-only: 93 of 93 asia-east2 holds, 0 director holds over
2026-09-20 14:00 to 2026-09-22 15:00 UTC; 0 from then to 2026-09-23 07:30.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-23 03:24:09 -04:00
Jinwoo Hong 1043dc5e1d fix(relay): stop rehoming hosts off Asia cells until the lock fix lands (#22443)
The source cell runs the rehome commit. An Asia source pays a cross-ocean
round trip per statement while holding relay_cells row locks every cell
needs, which convoys the fleet database. Selection now drops source cells
outside the director's region before building the decision window, so the
incumbent_region filter shrinks while Asia cells stay valid targets. The
preview counts the same hosts as source-outside-director-region and the
poll summary reports skippedOffRegionSourceCells. Temporary stopgap.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-23 03:15:52 -04:00
Jinwoo Hong 51c3434851 chore(relay): treat Asia cell c30 as a general cell now that it is promoted (#22439)
C30 was promoted to general on 2026-09-23 (selector generation 286). The
same-cap wave now rolls it as a general cell instead of handing it back
isolated, and the shadow gate reads its pool beside C27-C29. Follow-up to
#22375.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-23 02:55:27 -04:00
Jinwoo Hong 9fef7a0f04 fix(cloud): gate the Asia canary on its own cell's SQL failures, not the directors' (#22405)
* fix(cloud): gate the Asia canary on its own cell's SQL failures, not the directors'

The production canary summed sqlFailuresDelta over every director and the
canary cell and required zero. Directors log a steady baseline of
relay_cells NOWAIT and lock-timeout refusals unrelated to the canary cell,
so a C30 canary failed most attempts on that noise. The canary now requires
zero SQL failures from the canary cell's own metrics and records the
director sum as directorSqlFailures without gating on it. Directors keep
every other rule (unavailable regions, fallbacks, pool waiting, transient
waiter and wait-time bounds). Staging keeps the combined zero rule and its
evidence shape unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(cloud): gate the Asia canary's pool bounds on its own cell too

Directors also show a steady pool-wait baseline (waiting above zero and
waits over 50 ms in about 6 of every 60 minutes), so a five-minute canary
still failed about half the time on director pool pressure unrelated to
the canary cell. With gateDirectorDatabase off, the production canary now
applies databasePoolWaitingMax, databasePoolWaitersMax and
databasePoolWaitMsMax to the canary cell's metrics only and records the
director values under director-prefixed names. Directors still gate Asia
selections, region fallbacks and unavailable regions. The staging path
keeps its combined values, key order and validation order.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-23 01:43:26 -04:00
Jinwoo Hong 483fa0aca2 fix(cloud): compare the Asia topology budget gate against the measured 500-connection default (#22386)
The topology workflow's Cloud SQL gate carried a hard-coded 400 for the
instance's tier default while the consumer contract records the value
measured on the live instance (SHOW max_connections = 500, 2026-09-16,
#21163). The gate compares the two and the first production plan run
(35815654836) failed silently on that mismatch before Terraform ran.

The verified default now lives beside the tier and version it is verified
for, as VERIFIED_DEFAULT_MAX_CONNECTIONS, so the contract and the workflow
are two independent records of the same measurement and the gate keeps
its cross-check. The test pins the new source and forbids a bare literal.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-23 00:28:44 -04:00
Jinwoo Hong bf3f95245c feat(relay): declare Asia cell c30 at the c27 shape (#22375)
* feat(relay): declare Asia cell c30 at the c27 shape

Adds production-gce-c30 in asia-east2-a at the reviewed Asia shape (6,000
request units, 3,000/60 connection limits, 16-connection pool, disabled) and
the rehome trust the other Asia cells carry.

Every Asia enumeration now knows C30. The topology, admission, and director
tools treat it as its own reviewed wave so its plan and registration never
touch the live launch cells. C30 promotion requires C27 general and fresh
staging evidence. The topology and director validators now pin the committed
production pool of 16 instead of the stale 10, which had made the topology
workflow reject the committed launch cells.

* fix(relay): plan C30 at live images and prove it with its own canary

The shared URL map pulls every cell into the C30 topology plan, so the workflow
now plans each non-target cell at the image its live template serves, and the
validator names any change to a cell outside the wave. C30 promotion runs the
same five-minute production canary and automatic rollback C27 used, with the
load report proving the canary control was placed on C30, instead of relying
on staging evidence. C30 leaves the shadow gate's fleet pool list until it
serves, rollback rejects mixed partial sets, and a budget test pins the
mixed-Asia-pool refusal.

* fix(relay): pin C30 to the production director's live image digest

C30 promotion requires the director and C30 to report one digest, so C30
takes the director's sha256:4158d8a2 (read 2026-09-22). C27-C29 keep their
committed lines; every Asia check compares only the cells named in a run.

* fix(relay): read the committed cell map from a plan, not console

terraform console evaluates every output against state, and the Relay
deployments output indexes each cell's MIG, so it fails with Invalid index
while C30 is declared but not created. Read the map from a no-refresh,
unlocked plan over the same targets instead, and refuse empty overlay input.

* fix(relay): keep console readers working and C30 migration-only until promotion

relay_gce_cell_deployments indexed each cell's MIG, backend, and template,
so once C30 is declared but not applied every production terraform console
reader printed a warning to stdout and broke its jq parse. Wrap those six
lookups in try(..., null).

Same-cap listed C30 as general, so a rollback dispatch on a migration-only
C30 would restore it with activate and skip its canary. List it with the
migration-only cells until the promotion follow-up moves it.
2026-09-22 23:41:09 -04:00
Jinwoo Hong b013590363 fix(push): bound the delivery claim, delete finished batches, and keep a connection for requests (#22307)
* fix(push): bound the delivery claim and stop keeping finished batches

* fix(push): bound claim scans to the notification TTL and document the queue

* test(push): boot waits out a table writer for the queue indexes; queue stays correct without them

* fix(push): share the claim lock so a previous-revision claim cannot re-lease a delivery

During a deploy overlap the previous revision's claim scans under an exclusive
push-worker-claim lock and re-reads the row without checking lease_until, so it
could overwrite a lease this revision had just committed and send twice. The
new claim now takes the same key shared: new claimers never block each other,
and the previous claim waits until their leases commit before it scans.
Droppable one release after every worker runs this revision.

* fix(push): install one queue index at boot, not three

push_batches_leased_device indexed lease_until, so every lease, renew and finish
UPDATE lost heap-only eligibility and rewrote every index. push_batches_pending_due
was unused: on a synthetic 902k-row table the candidate scan plans onto the
existing (state, due_at) and expiry indexes with or without it. The per-device
pending index stays; the head check and the busy anti-join use it. Fewer
boot-time builds also shorten the SHARE lock the first boot takes on the table.

* test(push): pin the claim's TTL scan bound and the server's worker connection cap

Removing either guard left the suite green. The claim test captures every row
the candidate scan returns and plants one row that only the TTL term excludes;
the server test drives the real worker through createPushServer and fails when
the request-connection reservation is unwired (peak 4 instead of 2).

* fix(push): renew delivery leases outside the background connection cap

Renew shared the single background slot with claim retries and prune batches,
so a heartbeat could wait long enough for a lease to lapse and the delivery to
be re-leased mid-send. It is a keyed one-row UPDATE, so request traffic cannot
starve it on the ungated pool.

* docs(push): describe the shared claim lock for mixed-revision deploys
2026-09-22 15:31:51 -04:00
Jinwoo Hong 8568d77b06 fix(relay): let a same-cap cell plan delete the deposed template a failed wave left behind (#22170)
A same-cap wave applies with create_before_destroy. When run 35684694704 died
after the new template was made, the previous template stayed in Terraform state
as a deposed object, so every later plan for that cell carried its delete. The
plan validator only tolerated deposed deletes in same-cap-image mode, so the
recovery run 35698133226 was refused with "cell plan must change only the exact
instance template and MIG" and the cell was stranded.

Allow exactly one deposed delete at the cell's own template address in
same-cap-cell mode too, mirroring the one-deposed bound the convergence path
already applies to non-image modes. Report it as `obsoleteTemplates` rather than
in `changes`, the way the cell backend update is already split out, so the job's
`changes == 2` resume gate and `changes == 0` stranded roll keep reading the
template-and-MIG count.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-22 03:27:00 -04:00
Jinwoo Hong bf8d63bc67 fix(relay): keep the backend service out of the same-cap wave; the capacity role cannot update it (#22140)
The capacity role the same-cap wave authenticates as, orcaRelayProductionCapacity,
has no compute.backendServices.update. Since #21860 added
`google_compute_backend_service.relay_gce_cell["${TARGET_CELL_ID}"]` to both of the
job's plan invocations, every wave has therefore created the new instance template,
modified the MIG, and then failed 403 on the backend, leaving the cell isolated with
its trust probe, admission restore, and shadow gate all skipped. Run 35684694704 on
production-gce-c7 is the first one that hit it in production.

Drop the backend target from both plans and restore the resume gate to exactly
`.changes == 2` (the template-and-MIG rollback-image drift) or a converged plan,
removing the backend-only resume apply #21865 added on top. A resume applies nothing
again, which is what a resume means.

The validator keeps its bound on a cell backend update, so it still reports one and
refuses anything wider, but a wave plan can no longer contain one. The drain timeout
from #21848 and the log_config from #21860 need a root apply by a principal that holds
the permission; granting the capacity role that permission is itself a root apply, so
it can follow as its own change rather than blocking every wave in the meantime.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-22 02:50:32 -04:00
Jinwoo Hong c8a5580659 fix(relay): re-place hosts off a cell isolated for a roll (#21911)
* fix(relay): re-place hosts off a cell isolated for a roll

A roll isolates a cell by moving it out of the 'general' admission class; the
cell then refuses every attach with 4503. The director never noticed, because
the only liveness test it applies to a host's current cell reads
`relay_cell_runtime.ready` and the heartbeat, and an isolated cell keeps
heartbeating ready=1 for the whole drain. So every host on that cell was handed
its own dead cell, closed, and handed it back — 500-1,900 hosts looping for
13-16 minutes per cell roll, at ~6 dials each per minute, with no neighbour
absorbing anything.

The sticky lane now treats a live incumbent whose admission is 'migration-only'
— the state a roll's isolate step writes — the same way it treats a dead one:
it returns null, which means "fall through to placement". The placement lane
had the identical hole eleven lines further down, so it takes the same
predicate; without that second swap the sticky change is inert, because
placement would hand the pin straight back (a draining cell has more headroom
than anyone). An isolated incumbent skips the dead-cell fence branch: that
branch exists to prove an unreachable cell stopped serving a host, and this one
is reachable and enforces the epoch itself.

'existing-only' is deliberately untouched — those cells serve the hosts they
already hold, and only `assignmentStrandedOnUnservedCell` may release that pin.
A host with an open `relay_assignment_migrations` row keeps its pin too, so
this stays disjoint from the migration machinery.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): gate re-placement on a roll-isolation marker, not on admission

Review of the first commit found the predicate wrong. `migration-only` is an
admission class, not a drain signal: an Asia `--mode rollback`, an evacuation or
forward-recovery target awaiting a separate promote dispatch, a failed same-cap
wave's re-isolate, an abandoned migration retired on its target and a rehome
settlement all park loaded cells there durably, with no migration lease and no
open migration row. All five were indistinguishable from a roll's isolate, so
the first commit would have converted `operate-relay-asia-admission --mode
rollback` from a reversible admission flip into a mass move of ~4,000 hosts —
and, because `leastLoadedCell` treated region as a preference, into us-central1.

The signal is now an explicit stamp. `relay_cell_admission` gains a nullable
`roll_isolated_at`, added through the shared schema runner's catalog pre-check
so a migrated database takes no relation lock on boot and an un-migrated one
gets a catalog-only rewrite. The same-cap isolate step is its only writer, via a
new optional `rollIsolatedCells` on the selector apply; the same UPDATE that
writes the state clears the stamp whenever a cell leaves 'migration-only', so a
restore cannot leave one behind and a failed wave's re-isolate keeps the one it
has. Every other admission writer omits the field, so its cells stay unmarked
and their hosts stay pinned. Old directors ignore the field; old callers never
send it.

Region is now a constraint rather than a preference on this path only: a
re-placement must find a general, live cell with connection headroom in the
host's own region, or the pin is kept and one
`orca_relay_sticky_replacement_deferred` event is logged. Cross-region spill is
no longer reachable here.

The fence bypass is narrowed to a live incumbent. It was always a no-op for the
intended case, and for a stamped cell that stops heartbeating while still
holding sockets it reopened split-brain; that cell now takes the dead-cell path
unchanged.

Also: the hot-path admission reader no longer throws on an unrecognised state —
it sits on every sticky dial and the rule it feeds is "move the host", so an
unreadable row has to mean "don't". And the sticky lane reads the admission row
once for both the stranded rule and the stamp instead of twice.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): emit the re-placement events after the transaction commits

CodeRabbit on assignment-store.ts:1075. Both events were written where they are
decided, which is inside assignOnce's transaction. A reservation or lease write
failing after that point rolls the placement back, but a line already on stdout
cannot be rolled back with it — so the canary this PR asks an operator to read
would count re-placements that never happened, and a Postgres transaction retry
could leave a stale line behind as well.

The transaction now returns its events alongside the RelayAssignment and the
caller flushes them once it has resolved. Returning them rather than setting a
variable in the enclosing scope is what makes the retry case safe too: only the
attempt that committed can carry its events out. assign()'s signature is
unchanged; the extra shape lives entirely inside assignOnce.

orca_relay_sticky_replacement_deferred was moved the same way. It cost one more
push into the array that already existed, and it is decided inside the same
transaction, so leaving it behind would have been the odd case rather than the
cheap one.

The new test injects a failure on the first write after the decision, asserts no
event is emitted, and asserts the assignment is still on its original cell —
without that second assertion the absence would only prove the emit was early,
not that it would have been wrong. A control dial with nothing injected emits
exactly one event, so the case cannot pass on a broken harness. With the emit
put back inside the transaction, it fails.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): expire the roll stamp, correct the wire note, assert the stamp landed

Delta review findings B, D and E. A (the deferral path's cost) is deliberately
not implemented; it is now written up under Follow-ups in the PR body as
required before any Asia roll, because it cannot fire in a US canary.

B, which also closes C: the stamp was written, carried and never compared to
anything. A roll isolates and restores one cell inside ~15 minutes, so a stamp
older than two hours is not a roll in progress. It is a failed wave whose
failsafe re-isolated a possibly healthy cell and is waiting on an operator — the
postmortem in this tree records gaps of hours — or an orphan left by a director
rollback whose restore wrote 'general' without the clause that clears the stamp,
which the selector's 'keep' branch would then preserve until some later park
reactivated it. Both want the same answer and it is the pre-existing one: keep
the pin. One comparison against a value already on the row.

The bound takes the caller's `now` rather than reading the clock again, so one
assign reasons about one instant; the stamp's age is now a thing that decides
whether a host moves, and two clock reads could disagree across it.

D: the comment beside the new request field claimed an updated caller reaching
an older director "is simply ignored". The schema is .strict(), so it is a 400.
That fails closed — the isolate aborts before MUTATION_STARTED is set and
nothing is written — but it is a deploy ordering constraint, and it was
undocumented. The comment now says so and the PR body's rollout notes carry it.

E: nothing read the `rollIsolated` the script already prints, so an older script
against a newer director would silently produce today's behaviour and the canary
would read as "the fix did nothing" with no way to tell that from a wrong
premise. Both isolate steps now assert it, beside the generation they already
parse.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 04:18:57 -04:00
Jinwoo Hong e476193bf5 chore(relay): bound the shadow health gate and apply a pending backend update on resume (#21865)
* fix(relay): bound the same-cap shadow gate and apply a resumed backend update

Two findings both adversarial reviews of tonight's merged set agree on.

The report-only shadow health gate (#21849) had `continue-on-error: true` but
no step timeout. That bounds the step's contribution to the job outcome, not
its clock. Its reads are serialised, and a failure that answers nothing slowly
— an expired credential, a project-wide Logging 429 storm — makes every read
cost its full 3 x 60 s retry budget, so the cost scales with the roll window:
roughly 8S + 2 reads for S ten-minute sub-windows. A 40-minute window is about
34 reads, or 108 minutes, against the job's `timeout-minutes: 75`. A cancelled
job cannot be absorbed by continue-on-error, fires the failure-gated cleanup
isolation on an already-restored cell, and stops the strict next-cell chain.

Give the step `timeout-minutes: 5` and the artifact upload `timeout-minutes: 2`.
A timed-out step is a failed step, which continue-on-error covers, so the job
stays green. Inside the script, stop reading after an overall four-minute
deadline and report the remaining checks unverified, so the normal outcome is a
written verdict rather than a killed process; the step timeout is then only for
a hung process. The census test pins both timeouts and that the deadline leaves
the step time to write its verdict.

The resume branch (#21860) accepted `changes == 0` with a non-empty
`backendUpdate` as complete and applied nothing, so a resumed cell silently
kept the 300-second drain and no request logging behind a green resume. That
shape means the template and MIG are converged and only this cell's reviewed
backend update is left, so apply the saved resume plan — the validator has
already bounded it to this cell's backend and neither attribute restarts an
instance — then continue as converged. Template-and-MIG drift still applies
nothing, which is what a resume means, and a stranded cell's explicit MIG
replace is unchanged.

Claude-Session: relay-same-cap-gate-timeout-and-resume

* fix(relay): raise the shadow gate bounds clear of a healthy gate's read time

A healthy gate is already minutes of serial reads on the 2-vcpu runner, so a
four-minute deadline would report unverified tails on ordinary days and stop
the shadow roll measuring the comparison it exists for. Raise both together:
the step to eight minutes and the script's own deadline to seven, keeping the
census pin that the deadline leaves the step room to write its verdict. The
job budget is unaffected: a ~14-minute cell plus eight is well inside 75.

Claude-Session: relay-same-cap-gate-timeout-and-resume
2026-09-20 21:35:09 -04:00
Jinwoo Hong 2524737ef0 chore(relay): apply the cell backend drain and request-logging settings inside each same-cap wave (#21860)
* chore(relay): target each cell's backend service from the same-cap job

The 60 s connection drain timeout merged in #21848 has no safe apply path.
A root plan scoped to the backend services alone still pulls every
`google_compute_instance_template.relay_gce_cell` in as a dependency, and
standing image drift turns all 29 into replacements, so applying it would roll
the fleet at once.

Add `google_compute_backend_service.relay_gce_cell["${TARGET_CELL_ID}"]` to
both plan invocations in the per-cell same-cap job, next to the template and
MIG it already targets, and teach the reviewed plan validator to allow exactly
one extra change: an in-place update of that one cell's backend whose only
changed attribute is `connection_draining_timeout_sec`, landing on the
constant `validate-relay-asia-topology-plan.mjs` exports. Any other attribute,
any other resource, or a backend for another cell still fails the validator.

The accepted update is reported as `connectionDrainUpdate` and kept out of
`changes`, so the apply step's stranded branch and the resume step's drift
branch keep reading the template-and-MIG count they were written against; the
resume branch additionally accepts a plan whose only pending change is that
drain update, which restarts nothing.

Claude-Session: relay-same-cap-targets-cell-backend

* fix(relay): also let the same-cap wave apply this cell's LB request logging

A read-only production plan for production-gce-c7 showed the live US cell
backends carry no `log_config` at all, while relay-gce-cells.tf has declared
`log_config { enable = true, sample_rate = var.relay_gce_cell_log_sample_rate }`
on every cell backend since the Terraform root landed in 3eec77c11a (#18413).
Nothing has applied it because every production apply since is a per-cell
targeted plan that names only the template and the MIG.

So the real canary plan's backend moves two paths, not one:
`["connection_draining_timeout_sec", "log_config.0"]`. The drain-only validator
rejected exactly that plan, which would have stranded the cell mid-wave after
the drain had already started.

Accept both, each optional, for this cell's backend only: the drain landing on
RELAY_CELL_CONNECTION_DRAIN_SECONDS, and a log_config of exactly one block with
`enable = true` and `sample_rate` equal to RELAY_CELL_LOG_SAMPLE_RATE, the
declared default of a variable no environment file overrides. Any third path,
a different sample rate, disabled logging, another cell, or a replacement still
fails. The accepted paths are reported as `backendUpdate`, which the resume
branch now reads instead of the drain-only flag.

Verified against the real production plan: the masked, three-target plan for
production-gce-c7 contains exactly that cell's template, MIG, and backend
service and nothing else, and this validator returns
`{"changes":2,"backendUpdate":["connection_draining_timeout_sec","log_config.0"]}`.

Claude-Session: relay-same-cap-targets-cell-backend
2026-09-20 20:35:52 -04:00
Jinwoo Hong 5b8ac36f41 chore(relay): add a report-only post-wave health gate to the same-cap cell job (#21849)
* feat(relay): report a post-wave health verdict on each same-cap cell, without gating on it

After a same-cap cell finishes rolling, an operator reads five things by hand
before dispatching the next cell: director 503s against the same clock hour a day
and two days earlier, whether the cell's new container announced its listener and
has stayed up, the cell's own pool pressure, the asia-east2 pool trio, and Cloud
SQL FATALs. This runs those same reads automatically and records PASS / WARN /
WOULD_BLOCK with its numbers, so its calls can be compared with the operator's
over a full roll before it is ever allowed to stop one.

It cannot fail a cell in this change. The script exits 0 on every verdict, and
the step is continue-on-error, so even a crash stays off the job's outcome and the
failure failsafe cannot fire on anything it observes. It also runs after the
restore, so no cell waits on it to go back into admission.

Cloud Logging returns only --limit entries and says nothing when it truncates, so
every count is split into sub-windows of ten minutes and a sub-window that comes
back at the limit is reported unverified rather than as a count. Windows are
always explicitly bounded: --freshness does not bind on these logs.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): bound the shadow gate's cell reads at the apply start and cap every read

Four fixes from review, all in the report-only shadow health gate.

The boot search opened at apply-completed-at, which is stamped after
`terraform apply` and `wait-until --stable`. The new container announces its
listener while the MIG is still converging, so that bound is already past the
announcement it looks for and a healthy roll read as would-block. The job now
stamps apply-started-at immediately before the apply, and the boot search opens
there; apply-completed-at is kept, recorded rather than judged, so an operator
comparing verdicts can see apply time next to boot time.

The crash query started at the newest listener timestamp, which erased any crash
before it. A crash-restart loop ends with an announcement that looks like a clean
boot, so that is exactly the case it hid: against production, the 2026-09-20 c28
crash at 20:18:10 was dropped because the listener landed at 20:18:27. It now
runs from the apply start, still scoped to the instance id the listener
identified, and that crash is counted.

A runtime-metrics read that came back at its 500-entry limit fed judgePool as
though it were a complete sample run. A truncated run has holes and the
consecutive-sample rule reads a hole as a recovery, so it now reports unverified.

gcloud reads had no timeout. continue-on-error bounds the job's outcome but not
its clock, so a stalled read could have spent the rollout's remaining minutes.
Each read now gets 60 s and a timed-out read is just a failed read.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): require each shadow-gate stamp's presence before asserting its order

The ordering assertion used indexOf, which answers -1 for an absent stamp, and
-1 precedes every real offset. Deleting the apply-started-at line left the test
green, so the census could not see the fix it was written to pin.

Each stamp's presence is now asserted first, with a message naming the stamp and
the step, and presence is judged inside the step that owns the stamp rather than
anywhere in the file: a stamp written into a neighbouring step records the wrong
instant but would satisfy a whole-file match.

Control-run against a scratch copy of the job. Deleting drain-started-at,
apply-started-at, or apply-completed-at each reds with its own message, and
moving apply-started-at after terraform apply reds on the ordering assertion, so
presence and order both fail independently.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-20 19:10:35 -04:00
Jinwoo Hong 4f839cc8c9 chore(relay): cut the cell LB connection drain to 60 s and allow ten-cell same-cap batches (#21848)
* perf(relay): cut the cell LB drain to 60s and widen the same-cap batch to ten cells

Two independent sources of relay roll wall clock, neither of which protects a
host:

1. `connection_draining_timeout_sec` on the per-cell backend services was 300s.
   The same-cap job drains every host off the cell to a restart-safe condition
   before Terraform runs, so the LB drain only ever covers a host still
   mid-handshake. Measured 2026-09-16 over ten same-cap cell jobs, it sat as
   ~5m55s of dead time between `Apply complete` and the old VM powering off,
   inside an 8.5-minute `wait-until --stable` step. Now 60s, and pinned in the
   topology `check` block beside the other fixed-one invariants.

2. The same-cap wave capped a batch at four cells, so a 22-cell roll needed six
   batches, six single-use monitor gates, and a human handoff per batch. The
   wave workflow now declares cell_1..cell_10 with the identical serial shape
   and chaining, and the validator accepts two to ten.

The shared wave-index rule (`relay-monitor-evidence.mjs` and the relay-ops
preflight CLI) widens from 0-3 to 0-9 so the later cells can present the same
evidence; each job workflow keeps its own narrower range, so the capacity wave
stays at four. Cells remain strictly serial, one at a time behind the rollout
lease, each with its own live preflight.

Claude-Session: https://claude.ai/session/relay-roll-drain-timeout-and-batch-cap

* fix(relay): align the Asia topology plan validator with the 60s cell drain

`validate-relay-asia-topology-plan.mjs` rejected any Asia backend whose
`connection_draining_timeout_sec` was not 300, and
`cloud-deploy-relay-asia-topology.yml` targets
`google_compute_backend_service.relay_gce_cell["<cell>"]` per cell. With the
Terraform local at 60 that workflow would have failed its own plan review.

The validator's two restated topology values are now named exports, and a new
census test reads `relay-gce-cells.tf` and equates three statements of each:
the `relay_gce_topology` local, the topology `check` assert that pins it, and
the validator constant. Terraform cannot export a local to JS, so reading the
source is the only way to stop them drifting; the test was confirmed to fail
when the local alone is moved back to 300.

Repo-wide grep finds no other pin of the drain value.

Claude-Session: https://claude.ai/session/relay-roll-drain-timeout-and-batch-cap
2026-09-20 19:01:44 -04:00
Jinwoo Hong eb6068a434 fix(relay): stop a terminated checked-out PostgreSQL client from killing the cell (#21840)
pg-pool removes its own `error` listener when it hands a client out
(pg-pool@3.14.0 index.js:344) and only reattaches it in `_release`
(index.js:385). Between acquire and release the client therefore has no
`error` listener, so when Cloud SQL terminates that session mid-statement
the emit becomes an unhandled 'error' event and the process exits.
`absorbPostgresIdleClientErrors` cannot see it: pg-pool routes to
`pool.on('error')` only from the idle listener.

Attach a per-checkout `error` listener in the one seam every relay
checkout passes through, log a single warn line, and release the client
with the error so pg-pool destroys it instead of pooling a dead
connection. The listener is removed on release so it cannot accumulate.
The in-flight query still rejects, so existing failure reporting and the
transaction retry ladder are unchanged.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-20 17:46:25 -04:00
Jinwoo Hong 68b11282a5 fix(relay): let the rehome evidence parser read a line the director grew (#21823)
The enable workflow reads the director's `[orca-relay] regional rehome
inventory` line out of Cloud Logging and pins the whole line with one regex.
Adding `hostNotArrivedLast24Hours` in #21813 made every healthy line stop
matching, so "Read fresh aggregate completion and abort evidence" threw
"no aggregate regional rehome inventory evidence" and the fail-closed step
disabled the durable switch at control generation 26.

The parser now requires the six original fields and tolerates further ones
in any order. Extra fields stay fenced by value shape rather than by pinning
the whole line: a field must be a bare name and a non-negative integer or
`none`, so `hostId=someone` is still not a counter and cannot ride along.
An absent count reads as null, not zero, because an older director not
reporting leaks is not the same as reporting none.

`hostNotArrivedLast24Hours` and `oldestActiveAgeMs` now reach the evidence
JSON and the operator step summary.

Two guards close the chain, each verified to fail on the regression it
exists for: a census in the relay package feeds the real formatter's output
to the real parser, and a script-side test pins the parser's output to the
fields the workflow summary renders.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-20 15:50:43 -04:00
Jinwoo Hong fa0010e8d6 fix(relay): abort rehomes whose host never arrived, without disabling the switch (#21813)
A regional rehome whose host went offline right after accepting the move
left its migration row open forever: the target had registered it, the host
held nothing on the source, and the completion sweep could never finish it.
Eight such rows filled REGIONAL_REHOME_CONCURRENT_LIMIT and every later
candidate came back deferred, silently, for 21 hours.

The only sweep that touched them fires at 24 hours and also sets
enabled = 0 on the durable control, so the first leak to age out would have
turned rehoming off, repeatedly.

Adds a director sweep that rolls such an attempt back to its source after
one migration lease, with abort_reason = 'host_not_arrived', reusing the
existing rollback (assignment epoch bump back to the source, lease removal,
superseded target reservation release) and leaving the switch untouched.
The 24-hour sweep keeps its disable as a last-resort latch.

The source cell now names why it deferred, on a new optional response field,
and the director stops walking its candidate page on a deferral no later
candidate can pass. Each poll that dispatched logs one summary line.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-20 15:13:27 -04:00
Jinwoo Hong 030a1e0c77 fix(relay): stop holding a cell row across the whole control accept (#21563)
* fix(relay): stop holding a cell row across the whole control accept

The cell accept path took the host's cell row FOR UPDATE at its first
supersession statement and held it to COMMIT across a dozen round trips,
which capped a cell far from Postgres at a couple of accepts a second.
Fold every cell-row change on the path into one conditional delta write
issued last, so the contended row is held only across the commit.

* fix(relay): give relay_cells one global row lock order, taken last

Moving the accept's cell-row write to the end of its transaction put it
after the host's relay_control_connection_reservations rows, while every
director path that reads the inventory took those rows the other way
round. Pin one order for both roles -- host rows, then the shared cell
row -- by locking the host's reservation rows before the inventory in the
nine director paths that take both, document the tiers next to
CellInventoryLockMode, and add a census that fails on a new path taking
relay_cells first.
2026-09-18 23:16:15 -04:00
Jinwoo HongandClaude 7080eb0604 fix(relay): bound the idle-rehome candidate poll to a window of decisions (#21557)
* fix(relay): bound the idle-rehome candidate poll to a window of decisions

The director's idle-regional-rehome poll built every (eligible host x target
cell in its preferred region) pair, applied the cohort predicate downstream of
that fan-out, sorted the lot, and took LIMIT 100 OFFSET n. Its cost was set by
the size of the fleet and the width of the cohort, so raising the cohort from
10% to 100% pushed it past the serving pool's 5 s statement_timeout and the
rollout stalled at 0.37 hosts/min.

The poll now resolves the cell inventory once (tens of rows), takes a bounded
window of decision rows in primary-key order from a keyset cursor with the
cohort, freshness and cross-region predicates applied first, verifies only that
window against the host-side gates, and ranks targets in the process. Same
candidates in the same priority order; the work per poll no longer depends on
the cohort or the fleet.

Adds a once-a-minute aggregated poll summary so an operator can tell a poll
gated by the dispatch budget from one that found nobody to move.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(relay): pin the rehome verification to the window's exact keys

The window read and the verification read take separate snapshots. The
verification repeated the window's predicate with its own LIMIT, so a decision
that turned eligible between the two reads shifted that LIMIT and pushed the
window's last host out of it -- while the cursor still advanced past that host,
skipping it for a whole sweep.

The verification now names the keys the window returned. Its LIMIT stays as the
optimisation fence that stops Postgres flattening the subquery, but can no
longer truncate a key set that is at most one window long.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-18 21:16:19 -04:00
Jinwoo Hong 164c7140fd fix(relay): stop reporting an unavailable home cell as exhausted capacity (#21518)
A host whose home cell is not live — readiness false, drained, or inside a
boot window — is refused by the committed-fence branch in assignOnce()
without any capacity being consulted. It answered relay_capacity_exhausted,
so every cell boot and every readiness dip printed capacity rejections at
17% fleet utilisation and sent an investigation after headroom that was
never short.

The branch now raises RelayHomeCellUnavailableError, which carries the cell
id and which of cellIsLive()'s conditions failed (draining / booting /
unheard / not_ready). The director logs reason, cause and cell, and returns
the new reason in the same retryable 503. Nothing on the wire reads the
body: the desktop client discards it unread and branches on status only,
and no log-based metric or alert parses the reason. The load harness, the
only body-reading consumer, gets its own bucket so a home-cell rejection no
longer inflates the capacity count.

Hinted grants are now logged on whichever lane served them, so a host that
failed sticky verification and was rehomed by placement leaves a record of
where it landed. Unhinted placement grants stay silent.
2026-09-18 17:55:35 -04:00
Jinwoo Hong 3467e5f6b5 fix(relay): check the rehome dispatch budget before planning the candidate join (#21517)
* fix(relay): check the rehome dispatch budget before planning the candidate join

`selectIdleRegionalRehomeCandidates` read the enable control and the fleet
safety snapshot, then ran the twenty-table candidate join, then handed every
row to the worker, which POSTed each one to its source cell. Only there — in
`commitIdleRegionalRehome`, three statements into a write transaction that
takes `FOR UPDATE` on two global single-row tables — was the durable dispatch
budget consulted.

The budget is ten moves a minute (`next_dispatch_at = now + 6s`), and five
directors poll every six seconds, so most of that work was spent to be told
the budget was closed. A five-minute `paused_until` made every poll in the
window do it.

The gate is a single-row primary-key read, so it goes in front. An absent row
means the budget has never been spent and opens the gate, matching the
INSERT ... ON CONFLICT DO NOTHING the commit path already relies on.

* test(relay): assign the closed budget field once so the case runs on Postgres

The two gate cases zeroed both `next_dispatch_at` and `paused_until` and then
set the one under test, which names that column twice in a single `SET`. SQLite
accepts it; Postgres raises "multiple assignments to same column", so both cases
failed whenever `ORCA_IDLE_REHOME_POSTGRES_URL` pointed the suite at a real
server -- exactly the backend the gate has to hold on.

Setup already leaves both fields at 0, so naming the other one bought nothing.
2026-09-18 17:55:27 -04:00
Jinwoo Hong ce5d8c02d4 fix(relay): wait out a cold proxy at boot instead of exiting the cell (#21516)
* fix(relay): wait out a cold proxy at boot instead of exiting the cell

A cell container starts its relay process beside a cloud-sql-proxy that is
itself still dialling. The first pool acquire therefore competes with a proxy
cold start, and the 2s connect timeout that protects the request path fires
before the proxy is listening. `openRelayDatabase` rejects out of the region
backfill, the top-level await rejects, and the process exits; COS restarts the
container and the next boot succeeds 1-3s later. The 2026-09-18 fleet roll saw
0-7 of these per cell, including on cells with zero hosts, so it is a property
of the boot sequence rather than of database load.

The boot open now retries on transient errors only, inside a 45s wall-clock
window with exponential backoff from 250ms to 4s. The classifier is the one the
request path already uses, so a rejected credential or a bad URL still exits on
the first attempt. Each wait logs `orca_relay_boot_database_retry` and a
give-up logs `orca_relay_boot_database_failed`, both with the bounded error
category, so a rollout can tell a slow boot from a stuck one without reading
container exit codes.

The bounded startup retry is lifted out of `reconcileCellAdmissionAtStartup`,
which had the same loop; its attempt budget, flat delay, and both log events are
unchanged (a flat delay is a cap equal to the base).

* fix(relay): retry the boot open only when Postgres is unreachable

The boot open re-runs the schema apply, and applyPostgresSchema refuses to
repeat a DDL lock timeout on purpose: relation locks are granted in queue order,
so a repeat parks every writer behind the same statement again. Gating the boot
retry on the full request-path classifier would have re-queued it up to 16 times
in 45s on sustained 55P03 - the mechanism behind the 2026-09-16 outage.

The boot call site now has its own predicate: pool connect failures (both
connect-timeout messages and an acquire-marked early-ended socket) plus 08001
and 08006. Lock and overload SQLSTATEs - 55P03, 57014, 53300 - exit on the first
attempt. The retry predicate moves onto the policy because what a step re-runs,
not the request path, decides what it may repeat; the startup reconcile keeps
the full classifier, which is what lets it wait out 55P03.
2026-09-18 17:55:19 -04:00