Adds a sixth asia-east2 cell at the C31 shape (cap 3000, 6000 request
units, pool 16, e2-standard-4) in asia-east2-c, pinned to the f30b5cb1
cell image. It gets its own topology and registration wave but no
promotion wave, so the admission script and workflow refuse to promote
it; it stays a migration-only landing zone and out of the fleet pool list.
Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041
* refactor(relay): stop creating three tables nothing writes
relay_confirmable_splices, relay_cell_drain_attempts and
relay_migration_leases are created at every boot and referenced nowhere
else on main except tests. In production all three hold 0 rows and
pg_stat_user_tables shows 0 inserts, 0 sequential and 0 index scans since
the 2026-09-28 stats reset.
This removes them from the schema and the transaction-phase table, and
drops the test references. It does not drop them: an older image still
runs CREATE TABLE IF NOT EXISTS at boot, and a drop racing that create can
fail the older boot's schema step. A follow-up can drop them once no image
that creates them can be rolled back to.
* docs(relay): note the account-erasure prerequisite for dropping retired tables
* fix(relay): retain confirm results for a week and audit events for 90 days
relay_confirm_results (8.3M rows, 5.8 GB) and relay_audit_events (8.4M
rows, 3.6 GB) were never deleted. The director's credential cleanup now
reaps both after its existing passes:
- Confirm results older than 7 days. The only reader replays a stored
result for a retry of the same request on the same connection basis; a
different basis is already refused as a tuple mismatch. committed_at has
no index, so this uses the TID-window reaper from the reservation prune,
capped at 250 rows a tick (the 7.4M-row backlog drains over 2-3 days).
- Audit events older than 90 days, ordered by `at` so the batch walks
relay_audit_events_at. Nothing in the relay reads them back. The oldest
row is from 2026-07-14, so this deletes nothing until 2026-10-12.
reapBatch now deletes by `ctid = ANY(ARRAY(...))` on Postgres: with
`ctid IN (...)` the planner can choose a hash join over a sequential scan
of the whole table.
* chore(relay): record the decided audit retention
Both were promoted to general on 2026-10-01 (selector 344 shows them general),
but the same-cap job still classed them migration-only. Rollback mode on either
would isolate (general -> migration-only) before its no-op check failed, demoting
a live cell. They stay out of the fleet pool list (pool 10, not 16).
Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041
relay_control_connection_reservations was never deleted: 13.8M rows and
9.1 GB in production, 99.999% of them in state 'released' and ~400k more
a day. Nothing reads a released row (every reader filters it out by
state), but each placement still locks all of its host's rows, ~160 on
average.
A new director sweep step deletes released rows older than a day. It walks
the heap in TID ranges of 16 pages, because without an index on
released_at a `LIMIT n` delete plans as a sequential scan from page 0 that
gets slower as the head of the heap empties (production EXPLAIN). Each
statement selects its rows FOR UPDATE SKIP LOCKED, so a row a request holds
is skipped rather than waited on, and deletes them by `ctid = ANY(ARRAY(...))`
so the delete is always a TID scan. A tick stops at 400 rows, 128 pages or
250 ms, which drains the backlog over about three days at five directors.
evacuateDeadCells ran one query that joined every assignment to its cell's
liveness and fence state, ordered by primary key with LIMIT 100. In
production no row ever matched (the only dead cells are stale existing-only
cells without a committed fence), so the planner walked the whole
relay_assignments primary key every call: 551 ms mean over 85k calls,
~12% of relay database time.
A cell-level pre-check now evaluates the predicate's cell-only half over
the ~33-row cell tables (3 ms in production). It is a superset of the cells
the host query can act on, so the sweep returns early when it is empty and
otherwise restricts the host query to those cells. The host predicate is
unchanged.
Reuse the existing first-frame finish pattern to remove stage-owned timer/message/close callbacks on receipt, timeout or close; check OPEN after successful async assignment verification.
* Reuse the parsed notification when leasing a push delivery
Pass the notification already parsed for the dismissal check into the existing private delivery builder.
* Reuse the parsed notification when leasing a push delivery
Pass the notification already parsed for the dismissal check into the existing private delivery builder.
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run
A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed,
single-use, five-minute-fresh evidence. Each apply wave now samples fleet health
itself right before isolation, with the monitor's evaluator, thresholds, and
tolerances, for a window sized to the cell's host count (3/5/8 min), plus three
lookback rules: no cell container exit in 10 min, no minute over 500 director
503s in 10 min, and director concurrency p99 within the monitor bar over 4 min.
Removes the monitor-run inputs, the gate's consume/authorize steps, the
break-glass override, and the same-cap-only authorization shapes in
relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are
unchanged.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): bound the pre-drain sample overrun and keep the drain token fresh
Review follow-ups: alternating tolerated readings could hold the sample open
until its step timeout, so cap the overrun at three samples past the window;
record why a read failed; mint a fresh admin ID token for the drain after the
sample; raise the job timeout to 90 min so a long sample cannot cancel the
job past the failsafe.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule
The exit rule counted every relay container exit fleet-wide, so a cell that
crashes every few hours (c25, 12 a week) blocked the very roll that fixes it,
and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part
in. Exits are now grouped by instance, each instance is named by its own newest
runtime-metrics log line, and only exits on general or migration-only cells
other than the target count. An exit no configured cell can be named for trips
the rule; a failed lookup is a failed read. relay-observability.tf joins the
evidence-code set because the rule depends on its filter.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* feat(relay): declare US cells c32 and c33 at the 3,000-host shape
Declares two us-central1 cells at the Asia shape (cap 3000, 6000 request
units, e2-standard-4) with the US default pool of 10, and generalises the
Asia topology and admission ladder to derive each wave's region from its
reviewed zone, leaving every Asia wave's behaviour unchanged.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* docs(relay): note the US canary tie-break and leave the fleet pool list to promotion
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): plan C32 and C33 as one topology wave
The live-image overlay refuses a declared non-target cell with no template,
so a lone C32 plan would fail on C33. Registration and promotion stay
one cell at a time.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): admit drained hosts through their own lane and stagger their return
A same-cap roll's drain sends every host on the isolated cell back to the
director at once. Those hosts reconnect through the sticky lane (one slot per
director), and each one's re-placement holds that slot for most of a second
behind the region-wide inventory lock, so ordinary reconnects time out behind
them and the drained hosts retry every 2 s: ~30k 503s per drain.
The reconnect verification read now also says whether the host's home cell is
isolated for a roll right now (same predicate re-placement uses). Those hosts
release the sticky slot after the read and take a separate drain-return lane
(1 per director, matching the store's per-director placement serialization).
When that lane is full the host gets a Retry-After that reserves the next
free service slot, paced by the measured re-placement time and capped at 300 s.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): keep drain returns inside the placement pool budget and their own slot
Review follow-ups for the drain-return lane:
- The lane now borrows placement permits (never placement's last, never ahead
of a queued placement), so placement + sticky still bounds the database pool.
- A host's own early retry (row-busy redial, duplicate dial) gets the 2 s lane
interval instead of a fresh slot behind the cohort, and a host that returns
early to the same director keeps its reserved slot.
- Classification also excludes an open migration row whose lease counter
lapsed, matching the re-placement rule.
- The load test now runs five directors behind random routing with a shared
inventory lock.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG
The stranded-rollback recovery ran a gcloud rolling action, which renames the
MIG version outside Terraform. Every later plan for that cell then reverted the
label, and the capacity-plan validator refused the revert as an unreviewed MIG
change, so the cell could be neither rolled nor rolled back.
The validator now accepts a MIG field moving back to what relay-gce-cells.tf
declares (version name and update policy), in every mode, and a test pins those
values to the Terraform file. The stranded branch recreates the cell's single
instance with recreate-instances, which leaves the MIG untouched.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): let a label-only MIG plan through and recreate on it in a stranded rollback
A stranded rollback whose template is already in place plans only the version
name revert. The validator still required the MIG template to move, so that
plan was refused, and the recreate gate (changes == 0) would have skipped a
plan of one change and left the drain flag set. Require the template move only
when no declared field reconciles, and recreate whenever the template was not
replaced (changes < 2).
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup
The same-cap gate now verifies the dry-run on its own clock and records the
authorisation instant in the single-use consumed marker; each cell job checks
the evidence age at that instant and bounds its own start after it.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): refuse a re-run same-cap gate before it consumes evidence; tighten order tests
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): let a draining cell pass restart-safe through refused redials
A draining cell refuses every control and host proof, so once no session,
splice, or queued byte remains, an in-flight or reserved connection unit can
only belong to a dial the cell is about to refuse. The restart-safe wait no
longer resets its pace-window streak on those units, and its progress line
now prints every counter the gate reads.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* test(relay): pin fail-closed parsing of handshake counters on a draining cell
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
The parent workflow collapses canary-apply and batch-apply into the job mode
apply, so the headroom step's canary-apply/batch-apply condition never held and
the gate was skipped on every real roll. Run it wherever the drain runs (apply,
rollback before its restart) and in read-only verify; skip only a resumed
rollback, which drains nothing. A new workflow-shape test fails on any job
step comparing against a mode the parent cannot pass, and on a drain that can
run without the headroom check.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): gate the Asia canary on region fallbacks against a pre-canary baseline
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): gate the Asia canary on region fallbacks against a pre-canary baseline
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): let a same-cap drain finish when only unplaceable hosts remain
The c28 canary on 2026-10-01 drained the cell to zero live connections, but four
hosts with no free slot anywhere kept redialling and held director leases on it,
so the restart-safe wait timed out and left the cell isolated and empty.
The drain wait now also passes once the runtime has carried nothing for a
sustained quiet window while a small, capped number of leases remain, and logs
the escape. Apply modes also refuse a cell whose hosts exceed 80% of the free
slots on the other general cells, so a wave cannot strand hosts in the first place.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): define restart-safe by the cell runtime, not director leases
Replaces the opt-in stranded-host escape with a corrected definition. A
restart is safe when the cell runtime carries nothing live and no migration
is open, sustained for the drain pace window. Director activity leases lag
hosts that already left or cannot be placed, so they are reported in a
progress line and the verified result instead of blocking the restart.
The same-cap drain passes its existing pace window. The headroom script is
added to the trusted evidence code paths with the other production scripts.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): print stranded director counts on every restart-safe sample
Each restart-safe poll now prints its sample count and the director's
restart-blocking leases, request units, reserved remainder, and migrations
under `stranded`; the verified line carries the same object. Open migrations
still block because each is pinned to the cell incarnation a restart replaces.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): require the pace window for every live restart-safe wait
Pre-auth and total connections no longer reset the restart-safe window:
on drained c28 they flickered with unauthenticated redials in a third of
samples, which a restart does not lose. They stay in the progress output.
Every live restart-safe call must now pass --pace-window-ms. The capacity
job and staging proof drain unpaced, so they pass the production 300000 ms
window, and the calls that relied on the 180000 ms default get 480000 ms.
Headroom free slots now follow the director's placement rule: the admission
pause minus the larger of observed and enforced units, minus outstanding
control reservations.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): refuse a redial at once while the host's own release holds its row
During an Asia drain the host whose socket closes is the one that redials.
Its release on the draining cell locks its assignment row first, then waits
on the cell's busy row for up to the lock timeout. The director's sticky
and placement paths waited on that assignment row inside the single sticky
slot, and the sticky path then locked the busy cell row itself before it
checked isolation. The slot backed up and dials timed out fleet-wide.
Both paths now take the host's assignment row NOWAIT and throw
RelayAssignmentRowBusyError when it is held. /v1/assign answers that with
503, Retry-After 1 and error assignment_row_busy, logged with its own
reason. The sticky path decides isolation before it touches the pinned
cell row. An isolated retry keeps its own tier as its retry scope, so a
busy lock inside it no longer falls to the all-rows path. The local drain
arm drops its zero-release-failures bar, which the Asia arm never had.
The drain harness gains a departing-host arm: each host releases its own
lease, then redials after 150, 400 or 1000 ms on the desktop client's
5-5.5 s pacing. At 400 ms, main rejected 83 of 180 first dials by sticky
wait timeout, placed 11.6/s with 3.1 director backends lock-waiting, and
took 11.2 s at p95 from release to placed. Now: 16 fast refusals, 18/s,
no lock waits, 5.7 s at p95.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): wait briefly for a calm host's row and keep the dead-cell sweep going
The dead-cell sweep treated RelayAssignmentRowBusyError as fatal, so one
busy host ended the sweep for every later host each tick. It now skips
that host and carries on.
The sticky path refused a busy row at once for every host. A calm host
redialling after its own clean close often meets its own short release,
and a refusal costs it the client's 5 s assign gate. When the pinned cell
is general and live, the sticky path now waits up to 1 s for the row
before refusing. A roll-isolated, parked or dead cell still gets the
immediate refusal. Placement keeps NOWAIT, because it holds cell rows
while it would wait. A resume refused for a busy row now carries
Retry-After 1 as well.
The departing-host harness arm now bounds the busy refusals at 20% of
hosts and the p95 at 8 s, and counts unexpected errors apart from
retryable refusals.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): never wait on a host's row while the sticky retry holds its cell row
The inventory-first sticky retry takes the pinned cell row before the
assignment row. With the calm-host bounded wait it could then wait up to
1 s on the assignment row while holding the cell row, the reverse of the
ranked lock order, against this host's own release, which holds its row
and wants the cell's. The bounded wait now applies only when no cell row
is held; the retry stays NOWAIT.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* test(relay): reproduce drain-release contention against director placement
Adds a Postgres harness that drives releases from an isolated cell over a
171 ms per-statement pool while five directors re-place reconnecting hosts
through the sticky lane. At 18 releases/s placements fall from 20/s to about
5/s and every active director backend is blocked on relay_cells.
Moves the per-statement delay pool into a shared test fixture so the
rehome target-row test and this harness use one implementation.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): re-place hosts off a draining cell without locking its row
Sticky re-placement of a host whose cell is isolated for a roll locked
every relay_cells row. During an Asia drain the source row is held by the
cell's own releases for a round trip each, so the placement waited on it,
the single sticky slot backed up, and /v1/assign returned 503 fleet-wide.
The isolation decision now comes from an unlocked read, and the placement
locks only the same-region general rows it can move to. It no longer writes
the source row: the host's source leases stay, and each one's own release
or expiry takes its units back off the source. The assignment keeps its
activity counters and adds one control instead of resetting them. With no
same-region headroom the path falls back to the all-rows lock, as before.
Dormant hosts hold no units, so their placement also locks only the
general rows and skips the zero write to their old cell.
The drain harness now asserts the after picture: 20 placements/s at Asia
latency with no director lock waits, against 9.2/s and 4.8/s before.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): take a moved host's units off the cell that holds them
After a narrowed re-placement a host keeps leases on its old cell while
its assignment row names the new one. Two paths charged the row's whole
counted total to the row's cell: aggregate expiry, and the lease deletion
in dead-cell and stranded re-placement. Both over-charged the new cell and
left the old cell's units stranded.
Aggregate expiry now skips hosts that still hold any lease; the lease
sweep takes each lease's units off its own cell. Placement frees each
deleted lease's units on that lease's cell, charges the old cell only for
units no lease backs, and sets the counters from the leases it keeps plus
the new control. The narrowed path runs only when the counters already
match the leases, so it never needs to write the old cell's row.
The all-rows re-placement off an isolated cell follows the same rule, so
it no longer decrements the source at placement either.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): try other regions before the all-rows lock when re-placing off a roll
With every same-region neighbour at its connection cap, the narrowed path
found no target and fell back to the all-rows lock behind the busy source
row, which is the drain brownout again. It now tries a second tier, general
cells in every other region, in its own transaction over one ordered
lockCellRows, still never the source row. Only when no general cell in any
region has room does it fall back to the all-rows path, which keeps the pin.
This changes the policy from #21911, which refused to re-place an isolated
host across a region. The unit tests that encoded that rule now assert the
tier order instead.
The drain harness gains a US cell and an arm with every Asia neighbour
capped: 200 of 200 dials placed cross-region at 20/s with no lock waits,
against 0 placed and 144 sticky rejections on the previous head. Its pass
bars are now the rejection share and the lock-waiting share, not the
placement rate a slow runner's pacing can move.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): take the narrowed path for hosts whose counters sit below their leases
Main's old placement reset a moved host's counters while keeping its source
leases, and those leases' releases floored the counters at zero. Such hosts
hold fewer counted units than lease units, and requiring equality sent them
down the all-rows path behind the busy source row.
The narrowed path now requires only that the host holds no units no lease
backs, the one case that needs a write to the old row. Its placement
already rebuilds the counters from the kept leases plus the new control,
so a drifted host is healed by its next re-placement.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* test(relay): judge the local drain arm on completion and lock waits, not rate
The local-latency arm asserted at least 18 placements/s at a 20/s dial rate,
which a slow runner's pacing alone can miss. It now asserts what the Asia
arm does: no dial failures, sticky rejections under 10% of dials, every
other dial placed, and director lock-waiting under half a backend.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
Completes the first pass over every test area in the repository. Sweep over
`mobile/src`, `config/scripts`, `cloud/`, and `tests/` (1,494 files in scope, with
the 24 files under `mobile/src/test-support/rpc-recording/` deliberately excluded).
31 case declarations removed across 17 files, 2 test files deleted, 356 lines gone.
What went, by pattern:
- Cross-boundary replays of a shared helper. A whole mobile file re-ran
`extractPendingAsk`/`parseAskFromStatus`/`formatAskAnswer`, all owned by
`src/shared/native-chat-ask.test.ts`, `native-chat-ask-fifo.test.ts` and the
renderer's interactive-prompt suite — one case title was verbatim identical to the
owner's, and the owners' inputs are supersets. The mobile file imported the shared
module directly and exercised no mobile transport, lifecycle or rendering.
- A case whose input cannot reach the behavior its title names: "arms it on Android
while the drawer is open", where `use-back-claim.ts` has zero
Platform/OS references, so flipping the mocked OS changes only shadow styles.
- Identity copiers, including one asserting `prSidebarRenderBranch(state) ===
state.kind` against a production body that is `return state.kind`. The function
stays; it has three live callers.
- A test of the runtime rather than the product: a case asserting Node's own
`EventEmitter` crash contract on a bare emitter, with zero production code in the
path. The guard it documents is exercised behaviourally by the case after it.
- Duplicate invocations, one of them provable rather than eyeballed: with
`MODULE_SCOPE_ENV_WRITER_PIN = 0`, `files.size <= 0` is strictly implied by the
sibling's `expect(offenders).toEqual([])`, since a non-empty `offenders` forces
`files.size >= 1`. The pin's own doc says it may only ever be decreased from 0, so
it could never become a meaningful bound either. Its policy guidance survives as a
comment; the file's real ratchet and its regex self-test both stay.
- Expected values produced by the test's own arithmetic, and a p95 case strictly
implied by a sibling that already pins exact p95 and exact max over a wider range.
One production line goes: the `export` keyword on `assignmentCleanupSteps` in
`cloud/apps/relay/src/assignment-cleanup-steps.ts`. The function itself stays and is
still called internally; only the test-only export was orphaned.
Kept deliberately: everything a gate cites, checked by case title and not only by
file path; a gate-cited case that does not deliver its claim (reported instead — see
below); a cross-version wire cell whose ledger is never invoked, left under the
raised bar for wire coverage; and every limit, bound, quota and provenance guard.
Nothing under `mobile/src/test-support/rpc-recording/` or
`mobile/rpc-foundation/goldens/` was touched — those bytes feed a `recorderSha256`
digest pinning 398 golden recordings.
Verified: `mobile` vitest over the modified mobile files (8 files, 50 cases);
`mobile/scripts/check-tests-typecheck-ratchet.mjs` OK (898 files in program, 125
grandfathered, none @ts-nocheck); relay suite 799 passed; `check-reliability-gates.mjs`
140 gates; both deleted files confirmed absent from the gate manifest,
`cloud/package.json` and `mobile/tests-typecheck-baseline.txt`.
Seven local failures were investigated and none is caused by this change: five
`mobile-web-app-*-render` tests drive `playwright-core` chromium/webkit and need
browsers this machine lacks, `release-checkout.unit.test.ts` needs cross-version git
refs, and `e2e-worker-env-isolation.unit.test.ts` fails identically with its HEAD
content restored — it recurses `tests/e2e` with symlink-following `statSync` and no
depth guard.
* fix(push): size the claim-attempt budget from the drain count
With twelve drains, up to eleven peers can hold device heads, so a
four-attempt claim budget can run out while claimable rows remain and the
drain exits idle for a tick. Move the drain count into one module and derive
the attempt budget from it.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* style(push): keep the worker and store in repo formatting
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* style(push): drop the stray semicolon in the concurrency constant
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
Four drains, each holding one provider round trip of ~100 ms plus its
database statements, capped delivery near 30/s. Production inflow reached
35/s on 2026-09-30, so the backlog aged past the five-minute TTL and
notifications expired. Twelve drains lift the ceiling to roughly 90/s; the
pool is now six per instance, so the extra drains queue on connections
instead of starving the request path.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
Delivery collapsed on 2026-09-29 once send volume doubled: the worker, the
retention pruner and the request path share a two-connection pool, and the
database transaction rate pinned at ~140/s regardless of how many notifications
were delivered. Raise the pool to six so worker and pruner stop serialising on
one connection. The budget precondition stays satisfied (2 x 6 x 3 = 36 <= 64).
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
Third audit wave. The detector looked for production modules exporting three
or more symbols that no production file imports — only tests do. That shape is
the authoring gate's fourth question failing: a test needing a production seam
no caller needs belongs at the real boundary instead.
Most hits were detector false positives and were left alone; the scanner misses
re-export barrels and dynamic imports, so every module was re-verified with rg
before any edit. Where a private predicate's behavior was already covered
through the module's real entry point, the duplicate cases are gone and the
symbol is module-private again. Where it was NOT covered anywhere else, the test
stays — this audit removes tests, it does not author replacements.
Production code deleted where tests were its only callers: the superseded
`filesystem-directory-listing-limit` module, the unused
`format{Hourly,Daily,Adhoc}Version` helpers and their orphaned prerelease
identifiers, the dead `filterByAutomationListSearch*` family superseded by
`matchAutomationListSearchRowKeys`, and the dead
`getAiVaultResumeWorktreeTargetStatus` copy of the live workspace branch.
Also drops two call-shape source greps in `relay-sweep-schedule.test.ts` that
asserted `index.ts` spells `jitteredSweepIntervalMs(30_000)`; the jitter math
has a behavioral owner at the top of the same file. The structural census that
counts role-gated vs total `setInterval(` calls stays — an ungated sweep runs in
every cell, and nothing else can catch that.
Second audit wave, targeting three more junk patterns:
- assertion-free cases that run code and assert nothing, so they pass no
matter what the code does;
- inventory literals re-typed from a production declaration, where the only
way the assertion can fail is someone editing one of the two copies;
- export key-set and export-shape loops (`typeof x === 'function'` over every
export) that restate what TypeScript already enforces.
Yield is much smaller than wave 1 on purpose: the assertion-free scanner has
a high false-positive rate, because many flagged blocks assert through a
shared helper or their oracle is "this must not throw". Those were kept.
`mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is
de-exported — after the inventory comparison went away, nothing outside the
module read it.
Deletes 101 test files and trims 112 more, all matching documented junk
patterns: exact source/import/string greps, copied inventories and export
lists, duplicate invocations of a contract another test already owns,
typeof-shape checks TypeScript already enforces, and self-comparisons.
The largest group read a production `.ts` file and asserted on its text —
for example a TaskPage test that required the source to contain
`selectedRepos.find((r) => r.id === newIssueRepoId) ?? selectedRepos[0] ?? null`.
Any behavior-preserving rename broke it; no behavior change ever did.
Production-side follow-through: exports that only these tests imported are
de-exported or deleted, stale comments pointing at removed censuses are
dropped, and the reliability-gate registry, `cloud/package.json` test lists,
and orphaned source-reading helpers are updated so nothing references a
deleted file.
Two files kept their real coverage and lost only the census scaffolding:
`agent-status-producer-census.test.ts` now drives all five producers end to
end instead of grepping the source tree, and `config-toml-trust-stale-writes`
replaces an export-list parity check.
* perf(relay): release rejected first-frame connections [trade-off]
* fix(relay): bound the director's rejected and redirected first-frame closes too
The parent PR routed four cell-side first-frame rejections through
closeRelayWebSocket but left two raw socket.close() calls in the same
handler. A real-socket probe shows both still pin a connection unit for
ws's full 30s close timer when the peer ignores the close frame:
- 'invalid invite' is reachable by an unauthenticated peer with a
well-formed but bogus credential, so the exhaustion the parent PR
claims to prevent stayed reachable on the director;
- 'connect to assigned cell' is the happy path for every phone's first
director contact, so it is the highest-volume unbounded close here.
closeWithDrain gets the same treatment; host-session-registry already
closes the identical drain through the helper.
Also records that the bounded close is not a user-facing trade-off: the
close frame is written before the force-close timer can fire and TCP
delivers it ahead of the FIN, so an abandoned peer still reads code and
reason over a graceful close. The new regression asserts that, plus
exactly-once release across concurrent bursts and rejection racing the
peer's own disconnect (the ledger does not clamp at zero, so a double
release would surface as a negative count).
* refactor(relay): drop the unused closeWithDrain helper
`closeWithDrain` has no callers anywhere in the repo, and its `graceMs`
parameter promised a caller-supplied drain window that the body no longer
honours: routing it through `closeRelayWebSocket` force-terminates after 1s
regardless, so a future caller passing `graceMs: 30_000` would have had its
drain silently cut short while the signature still claimed otherwise.
The real drain path is `host-session-registry`, which sends the same
`resolve-director` drain and closes it there. Delete the dead duplicate
rather than bound a helper whose contract says "graceful".
Co-Authored-By: Claude <noreply@anthropic.com>
* docs(relay): call the bounded close's rejection delivery best effort, not guaranteed
The helper claimed the forced terminate costs an abandoned peer nothing it can
observe, and the test presented its fast-peer assertion as proof. Neither holds in
general: `ws` writes the close frame to the socket and `terminate()` destroys that
socket a second later, so under backpressure the frame — and any `relay-moved`
message queued ahead of it — can go unsent even to a peer that never stopped
reading.
Qualifies both comments to describe delivery as best effort and name the
backpressure case. The fast-peer assertion is valid and stays exactly as it was;
only its stated scope narrows, and the test is renamed to say which peer it speaks
for. What the bound actually buys — a stalled peer cannot hold admission — is now
stated on its own rather than resting on a delivery claim.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* perf(push): keep retention sweeps from overlapping
* perf(push): drain a saturated retention sweep instead of idling out the tick
The overlap guard on the shared prune timer removed a side effect the sweeper had
been relying on: overlap was the only thing that let a backlog exceed the
50-batch-per-call cap inside one 60-second tick. With the guard, a sweep that
spent its whole budget went idle for the rest of the interval, so a large backlog
drained far slower exactly when retention matters most.
The timer is now a chained setTimeout rather than an interval. `deleteInBatches`
reports whether it exhausted its batch budget, `prune()` returns
`{ deleted, saturated }`, and a saturated sweep is rescheduled immediately. The
connection gate hands a freed slot to the longest waiter, so one serial sweeper
looping back to back still parks a single statement ahead of a worker claim: claim
latency keeps the value the guard bought while the maximum drain rate returns to
what it was before. The loop is self-limiting and stops once the backlog clears.
A sweep that has not settled a full interval after it started now logs
`orca_push_prune_overdue` with its target. Admission waits have no timeout, so a
lost slot release could previously wedge retention permanently and silently.
Chaining makes the overlap guard structural, so there is no flag to scope. The two
single-DELETE sweeps state through `unbatchedSweep` that they have no batch budget
to exhaust, which keeps the immediate-resume path readable as delivery-only.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* perf: scope relay capacity checks to the requested cell
* fix: restore catalog entries required by the current CI baseline
* fix(i18n): make AI the recipient of diff notes
* fix(relay): drain a same-cap cell over 5 minutes, not 2
The 2026-09-23 c27 roll drained 2,145 hosts over the 2-minute window,
about 18 re-dials/s, while the director re-places roughly 8/s through
its single-slot sticky lane. The overflow queued behind slow
re-placements and timed out, so /v1/assign returned 503 fleet-wide for
about 5 minutes. 5 minutes is the cell's maximum pace window and keeps
the remaining 2,650-host cells near lane capacity. The transition wait
already outlasts a 5-minute window (17 min).
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* test(relay): pin the same-cap drain window contract at 5 minutes
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): keep the drain wait at lease plus the 5-minute window
The transition wait after a drain was set to the 15-minute migration
lease plus the pace window. Widening the window to 5 minutes without
moving the wait left 12 minutes for a migration that can hold for 15,
so a late-window migration would time the wave out into rollback.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
The selector membership check capped cell ids at c29, so enable failed
closed with "selector membership is invalid" once c30 went general.
Accept c1-c99 with no leading zero.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): lock only the target cell row, last and NOWAIT, in the rehome commit
The idle-rehome commit runs on the source cell. From an Asia cell each
statement is a cross-region round trip, and the transaction locked every
relay_cells row plus every runtime, capability and safety row before about
twenty more statements, so each Asia-source rehome held the whole fleet's
cell rows for ~3.6 s and every reconnect, renewal and placement queued or
timed out behind it.
The commit now reads the cell inventory and the runtime, capability and
safety tables unlocked, keeps the control and worker rows locked (now
NOWAIT), and takes one cell lock: the target row, in a single statement
that locks it NOWAIT, re-checks enabled, general admission and capacity,
and reserves the units, issued as the last statement before COMMIT. A
target that changed admission, filled up, or is locked by another writer
rolls the whole commit back and answers deferred (candidate-ineligible).
The hold is sampled under a site label, so cellInventoryHoldMsMax still
sees rehome holds and rehomeTargetRowHoldMsMax reports them apart.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* fix(relay): fail closed on the rehome target-row lock clause
The target-row statement now carries FOR UPDATE ... NOWAIT unless the
dialect is explicitly SQLite, so a wrapper that omits the optional dialect
can no longer run the reservation unlocked. Test wrappers and the fault
injection entry forward the dialect they wrap.
The latency test also probes the admission and region tables at every
round trip; only the target's admission row may be locked, and only
before COMMIT. The runbook notes that an Asia-sourced commit holds the
rehome control row for about 6.5 s, so a pause that fails once is retried.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* feat(relay): alert on relay cell table lock convoys
Adds a log-based metric and alert for cell-inventory lock holds of at
least 1,000 ms, and a Cloud SQL log metric and alert for relay-only lock
timeout cancels at 20 or more per minute. NOWAIT refusals are excluded:
background sweeps produce about 160 per minute even with rehome paused.
Replayed over 2026-09-20 14:00 to 2026-09-22 15:00 UTC: the hold filter
matches all 93 asia-east2 rehome holds plus 9 director holds, and every
one of the 88 cancel burst minutes overlaps an asia-east2 hold.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): page only on cell lock holds; director holds stay visible
Director holds of 1-2.5 s recur several times a day with rehoming paused,
and pausing rehome does not stop them. The paging hold policy now selects
role=cell samples only; a separate policy with no notification channel
keeps director holds visible. The burst documentation no longer claims no
burst happens while paused, and the runbook points a burst with no cell
hold at the director policy.
Replayed cell-only: 93 of 93 asia-east2 holds, 0 director holds over
2026-09-20 14:00 to 2026-09-22 15:00 UTC; 0 from then to 2026-09-23 07:30.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
The source cell runs the rehome commit. An Asia source pays a cross-ocean
round trip per statement while holding relay_cells row locks every cell
needs, which convoys the fleet database. Selection now drops source cells
outside the director's region before building the decision window, so the
incumbent_region filter shrinks while Asia cells stay valid targets. The
preview counts the same hosts as source-outside-director-region and the
poll summary reports skippedOffRegionSourceCells. Temporary stopgap.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
C30 was promoted to general on 2026-09-23 (selector generation 286). The
same-cap wave now rolls it as a general cell instead of handing it back
isolated, and the shadow gate reads its pool beside C27-C29. Follow-up to
#22375.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* fix(cloud): gate the Asia canary on its own cell's SQL failures, not the directors'
The production canary summed sqlFailuresDelta over every director and the
canary cell and required zero. Directors log a steady baseline of
relay_cells NOWAIT and lock-timeout refusals unrelated to the canary cell,
so a C30 canary failed most attempts on that noise. The canary now requires
zero SQL failures from the canary cell's own metrics and records the
director sum as directorSqlFailures without gating on it. Directors keep
every other rule (unavailable regions, fallbacks, pool waiting, transient
waiter and wait-time bounds). Staging keeps the combined zero rule and its
evidence shape unchanged.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* fix(cloud): gate the Asia canary's pool bounds on its own cell too
Directors also show a steady pool-wait baseline (waiting above zero and
waits over 50 ms in about 6 of every 60 minutes), so a five-minute canary
still failed about half the time on director pool pressure unrelated to
the canary cell. With gateDirectorDatabase off, the production canary now
applies databasePoolWaitingMax, databasePoolWaitersMax and
databasePoolWaitMsMax to the canary cell's metrics only and records the
director values under director-prefixed names. Directors still gate Asia
selections, region fallbacks and unavailable regions. The staging path
keeps its combined values, key order and validation order.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The topology workflow's Cloud SQL gate carried a hard-coded 400 for the
instance's tier default while the consumer contract records the value
measured on the live instance (SHOW max_connections = 500, 2026-09-16,
#21163). The gate compares the two and the first production plan run
(35815654836) failed silently on that mismatch before Terraform ran.
The verified default now lives beside the tier and version it is verified
for, as VERIFIED_DEFAULT_MAX_CONNECTIONS, so the contract and the workflow
are two independent records of the same measurement and the gate keeps
its cross-check. The test pins the new source and forbids a bare literal.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* feat(relay): declare Asia cell c30 at the c27 shape
Adds production-gce-c30 in asia-east2-a at the reviewed Asia shape (6,000
request units, 3,000/60 connection limits, 16-connection pool, disabled) and
the rehome trust the other Asia cells carry.
Every Asia enumeration now knows C30. The topology, admission, and director
tools treat it as its own reviewed wave so its plan and registration never
touch the live launch cells. C30 promotion requires C27 general and fresh
staging evidence. The topology and director validators now pin the committed
production pool of 16 instead of the stale 10, which had made the topology
workflow reject the committed launch cells.
* fix(relay): plan C30 at live images and prove it with its own canary
The shared URL map pulls every cell into the C30 topology plan, so the workflow
now plans each non-target cell at the image its live template serves, and the
validator names any change to a cell outside the wave. C30 promotion runs the
same five-minute production canary and automatic rollback C27 used, with the
load report proving the canary control was placed on C30, instead of relying
on staging evidence. C30 leaves the shadow gate's fleet pool list until it
serves, rollback rejects mixed partial sets, and a budget test pins the
mixed-Asia-pool refusal.
* fix(relay): pin C30 to the production director's live image digest
C30 promotion requires the director and C30 to report one digest, so C30
takes the director's sha256:4158d8a2 (read 2026-09-22). C27-C29 keep their
committed lines; every Asia check compares only the cells named in a run.
* fix(relay): read the committed cell map from a plan, not console
terraform console evaluates every output against state, and the Relay
deployments output indexes each cell's MIG, so it fails with Invalid index
while C30 is declared but not created. Read the map from a no-refresh,
unlocked plan over the same targets instead, and refuse empty overlay input.
* fix(relay): keep console readers working and C30 migration-only until promotion
relay_gce_cell_deployments indexed each cell's MIG, backend, and template,
so once C30 is declared but not applied every production terraform console
reader printed a warning to stdout and broke its jq parse. Wrap those six
lookups in try(..., null).
Same-cap listed C30 as general, so a rollback dispatch on a migration-only
C30 would restore it with activate and skip its canary. List it with the
migration-only cells until the promotion follow-up moves it.
* fix(push): bound the delivery claim and stop keeping finished batches
* fix(push): bound claim scans to the notification TTL and document the queue
* test(push): boot waits out a table writer for the queue indexes; queue stays correct without them
* fix(push): share the claim lock so a previous-revision claim cannot re-lease a delivery
During a deploy overlap the previous revision's claim scans under an exclusive
push-worker-claim lock and re-reads the row without checking lease_until, so it
could overwrite a lease this revision had just committed and send twice. The
new claim now takes the same key shared: new claimers never block each other,
and the previous claim waits until their leases commit before it scans.
Droppable one release after every worker runs this revision.
* fix(push): install one queue index at boot, not three
push_batches_leased_device indexed lease_until, so every lease, renew and finish
UPDATE lost heap-only eligibility and rewrote every index. push_batches_pending_due
was unused: on a synthetic 902k-row table the candidate scan plans onto the
existing (state, due_at) and expiry indexes with or without it. The per-device
pending index stays; the head check and the busy anti-join use it. Fewer
boot-time builds also shorten the SHARE lock the first boot takes on the table.
* test(push): pin the claim's TTL scan bound and the server's worker connection cap
Removing either guard left the suite green. The claim test captures every row
the candidate scan returns and plants one row that only the TTL term excludes;
the server test drives the real worker through createPushServer and fails when
the request-connection reservation is unwired (peak 4 instead of 2).
* fix(push): renew delivery leases outside the background connection cap
Renew shared the single background slot with claim retries and prune batches,
so a heartbeat could wait long enough for a lease to lapse and the delivery to
be re-leased mid-send. It is a keyed one-row UPDATE, so request traffic cannot
starve it on the ungated pool.
* docs(push): describe the shared claim lock for mixed-revision deploys
A same-cap wave applies with create_before_destroy. When run 35684694704 died
after the new template was made, the previous template stayed in Terraform state
as a deposed object, so every later plan for that cell carried its delete. The
plan validator only tolerated deposed deletes in same-cap-image mode, so the
recovery run 35698133226 was refused with "cell plan must change only the exact
instance template and MIG" and the cell was stranded.
Allow exactly one deposed delete at the cell's own template address in
same-cap-cell mode too, mirroring the one-deposed bound the convergence path
already applies to non-image modes. Report it as `obsoleteTemplates` rather than
in `changes`, the way the cell backend update is already split out, so the job's
`changes == 2` resume gate and `changes == 0` stranded roll keep reading the
template-and-MIG count.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
The capacity role the same-cap wave authenticates as, orcaRelayProductionCapacity,
has no compute.backendServices.update. Since #21860 added
`google_compute_backend_service.relay_gce_cell["${TARGET_CELL_ID}"]` to both of the
job's plan invocations, every wave has therefore created the new instance template,
modified the MIG, and then failed 403 on the backend, leaving the cell isolated with
its trust probe, admission restore, and shadow gate all skipped. Run 35684694704 on
production-gce-c7 is the first one that hit it in production.
Drop the backend target from both plans and restore the resume gate to exactly
`.changes == 2` (the template-and-MIG rollback-image drift) or a converged plan,
removing the backend-only resume apply #21865 added on top. A resume applies nothing
again, which is what a resume means.
The validator keeps its bound on a cell backend update, so it still reports one and
refuses anything wider, but a wave plan can no longer contain one. The drain timeout
from #21848 and the log_config from #21860 need a root apply by a principal that holds
the permission; granting the capacity role that permission is itself a root apply, so
it can follow as its own change rather than blocking every wave in the meantime.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): re-place hosts off a cell isolated for a roll
A roll isolates a cell by moving it out of the 'general' admission class; the
cell then refuses every attach with 4503. The director never noticed, because
the only liveness test it applies to a host's current cell reads
`relay_cell_runtime.ready` and the heartbeat, and an isolated cell keeps
heartbeating ready=1 for the whole drain. So every host on that cell was handed
its own dead cell, closed, and handed it back — 500-1,900 hosts looping for
13-16 minutes per cell roll, at ~6 dials each per minute, with no neighbour
absorbing anything.
The sticky lane now treats a live incumbent whose admission is 'migration-only'
— the state a roll's isolate step writes — the same way it treats a dead one:
it returns null, which means "fall through to placement". The placement lane
had the identical hole eleven lines further down, so it takes the same
predicate; without that second swap the sticky change is inert, because
placement would hand the pin straight back (a draining cell has more headroom
than anyone). An isolated incumbent skips the dead-cell fence branch: that
branch exists to prove an unreachable cell stopped serving a host, and this one
is reachable and enforces the epoch itself.
'existing-only' is deliberately untouched — those cells serve the hosts they
already hold, and only `assignmentStrandedOnUnservedCell` may release that pin.
A host with an open `relay_assignment_migrations` row keeps its pin too, so
this stays disjoint from the migration machinery.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* fix(relay): gate re-placement on a roll-isolation marker, not on admission
Review of the first commit found the predicate wrong. `migration-only` is an
admission class, not a drain signal: an Asia `--mode rollback`, an evacuation or
forward-recovery target awaiting a separate promote dispatch, a failed same-cap
wave's re-isolate, an abandoned migration retired on its target and a rehome
settlement all park loaded cells there durably, with no migration lease and no
open migration row. All five were indistinguishable from a roll's isolate, so
the first commit would have converted `operate-relay-asia-admission --mode
rollback` from a reversible admission flip into a mass move of ~4,000 hosts —
and, because `leastLoadedCell` treated region as a preference, into us-central1.
The signal is now an explicit stamp. `relay_cell_admission` gains a nullable
`roll_isolated_at`, added through the shared schema runner's catalog pre-check
so a migrated database takes no relation lock on boot and an un-migrated one
gets a catalog-only rewrite. The same-cap isolate step is its only writer, via a
new optional `rollIsolatedCells` on the selector apply; the same UPDATE that
writes the state clears the stamp whenever a cell leaves 'migration-only', so a
restore cannot leave one behind and a failed wave's re-isolate keeps the one it
has. Every other admission writer omits the field, so its cells stay unmarked
and their hosts stay pinned. Old directors ignore the field; old callers never
send it.
Region is now a constraint rather than a preference on this path only: a
re-placement must find a general, live cell with connection headroom in the
host's own region, or the pin is kept and one
`orca_relay_sticky_replacement_deferred` event is logged. Cross-region spill is
no longer reachable here.
The fence bypass is narrowed to a live incumbent. It was always a no-op for the
intended case, and for a stamped cell that stops heartbeating while still
holding sockets it reopened split-brain; that cell now takes the dead-cell path
unchanged.
Also: the hot-path admission reader no longer throws on an unrecognised state —
it sits on every sticky dial and the rule it feeds is "move the host", so an
unreadable row has to mean "don't". And the sticky lane reads the admission row
once for both the stranded rule and the stamp instead of twice.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* fix(relay): emit the re-placement events after the transaction commits
CodeRabbit on assignment-store.ts:1075. Both events were written where they are
decided, which is inside assignOnce's transaction. A reservation or lease write
failing after that point rolls the placement back, but a line already on stdout
cannot be rolled back with it — so the canary this PR asks an operator to read
would count re-placements that never happened, and a Postgres transaction retry
could leave a stale line behind as well.
The transaction now returns its events alongside the RelayAssignment and the
caller flushes them once it has resolved. Returning them rather than setting a
variable in the enclosing scope is what makes the retry case safe too: only the
attempt that committed can carry its events out. assign()'s signature is
unchanged; the extra shape lives entirely inside assignOnce.
orca_relay_sticky_replacement_deferred was moved the same way. It cost one more
push into the array that already existed, and it is decided inside the same
transaction, so leaving it behind would have been the odd case rather than the
cheap one.
The new test injects a failure on the first write after the decision, asserts no
event is emitted, and asserts the assignment is still on its original cell —
without that second assertion the absence would only prove the emit was early,
not that it would have been wrong. A control dial with nothing injected emits
exactly one event, so the case cannot pass on a broken harness. With the emit
put back inside the transaction, it fails.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* fix(relay): expire the roll stamp, correct the wire note, assert the stamp landed
Delta review findings B, D and E. A (the deferral path's cost) is deliberately
not implemented; it is now written up under Follow-ups in the PR body as
required before any Asia roll, because it cannot fire in a US canary.
B, which also closes C: the stamp was written, carried and never compared to
anything. A roll isolates and restores one cell inside ~15 minutes, so a stamp
older than two hours is not a roll in progress. It is a failed wave whose
failsafe re-isolated a possibly healthy cell and is waiting on an operator — the
postmortem in this tree records gaps of hours — or an orphan left by a director
rollback whose restore wrote 'general' without the clause that clears the stamp,
which the selector's 'keep' branch would then preserve until some later park
reactivated it. Both want the same answer and it is the pre-existing one: keep
the pin. One comparison against a value already on the row.
The bound takes the caller's `now` rather than reading the clock again, so one
assign reasons about one instant; the stamp's age is now a thing that decides
whether a host moves, and two clock reads could disagree across it.
D: the comment beside the new request field claimed an updated caller reaching
an older director "is simply ignored". The schema is .strict(), so it is a 400.
That fails closed — the isolate aborts before MUTATION_STARTED is set and
nothing is written — but it is a deploy ordering constraint, and it was
undocumented. The comment now says so and the PR body's rollout notes carry it.
E: nothing read the `rollIsolated` the script already prints, so an older script
against a newer director would silently produce today's behaviour and the canary
would read as "the fix did nothing" with no way to tell that from a wrong
premise. Both isolate steps now assert it, beside the generation they already
parse.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb