federation-release-recovery-scenarios, the federation test runtime and the
federation start-request builder are imported only by tests; name them
*.test-support.ts beside them, as worker-release.test-support.ts already does.
The branch's rewrite dropped main's verbatim phrases ("hand off", "handoff",
"handover", "give this to another agent", "another worktree", threaded
messages, worker_done/escalation waits, decision gates, reading or waiting on
terminals) — the only text a model sees when choosing this skill. Restored in
both the kernel frontmatter and the identical stub, still shorter than main's,
and pinned by a routing test.
worker-list --include-remote issued one federated-dispatch lookup and one
observation-fence capture per row, so a 40-row page prepared 46 statements
against a budget of 7. Both are now one statement per page/host group.
retry_of_dispatch_id, creator_role, endpoint_id, endpoint_incarnation,
attachment_kind and resource_id were written and never read back;
worker_terminal_resources already owns the endpoint/resource identity. Only
creator_dispatch_id (joined by task-store) and host_scope (read by the worker
liveness fallback) stay. v35 drops the columns and their two indexes.
Note: --retry-of still gates the retry transition, it just no longer stamps a
lineage column nothing reads.
The guide still documented queued_pending_turn and submission_observed, and the
send-submit repro gated on the removed submission_observed literal, so its
unsubmitted assertion passed vacuously. Both now use input_accepted then
turn_started.
check --json shipped raw mailbox rows, so agents read delivery plumbing
(pointer_*, sender_pane_key, read, sequence) as mailbox truth, and worker-show
shipped the snake_case dispatch row beside a camelCase worker. Both now go
through presenters.
requestRemoteAttachmentTerminalRelease carried a verbatim copy of the pre-table
ownership ladder and its own release_state list; both now come from
decideWorkerTerminalRelease and WORKER_TERMINAL_RELEASABLE_ROW_SQL.
A pre-rowid cursor resolved its order key through a subquery, so once a reset
deleted the anchor dispatch `rowid > NULL` excluded every row and the client
read a finished, empty inventory; it now expires the cursor.
A pinned filtered page also took its total from the snapshot's membership and
its counts from a live scan, so a pinned row leaving the filter left a total no
count could reach. Both now come from the pinned row set.
selector_not_found on an orchestration mutation already carried
orchestrationRequestId, so `data ?? selector` and the RuntimeClientError
short-circuit dropped the selector grammar from --json and its next steps from
the text message. Both recoveries now merge.
The four prompt-delivery warnings were built inside the text formatter, so
--json callers (every agent) saw none of them. The receipt now carries a
warnings array built from the same function, and the swallowed-Enter warning is
ordered ahead of the unsupported-observation arm so an agent provider always
gets the recovery command; a plain shell keeps its cannot-report-delivery text.
lifecycle_transition_receipts had no production reader: the append, the getter,
the delete triggers, the reset paths and the bounded recovery retention all fed
a table only tests read. The transition graph and its guards stay; v35 sheds the
table and rebuilds the two delete triggers that survive it.
Tests that read the ledger now assert the observable state change, and the
atomicity tests inject their failure on the last real projection instead.
The line printed the PTY verdict with the same word the fleet verdict uses, so a
live pane holding a dead agent read as "Liveness: live". It is now
"Terminal liveness", with "Agent liveness" printed from the fleet projection
when the receipt carries one.
A close that throws terminal_handle_stale on a host-certified exit was
classified permanent, so an exited remote worker could never reach released and
the recovery text told the agent to retry into the same stale handle. Both the
local and federated paths now treat 'nothing left to close' as the close
succeeding; a lost endpoint still parks as release_pending for recovery.
The kernel's stall exit fired on "not live", which includes every unverifiable
arm — all of which are absence — contradicting the safety floor and the recovery
table. It now names the positive signals (exited liveness, the worker's own
observation of exit, a final agent turn with no worker_done) and states that
unverifiable never authorizes stop, abandon, retry, or release. Also corrects the
projection.* field paths the worker-list row actually nests.
Settlement always requires an archive, but the archive is only written while
release_state is 'requested' — a state a stopped or abandoned worker never
reaches. Its owned pane was retained forever. An owner releasing a
process-proven-exited pane that never recorded release intent now settles with
archive_status 'unavailable'; the archive stays mandatory everywhere it is
still reachable, and user_owned/external/transferred stay retained.
A stale or unknown --terminal fell through to the direct mailbox and returned
ok:true with an empty inbox forever, so a worker read absence as "no mail yet"
while its Run delivery sat unread. Consuming checks now fail with
stable_pane_required and the run-use / --run next action; --peek and --all still
inspect.
worker-start --terminal already rejected coordinator self-adoption; manual
dispatch never compared --to against the caller, so --inject delivered the
worker preamble into the coordinator itself.
v34 early-returns at >= 34 and every index probe uses IF NOT EXISTS, so a
database stamped by the pre-fix build kept the NOT NULL mailbox_handle with no
DEFAULT and the old index predicates. v35 re-applies both against the stored
SQL, and the skew lists now gate the new invariants.
An operator close settles the worker as failed, so the fleet projection fell
through to missing_status and reported a proven-dead worker as absence in the
same receipt that carried its exit. A recorded process exit is now the verdict,
and a proven exit with no worker outcome routes to worker-read instead of the
worker-show self-loop. `unverifiable` still never authorizes stop or abandon.
`isReapableRelayHusk` required `childCount === 0`, where `childCount` came from
`pgrep -P <relay> | grep -c .`. But the relay forks service children of its own,
and `relay-ai-vault-service.js` never exits once spawned. Any relay that had
served a single AI Vault request therefore reported a non-zero child count
forever, so the sweep answered `retained-live-work` for a superseded,
disconnected relay holding no user work at all — and its version directory
stayed pinned against GC by its own live socket.
The probe now censuses each direct child instead of counting them, and the reap
gate reads the count of children it could *not* positively identify as relay
infrastructure. The asymmetry is the safety argument
(docs/reference/ssh-execution-boundary.md): subtracting a child we can name is
positive knowledge, assuming about one we cannot is not. An unrecognised argv,
an argv `ps` would not print, and a host without `pgrep` all keep the relay
unreapable. `reapEmptyRelayHuskCommand` re-runs the same census on the host
immediately before signalling.
Fixes#13614
The shim parses its own argv, so a shell-emptied --retry-request parsed as
`true` and fell through to undefined, minting a fresh mutation identity for
orchestration send/check/ask over SSH (#15180).
The PTY-exit path failed the dispatch out from under an in-flight worker-stop,
so a stop that worked returned dispatch_inactive and left the worker reading
failed/process_exited. An exit that lands while the worker is `stopping` is
that stop's outcome, so settle it through the stop path.
The fleet liveness projection reads getStatusSnapshot(), whose sole builder
dropped the observation clock, so every worker fell back to the delivery clock
and an hour-stale replayed agent read live.
A WSL pane's PTY is local (connectionId null), so every WSL hook status was
filtered out and worker-read always fell back to screen scraping. Pass the same
distro expression the headless terminal state already uses.
Also deletes the dead v1 transcript_pin archive path (nothing has written
version 1, and its reader read a possibly-remote path from the local
filesystem) along with the endOffset thread it was the only caller of, and
adds Archived:/Liveness:/Worker: lines so a released archive read no longer
prints identically to a live one.
- close an exited remote worker's terminal before labelling the release
closed_exited_terminal; a host-certified exit keeps the verdict when the
kill stops nothing
- replace the structured-read and fleet-snapshot capability probes with the
optimistic call plus method_not_found, so hosts that serve
federationReadOutput without advertising it stop downgrading to a scrape
- drop forceProbe so an unchanged peer at an unchanged epoch probes once
- distinct fleet reasons for home-side budget exhaustion and peer_changed;
keep a host-supplied unverifiable reason and lastObservedAt for exited
- delete three compile-time-true self-capability checks, the advertised but
never-read federation-release capability, and decode the pull page with zod
The harness is test-only code sitting among the rpc/methods production modules;
*.test-support.ts is the repo's existing marker for that. The five folds the
review asked for are not possible: every merge target is already at 292-300
effective lines against the 300 cap.
Between worker_done and release the pane still holds a resumable provider
session, but listLegacyWorkerTerminalRecoveryRows selected live worker states
only, so nothing fenced it: reopening the workspace after a restart re-ran
`codex resume <session>` on a finished worker.
Settled workers whose terminal resource is still owned and neither released nor
retained now join the recovery rows, the plan marks those panes settled so they
never compete for adoption, and any fence the plan no longer claims is lifted —
on release, retain, user takeover and dispatch prune. An unreadable plan fails
closed and lifts nothing.
Ported from PR #17651, adapted to this branch's worker_terminal_resources
ownership model.
worker-show spread the raw worker row beside its parsed copies, so a reader got
residual_resources (a JSON string) next to residualResources (an array), plus
host_scope as JSON-inside-JSON and two authority hashes with no consumer. Parse
once, emit camelCase once, and withhold the hashes.
worker-show also published only PTY liveness, so an agent that died at a trust
prompt read live there while worker-list called it unverifiable -- and
worker-list's nextAction pointed back at worker-show. Both now publish the same
fleet projection.
worker-list's projection.resource restated fields the row already carried, and
the unconfirmed-stop sentence doubled a terminator on an already-punctuated
reason.
Adds regression coverage that fails on ab6ee61aac: staging-failure and
DB-error exits leaving an active watermark, the restart scan skipping
pointer-pending and dispatch: mailboxes, the caller-edge walk over the
lifecycle graph, stopping -> failed, Task reopen/overturn, the v34
downgrade insert, the pointer-enter index predicate, the v33 skew column,
and in-memory/sqlite parity for pointer reservations and batch exclusion.
Fleet liveness measured staleness against status.receivedAt, the delivery
clock a relay reconnect restamps, so an hour-stale agent read live after every
reconnect; it now uses evidenceObservedAt when the host supplies it.
worker-attention-context kept a second, divergent copy of that projection which
ignored host scope entirely and could not say exited. It calls the shared
projectLiveness now, with hostScope carried on WorkerAttentionFacts.
refreshOrchestrationFleetLivenessAttention re-derived categories from liveness
alone and dropped the unverifiable an unproven outcome contributed; it re-runs
the one attention projection instead, using the outcome now exposed on the
worker. The byte-identical active-sibling predicate is one shared fragment.
- compare the recorded terminal binding on every prompt replay, not only the
--wait-submit ones, so a replayed receipt cannot claim observation:supported
for an incarnation that is gone
- normalise `terminal` to a pane key in replayStableCallerParams so a re-minted
handle replays instead of failing request_mismatch
- warn (exit 0) when a supported send never reached turn_started, naming
--retry-request <id> --wait-submit <seconds>
- collapse the write-only stages to input_accepted | turn_started
- move the PTY-keyed correlation state into AgentPromptRequestCorrelation
(pty-keyed arrays, not NUL-joined string keys); the abandoned scan in
lifecycle claim allocation now continues instead of returning
A bare repo id passed to --worktree returned a content-free `selector_not_found`
with no value and no grammar, so a caller could not tell what was wrong. Shape it
at the CLI boundary — the only layer that still knows what was typed — in the same
selector/suggestions/nextSteps shape as an unknown-flag error, and point `--from`
on `orchestration check` at `--terminal`, which edit distance cannot reach.
The watermark was set before stageMailboxPointerEnter, and neither the
staging-failure nor the DB-error exit cleared it, so a mailbox whose claim
was stolen parked every later delivery forever.
Also restores the restart scan's pointer-pending union and dispatch: handles,
aligns the pointer-enter index predicate with its query, adds
messages.pointer_enter_pending to the v33 skew columns, makes v34's
mailbox_handle downgrade-safe, and completes the lifecycle graph.
A coordinator adopted as its own worker answers its own dispatch preamble; the
--terminal target is now rejected by handle and by resolved pane key with a
typed terminal_is_coordinator error naming the next action.
The dispatch address is durable but never interrupts a worker, so a preamble
that only listed `check --terminal` produced workers that never read one
coordinator follow-up. Name the checkpoints and the pre-worker_done read.
D1: name the two liveness layers (worker-list projection.liveness is the fleet
verdict, worker-show observation.status is PTY-only) and give the supervised
loop a bounded stall procedure instead of an unbounded wait.
D2: put worker-list, attention, requiresAction and nextAction in the loop and in
completion accounting.
D4: document request-show / --retry-request / terminal send --wait-submit.
D5: check names its caller with --terminal, never --from.
D7: give a dispatched worker a concrete follow-up read cadence.
D10: document the real folder-workspace route (project setup-existing-folder).
worker-start --spec is now the canonical loop's default.
Guidance pins are contracts via squash() instead of reflow-fragile prose.
The SQL CASE in worker-terminal-inventory-counts duplicated
deriveWorkerTerminalListState and diverged on an owned, unreleased resource
with no worker_dispatches row: SQL returned NULL where TS returns 'active', so
worker-list --terminal-state active dropped a row it labelled active. The SQL
copy is gone; filtering and counting now read the one TS projection.
worker-list also ordered by COALESCE(w.created_at, d.created_at) while the
fence pinned d.rowid, so a worker registering between pages re-emitted its row.
Order and fence now share d.rowid, and a pinned filtered cursor counts over its
own row extent instead of a live scan.
A transport timeout is not evidence that another runtime answered, so the
preflight-attested host still owns the durable pending receipt. Strip the
request ID only when a different runtime handled the send, and drop the
false 'update Orca on the execution host' advice from that path.
worker-release's dead-process shortcut called settleDeadWorkerTerminalRelease
without requireArchive, and the guard only excluded ownership_state='released',
so a user takeover, an external terminal, or a transferred resource was marked
released and its output archive was lost.
Both release guards now read one (ownership_state, release_state) -> action
table, and settlement always requires the durable archive.
The merge kept only the branch's tombstone reseed and dropped main's guard, so an
explicitly promised surface raced a fallback terminal after the inventory gate.
Fold both gated callbacks into one seeding helper.
The global relay_cells FOR UPDATE lock made successful retries a
steady-state rate: fleet-wide p50 430 / p90 924 / p99 1320 / max 1504
per five minutes over the last 24 h, 55% of windows over the 300 bar,
only 22% of 15-minute gates clean. Three read-only dry-runs on
2026-09-04 froze on it, blocking the same-cap roll that carries #18521
and the beginProof crash guard to the 23 cells. 2000 clears every
measured healthy gate; the exhausted-retry, director concurrency, and
pool bars keep the incident discriminator role.
* perf(ipc): build the filesystem allowed-root list once per authorization
* perf(ipc): keep the allowed-root snapshot lazy so granted external paths build nothing
Hoisting getAllowedRoots to the top of resolveAuthorizedPath made every read of a
path covered by an external grant build the full root list, where main built none
(the grant answered before isPathAllowed reached the roots). Build on first use
instead: still one build per authorization, zero when a grant already answers.
* test(ipc): skip the allowed-root symlink escapes on Windows
Unprivileged Windows cannot create symlinks (EPERM), so both cases failed in
setup instead of exercising the escape check.