setSleepingAgentAutomaticResumeBlocked had no production caller: the stamp only
ran at startup and the sweep only after release/retain/takeover, so a worker
settled by worker_done, stop, abandon or its own process exit left its pane
unfenced and reopening it in the same session respawned the agent. Every
settlement now sweeps, and the sweep pushes each fence change to the live
renderer over agentStatus:legacyWorkerTerminalResumeFence.
worker-observation's dispatch receipt publishes it as retryOfDispatchId, and
--retry-of validates the prior attempt and gates failed -> dispatched. Only the
index stays dropped: nothing queries by the column. The other five v31 columns
remain deleted.
Nothing ever stamps AgentStatusIpcPayload.processIncarnation/endpointId/
endpointIncarnation, so matchesOptionalIdentity could only reject a status the
Dispatch had already claimed. The reminted-pane branch still requires the
terminal handle.
WorkerTerminalTailArchive.sourceExact/contentComplete were written as constant
false and read with ?? false; derive them at the read boundary instead.
federation-release-recovery-scenarios, the federation test runtime and the
federation start-request builder are imported only by tests; name them
*.test-support.ts beside them, as worker-release.test-support.ts already does.
The branch's rewrite dropped main's verbatim phrases ("hand off", "handoff",
"handover", "give this to another agent", "another worktree", threaded
messages, worker_done/escalation waits, decision gates, reading or waiting on
terminals) — the only text a model sees when choosing this skill. Restored in
both the kernel frontmatter and the identical stub, still shorter than main's,
and pinned by a routing test.
worker-list --include-remote issued one federated-dispatch lookup and one
observation-fence capture per row, so a 40-row page prepared 46 statements
against a budget of 7. Both are now one statement per page/host group.
retry_of_dispatch_id, creator_role, endpoint_id, endpoint_incarnation,
attachment_kind and resource_id were written and never read back;
worker_terminal_resources already owns the endpoint/resource identity. Only
creator_dispatch_id (joined by task-store) and host_scope (read by the worker
liveness fallback) stay. v35 drops the columns and their two indexes.
Note: --retry-of still gates the retry transition, it just no longer stamps a
lineage column nothing reads.
The guide still documented queued_pending_turn and submission_observed, and the
send-submit repro gated on the removed submission_observed literal, so its
unsubmitted assertion passed vacuously. Both now use input_accepted then
turn_started.
check --json shipped raw mailbox rows, so agents read delivery plumbing
(pointer_*, sender_pane_key, read, sequence) as mailbox truth, and worker-show
shipped the snake_case dispatch row beside a camelCase worker. Both now go
through presenters.
requestRemoteAttachmentTerminalRelease carried a verbatim copy of the pre-table
ownership ladder and its own release_state list; both now come from
decideWorkerTerminalRelease and WORKER_TERMINAL_RELEASABLE_ROW_SQL.
A pre-rowid cursor resolved its order key through a subquery, so once a reset
deleted the anchor dispatch `rowid > NULL` excluded every row and the client
read a finished, empty inventory; it now expires the cursor.
A pinned filtered page also took its total from the snapshot's membership and
its counts from a live scan, so a pinned row leaving the filter left a total no
count could reach. Both now come from the pinned row set.
selector_not_found on an orchestration mutation already carried
orchestrationRequestId, so `data ?? selector` and the RuntimeClientError
short-circuit dropped the selector grammar from --json and its next steps from
the text message. Both recoveries now merge.
The four prompt-delivery warnings were built inside the text formatter, so
--json callers (every agent) saw none of them. The receipt now carries a
warnings array built from the same function, and the swallowed-Enter warning is
ordered ahead of the unsupported-observation arm so an agent provider always
gets the recovery command; a plain shell keeps its cannot-report-delivery text.
lifecycle_transition_receipts had no production reader: the append, the getter,
the delete triggers, the reset paths and the bounded recovery retention all fed
a table only tests read. The transition graph and its guards stay; v35 sheds the
table and rebuilds the two delete triggers that survive it.
Tests that read the ledger now assert the observable state change, and the
atomicity tests inject their failure on the last real projection instead.
The line printed the PTY verdict with the same word the fleet verdict uses, so a
live pane holding a dead agent read as "Liveness: live". It is now
"Terminal liveness", with "Agent liveness" printed from the fleet projection
when the receipt carries one.
A close that throws terminal_handle_stale on a host-certified exit was
classified permanent, so an exited remote worker could never reach released and
the recovery text told the agent to retry into the same stale handle. Both the
local and federated paths now treat 'nothing left to close' as the close
succeeding; a lost endpoint still parks as release_pending for recovery.
The kernel's stall exit fired on "not live", which includes every unverifiable
arm — all of which are absence — contradicting the safety floor and the recovery
table. It now names the positive signals (exited liveness, the worker's own
observation of exit, a final agent turn with no worker_done) and states that
unverifiable never authorizes stop, abandon, retry, or release. Also corrects the
projection.* field paths the worker-list row actually nests.
Settlement always requires an archive, but the archive is only written while
release_state is 'requested' — a state a stopped or abandoned worker never
reaches. Its owned pane was retained forever. An owner releasing a
process-proven-exited pane that never recorded release intent now settles with
archive_status 'unavailable'; the archive stays mandatory everywhere it is
still reachable, and user_owned/external/transferred stay retained.
A stale or unknown --terminal fell through to the direct mailbox and returned
ok:true with an empty inbox forever, so a worker read absence as "no mail yet"
while its Run delivery sat unread. Consuming checks now fail with
stable_pane_required and the run-use / --run next action; --peek and --all still
inspect.
worker-start --terminal already rejected coordinator self-adoption; manual
dispatch never compared --to against the caller, so --inject delivered the
worker preamble into the coordinator itself.
v34 early-returns at >= 34 and every index probe uses IF NOT EXISTS, so a
database stamped by the pre-fix build kept the NOT NULL mailbox_handle with no
DEFAULT and the old index predicates. v35 re-applies both against the stored
SQL, and the skew lists now gate the new invariants.
An operator close settles the worker as failed, so the fleet projection fell
through to missing_status and reported a proven-dead worker as absence in the
same receipt that carried its exit. A recorded process exit is now the verdict,
and a proven exit with no worker outcome routes to worker-read instead of the
worker-show self-loop. `unverifiable` still never authorizes stop or abandon.
The shim parses its own argv, so a shell-emptied --retry-request parsed as
`true` and fell through to undefined, minting a fresh mutation identity for
orchestration send/check/ask over SSH (#15180).
The PTY-exit path failed the dispatch out from under an in-flight worker-stop,
so a stop that worked returned dispatch_inactive and left the worker reading
failed/process_exited. An exit that lands while the worker is `stopping` is
that stop's outcome, so settle it through the stop path.
The fleet liveness projection reads getStatusSnapshot(), whose sole builder
dropped the observation clock, so every worker fell back to the delivery clock
and an hour-stale replayed agent read live.
A WSL pane's PTY is local (connectionId null), so every WSL hook status was
filtered out and worker-read always fell back to screen scraping. Pass the same
distro expression the headless terminal state already uses.
Also deletes the dead v1 transcript_pin archive path (nothing has written
version 1, and its reader read a possibly-remote path from the local
filesystem) along with the endOffset thread it was the only caller of, and
adds Archived:/Liveness:/Worker: lines so a released archive read no longer
prints identically to a live one.
- close an exited remote worker's terminal before labelling the release
closed_exited_terminal; a host-certified exit keeps the verdict when the
kill stops nothing
- replace the structured-read and fleet-snapshot capability probes with the
optimistic call plus method_not_found, so hosts that serve
federationReadOutput without advertising it stop downgrading to a scrape
- drop forceProbe so an unchanged peer at an unchanged epoch probes once
- distinct fleet reasons for home-side budget exhaustion and peer_changed;
keep a host-supplied unverifiable reason and lastObservedAt for exited
- delete three compile-time-true self-capability checks, the advertised but
never-read federation-release capability, and decode the pull page with zod
The harness is test-only code sitting among the rpc/methods production modules;
*.test-support.ts is the repo's existing marker for that. The five folds the
review asked for are not possible: every merge target is already at 292-300
effective lines against the 300 cap.
Between worker_done and release the pane still holds a resumable provider
session, but listLegacyWorkerTerminalRecoveryRows selected live worker states
only, so nothing fenced it: reopening the workspace after a restart re-ran
`codex resume <session>` on a finished worker.
Settled workers whose terminal resource is still owned and neither released nor
retained now join the recovery rows, the plan marks those panes settled so they
never compete for adoption, and any fence the plan no longer claims is lifted —
on release, retain, user takeover and dispatch prune. An unreadable plan fails
closed and lifts nothing.
Ported from PR #17651, adapted to this branch's worker_terminal_resources
ownership model.
worker-show spread the raw worker row beside its parsed copies, so a reader got
residual_resources (a JSON string) next to residualResources (an array), plus
host_scope as JSON-inside-JSON and two authority hashes with no consumer. Parse
once, emit camelCase once, and withhold the hashes.
worker-show also published only PTY liveness, so an agent that died at a trust
prompt read live there while worker-list called it unverifiable -- and
worker-list's nextAction pointed back at worker-show. Both now publish the same
fleet projection.
worker-list's projection.resource restated fields the row already carried, and
the unconfirmed-stop sentence doubled a terminator on an already-punctuated
reason.
Adds regression coverage that fails on ab6ee61aac: staging-failure and
DB-error exits leaving an active watermark, the restart scan skipping
pointer-pending and dispatch: mailboxes, the caller-edge walk over the
lifecycle graph, stopping -> failed, Task reopen/overturn, the v34
downgrade insert, the pointer-enter index predicate, the v33 skew column,
and in-memory/sqlite parity for pointer reservations and batch exclusion.
Fleet liveness measured staleness against status.receivedAt, the delivery
clock a relay reconnect restamps, so an hour-stale agent read live after every
reconnect; it now uses evidenceObservedAt when the host supplies it.
worker-attention-context kept a second, divergent copy of that projection which
ignored host scope entirely and could not say exited. It calls the shared
projectLiveness now, with hostScope carried on WorkerAttentionFacts.
refreshOrchestrationFleetLivenessAttention re-derived categories from liveness
alone and dropped the unverifiable an unproven outcome contributed; it re-runs
the one attention projection instead, using the outcome now exposed on the
worker. The byte-identical active-sibling predicate is one shared fragment.
- compare the recorded terminal binding on every prompt replay, not only the
--wait-submit ones, so a replayed receipt cannot claim observation:supported
for an incarnation that is gone
- normalise `terminal` to a pane key in replayStableCallerParams so a re-minted
handle replays instead of failing request_mismatch
- warn (exit 0) when a supported send never reached turn_started, naming
--retry-request <id> --wait-submit <seconds>
- collapse the write-only stages to input_accepted | turn_started
- move the PTY-keyed correlation state into AgentPromptRequestCorrelation
(pty-keyed arrays, not NUL-joined string keys); the abandoned scan in
lifecycle claim allocation now continues instead of returning
A bare repo id passed to --worktree returned a content-free `selector_not_found`
with no value and no grammar, so a caller could not tell what was wrong. Shape it
at the CLI boundary — the only layer that still knows what was typed — in the same
selector/suggestions/nextSteps shape as an unknown-flag error, and point `--from`
on `orchestration check` at `--terminal`, which edit distance cannot reach.