Commit Graph
10202 Commits
Author SHA1 Message Date
Jinwoo-H 0df053b20b refactor(rpc): move federation test harnesses out of the production directory
federation-release-recovery-scenarios, the federation test runtime and the
federation start-request builder are imported only by tests; name them
*.test-support.ts beside them, as worker-release.test-support.ts already does.
2026-09-04 03:24:03 -04:00
Jinwoo-H 5ab25102e5 docs(orchestration): restore the routing triggers to the skill description
The branch's rewrite dropped main's verbatim phrases ("hand off", "handoff",
"handover", "give this to another agent", "another worktree", threaded
messages, worker_done/escalation waits, decision gates, reading or waiting on
terminals) — the only text a model sees when choosing this skill. Restored in
both the kernel frontmatter and the identical stub, still shorter than main's,
and pinned by a routing test.
2026-09-04 03:23:41 -04:00
Jinwoo-H 5e1f38120c perf(orchestration): batch the federated fleet lookups per page
worker-list --include-remote issued one federated-dispatch lookup and one
observation-fence capture per row, so a 40-row page prepared 46 statements
against a budget of 7. Both are now one statement per page/host group.
2026-09-04 03:23:06 -04:00
Jinwoo-H aebe987685 refactor(orchestration): drop the v31 dispatch identity columns no reader consumes
retry_of_dispatch_id, creator_role, endpoint_id, endpoint_incarnation,
attachment_kind and resource_id were written and never read back;
worker_terminal_resources already owns the endpoint/resource identity. Only
creator_dispatch_id (joined by task-store) and host_scope (read by the worker
liveness fallback) stay. v35 drops the columns and their two indexes.

Note: --retry-of still gates the retry transition, it just no longer stamps a
lineage column nothing reads.
2026-09-04 03:22:35 -04:00
Jinwoo-H e484d6a1a7 docs(cli): name the prompt stages the runtime actually reports
The guide still documented queued_pending_turn and submission_observed, and the
send-submit repro gated on the removed submission_observed literal, so its
unsubmitted assertion passed vacuously. Both now use input_accepted then
turn_started.
2026-09-04 03:22:02 -04:00
Jinwoo-H 094edf1976 refactor(orchestration): present check messages and the dispatch row as receipts
check --json shipped raw mailbox rows, so agents read delivery plumbing
(pointer_*, sender_pane_key, read, sequence) as mailbox truth, and worker-show
shipped the snake_case dispatch row beside a camelCase worker. Both now go
through presenters.
2026-09-04 03:21:29 -04:00
Jinwoo-H 3aadb6ea3b refactor(orchestration): read the federated release guard from the one table
requestRemoteAttachmentTerminalRelease carried a verbatim copy of the pre-table
ownership ladder and its own release_state list; both now come from
decideWorkerTerminalRelease and WORKER_TERMINAL_RELEASABLE_ROW_SQL.
2026-09-04 03:19:32 -04:00
Jinwoo-H 1d5bc6c893 fix(orchestration): make worker-list paging report what the cursor can reach
A pre-rowid cursor resolved its order key through a subquery, so once a reset
deleted the anchor dispatch `rowid > NULL` excluded every row and the client
read a finished, empty inventory; it now expires the cursor.

A pinned filtered page also took its total from the snapshot's membership and
its counts from a live scan, so a pinned row leaving the filter left a total no
count could reach. Both now come from the pinned row set.
2026-09-04 03:18:15 -04:00
Jinwoo-H 6343053732 fix(cli): merge worktree-selector recovery into an error that already has data
selector_not_found on an orchestration mutation already carried
orchestrationRequestId, so `data ?? selector` and the RuntimeClientError
short-circuit dropped the selector grammar from --json and its next steps from
the text message. Both recoveries now merge.
2026-09-04 03:17:40 -04:00
Jinwoo-H 31c3f74f79 fix(cli): publish terminal send delivery warnings in --json
The four prompt-delivery warnings were built inside the text formatter, so
--json callers (every agent) saw none of them. The receipt now carries a
warnings array built from the same function, and the swallowed-Enter warning is
ordered ahead of the unsupported-observation arm so an agent provider always
gets the recovery command; a plain shell keeps its cannot-report-delivery text.
2026-09-04 03:16:54 -04:00
Jinwoo-H 63b919bde2 refactor(orchestration): delete the write-only lifecycle transition ledger
lifecycle_transition_receipts had no production reader: the append, the getter,
the delete triggers, the reset paths and the bounded recovery retention all fed
a table only tests read. The transition graph and its guards stay; v35 sheds the
table and rebuilds the two delete triggers that survive it.

Tests that read the ledger now assert the observable state change, and the
atomicity tests inject their failure on the last real projection instead.
2026-09-04 03:16:25 -04:00
Jinwoo-H b49bcd4307 fix(cli): label the worker-read PTY verdict as terminal liveness
The line printed the PTY verdict with the same word the fleet verdict uses, so a
live pane holding a dead agent read as "Liveness: live". It is now
"Terminal liveness", with "Agent liveness" printed from the fleet projection
when the receipt carries one.
2026-09-04 03:15:26 -04:00
Jinwoo-H 5b1705806a fix(orchestration): let an already-gone terminal close finish a release
A close that throws terminal_handle_stale on a host-certified exit was
classified permanent, so an exited remote worker could never reach released and
the recovery text told the agent to retry into the same stale handle. Both the
local and federated paths now treat 'nothing left to close' as the close
succeeding; a lost endpoint still parks as release_pending for recovery.
2026-09-04 03:14:52 -04:00
Jinwoo-H c78f40fdd0 docs(orchestration): require positive evidence of exit before ending a wait
The kernel's stall exit fired on "not live", which includes every unverifiable
arm — all of which are absence — contradicting the safety floor and the recovery
table. It now names the positive signals (exited liveness, the worker's own
observation of exit, a final agent turn with no worker_done) and states that
unverifiable never authorizes stop, abandon, retry, or release. Also corrects the
projection.* field paths the worker-list row actually nests.
2026-09-04 03:14:51 -04:00
Jinwoo-H 797c070baa fix(orchestration): give an owned pane on a settled worker a way out of retained
Settlement always requires an archive, but the archive is only written while
release_state is 'requested' — a state a stopped or abandoned worker never
reaches. Its owned pane was retained forever. An owner releasing a
process-proven-exited pane that never recorded release intent now settles with
archive_status 'unavailable'; the archive stays mandatory everywhere it is
still reachable, and user_owned/external/transferred stay retained.
2026-09-04 03:12:05 -04:00
Jinwoo-H 1ed068d304 fix(orchestration): refuse a consuming check on a handle with no live pane
A stale or unknown --terminal fell through to the direct mailbox and returned
ok:true with an empty inbox forever, so a worker read absence as "no mail yet"
while its Run delivery sat unread. Consuming checks now fail with
stable_pane_required and the run-use / --run next action; --peek and --all still
inspect.
2026-09-04 03:10:58 -04:00
Jinwoo-H 3e3b0617d1 fix(orchestration): refuse a dispatch aimed at the caller's own pane
worker-start --terminal already rejected coordinator self-adoption; manual
dispatch never compared --to against the caller, so --inject delivered the
worker preamble into the coordinator itself.
2026-09-04 03:08:50 -04:00
Jinwoo-H d764e4b791 refactor(orchestration): drop the dead remote-transcript endOffset window
Nothing has produced an endOffset since the v1 pin archive reader was removed,
so both remote readers carried an unreachable bound.
2026-09-04 03:08:49 -04:00
Jinwoo-H c1cbdc31b0 fix(orchestration): repair v34 databases the pre-fix build already stamped
v34 early-returns at >= 34 and every index probe uses IF NOT EXISTS, so a
database stamped by the pre-fix build kept the NOT NULL mailbox_handle with no
DEFAULT and the old index predicates. v35 re-applies both against the stored
SQL, and the skew lists now gate the new invariants.
2026-09-04 03:08:18 -04:00
Jinwoo-H 5d19a53367 fix(orchestration): let a certified exit outrank the worker's settled state
An operator close settles the worker as failed, so the fleet projection fell
through to missing_status and reported a proven-dead worker as absence in the
same receipt that carried its exit. A recorded process exit is now the verdict,
and a proven exit with no worker outcome routes to worker-read instead of the
worker-show self-loop. `unverifiable` still never authorizes stop or abandon.
2026-09-04 03:07:23 -04:00
Neil 561a94038c fix(ssh): stop the daemon's own services from blocking the superseded-relay reap (#18586)
`isReapableRelayHusk` required `childCount === 0`, where `childCount` came from
`pgrep -P <relay> | grep -c .`. But the relay forks service children of its own,
and `relay-ai-vault-service.js` never exits once spawned. Any relay that had
served a single AI Vault request therefore reported a non-zero child count
forever, so the sweep answered `retained-live-work` for a superseded,
disconnected relay holding no user work at all — and its version directory
stayed pinned against GC by its own live socket.

The probe now censuses each direct child instead of counting them, and the reap
gate reads the count of children it could *not* positively identify as relay
infrastructure. The asymmetry is the safety argument
(docs/reference/ssh-execution-boundary.md): subtracting a child we can name is
positive knowledge, assuming about one we cannot is not. An unrecognised argv,
an argv `ps` would not print, and a host without `pgrep` all keep the relay
unreapable. `reapEmptyRelayHuskCommand` re-runs the same census on the host
immediately before signalling.

Fixes #13614
2026-09-04 00:06:09 -07:00
Jinwoo-H 668a3ec278 fix(ssh): validate --retry-request in the relay CLI shim
The shim parses its own argv, so a shell-emptied --retry-request parsed as
`true` and fell through to undefined, minting a fresh mutation identity for
orchestration send/check/ask over SSH (#15180).
2026-09-04 03:04:46 -04:00
Jinwoo-H 8f9a30cfc3 fix(orchestration): treat a stop-time process exit as the stop succeeding
The PTY-exit path failed the dispatch out from under an in-flight worker-stop,
so a stop that worked returned dispatch_inactive and left the worker reading
failed/process_exited. An exit that lands while the worker is `stopping` is
that stop's outcome, so settle it through the stop path.
2026-09-04 03:04:23 -04:00
Jinwoo-H f067ab45e4 fix(agent-hooks): carry evidenceObservedAt into the status snapshot
The fleet liveness projection reads getStatusSnapshot(), whose sole builder
dropped the observation clock, so every worker fell back to the delivery clock
and an hour-stale replayed agent read live.
2026-09-04 03:00:28 -04:00
Jinwoo-H e426592e5d refactor(rpc): group orchestration method files by verb cluster 2026-09-04 02:45:04 -04:00
Jinwoo-H caa3a72988 Merge fix-send-federation-transcript into integrate-fixes 2026-09-04 02:35:09 -04:00
Jinwoo-H 58adae375e fix(transcript): wire the attested WSL distro into exact worker session selection
A WSL pane's PTY is local (connectionId null), so every WSL hook status was
filtered out and worker-read always fell back to screen scraping. Pass the same
distro expression the headless terminal state already uses.

Also deletes the dead v1 transcript_pin archive path (nothing has written
version 1, and its reader read a possibly-remote path from the local
filesystem) along with the endOffset thread it was the only caller of, and
adds Archived:/Liveness:/Worker: lines so a released archive read no longer
prints identically to a live one.
2026-09-04 02:30:46 -04:00
Jinwoo-H 58efd45c33 fix(federation): negotiate structured read by method, not advertisement
- close an exited remote worker's terminal before labelling the release
  closed_exited_terminal; a host-certified exit keeps the verdict when the
  kill stops nothing
- replace the structured-read and fleet-snapshot capability probes with the
  optimistic call plus method_not_found, so hosts that serve
  federationReadOutput without advertising it stop downgrading to a scrape
- drop forceProbe so an unchanged peer at an unchanged epoch probes once
- distinct fleet reasons for home-side budget exhaustion and peer_changed;
  keep a host-supplied unverifiable reason and lastObservedAt for exited
- delete three compile-time-true self-capability checks, the advertised but
  never-read federation-release capability, and decode the pull page with zod
2026-09-04 02:24:03 -04:00
Jinwoo-H 92e1124c78 Merge branch 'fix-send-federation-transcript' into integrate-fixes 2026-09-04 02:23:26 -04:00
Jinwoo-H 994f01c20a Merge branch 'fix-ergonomics' (early part) into integrate-fixes 2026-09-04 02:23:18 -04:00
Jinwoo-H 80024119a9 Merge fix-worker into integrate-fixes 2026-09-04 02:20:19 -04:00
Jinwoo-H ca53fff0b7 test(orchestration): pin the settled arm of the legacy recovery plan 2026-09-04 02:19:00 -04:00
Jinwoo-H 6f20b25663 refactor(orchestration): name the worker-release harness as test support
The harness is test-only code sitting among the rpc/methods production modules;
*.test-support.ts is the repo's existing marker for that. The five folds the
review asked for are not possible: every merge target is already at 292-300
effective lines against the 300 cap.
2026-09-04 02:15:11 -04:00
Jinwoo-H 6d91a1c438 fix(orchestration): fence automatic resume for a settled, unreleased worker pane
Between worker_done and release the pane still holds a resumable provider
session, but listLegacyWorkerTerminalRecoveryRows selected live worker states
only, so nothing fenced it: reopening the workspace after a restart re-ran
`codex resume <session>` on a finished worker.

Settled workers whose terminal resource is still owned and neither released nor
retained now join the recovery rows, the plan marks those panes settled so they
never compete for adoption, and any fence the plan no longer claims is lifted —
on release, retain, user takeover and dispatch prune. An unreadable plan fails
closed and lifts nothing.

Ported from PR #17651, adapted to this branch's worker_terminal_resources
ownership model.
2026-09-04 02:12:47 -04:00
Jinwoo-H b1c56e2cbe fix(orchestration): stop worker receipts from contradicting themselves
worker-show spread the raw worker row beside its parsed copies, so a reader got
residual_resources (a JSON string) next to residualResources (an array), plus
host_scope as JSON-inside-JSON and two authority hashes with no consumer. Parse
once, emit camelCase once, and withhold the hashes.

worker-show also published only PTY liveness, so an agent that died at a trust
prompt read live there while worker-list called it unverifiable -- and
worker-list's nextAction pointed back at worker-show. Both now publish the same
fleet projection.

worker-list's projection.resource restated fields the row already carried, and
the unconfirmed-stop sentence doubled a terminator on an already-punctuated
reason.
2026-09-04 02:11:38 -04:00
Jinwoo-H 3053efe9bd test(orchestration): pin the watermark, restart-scan, lifecycle and downgrade fixes
Adds regression coverage that fails on ab6ee61aac: staging-failure and
DB-error exits leaving an active watermark, the restart scan skipping
pointer-pending and dispatch: mailboxes, the caller-edge walk over the
lifecycle graph, stopping -> failed, Task reopen/overturn, the v34
downgrade insert, the pointer-enter index predicate, the v33 skew column,
and in-memory/sqlite parity for pointer reservations and batch exclusion.
2026-09-04 02:08:17 -04:00
Jinwoo-H f9d95574e3 fix(orchestration): project fleet liveness once, on the evidence clock
Fleet liveness measured staleness against status.receivedAt, the delivery
clock a relay reconnect restamps, so an hour-stale agent read live after every
reconnect; it now uses evidenceObservedAt when the host supplies it.

worker-attention-context kept a second, divergent copy of that projection which
ignored host scope entirely and could not say exited. It calls the shared
projectLiveness now, with hostScope carried on WorkerAttentionFacts.

refreshOrchestrationFleetLivenessAttention re-derived categories from liveness
alone and dropped the unverifiable an unproven outcome contributed; it re-runs
the one attention projection instead, using the outcome now exposed on the
worker. The byte-identical active-sibling predicate is one shared fragment.
2026-09-04 02:06:58 -04:00
Jinwoo-H 4da7866b41 fix(runtime): re-check the prompt binding on every replay and warn on a swallowed Enter
- compare the recorded terminal binding on every prompt replay, not only the
  --wait-submit ones, so a replayed receipt cannot claim observation:supported
  for an incarnation that is gone
- normalise `terminal` to a pane key in replayStableCallerParams so a re-minted
  handle replays instead of failing request_mismatch
- warn (exit 0) when a supported send never reached turn_started, naming
  --retry-request <id> --wait-submit <seconds>
- collapse the write-only stages to input_accepted | turn_started
- move the PTY-keyed correlation state into AgentPromptRequestCorrelation
  (pty-keyed arrays, not NUL-joined string keys); the abandoned scan in
  lifecycle claim allocation now continues instead of returning
2026-09-04 02:00:46 -04:00
Jinwoo-H 8d142d2986 fix(cli): make selector_not_found name the offending worktree selector
A bare repo id passed to --worktree returned a content-free `selector_not_found`
with no value and no grammar, so a caller could not tell what was wrong. Shape it
at the CLI boundary — the only layer that still knows what was typed — in the same
selector/suggestions/nextSteps shape as an unknown-flag error, and point `--from`
on `orchestration check` at `--terminal`, which edit distance cannot reach.
2026-09-04 02:00:36 -04:00
Jinwoo-H 5974c500e0 fix(orchestration): release the mailbox watermark with the DB reservation
The watermark was set before stageMailboxPointerEnter, and neither the
staging-failure nor the DB-error exit cleared it, so a mailbox whose claim
was stolen parked every later delivery forever.

Also restores the restart scan's pointer-pending union and dispatch: handles,
aligns the pointer-enter index predicate with its query, adds
messages.pointer_enter_pending to the v33 skew columns, makes v34's
mailbox_handle downgrade-safe, and completes the lifecycle graph.
2026-09-04 01:59:14 -04:00
Jinwoo-H 0719a459f9 fix(orchestration): reject worker-start --terminal on the coordinator's own pane
A coordinator adopted as its own worker answers its own dispatch preamble; the
--terminal target is now rejected by handle and by resolved pane key with a
typed terminal_is_coordinator error naming the next action.
2026-09-04 01:58:18 -04:00
Jinwoo-H dcc540102d fix(orchestration): give a dispatched worker a concrete follow-up read cadence
The dispatch address is durable but never interrupts a worker, so a preamble
that only listed `check --terminal` produced workers that never read one
coordinator follow-up. Name the checkpoints and the pre-worker_done read.
2026-09-04 01:57:37 -04:00
Jinwoo-H dc6e0f6e47 docs(orchestration): give the kernel loop an exit condition and the missing commands
D1: name the two liveness layers (worker-list projection.liveness is the fleet
verdict, worker-show observation.status is PTY-only) and give the supervised
loop a bounded stall procedure instead of an unbounded wait.
D2: put worker-list, attention, requiresAction and nextAction in the loop and in
completion accounting.
D4: document request-show / --retry-request / terminal send --wait-submit.
D5: check names its caller with --terminal, never --from.
D7: give a dispatched worker a concrete follow-up read cadence.
D10: document the real folder-workspace route (project setup-existing-folder).
worker-start --spec is now the canonical loop's default.
Guidance pins are contracts via squash() instead of reflow-fragile prose.
2026-09-04 01:56:59 -04:00
Jinwoo-H 233fb0e92e fix(orchestration): derive worker terminal state once and page by rowid
The SQL CASE in worker-terminal-inventory-counts duplicated
deriveWorkerTerminalListState and diverged on an owned, unreleased resource
with no worker_dispatches row: SQL returned NULL where TS returns 'active', so
worker-list --terminal-state active dropped a row it labelled active. The SQL
copy is gone; filtering and counting now read the one TS projection.

worker-list also ordered by COALESCE(w.created_at, d.created_at) while the
fence pinned d.rowid, so a worker registering between pages re-emitted its row.
Order and fence now share d.rowid, and a pinned filtered cursor counts over its
own row extent instead of a live scan.
2026-09-04 01:56:41 -04:00
Jinwoo-H 482766caaa fix(cli): keep the prompt retry ID when only transport fails
A transport timeout is not evidence that another runtime answered, so the
preflight-attested host still owns the durable pending receipt. Strip the
request ID only when a different runtime handled the send, and drop the
false 'update Orca on the execution host' advice from that path.
2026-09-04 01:51:37 -04:00
Jinwoo-H 08bea361f9 fix(orchestration): never settle a worker terminal the dispatch no longer owns
worker-release's dead-process shortcut called settleDeadWorkerTerminalRelease
without requireArchive, and the guard only excluded ownership_state='released',
so a user takeover, an external terminal, or a transferred resource was marked
released and its output archive was lost.

Both release guards now read one (ownership_state, release_state) -> action
table, and settlement always requires the durable archive.
2026-09-04 01:51:14 -04:00
Jinwoo Hong b378101901 docs(cloud): reconcile the 2026-08-23 retry figure with the gate metric (#18581) 2026-09-04 01:40:55 -04:00
Jinwoo-H ab6ee61aac Keep main's providesInitialSurface guard on gated activation callbacks
The merge kept only the branch's tombstone reseed and dropped main's guard, so an
explicitly promised surface raced a fallback terminal after the inventory gate.
Fold both gated callbacks into one seeding helper.
2026-09-04 01:40:04 -04:00
Jinwoo Hong 79d5fb469a fix(cloud): recalibrate the relay monitor's postgres-retry freeze to a measured bar (#18580)
The global relay_cells FOR UPDATE lock made successful retries a
steady-state rate: fleet-wide p50 430 / p90 924 / p99 1320 / max 1504
per five minutes over the last 24 h, 55% of windows over the 300 bar,
only 22% of 15-minute gates clean. Three read-only dry-runs on
2026-09-04 froze on it, blocking the same-cap roll that carries #18521
and the beginProof crash guard to the 23 cells. 2000 clears every
measured healthy gate; the exhausted-retry, director concurrency, and
pool bars keep the incident discriminator role.
2026-09-04 01:23:54 -04:00
Neil 3941edd4b6 perf(ipc): build the filesystem allowed-root list once per authorization (#18423)
* perf(ipc): build the filesystem allowed-root list once per authorization

* perf(ipc): keep the allowed-root snapshot lazy so granted external paths build nothing

Hoisting getAllowedRoots to the top of resolveAuthorizedPath made every read of a
path covered by an external grant build the full root list, where main built none
(the grant answered before isPathAllowed reached the roots). Build on first use
instead: still one build per authorization, zero when a grant already answers.

* test(ipc): skip the allowed-root symlink escapes on Windows

Unprivileged Windows cannot create symlinks (EPERM), so both cases failed in
setup instead of exercising the escape check.
2026-09-03 22:20:35 -07:00