chore: keep orchestration audit artifacts out of PR root

This commit is contained in:
Jinwoo-H
2026-08-31 15:49:26 -04:00
parent 419ac29149
commit 70aa921a01
28 changed files with 0 additions and 6306 deletions
-105
View File
@@ -1,105 +0,0 @@
# Federation / SSH / mixed-version complexity audit
Scope: read-only review of federation and SSH execution-boundary paths in
`4a32d073fa` and their focused tests. No tracked source files were changed.
## Findings
### F1 — Keep: execution-host authority and peer pinning are implemented
Federated reads, stop, release, fleet snapshots, and relay pull/ack/import all
resolve the saved peer fingerprint and pass the resolved environment's pairing
revision as a fence (`src/main/runtime/rpc/methods/orchestration-worker-observation.ts:172-194`,
`src/main/runtime/rpc/methods/orchestration-federated-worker-read.ts:32-68`,
`src/main/runtime/orchestration/federation-sync.ts:63-90,156-199`). A re-paired
environment is rejected before an operation can target the wrong execution host.
Host-side attachment methods authenticate the Run-home fingerprint before
touching a terminal (`src/main/runtime/rpc/methods/orchestration-federation-attachment-observation.ts:4-17`),
and process identity is checked against the persisted pane/incarnation before
observation or close (`.../orchestration-federation-attachment-observation.ts:37-51`).
Focused regressions cover changed peer, pairing fences on every call, and stale
home observation fences (`orchestration-federated-transport-safety.test.ts:28-137`,
`orchestration-federated-fleet-snapshot.test.ts:62-145`). Keep this complexity;
these are concrete wrong-host and wrong-process failure modes.
### F2 — Keep: SSH contact loss remains `unverifiable`
`inspectRemoteAttachment` treats a missing/unknown host liveness verdict as
`unverifiable` and only an explicit host verdict yields `exited`
(`src/main/runtime/rpc/methods/orchestration-federation-attachment-observation.ts:53-75`).
The host release/stop code refuses to report a settled close when the PTY kill
was not confirmed (`src/main/runtime/rpc/methods/orchestration-federation-control.ts:215-264`).
The real-host tests exercise provider loss, old peers without verdict fields,
natural host exit, and unconfirmed close (`orchestration-federation-liveness-verdict.test.ts:145-355`).
This directly satisfies `docs/reference/ssh-execution-boundary.md` and should not
be simplified into `connected === false` heuristics.
### F3 — Keep: remote transcript reads do not fall through to desktop files
`readExactWorkerOutput` requires a registered SSH filesystem provider and returns
`remote_capability_unavailable` when the provider or attested path is absent
(`src/main/runtime/rpc/methods/orchestration-worker-output.ts:39-55`). The remote
reader only calls the provider with the hook-attested path and has a bounded
range/whole-file fallback (`src/main/runtime/orchestration/worker-transcript-read.ts:46-79`,
`worker-transcript-remote-read.ts:17-116`). Federated old peers are fenced to
terminal-only output and reject transcript cursors rather than silently changing
source (`orchestration-worker-legacy-federated-read.ts:17-47`). Keep; this is an
authority/security fix for same-path local lookalikes, not speculative handling.
### F4 — Simplify timing: fleet's five-second budget is not hard
Fleet sets a 5,000 ms total deadline and 3,000 ms per-host timeout
(`src/main/runtime/rpc/methods/orchestration-federated-fleet-snapshot.ts:14-17,40-57`),
but `OrchestrationPeerCapabilityCache.resolve` may retry once when the probe
observes a changed runtime epoch (`src/main/runtime/orchestration/orchestration-peer-capability-cache.ts:58-112`).
The retry receives the same `timeoutMs` computed before the first probe, so two
3-second status probes can consume ~6 seconds before the snapshot call (which is
then given a 1 ms minimum timeout). This violates the advertised total budget
under a fast peer restart. Pass the remaining deadline into capability probes or
disable stale-epoch retry for fleet reads; add a fake-timer regression asserting
the host call sequence stays within `FLEET_TOTAL_TIMEOUT_MS`.
### F5 — Simplify timing: relay page recursion has no total deadline
`syncFederatedDispatchPages` allows six 50-item pages
(`src/main/runtime/orchestration/federation-sync.ts:23-24,194-207`), each with
15-second pull/ack/import timeouts (`.../federation-sync.ts:76-90,156-199`). A
large backlog can therefore keep one dispatch sync in flight for roughly 90
seconds; overlapping syncs are coalesced by `OrcaRuntimeService`, so this can
delay newer lifecycle reports. The page cap is useful back-pressure, but carry a
single deadline through recursive pages (or lower per-page timeout) and add a
test that a six-page slow backlog respects that deadline. This is a timing
simplification, not an authority blocker.
### F6 — Defer (compatibility infrastructure): no real old/new federation wire harness
The focused federation tests model old peers by deleting capability strings or
mocking `method_not_found` (`orchestration-federation-output.test.ts:699-748`,
`orchestration-federation-liveness-verdict.test.ts:272-281`), while the actual
cross-version harness only covers terminal stream and structured agent-session
surfaces (`docs/reference/remote-wire-compatibility.md:45-63`). It does not run
released and current federation RPC/relay implementations against each other,
so required-field changes, zod stripping, relay payload semantics, or host
published-content drift could pass current tests. Keep the explicit capability
fallbacks now, but defer removal/expansion until a two-build federation/relay
journey covers start, pull/ack/import, worker read, fleet, and release; this is
an infrastructure gap rather than evidence that the current branch targets a
speculative user edge.
## Regression coverage run
```
pnpm exec vitest run --config config/vitest.config.ts \
src/main/runtime/rpc/methods/orchestration-federated-transport-safety.test.ts \
src/main/runtime/rpc/methods/orchestration-federated-fleet-snapshot.test.ts \
src/main/runtime/orchestration/orchestration-peer-capability-cache.test.ts \
src/main/runtime/rpc/methods/orchestration-federation-output.test.ts
```
Result: 4 files, 35 tests passed (66.8 s). No tracked edits were made.
## Classification
Keep F1–F3 (proven execution-boundary and mixed-version safety); simplify F4–F5
(timing budgets); defer F6 (real cross-version federation harness). No blocking
authority defect was found in the reviewed federation/SSH paths.
@@ -1,105 +0,0 @@
# Lifecycle/identity/liveness integration audit
Scope: independent review of the current worktree diff against
`orchestration-vnext-implementation-audit.md` and the accepted lifecycle/R3c
reports. No source files were changed.
## Findings
### L1 — High: lifecycle writer centralization is incomplete and transitions are unconstrained
The new `transitionLifecycleWithDb` helper only checks the current state against
the caller's `from` list; it does not enforce a legal transition graph. The
public `recordWorkerStage` therefore accepts arbitrary `state` values and can
move a terminal worker backwards (for example `succeeded -> ready`) while
appending a plausible receipt (`src/main/runtime/orchestration/db/lifecycle-transition.ts:66-113`,
`db/worker-dispatch/worker-dispatch-stage.ts:6-59`). More importantly, several
production lifecycle writers still bypass the helper: `markWorkerStartUnknown`
directly updates worker/dispatch/task state (`db/worker-dispatch/worker-dispatch-outcome.ts:102-130`),
`settleActiveDispatchesForTask` and `failDispatch` directly update dispatch/task
projections (`db/dispatch-context/dispatch-completion.ts:40-56,105-177`), and
federated start reconciliation has direct updates (`db/worker-dispatch/federated-worker-start-reconcile.ts:28-105`).
These paths produce missing or synthetic `from_state` receipts and leave
projection/receipt divergence, violating P3's “every lifecycle writer” gate.
**Impact:** stale/late stop, retry, or remote-start events can revive or settle
the wrong projection; replay/audit cannot reconstruct the actual transition.
Block until all writers use a guarded legal-transition primitive (or are
explicitly schema/migration exceptions), with rollback and direct-update ratchet
tests.
### L2 — High: local fleet rows without `host_scope` are classified as remote
`projectHost` returns `{kind: 'remote', id: 'unknown'}` whenever a worker has a
resource but no `hostScope` (`src/shared/orchestration-fleet-projection.ts:197-207`).
Normal local worker authority calls leave `host_scope` null unless a caller
supplies it (`db/worker-dispatch/worker-dispatch-authority.ts:50-61,120-151`).
`projectLiveness` then treats that row as remote and rejects any status without a
`connectionId` as `unverifiable` (`orchestration-fleet-projection.ts:145-153`).
Thus ordinary local workers can show `host=remote/unknown` and lose fresh local
liveness evidence. The projection tests only cover an explicit local hostScope,
not the production-null case.
**Impact:** false `unverifiable`/stale attention and incorrect host routing for
local and folder-workspace workers. Block fleet promotion until null host scope
is represented as local (or every local authority write always stamps a typed
local scope) and a regression test covers it.
### L3 — High: remote attachment observation can infer exit from relay loss
`inspectRemoteAttachment` documents that a dropped relay is not a death
certificate, but when `getTerminalLivenessVerdict()` is absent/null it returns
`exited` solely from `terminal.connected === false`
(`src/main/runtime/rpc/methods/orchestration-federation-control.ts:281-295`).
The same fallback exists in local observation, but federation is the execution
host boundary where contact loss must remain `unverifiable`. An older/mixed
runtime or a provider that cannot produce the verdict can therefore make
`federationShow`, fleet projection, archive reads, and release/stop treat relay
loss as process exit.
**Impact:** unsafe remote stop/release and false completion/liveness after SSH or
relay loss. Block remote promotion until null/unknown verdict maps to
`unverifiable` (only an explicit host exit maps to `exited`), with an old-peer
fixture.
### L4 — Medium: federated release result is not reflected in the home projection
The new home-side `releaseFederatedWorker` returns the execution-host receipt
but never updates the home `worker_dispatches`, `dispatch_contexts`, or
`federated_dispatches` rows (`src/main/runtime/rpc/methods/orchestration-federated-worker-release.ts:14-91`).
The remote handler only marks its own attachment stage (`orchestration-federated-worker-release-host.ts:65-74`).
After a confirmed remote release, the home worker-list still has the prior
state/terminal projection (often a handle with no Resource and `retained`), so a
subsequent query can prescribe inspection/release again despite a successful
release receipt.
**Impact:** split-brain cleanup state and repeated operator actions; at minimum
the result must carry a durable home receipt or update the federated projection
idempotently, preserving `release_unknown` on ambiguous relay failure.
## Verification
Passed focused tests:
```
pnpm test src/main/runtime/orchestration/db/attempt-outcome-projection.test.ts \
src/main/runtime/orchestration/db/lifecycle-transition.test.ts \
src/main/runtime/orchestration/r1-identity-migration.test.ts \
src/main/runtime/orchestration/orchestration-worker-dispatch-db.test.ts
# 4 files, 29 tests passed
pnpm test src/shared/orchestration-fleet-projection.test.ts \
src/main/runtime/rpc/methods/orchestration-federation-liveness-verdict.test.ts \
src/main/runtime/orchestration/orchestration-db-retention-pagination.test.ts
# 3 files, 26 tests passed
pnpm tc:node
# passed
```
## Verdict
**BLOCK.** The new identity/observation fixtures are useful and the focused
tests pass, but L1–L3 violate explicit vNext safety gates and can produce
incorrect lifecycle or remote liveness facts in production; L4 leaves remote
cleanup projection inconsistent.
-178
View File
@@ -1,178 +0,0 @@
# Orca main orchestration identity/lifecycle audit
## Executive summary
The current main-process orchestration kernel has a durable SQLite model for Runs, Tasks, Dispatch contexts, worker state, messages, and worker-terminal resources. It already fences most stale or foreign lifecycle messages, serializes Task/Dispatch races, accumulates retry failures into a circuit breaker, and makes terminal cleanup an explicit request/archive/close/reconcile flow. The remaining reliability boundary is that these are still several projections and heuristics rather than one Attempt event log: startup `ready` means input was accepted by Orca, `worker_done` is the only positive completion fast path, Task status can be edited through multiple APIs, and remote/SSH observation plus unsupervised lanes remain explicitly incomplete.
This audit maps observed behavior to implementation slices, non-goals, dependencies, and independently verifiable tests. It intentionally proposes contract/test work only; no source or `orchestration-issues.md` changes are included.
## Scope and evidence
Primary implementation inspected:
- `src/main/runtime/orchestration/db/schema/create-core-tables-sql.ts` and `create-graph-tables-sql.ts` (schema, constraints, indexes).
- `db/dispatch-context/*` (capabilities, lookup, completion, worker-report settlement, Task reconciliation).
- `db/worker-dispatch/*` (start, authority, readiness, stop, abandon, missing-terminal recovery, federated start/stop).
- `db/worker-terminal/*` plus `worker-terminal-ownership.ts` and `worker-terminal-release-reconciliation.ts` (lease ownership, transfer, archive, release, user takeover).
- `lifecycle-reconciliation.ts`, `coordinator.ts`, `coordinator-task-dispatch.ts` (message authority and coordinator projection).
Representative tests include `lifecycle-reconciliation.test.ts`, `orchestration-worker-dispatch-db.test.ts`, `db-task-dispatch-{invariant,races,lifecycle-guards}.test.ts`, `db-heartbeat-straggler-guard.test.ts`, `coordinator.test.ts`, `orchestration-worker-release-recovery.test.ts`, `orchestration-worker-stop-liveness-verdict.test.ts`, federation tests, migration tests, and RPC delivery/receipt tests.
## Entity and identity map
| Entity | Current durable identity/evidence | Current authority rule | Audit implication |
| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| Run/coordinator | `runs.id`, coordinator handle/pane; handle history in `run_coordinator_handles`; `coordinator_runs` scheduler row | Run mailbox routing and coordinator ownership are DB-derived; legacy adoption has separate compatibility principals | Good Run-level routing, but scheduler/Run/role identity is not one immutable actor record. |
| Task | `tasks.id`, `run_id`, parent, dependency JSON, creator handle/pane/process/generation | Task status transitions require/forbid active Dispatch depending on target state; dependency completion promotes children | Status is a projection maintained by several writers, not an append-only Attempt fact. |
| Dispatch/Attempt | `dispatch_contexts.id`, task/run, contract version, assignee handle, pane key, process incarnation, capability hash, depth, status, failure count | Capability + pane/process and status gates lifecycle writes; retry must name current terminal Dispatch | This is the strongest current identity boundary, but capability is mutable and no immutable Attempt event stream exists. |
| Worker execution | `worker_dispatches.dispatch_id`, runtime epoch, state/stage, setup/effects/residual resources, terminal handle | Starting/ready/stopping/settled state machine; process exit can force-fail | `ready/input_accepted` does not prove provider turn-start or prompt submission. |
| Terminal resource lease | `worker_terminal_resources.id`, owner/origin/prior owners, pane/process/host identity, ownership/release state | Exact identity re-proven before close; external/user-owned/transferred resources retained | Lease state is durable and conservative, but there is no TTL/bulk cleanup and federated release is unsupported. |
| Message/lifecycle evidence | Immutable message id/sequence, type, sender pane, payload; rejection marker persisted | `worker_done`/heartbeat require dispatch id and assignee pane (or exact legacy handle) | Rejection and duplicate suppression work; generic status/progress remains non-authoritative and no unified receipt cursor exists. |
## Current state and transition behavior
### Identity and authority
- Pane keys are remint-stable by parsed leaf UUID; opaque keys require exact equality (`lifecycle-reconciliation.ts:5-27`). A pane-bound Dispatch ignores handle churn, while a legacy row without pane identity accepts only exact handle equality.
- Dispatch capabilities are hashed at rest, timing-safe compared, revoked on stop/failure/completion, and require equivalent pane plus matching process incarnation (`db/dispatch-context/dispatch-capability.ts`). `prepareStartingWorkerAuthority` atomically records assignee/process/capability and rejects a second active Dispatch for the same handle/pane.
- Dispatch occupancy is checked by handle, exact pane, then parsed pane suffix. The writer boundary centralizes all live-worker inserts and stamps nesting depth (`db/dispatch-row-writer.ts`); migrations default unknown depth to 1 (fail-closed for child spawning).
- Canonical `dispatch:<id>` senders are permitted for imported/federated paths where explicitly enabled; ordinary lifecycle reconciliation uses sender pane/legacy handle, not payload claims.
Implementation work (I1 — stable actor/attempt identity):
1. Add immutable Agent, Endpoint, Attempt, and Resource IDs and an actor/role binding table; treat terminal handles and pane keys as ephemeral addresses.
2. Make capability issuance/refresh an explicit Attempt operation; bind every completion, heartbeat, question, cleanup, and nested dispatch to Attempt + process incarnation + host scope.
3. Expose role/parent/depth and an `unsupervised` state in every worker/show/list projection, including adopted/context-only rows.
Non-goals: do not remove pane-key compatibility, rotate all IDs in one migration, or infer stable identity from terminal title, provider name, or model-copied text. Do not make a foreign handle authoritative merely because it claims a valid Dispatch id.
Dependencies: vocabulary/invariant contract (I0) before schema/API changes; runtime/IPC caller identity; remote wire optional fields and capability negotiation; migration/backfill strategy for v30 and legacy principals.
Tests: extend `lifecycle-reconciliation.test.ts` for capability refresh, process-incarnation mismatch, nested parent/child completion, and stale Attempt after retry; extend `db-task-dispatch-races.test.ts` for concurrent remint claims; add a migration test proving old rows remain fail-closed; add an RPC integration test that caller identity cannot be supplied in payload text.
### Dispatch creation and startup
- Normal `createDispatchContext` atomically claims a ready Task, records assignee identity/depth, and marks the Task dispatched. Composed `createStartingWorkerDispatch` atomically records mutation receipt + pending Dispatch + worker `starting` row (+ federation attachment) and clears prior Task result.
- Authority attachment records terminal/worktree/setup/effect evidence and creates or transfers a terminal lease. `markWorkerDispatchReady` changes context to dispatched and worker to `ready/input_accepted`; `markWorkerStartUnknown` revokes capability and blocks the Task without asserting process death. Start failure settles Dispatch/worker and fails Task when no sibling remains.
- The coordinator creates at most one terminal per tick and polls pre-created Tasks; `decompose()` explicitly does not create a DAG.
Implementation work (I2 — honest delivery/readiness and atomic start):
1. Add a receipt/state vocabulary (`recorded`, `routed`, `endpoint_delivered`, `terminal_queued`, `provider_submitted`, `turn_started`, `report_persisted`, `settled`, `outcome_unknown`) instead of overloading `ready/input_accepted`.
2. Make common `task-create + worker-start` idempotent and atomic while preserving standalone Task creation for planned fan-out; persist partial effects and exact retry command.
3. Give providers a readiness capability and distinguish slow/unavailable endpoint from process exit; on input failure mark the endpoint unusable or roll back, never leave a silently live lane.
Non-goals: do not claim provider turn-start from a PTY write, title scrape, or successful local enqueue; do not auto-fail a worker solely because a readiness observer timed out; do not make AI decomposition part of this kernel slice.
Dependencies: I0 receipt definitions; provider adapter contract; mutation receipt capacity/idempotency; execution-host runtime (for SSH, WSL, and remote server evidence).
Tests: add unit tests for each readiness stage and duplicate request receipt; integration tests for write accepted-but-provider-not-started, slow boot, truncated prompt, and reconnect replay; retain `orchestration-worker-dispatch-db.test.ts` transactional acceptance/rollback tests and add cross-version optional-field coverage.
### Completion, heartbeat, and Task/Dispatch projection
- `worker_done` parsing requires object payload, `taskId`, `dispatchId`, and `outcome`; it verifies Task existence, Dispatch ownership/task match, sender authority, then calls transactional `settleWorkerReport`.
- Settlement rejects inactive/stale Dispatches, refuses completion while another supervised sibling is active, atomically updates Task + Dispatch + worker state, closes questions, promotes dependency-ready children, and treats identical settled reports as duplicates. Earlier same-Dispatch heartbeats are marked read; late heartbeats on inactive Dispatches are suppressed.
- Generic `completeDispatch`, `updateTaskStatus`, escalation failure, process-exit failure, and missing-terminal recovery are additional writers. `failDispatch` increments a three-failure circuit breaker, requeues Task to ready unless circuit-broken, and refuses failure while a supervised worker remains active unless process exit is proven.
- Coordinator processing is a polling loop; it records completed/failed task IDs in in-memory state and marks all read messages after per-message handling.
Implementation work (I3 — one Attempt lifecycle and completion truth):
1. Introduce an append-only Attempt event/receipt table and derive Task/Dispatch projections from it; fence late/reordered/duplicate events by Attempt sequence/causation id.
2. Separate process/turn observation, artifact/Git evidence, worker report, and coordinator acknowledgment; add explicit `finished_unverified`/`outcome_unknown` rather than treating `ready` or missing report as healthy.
3. Make heartbeat age host-computed and expose `never`, age, and `unverifiable`; keep `live`/`unverifiable`/`exited` distinct on SSH/remote loss.
4. Return typed rejection/failure receipts with whether mutation may have happened and an idempotent next command; wake coordinator from durable writes instead of relying only on polling.
Non-goals: `worker_done` remains a useful fast path; this work does not require trusting prose over host evidence, automatic success from Git cleanliness, or removal of heartbeat suppression. Do not collapse remote transport loss into `exited`.
Dependencies: I1 Attempt identity; I2 delivery receipts/provider observation; SSH execution boundary; durable mailbox/ack cursor; existing circuit-breaker and Task dependency semantics.
Tests: preserve and extend `lifecycle-reconciliation.test.ts` (foreign pane, handle remint, duplicate, heartbeat suppression); `db-task-dispatch-lifecycle-guards.test.ts` (active sibling, process exit, stop race); `db-task-dispatch-races.test.ts` (late failure vs report); add event-order/replay tests, missing-report observation tests, and SSH `live/unverifiable/exited` integration tests.
### Stop, retry, abandon, and recovery
- `beginWorkerStop` accepts `ready` or `start_unknown`, revokes capability, reblocks Task when no sibling remains, closes questions, and moves worker to `stopping`; `settleWorkerStop`/federated reconcile moves to stopped and Dispatch failed. `stop_unknown` intentionally remains potentially live until reconciled.
- `abandonWorkerDispatch` is idempotent for abandoned rows, a no-op for superseded Dispatches, and handles context-only legacy rows; it refuses stopping or succeeded workers. `reconcileMissingWorkerTerminal` turns active rows into failed/ready (or circuit-broken/failed), marks worker abandoned/stopped, and is idempotent.
- Retry is allowed only from the Task's current terminal Dispatch and failed/stopped/abandoned worker state; failure count carries across attempts. Legacy recovery lists potentially live worker rows for startup reconciliation.
Implementation work (I4 — first-class recovery supervisor):
1. Persist transition receipts for stop/cancel/supersede/retry/abandon/recover/takeover, including partial effects and safe next action.
2. Add a mechanical restart/reconnect supervisor with bounded retries and one convergence result; never use missing client inventory as proof of remote exit.
3. Distinguish `cancelled`, `superseded`, `abandoned`, `failed`, and `outcome_unknown` in projections and CLI/API; make stale detector output actionable instead of warning-only.
Non-goals: no broad terminal close, no automatic retry that can duplicate worktree/process ownership, and no coordinator policy that kills an uncertain SSH process. Preserve current stop fence ordering and circuit-breaker threshold until a versioned policy changes it.
Dependencies: I1 identity, I2 receipts, I3 Attempt events; host-side process liveness and remote relay; mutation receipts.
Tests: extend `orchestration-worker-dispatch-db.test.ts` for every transition and idempotent replay; `db-task-dispatch-lifecycle-guards.test.ts` for sibling/reblock behavior; `federation-terminal-recovery.test.ts` for bounded remote retry; add restart/reconnect integration tests that verify no duplicate Dispatch and no false exit.
### Worker-terminal ownership and release leases
- Lease rows capture origin/current/prior owner, worktree, terminal/pane/process/host identity, ownership (`owned`, `external`, `user_owned`, `transferred`, `released`), release intent/state, retention reason, archive metadata, and timestamps.
- Newly created terminals are `owned`; explicit external reuse creates `external`; exact settled owned resources may transfer to a retry and record prior owner. Legacy backfill is always `external + retained + legacy_ambiguous`.
- Release is post-settlement only. `requestWorkerTerminalRelease` records durable intent, but retains stopped/abandoned, external, user-owned, transferred, no-resource, and federated resources. Completion re-proves handle/pane/process/host identity, captures transcript/tail archive, closes only the exact terminal, and marks `released`; missing/unavailable/failed close yields `release_pending`, `release_unknown`, or retained. Startup reconciliation retries only requested/releasing backlog and coalesces concurrent passes.
- Real user input marks a matching owned resource `user_owned + retained(user_takeover)` and deletes its release archive. Worker list derives process accounting separately from Task/Dispatch outcome (`active`, `reclaimable`, `retained`, `release_pending`, `release_unknown`, `released`).
Implementation work (I5 — complete lease lifecycle):
1. Make lease owner a first-class projection for process, terminal/pane, worktree, and setup surface; settle exactly once into release/retain/suspend/takeover with reason and optional TTL.
2. Add bulk and per-row idempotent “clear Done”/release, retention expiry, and explicit historical-vs-live listing; keep archives queryable after process release.
3. Implement worker-server/federated release protocol (capability-negotiated); until then keep `federation_unsupported` visibly retained.
4. Make missing-tab bookkeeping idempotently settle only with positive host `exited` evidence; retain `unverifiable` otherwise.
Non-goals: never close a broad pane by handle/title, never release external/user-owned/transferred resources automatically, and never delete evidence as part of a normal list operation. No TTL should override user takeover or an unresolved identity conflict.
Dependencies: I1 stable Resource/Attempt identity; I3 completion evidence; I4 recovery supervisor; transcript/archive API; remote wire compatibility and host-side close endpoint.
Tests: `orchestration-worker-release-recovery.test.ts` already covers requested/releasing restart recovery, unavailable terminal, untouched backlog, concurrency, and list counts; add tests for TTL/retention, transfer fencing, identity conflict, user input racing release, archive durability, bulk idempotency, and federated capability negotiation. Keep `orchestration-worker-stop-liveness-verdict.test.ts` assertions that unknown is not exited.
### DAG and coordinator behavior
- Tasks store `deps` JSON and are initially `pending` unless every dependency is completed; `promoteReadyTasks` runs in completion transactions and marks dependents ready. Parent/run IDs are checked for same-Run ownership. Coordinator dispatches ready tasks up to `maxConcurrent`, one terminal creation per tick, and converges only when all Tasks are terminal.
- There is no cycle detection, immutable DAG snapshot, attempt-level graph, or atomic fan-out/create+dispatch. `decompose()` requires pre-created Tasks and explicitly leaves AI decomposition for a future phase. Task status is still independently writable (guarded, but not event-derived), and a single Run query is not yet a fleet/agent graph.
Implementation work (I6 — DAG projection and fleet query):
1. Validate DAG acyclicity and same-Run dependency closure at creation; persist a Run-scoped graph/version so retries do not rewrite topology.
2. Derive readiness/convergence from Attempt events and dependency facts; expose blocked reason and unsatisfied dependency IDs.
3. Add one bounded fleet query returning Run/Task/Attempt/role/provider/host/worktree/stage/heartbeat/exit/evidence/lease/next action; use it for coordinator and UI instead of per-lane polling.
Non-goals: no AI decomposition in the reliability kernel, no merge queue or cost budget in this slice, and no second graph/state store in the UI. Existing independent Task creation remains supported.
Dependencies: I1 Attempt/role identity, I3 event-derived completion, I4 recovery states, I5 lease projection; query pagination and remote host routing.
Tests: extend `db-task-create-readiness.test.ts` for cycles/cross-run deps and interleaved completion; `db-task-dispatch-invariant.test.ts` for atomic fan-out races; `coordinator.test.ts` for max-concurrency/convergence and one-wave behavior; add fleet-query integration tests across local, folder workspace, SSH, WSL, and federated rows.
## Explicit non-goals and compatibility constraints
- Do not rewrite or delete legacy compatibility tables/routes in the first slice; dual-write/dual-read with parity tests and bounded migration is safer.
- Do not conflate coordination hierarchy, filesystem/worktree isolation, Git lineage, execution host, UI grouping, or notification audience.
- Do not infer process death from client inventory absence, timeout, socket loss, or a quiet terminal. SSH execution belongs to the execution host; use `live`/`unverifiable`/`exited` only.
- Do not add a new remote stream opcode without capability negotiation. New lifecycle receipt fields should be optional and old readers must remain safe.
- Do not expand this audit into renderer UX, notification policy, Git merge behavior, or provider implementation; those consume the contracts above and need their own audits.
## Dependency DAG and safe parallelization
```text
I0 vocabulary + invariants
├── I1 stable identity/authority ───────┐
├── I2 delivery/readiness receipts ─────┼── I3 Attempt events + completion truth
└── mailbox ack/wakeup contract ────────┘ ├── I4 recovery supervisor
├── I5 resource leases
└── I6 DAG projection/fleet query
I1 + I2 + I3 + I4 + I5 ──> attention/UI projections and provider/remote adapters
```
Safe parallel work after I0 (each with isolated schema/API ownership):
1. I1 identity schema/capability tests and I2 provider receipt/readiness tests can proceed in parallel if both use additive fields and share only the I0 vocabulary.
2. Mailbox ack/wakeup work can proceed beside I1/I2; it must consume immutable Attempt/message IDs rather than invent another identity.
3. I4 recovery tests can begin against current states while I3 event storage is designed, but implementation should land after the Attempt event contract is fixed.
4. I5 lease/archive work can proceed beside I3 once its owner key is stable; federated release adapter is independently parallel after remote capability negotiation is specified.
5. I6 DAG validation/readiness tests can proceed independently of terminal cleanup, then integrate with I3-derived completion and I5 lease fields.
Unsafe to parallelize: changing status enums, capability authority, or remote verdict vocabulary independently; changing worker-list projections before lease semantics are fixed; adding provider-specific lifecycle signals before receipt stages are agreed; or modifying legacy migration and reset behavior without compatibility fixtures.
## Definition of independently verifiable completion
Each implementation slice is complete only when its unit tests cover valid, duplicate, late, reordered, unauthorized, and restart/reconnect paths, and an integration test proves the durable DB projection plus runtime-visible receipt. A release is not complete if it merely closes a PTY; a Dispatch is not complete if only a message or UI status says done; and a remote process is not exited unless the owning host positively proves it.
-118
View File
@@ -1,118 +0,0 @@
# Mailbox, wakeup, pointer, and provider-evidence audit
## Scope and executive summary
This audit follows a message from durable insert through `notifyMessageArrived`, waiter wakeup, PTY pointer injection, explicit `orchestration.check`, and (for worker prompts) provider lifecycle verification. It covers local, folder-workspace, WSL/SSH/federated routing, restart behavior, and the recent Windows reports (`agent_prompt_stalled` while Claude/Codex had queued input; Claude reads exposing only a short visible screen). No source or issue-ledger files were changed.
The durable mailbox is substantially stronger than a best-effort notification: rows are immutable, ordered, replayed until acknowledgement, fenced by Run consumer generation, and routed in bounded pages. The weak boundary is between “bytes accepted by a PTY” and “provider accepted/submitted the prompt”: the current API reports a single success/failure and can mark a real queued prompt as a hard worker failure. Similar boundary ambiguity exists for pointer Enter, waiter registration, ordinary replies, restart repair, and terminal scrollback evidence.
## Current sequence
```text
send/reply/ask
-> SQLite message/question insert (durable id + sequence)
-> canonical direct/group/foreign routing
-> notifyMessageArrived(handle,type)
-> matching in-memory waiter resolves, OR microtask pointer delivery
-> idle/live/owner checks; outstanding Run delivery suppresses pointer
-> PTY write settles; rows get delivered_at and watermark
-> 500 ms recheck; optional '\r' write; pointer flight retires/redrives
-> explicit check reads durable rows / creates Run Delivery / marks read on ack
-> worker prompt path separately waits for workingSequence (provider effect)
```
`messageWaitersByHandle` and pointer flight state are process-memory state. SQLite `read`, `delivered_at`, delivery contract, and Run Delivery rows survive restart; waiter promises and in-flight Enter timers do not.
## What users and agents see today
| Operation | Observable behavior | Important edge |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `orchestration.send` | Persists one row per recipient (group sends share a thread id but have independent read state), returns message/sequence; then wakes waiters or causes an idle Run PTY pointer. Bare handles are canonicalized to Run/Dispatch; foreign targets enqueue a relay. | `worker_done` and heartbeat are authority-checked. Invalid lifecycle mail is converted to a high-priority rejection row. A suppressed heartbeat is read and does not wake a waiter. |
| `orchestration.check --run` | Routes direct snapshots, validates current consumer generation, then returns a replayable FIFO Delivery (up to 50). Existing outstanding Delivery is replayed until `--ack`; `--peek` and `--all` are read-only. `--wait` is an exclusive in-memory long poll with optional type filter. | Wake filter only decides when to wake; the subsequent Delivery contains the full oldest batch, including other types. No waiter survives process restart. |
| `orchestration.check` on active Dispatch | Reads Dispatch mailbox and marks returned unread rows read immediately (unless peek/all); `--wait` resolves on a matching notification, then reads/marks all currently unread matching rows. Inactive Dispatch unread rows are migrated to its Run in bounded pages and Run waiters are notified. | A worker losing ownership is reported as `dispatch_inactive`; direct-mail cancellation can be interpreted as Run fencing. |
| Direct `check` | Unread rows are consumed and lifecycle rows reconciled; `--peek`/`--all` inspect without consuming. A non-signal cancellation while waiting is treated as ownership fencing. | Legacy-run rows are inspect-only and cannot be acknowledged by this path. |
| `orchestration.reply` | Question replies are transactional/idempotent: same body repeats return `duplicate`, a different body conflicts, and the answer row is immediately read. Ordinary replies mark the original read and insert a `Re: ...` row to `from_handle`. | Ordinary replies have no idempotency key or Delivery object; repeated calls create repeated unread rows. |
| `orchestration.ask` | Only an active supervised Dispatch may ask. A durable question is inserted, then the sender waits on `dispatch:<id>`; answer/close, timeout, or abort returns typed status. Timeout/abort leaves the question pending for explicit `--resume`; resume does not create a duplicate. | Capability and current-dispatch checks fence stale workers. |
| Automatic pointer | Only Run mailboxes are injected. Owner must resolve to a live, writable, idle leaf with live evidence. Unfiltered waiters suppress pointers; filtered waiters reserve their types so unrelated types can be pushed. Pointer text is a literal `orca orchestration check [--run <id>]`. | PTY transport settlement precedes `delivered_at`; a 500 ms idle/ownership recheck precedes Enter. If Enter is rejected/uncertain after staging, rows can remain delivered and no automatic replay occurs. Cursor-title path redrives rather than submitting. |
| Notification | `notifyMessageArrived` canonicalizes direct/foreign destinations, wakes matching waiters, and queues pointer delivery when no applicable waiter owns the type. A two-second unref'd repoint timer repairs Run pointers after restart/rebind. | There is no native OS/toast attention projection tied to orchestration mail; PTY pointer is the primary automatic user notice. |
| Restart/reconnect | Durable undelivered unread Run rows are scanned and repointed after handles return; detached direct rows are routed to their current owner. In-flight memory flights/waiters disappear. Dispatch handles are skipped by restored pointer scheduling. | Repair is timer/idle dependent; a non-idle or unavailable PTY requires explicit check. |
| Worker prompt send | Prompt text is bracketed-paste/chunk written through a per-PTY-generation serializer. Claude/Codex use render/quiet gates; other agents use 500 ms (1.5 s on Windows) delay before Enter. Verification waits up to 5 s for `workingSequence` to increment and otherwise throws `agent_prompt_stalled`. | This proves an observed lifecycle transition, not PTY acceptance, provider submission, or whether text is queued in a TUI. Worker-start contract currently turns this error into `ok:true` with failed Dispatch/task and revokes capability. |
| Terminal/worker read | `terminal.read` without `--screen` returns bounded tail (runtime caps 2,000 lines/256 KiB; default limit 120). `--screen` returns a visible frame; cursor reads completed transcript lines and bypass visible/provider fallback. `worker-read` prefers an exact provider transcript and otherwise returns terminal fallback with source/warning metadata. | Claude’s TUI visible/terminal tail is not guaranteed full scrollback. A terminal cursor pins subsequent reads to terminal source; screen is advisory, not provider truth. |
## Guarantees that already exist
- `insertMessage(s)` commits immutable IDs and monotonic sequences; default delivery contract is `current_delivery` (legacy Run is explicit).
- Run Delivery creation/ack is transactional, one outstanding delivery per Run, FIFO capped at 50, and replay/idempotent under duplicate ack. Consumer generation fences stale readers.
- Direct routing is paged with a through-sequence snapshot, so arrivals after a check starts are not stolen into that check. Active Dispatch ownership is preserved.
- Pointer state tracks PTY flight, mailbox sequence watermark, parked deliveries, reservation merging, and same-leaf/PTY checks; failed asynchronous transport does not mark rows delivered. Restart repoint scans durable `delivered_at IS NULL` rows.
- Question creation/answer and worker lifecycle authority are durable and capability/process-incarnation checked. Federated sends/answers use relay records and protocol gates.
- Existing tests exercise replay, fencing, routing races, pointer transport settlement, filtered waiters, restart repair, duplicate notifications, ask idempotency, Windows delay, prompt serialization, transcript source changes, and screen/cursor schema.
## Missing guarantees and risk
1. **Prompt receipt conflates stages (high, Windows-visible).** A PTY write accepted plus queued TUI input can still yield `agent_prompt_stalled`; worker start then fails the task and revokes capability even though the provider may consume that same text later. Retrying blindly can duplicate work. There is no durable operation/message id that lets a caller query or safely retry.
2. **Potential lost wakeup (high).** `check` performs a durable read/routing pass, then registers an in-memory waiter; an insertion/notification between those points can be missed, especially after paged routing yields. Registration is not an atomic “subscribe + recheck durable epoch.”
3. **Pointer delivery is not provider submission (high).** `delivered_at` means pointer bytes settled, not Enter accepted or check executed. If the delayed Enter write is rejected/uncertain, staged rows remain delivered and are not automatically redriven; users may see no further nudge while mail remains unread.
4. **Coarse Run acknowledgement (medium).** Ack marks an entire Delivery batch read. Newer mail is held behind the outstanding Delivery for automatic pointer suppression; there is no per-message ack cursor or partial-ack recovery.
5. **Ordinary reply duplication (medium).** Unlike question answers, ordinary reply has no idempotency/delivery receipt. Network retries can create multiple `Re:` rows and unread fan-out.
6. **Restart/waiter convergence (medium).** Waiters are memory-only; restored repoint uses a 2 s timer, idle gating, and skips `dispatch:` pointers. A runtime crash during pointer/Enter loses the in-flight intent, requiring a later idle edge or explicit check.
7. **Platform command mismatch (medium).** `formatMessagePointer` emits literal `orca orchestration check`; the CLI resolver knows `orca` versus `orca-ide` and WSL/remote command context, but pointer formatting does not select it. A pasted pointer can fail or invoke the wrong binary.
8. **No orchestration attention policy (medium).** There is no durable native notification/toast/badge projection, nor a policy for muted/background/remote work. PTY injection is unavailable when a pane is not writable or idle.
9. **Provider output evidence is bounded/advisory (high for completion evidence).** Claude screen/tail reads expose only the visible/limited PTY buffer; cursor mode intentionally avoids snapshots. Worker-read fallback is explicitly terminal source, while exact transcript availability varies by provider/session. A raw TUI frame cannot prove completion or full scrollback.
10. **Remote uncertainty vocabulary needs enforcement (cross-cutting).** Loss of SSH/relay contact is not process death; status must remain `live`, `unverifiable`, or `exited`, and prompt/mail receipts must not infer `exited` from a timeout.
## Focused implementation slices
### A. Atomic wake correctness
Change `src/main/runtime/orca-runtime.ts` (`waitForMessage`), `src/main/runtime/rpc/methods/orchestration.ts`, and `mailbox-notification-coordinator.ts` to register a waiter with a durable mailbox epoch (or insert a register-then-recheck transaction). Re-read unread state immediately after registration; resolve synchronously if the epoch advanced. Preserve exclusive/type-filter semantics. Add deterministic insertion-at-registration tests to `orchestration-mailbox-routing-races.test.ts`, `orchestration-mailbox-notification-consistency.test.ts`, and `orchestration-check.test.ts`, including a process/restart case that falls back to explicit check.
### B. Per-message Run ack and replay
Extend `db/schema/create-core-tables-sql.ts` and `db/runs/run-delivery.ts` with an ack cursor/message set while retaining one outstanding delivery for old clients. Define partial-ack and newer-arrival behavior; make duplicate/old cursor a no-op. Test FIFO partial ack, replay after crash, concurrent consumer fencing, and mixed wake types in `orchestration-run-delivery-db.test.ts` and `orchestration-message-delivery-identity.test.ts`.
### C. Staged prompt receipt and idempotent retry
Split `sendTerminalAgentPrompt`/`writeTerminalAgentPrompt` and `agent-prompt-submission-verification.ts` into observable stages (`recorded`, `terminal_queued`, `provider_submitted`, `turn_started`, `outcome_unknown`). Persist an operation/request id (including worker-start mutation receipt), return it on timeout, and make retry-by-id query status rather than resend. Do not convert `agent_prompt_stalled` directly to task failure when queued/provider state is unknown; require explicit reconciliation. Cover Windows delay, swallowed Enter, Claude/Codex queued TUI, generation/permission races, cancellation, and duplicate retry in `agent-prompt-submission-runtime.test.ts`, `agent-prompt-submission-windows-submit-delay.test.ts`, and `orchestration-worker-start-prompt-contract.test.ts`.
### D. Provider-backed output evidence
In `orchestration-worker-output.ts`, `worker-transcript-read.ts`, runtime read methods, and provider adapters, expose capability/source/coverage (`transcript`, bounded terminal tail, visible screen) and a stable cursor. Add explicit “transcript required” and “coverage incomplete” assertions without pretending a screen is full scrollback. Extend `orchestration-worker-output.test.ts`, `worker-transcript-read.test.ts`, and `terminal-read-screen-cursor.test.ts` with Claude short-screen, cursor pinning, source-change, and SSH-unverifiable cases.
### E. Pointer command and attention projections
Have `formatter.ts` ask `cli-command.ts` for the platform/connection-specific command (including WSL/SSH), and add an explicit pointer receipt state separate from `delivered_at`/Enter. Add durable attention events plus renderer/native notification adapters only after policy (background, mute, remote) is agreed. Test command selection and no-focus/remote behavior in formatter, mailbox notification, and runtime integration tests.
## DAG contribution
```text
contract vocabulary + regression fixtures
├─ A atomic wake registration/recheck
├─ B per-message ack/replay
├─ C staged prompt receipt + idempotent retry
└─ D provider transcript/coverage contract
A + B ──> restart/reconnect convergence and pointer receipts
C + D ──> operator evidence and safe worker-start settlement
foundations ──> E platform pointer command and native attention projections
```
A–D can proceed in parallel after the shared status vocabulary/fixtures; restart convergence depends on A+B; operator evidence depends on C+D; E consumes all foundations. This keeps the DAG depth to three levels and avoids a scheduler rewrite.
## Testing and acceptance matrix
| Slice | Existing safety net | New acceptance |
| ----------------- | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| Wake | `orchestration-mailbox-routing-races`, notification consistency, check tests | No insertion between initial read and waiter registration is lost; filtered and exclusive waits remain correct. |
| Ack | Run-delivery DB and identity tests | Partial ack/replay is monotonic and generation-fenced; old clients still replay whole batch. |
| Prompt | submission runtime/verification/Windows-delay and worker-start contract tests | Every receipt names the furthest observed stage; queued/unknown never auto-fails or duplicates on retry-by-id. |
| Output | worker-output, transcript-read, terminal screen/cursor tests | Source and coverage are explicit; Claude screen cannot be mistaken for full transcript; remote contact loss is `unverifiable`. |
| Pointer/attention | notification consistency, transport settlement, runtime pointer tests | `delivered_at`, Enter, and provider stages are distinct; restart/redrive and platform command selection are deterministic. |
## Explicit non-goals
- This document does not implement source changes and does not edit `orchestration-issues.md`.
- No flag-day rewrite of orchestration, scheduler/DAG engine, or provider TUIs; slices must preserve mixed-version wire compatibility.
- No automatic prompt resend solely because of timeout/stall, and no inference of remote process death from lost contact.
- No promise that a raw PTY/TUI screen equals provider truth or complete scrollback; screen remains advisory.
- No OS notification/focus stealing before a durable, user-configurable attention policy; no bespoke semantics outside a provider capability contract.
- No assumption that every workspace is a git worktree or every execution is local; execution-host ownership and `live`/`unverifiable`/`exited` vocabulary remain in force.
@@ -1,26 +0,0 @@
# Transcript / terminal-read audit
Scope: current `orchestration-v3` HEAD; read-only audit, no production edits.
## Keep
- **Execution-host authority.** `readWorkerTranscript` accepts a remote filesystem provider only with the hook-attested `transcriptPath`; it explicitly refuses local provider-root search (`src/main/runtime/orchestration/worker-transcript-read.ts:62-78`). WSL sessions stay on the local guarded resolver with an attested distro, while SSH sessions are routed through `getSshFilesystemProvider` (`src/main/runtime/rpc/methods/orchestration-worker-output.ts:39-71`).
- **Race fences.** Local reads capture file identity before/after and hash a 64-byte boundary checkpoint; forward reads revalidate path and open-handle identity (`worker-transcript-local-read.ts:46-75, 87-145`; `worker-transcript-source-identity.ts:15-58`). Remote ranged reads re-stat, verify boundary bytes, and reject changed identity; unsupported range capability degrades once to the bounded SSH whole-file path (`worker-transcript-remote-range-read.ts:49-112`; `worker-transcript-remote-read.ts:82-111`).
- **Source-pinned worker output.** Exact worker reads validate provider session before and after transcript reads and pin opaque cursors to process/session/path identity (`rpc/methods/orchestration-worker-output.ts:39-139`). Fallback terminal reads are labeled `source: terminal`, `sourceExact: false`, and preserve fallback reasons (`:142-220`).
- **Terminal screen/stream distinction.** `terminal.read` rejects cursor+screen at the RPC schema (`rpc/methods/terminal/unary-schemas.ts:40-76`), labels every response source, and CLI refuses an old host that silently drops `screen` (`cli/handlers/terminal.ts:72-105`). Provider snapshots are timeout-bounded/shared and stale generations or output sequences are discarded (`orca-runtime.test.ts:19191-19207, 19241-19257, 19284-19309, 19358-19389`).
## Simplify
- Local initial transcript reads use the tail reader and return `limited: page.hasMore` with `nextOffset: consumedTo` (normally EOF), but emit no explicit warning that omitted older records cannot be paged (`worker-transcript-local-read.ts:53-75`). The remote implementation does emit that warning when its bounded EOF window clips history (`worker-transcript-remote-read.ts:205-213`), and the formatter test expects such wording (`cli/handlers/orchestration/worker-output.test.ts:5-53`). Add a local regression asserting `contentComplete`, `clipping`, and warning semantics, or make local and remote clipping metadata consistent.
- Add CLI regression coverage for `terminal read --screen`: current tests cover only schema rejection/acceptance (`rpc/methods/terminal-read-screen-cursor.test.ts:12-42`), not old-host source absence or `screen-unavailable` formatting.
## Defer (not block)
- Remote source identity fingerprints only `(dev, ino)` and same-size changes rely on mtime plus a 64-byte boundary hash (`worker-transcript-source-identity.ts:15-58`). A same-size rewrite outside that boundary with coarse/unchanged remote mtime is theoretically undetectable; add a test if providers ever compact transcripts in place.
- Local initial tailing inherits the native reader's unbounded backward scan when malformed records never increment the valid-message count (`native-chat/transcript-tail-reader.ts:86-153`); a very large malformed file can cost a full-file scan, but no user-impacting case is evidenced in this PR.
## Evidence
Targeted suite passes: `pnpm test src/main/runtime/orchestration/worker-transcript-read.test.ts src/main/runtime/orchestration/worker-transcript-remote-read.test.ts src/main/runtime/rpc/methods/orchestration-worker-output.test.ts src/main/runtime/rpc/methods/terminal-read-screen-cursor.test.ts` (4 files, 34 tests).
Existing tests cover split EOF append, bounded >8 MiB remote tails, legacy SSH fallback, same-inode replacement, WSL/SSH routing, payload redaction, source-pinned cursors, and stale provider snapshots (`worker-transcript-read.test.ts:37-274`; `worker-transcript-remote-read.test.ts:67-359`; `orchestration-worker-output.test.ts:37-357`; `orca-runtime.test.ts:19191-19389`). No correctness blocker found.
-256
View File
@@ -1,256 +0,0 @@
# Transcript and terminal-read architecture audit
This audit covers the Orca worker-output and native-chat readers, provider/session
adapters, terminal RPC/CLI, SSH federation and WSL handling, fleet observability,
and the analogous surfaces in Overstory, Paperclip, Herdr, and Gas Town. It is an
architecture note only; `orchestration-issues.md` and source code were not changed.
## Executive findings
- Orca has two deliberately different read products. Native Chat reads
provider-authored JSONL (`src/main/native-chat/*`), while orchestration worker
reads prefer that transcript and fall back to an exact PTY tail. A source identity
and process-incarnation fence prevents a cursor from silently crossing a new
process or transcript.
- The local, SSH-host, and federated paths now execute reads on the execution host.
`orchestration.federationReadOutput` calls the remote runtime; old peers are
fenced to terminal-only `federationRead`. Losing an SSH relay is
`unverifiable`, never `exited`, and does not make a stop successful.
- Transcript support is intentionally provider-specific at the decode/resolution
boundary (Claude/OpenClaude, Codex, Grok, OMP). Providers without a trustworthy
path report explicit fallback reasons; `auto` may use PTY output. This is safer
than treating every agent's JSONL as interchangeable, but the supported-agent
registry and provider hooks remain a growing maintenance seam.
- The strongest reference pattern is a normalized event stream plus a durable,
bounded log cursor. Overstory's adapter interface makes provider differences
explicit; Paperclip's run-log store survives pod loss with a transparent object
storage mirror; Herdr separates visible/recent/detection reads and has passive
subscriptions; Gas Town combines a Claude JSONL watcher (opt-in) with a
feed/audit event log and restart-first lifecycle. None provides Orca's
cross-provider transcript/PTY cursor contract out of the box.
## Orca architecture (current)
### Provider transcript and native chat
`src/shared/native-chat-agent-support.ts` maps `claude` and `openclaude` to the
Claude decoder and exposes Codex, Grok, and OMP as separate transcript agents.
`nativeChatRequiresLocalTranscript` marks Grok/OMP as unable to work from a
remote Model-A SSH main because their hooks do not disclose a transcript path.
`src/shared/agent-session-resume.ts` is the provider-session authority: it
normalizes session ids/paths, captures Claude/Codex `transcript_path`, captures
Pi's `session_file`, and defines provider-specific resume argv.
`src/main/native-chat/session-file-resolver.ts` first probes an authoritative hook
path, then provider-specific roots (Claude project slugs; managed Codex home then
`CODEX_HOME`; Grok/OMP roots). Windows WSL paths are translated to a ranked
`\\wsl.localhost` path in `host-readable-transcript-path.ts`; a stopped distro is
reported as an FS-gate refusal, not a missing transcript. `transcript-tail-reader.ts`
does bounded tail reads, complete-line boundaries, 2 MiB record caps, malformed/
oversized counters, lifecycle decoding, and forward cursor paging.
`src/main/ipc/native-chat.ts` exposes desktop IPC read/subscribe. Subscriptions
have sender-scoped cleanup, generation-safe unsubscribe, an unflushed-session
pending frame, and a retry/poll loop (`transcript-watch.ts`) that handles delayed
first JSONL flushes, replacement, and WSL failures. Runtime RPC
`src/main/runtime/rpc/methods/native-chat.ts` reuses the same readers for mobile/web,
with a 40-message initial / 2,000-message maximum window and client-kind payload
caps. Thus native chat is agent-specific in decoding, but transport and lifecycle
semantics are agent-agnostic.
### Worker transcript-first output
`src/main/runtime/orchestration/worker-transcript-read.ts` resolves the exact
provider session captured after dispatch attach, reads an initial tail or an
8 MiB-bounded forward page, and bounds response messages through
`worker-transcript-payload.ts` (40 default/50 max messages, block/input/response
limits, local-image omission, dispatch-capability redaction). Failure reasons are
typed: `provider_unsupported`, `session_not_reported`, `transcript_missing`,
`transcript_unreadable`, or `transcript_parse_failed`.
`readExactWorkerOutput` (`orchestration-worker-output.ts`) defaults to transcript
when the session is exact, but `source: terminal` selects the PTY directly and
`auto` falls back with a reason. Cursors encode dispatch id, source, source identity,
and byte/stream position. Before and after the read, process incarnation, agent,
session key/id, and transcript path are compared; mismatches return
`source_changed`/`worker_identity_changed`. PTY fallback uses the bounded recent
buffer, redacts dispatch capability tokens, and preserves terminal status.
Released workers are read from the archive path (`orchestration-worker-archive-read.ts`)
instead of a dead PTY. This makes release/read a durable handoff, not a best-effort
screen scrape.
### Terminal RPC and CLI
`terminal.read` (`terminal-query-methods.ts`) is a unary RPC over
`runtime.readTerminal`, supporting a cursor, bounded limit, and a `screen` flag.
The CLI (`src/cli/handlers/terminal.ts`) rejects `--screen` + `--cursor` and checks
the response `source` so an older host that silently drops `screen` cannot return
the wrong question's accumulated output. `terminal.read` is provider-neutral: it
reads the runtime PTY tail/current screen, not a provider transcript. Recent output
is retained by `RecentPtyOutputBuffer` (64 KiB default), with larger durable
scrollback snapshots handled by provider/renderer serializers.
### SSH, federation, and WSL boundaries
`SshPtyProvider` and `SshPtyProviderOutputState` proxy PTY lifecycle/data through a
relay, track provider generations/incarnations, and expose capability probes for
agent-session claims and idempotent creates. Snapshot authority is explicitly false
for this provider; reconnect attaches and replays through source-activation leases.
`orchestration.federationReadOutput` executes the full transcript/PTY decision on
the home (execution) host and returns its runtime epoch. A pre-structured peer is
handled by `orchestration-worker-legacy-federated-read.ts`, which permits only
terminal reads and returns `remote_capability_unavailable`; a transcript cursor is
rejected rather than silently downgraded.
`inspectWorkerTerminal` and federation control use the liveness vocabulary
`live`/`unverifiable`/`exited`; unregistered providers or lost relay contact are
never evidence of death. This matches `docs/reference/ssh-execution-boundary.md`.
WSL transcript access is separately gated (`wsl-transcript-fs-*`): guest paths are
never opened as `C:\home\...`, distro homes are cached/ranked, and refusal is
retryable. The same resolver runs in the remote main, so SSH reads use the remote
home rather than the desktop user's home.
### Fleet observability and notifications
PTY data notifications are sequenced and flow-controlled in the relay; gaps trigger
provider snapshots where available. Runtime orchestration status is persisted in
the dispatch DB, and setup/status changes enqueue federation relay messages. Worker
show/read responses carry exactness, liveness reason, `agentWait`, source identity,
runtime epoch, fallback reason, and warnings. Desktop/mobile unread/notification
policy is intentionally downstream of terminal side-effect facts, not embedded in
the PTY provider.
## Reference products
| Product | Transcript/output source | Agent scope | Default/fallback behavior | Remote/fleet observability |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Overstory** | `AgentRuntime` adapters; Claude hooks/JSONL today, planned headless `stream-json`; `metrics/transcript.ts` parses usage | Provider-specific adapter (`claude`, `codex`, `pi`, `copilot`, `cursor`, etc.); interface is agent-agnostic | Interactive tmux is legacy/default for some flows; headless stream-json is opt-in by runtime; tmux capture is a fallback/visibility path, not truth | No built-in SSH/federation in the audited tree; event/log files (`session.log`, `events.ndjson`) are local and redacted |
| **Paperclip** | Adapter `onLog` emits stdout/stderr/system chunks into `RunLogStore` NDJSON; result JSON and heartbeat events | Process/HTTP/plugin adapters are provider-specific behind one adapter contract; logs are provider-agnostic chunks | Local-file log is default; optional S3 mirror is transparent fallback on local loss; throttled in-flight mirror is opt-in | API/CLI heartbeat streams live logs and events; S3 finalize/in-flight mirror supports pod restart, but remote workspace preview is explicitly refused |
| **Herdr** | In-memory Ghostty terminal state; `pane.read` sources `visible`, `recent`, `recent-unwrapped`, `detection`; alt-screen reads scroll only while agent idle | Read API is agent-agnostic; agent metadata/session hooks are provider-aware | `pane read` defaults to recent text; passive subscriptions poll and emit output/agent-status/scroll events; tmux/terminal attach remains transport | SSH `--remote` runs a Herdr server on the target and forwards the local socket; no transcript-file protocol, so reads stay execution-host terminal state |
| **Gas Town** | tmux is universal transport; optional `gt agent-log` tails Claude `~/.claude/projects/*/*.jsonl` into OTel `agent.event`; `.events.jsonl` records lifecycle/feed audit | Transcript watcher is Claude-specific; runtime/provider integration is generic at config level and defaults to tmux | `GT_LOG_AGENT_OUTPUT=true` + OTel URL opt-in; otherwise screen/tmux and lifecycle hooks; restart-first session management and `gt prime` handoff are fallbacks | Feed/audit NDJSON has visibility levels and run-id correlation; no general SSH transcript reader (remote use is via tmux/host process) |
### Actionable patterns from references
1. **Make provider variation a narrow adapter contract (Overstory).** Keep
`resolveSessionFilePath`, line decoding, resume argv, and capability claims in a
registry; do not broaden a generic reader to guess foreign JSONL. Add a provider
conformance fixture whenever a new decoder is admitted.
2. **Persist a bounded, cursorable canonical log (Paperclip).** Orca's released
archive already does this for workers. Extend the same contract to optional
in-flight mirroring (local append remains hot; mirror at a cadence and flush on
graceful shutdown) so a runtime/relay crash does not erase the tail before
release. Preserve `sourceIdentity` and `nextCursor` across object-store reads.
3. **Separate read intent/source and passive subscriptions (Herdr).** Keep
`terminal.read` (current screen vs accumulated tail) distinct from transcript
`workerRead`; expose an explicit passive/watch mode for fleet dashboards rather
than polling an interactive endpoint. Consider an `outputMatched` subscription
for readiness/blocker detection, with provider status as a separate stream.
4. **Correlate lifecycle and content with a run id (Gas Town).** Add a stable
dispatch/run correlation to archive records and observability events, and expose
feed-vs-audit visibility. Keep agent output opt-in/redacted; never put capability
tokens or secrets in logs.
5. **Treat execution-host reads as authoritative.** Herdr's remote mode and Orca's
federation both forward the operation to the host owning the PTY/filesystem. Do
not attempt desktop-side transcript path reads for SSH or WSL and do not infer
process death from a broken transport.
## Proposed implementation slices (DAG)
The slices are deliberately small and can land independently. Dependencies are
listed as `A -> B` (A must land first).
- **A — Contract inventory and fixtures.** Freeze the current source/fallback
vocabulary, cursor envelope, provider-session metadata, and archive record shape;
add one fixture per supported transcript agent plus an unsupported-agent fixture.
**Depends on:** none.
- **B — Provider adapter registry.** Introduce a typed registry that owns transcript
decoder, resolver roots, lifecycle decoder, and `requiresLocalTranscript`; wire
existing functions through it without changing behavior. **A -> B.**
- **C — Durable worker-log mirror.** Add an optional local-file/object-store mirror
for in-flight worker archives, with cadence, graceful flush, stale-upload fencing,
bounded range reads, and the existing source identity in the cursor. **A -> C.**
- **D — Passive output watch RPC.** Add a capability-negotiated stream/subscription
for worker output/status (append frames + gap/replacement + liveness reason),
keeping `workerRead`/`terminal.read` unary semantics unchanged. **A, B -> D.**
- **E — Federation compatibility matrix.** Add a host capability bit for structured
worker output and tests covering new peer, old peer, relay loss (`unverifiable`),
WSL refusal/recovery, and source/cursor changes. **A, D -> E.**
- **F — Fleet observability projection.** Project dispatch/run id, host id/epoch,
source, cursor range, fallback reason, warnings, and liveness into redacted
diagnostics/notifications with feed-vs-audit policy. **C, D, E -> F.**
## Non-goals and guardrails
- Do not make every provider's transcript format look identical, or infer a
transcript path from an untrusted session id when the provider does not publish
one.
- Do not replace PTY reads with transcript reads for interactive terminal UX;
alternate-screen/current-frame semantics and ANSI output remain terminal-owned.
- Do not read SSH/WSL files from the desktop process, copy whole transcripts over
the wire, or treat relay loss as process exit.
- Do not add a new stream opcode without capability negotiation and an old-client
fallback; optional fields are preferred for mixed-version peers.
- Do not log raw prompts, tool payloads, image paths, dispatch capabilities, or
provider credentials. Keep limits and redaction at the execution boundary.
- Do not require Git/tmux or a git worktree for folder workspaces; transcript
resolution is runtime-relative.
## Verification tests
Existing high-value coverage includes:
- `src/main/native-chat/transcript-reader.test.ts`, `transcript-tail-reader-cancellation.test.ts`,
`transcript-watch-{resolve-poll,unflushed-settle,liveness,error,unsubscribe-race}.test.ts`;
`session-file-resolver-{codex-roots,wsl,wsl-scan-gate}.test.ts`; and
`host-readable-transcript-path*.test.ts` for bounded parsing, delayed flush,
replacement, cancellation, WSL ranking, and refusal recovery.
- `src/main/runtime/rpc/methods/orchestration-worker-output.test.ts`,
`orchestration-federation-output.test.ts`, `orchestration-federation.test.ts`,
and release/archive tests for source selection, redaction, cursor fencing,
old-peer fallback, archive reads, and liveness semantics.
- `src/main/runtime/rpc/methods/terminal-read-screen-cursor.test.ts` and
`src/cli/handlers/terminal.ts` characterization tests for screen/cursor
validation and mixed-version `source` detection.
- `src/main/providers/ssh-pty-provider*.test.ts` and relay PTY publication/
differential tests for incarnation, reconnect, notification gaps, flow control,
and provider-generation teardown.
Tests required for the proposed slices:
1. Registry conformance: each supported decoder parses fixture lines; unsupported
agents produce `provider_unsupported`; OpenClaude shares Claude format without
changing identity; malformed/oversized lines are bounded and warned.
2. Cursor/source races: append while reading, truncation/rotation, process
replacement, provider-session path change, archive handoff, and old cursor reuse
must return `source_changed` rather than duplicate or cross-lane data.
3. Cross-host matrix: local, SSH, federated-new, federated-old, WSL guest path,
cold/stopped distro, relay disconnect/reconnect; assert execution-host reads,
`live`/`unverifiable`/`exited`, and no false stop success.
4. Durability: in-flight mirror cadence, failed upload retry, graceful flush,
finalize-vs-stale-upload ordering, local-missing object fallback, range cursor,
and response-size/redaction limits.
5. Passive watch: initial snapshot, append, gap/replacement, unsubscribe and sender
teardown, capability-negotiated downgrade, and notification backpressure.
6. Fleet projection: stable dispatch/run/host correlation, feed-vs-audit filtering,
no secrets/capabilities in diagnostics, and dashboard behavior when liveness is
`unverifiable`.
## Reference paths consulted
- Orca: `src/main/native-chat/*`, `src/main/runtime/orchestration/worker-transcript-*`,
`src/main/runtime/rpc/methods/{orchestration-worker-output,orchestration-worker-control,orchestration-federation-control,terminal/terminal-query-methods}.ts`,
`src/cli/handlers/terminal.ts`, `src/main/providers/ssh-pty-provider*.ts`,
`src/shared/{native-chat-agent-support,agent-session-resume,pty-liveness-verdict}.ts`,
and `docs/reference/{ssh-execution-boundary,remote-wire-compatibility,wsl-command-execution}.md`.
- Overstory: `docs/runtime-adapters.md`, `docs/runtime-abstraction.md`,
`docs/direction-ui-and-ipc.md`, `docs/headless-hooks-design.md`,
`src/runtimes/*.ts`, `src/metrics/transcript.ts`, `src/events/*`, `src/logging/*`.
- Paperclip: `server/src/services/run-log-store.ts`, `run-liveness.ts`,
`heartbeat-run-summary.ts`, `adapters/process/execute.ts`, `adapters/types.ts`,
`cli/src/commands/heartbeat-run.ts`, and `server/src/__tests__/heartbeat-*.test.ts`.
- Herdr: `src/api/schema/{panes,agents,common}.rs`, `src/server/{headless,alt_screen_read}.rs`,
`src/api/subscriptions.rs`, `src/terminal/history_read.rs`, `src/remote/attach.rs`.
- Gas Town: `internal/session/{lifecycle,agent_logging_unix}.go`,
`internal/events/events.go`, `docs/agent-provider-integration.md`,
`docs/design/polecat-lifecycle.md`, and `docs/otel-data-model.md`.
-103
View File
@@ -1,103 +0,0 @@
# Orchestration v3 lifecycle and persistence complexity audit
## Verdict
**Block on F1.** The new centralized lifecycle graph misses a production transition reached when a delayed PTY exit follows an uncertain stop. The rest of the audited design is conservative and generally transaction-safe, but worker composition, mutation idempotency, release convergence, and receipt correlation still carry avoidable branch/ledger complexity.
Scope was the current `orchestration-v3` working tree against `origin/main`, emphasizing lifecycle, composed worker start, SQLite persistence, Task/Dispatch/worker projection, receipts, release, and recovery. No tracked files were edited.
## Findings
### F1 — High — block: delayed exit after `stop_unknown` is rejected by the new lifecycle graph
Evidence:
- `src/main/runtime/orchestration/db/lifecycle-transition.ts:89-99` permits `stop_unknown -> stopped|abandoned`, but not `stop_unknown -> failed`.
- `src/main/runtime/orchestration/db/dispatch-context/dispatch-completion.ts:173-186` handles every positively observed worker process exit by transitioning the current worker state to `failed`.
- `src/main/runtime/orca-runtime.ts:19147-19175` routes a PTY exit for any still-active Dispatch through `failDispatch(..., { workerProcessExited: true })`.
- `src/main/runtime/orchestration/db/worker-dispatch/worker-dispatch-stop.ts:234-263` leaves the Dispatch active when stop outcome is unknown, so the later PTY-exit hook can reach exactly this combination.
- `src/main/runtime/orchestration/db-task-dispatch-lifecycle-guards.test.ts:132-147` tests process exit only from `ready`; the stop-unknown tests do not deliver a later positive exit.
Impact: after a lost stop response, a later real PTY exit throws `lifecycle_conflict` while the runtime is processing process death. The transaction rolls back, leaving the worker `stop_unknown`, the Dispatch active, and the Task blocked; it can also interrupt the remainder of exit cleanup because the authoritative `failDispatch` call is outside the best-effort notification `try` in `orca-runtime.ts:19177-19188`. This is a regression from `origin/main`, whose direct worker update allowed any nonterminal worker state to become `failed`.
Safe simplification: decide one canonical meaning and encode it once: either allow `stop_unknown -> failed` for the generic exit path, or special-case a positive exit after a stop request as `stop_unknown -> stopped` while settling the Dispatch/Task consistently. Add an integration test `ready -> stopping -> stop_unknown -> PTY exited` asserting no throw, revoked authority, terminal worker state, settled Dispatch, blocked/retryable Task, and one receipt per projection.
### F2 — Medium — simplify: concurrent identical `workerStart` calls bypass in-process coalescing
Evidence:
- `src/main/runtime/rpc/orchestration-mutation-executor.ts:82-101` treats worker start specially and performs only a DB lookup instead of atomically beginning the receipt.
- The existing `inFlight` promise is consulted only after a durable row is observed as `completed` or `pending` (`:109-146`). A call with no row proceeds without checking it.
- The promise is not installed until `:179-194`.
- Local composition performs asynchronous terminal/worktree validation before `createStartingWorkerDispatch` atomically inserts the acceptance receipt (`src/main/runtime/rpc/methods/orchestration-local-worker-start.ts:47-114`). Two same-request calls can therefore both classify themselves as `started`; one later wins DB acceptance and the other returns `operation_unknown` instead of coalescing on the first result.
- `src/main/runtime/rpc/orchestration-mutation-executor.test.ts:38-118` covers prompt receipt recovery, not concurrent atomic worker acceptance.
Impact: effects remain fenced by the DB receipt, so this is not a duplicate-worker bug, but an identical concurrent retry can receive an ambiguous failure while the original succeeds. That weakens the advertised idempotent composition contract and pushes avoidable inspection/recovery onto coordinators.
Safe simplification: consult/install `inFlight` before the special atomic-acceptance branch, or add an atomic acceptance-claim state that is created before asynchronous topology checks and can be safely discarded before effects. Add a `Promise.all` test for identical local and federated starts that expects one invocation/effect set and the same receipt with `replayed` metadata.
### F3 — Medium — simplify: release cannot converge automatically across close-success / DB-settlement loss
Evidence:
- Local release commits the archive/`releasing` state, closes the terminal, and only afterward marks the resource released (`src/main/runtime/rpc/methods/orchestration-worker-release-completion.ts:178-245`). A crash or SQLite error between `closeTerminal` and `settleWorkerTerminalRelease` leaves durable `releasing` state after the process is gone.
- Startup reconciliation selects `requested` and `releasing` rows (`src/main/runtime/orchestration/db/worker-terminal/worker-terminal-listing.ts:34-43`), but a missing/unattached terminal in recovery mode is returned as `release_pending` before process-incarnation liveness is consulted (`orchestration-worker-release-completion.ts:131-141`).
- The positive-exit convergence helper excludes the two backlog states: `settleDeadWorkerTerminalRelease` accepts only `not_requested|retained|unknown` (`src/main/runtime/orchestration/db/worker-terminal/worker-terminal-release.ts:133-152`).
- The interactive handler performs process-incarnation liveness convergence only for a `retained` disposition (`src/main/runtime/rpc/methods/orchestration-worker-release.ts:46-70`), not for a requested/releasing backlog row.
Impact: the safe close already happened, but the durable resource can remain pending/unknown indefinitely. The design correctly refuses to infer death from missing inventory; the gap is that it also fails to consume a later *positive* `exited` verdict for the exact process incarnation.
Safe simplification: on missing/unattached recovery, query exact process-incarnation liveness; when it is positively `exited`, allow an atomic `requested|releasing -> released` settlement using the already committed archive. Retain current pending behavior for `live`/`unverifiable`. Add a crash-seam test that injects failure immediately after successful close, reopens the DB, supplies `exited`, and converges without a second close.
### F4 — Medium — simplify: local and federated worker composition duplicate the same state machine
Evidence:
- Dependency parsing is duplicated verbatim in `src/main/runtime/rpc/methods/orchestration-local-worker-start.ts:274-289` and `orchestration-federated-worker-start.ts:35-51`.
- Local start separately assembles start options, creates the pending Task/Dispatch/worker transaction, records topology effects, waits for readiness, attaches authority, sends the preamble, and marks input accepted (`orchestration-local-worker-start.ts:79-270`).
- Home-side federated start repeats task/start-option normalization and pending acceptance (`orchestration-federated-worker-start.ts:135-172`), while the execution host repeats setup/readiness, authority, preamble delivery, effect recording, and failure classification (`src/main/runtime/rpc/methods/orchestration-federation.ts:198-297`).
Impact: receipt vocabulary and legal stages must stay aligned across three orchestration paths. The code already shows patch pressure around prompt budgets, setup evidence, launch preferences, and unknown outcomes; another stage or receipt field currently requires multi-file synchronization and parallel tests.
Safe simplification: keep host-specific effects separate, but extract one pure normalized start plan and one host-side `materialize -> ready -> attach authority -> deliver -> settle` driver parameterized by placement/transport adapters. Do not merge local and remote liveness policies or infer remote exit from transport loss.
### F5 — Low/Medium — defer then simplify: the lifecycle boundary contains a delivery-error exception
Evidence:
- The generic lifecycle primitive exposes `correction: 'unobserved_prompt_report'` (`src/main/runtime/orchestration/db/lifecycle-transition.ts:51-62`) and bypasses terminal-state legality for failed Task/Dispatch/worker rows (`:167-179`).
- Worker settlement imports `AGENT_PROMPT_STALLED_ERROR` from prompt-submission verification and has a large dedicated reopen branch (`src/main/runtime/orchestration/db/dispatch-context/worker-report-settlement.ts:3,90-121,165-254`).
- Current composed worker start uses queued acceptance with zero observation wait (`src/main/runtime/rpc/methods/orchestration-local-worker-start.ts:226-237`), and the runtime returns an honest input-accepted receipt rather than throwing when turn start is not observed (`src/main/runtime/orca-runtime.ts:21815-21894`).
Impact: the core state graph knows about one upper-layer delivery implementation error and can reopen otherwise terminal states. The branch is needed for legacy/pre-update rows, but it should not be the long-term normal lifecycle model.
Safe simplification: keep this compatibility path for existing failed-stall rows, but isolate it in a named legacy-repair operation with explicit provenance. New writes should represent uncertainty as `start_unknown/outcome_unknown`, which ordinary authenticated reports can settle without reopening terminal states. Remove the exception only after migration/compatibility coverage proves no supported row still depends on it.
### F6 — Low — keep semantics, simplify audit correlation later: four durable ledgers describe one worker report
Evidence:
- `mutation_receipts`, `lifecycle_transition_receipts`, and `attempt_observation_facts` are separate tables (`src/main/runtime/orchestration/db/schema/create-core-tables-sql.ts:95-147`), in addition to the durable message row/delivery.
- A worker report writes Task, Dispatch, and worker lifecycle receipts plus a worker-report observation (`src/main/runtime/orchestration/db/dispatch-context/worker-report-settlement.ts:167-294`). Lifecycle receipts have no causation/message-id column (`src/main/runtime/orchestration/db/lifecycle-transition.ts:122-130,219-247`).
- Atomic rollback and replay are well covered in `src/main/runtime/rpc/orchestration-commit-notify-characterization.test.ts:193-440`; the complexity is operability/correlation, not transaction safety.
Disposition: **keep** the separate purposes and existing atomic transaction. Later, add an optional causation ID to lifecycle receipts and populate it with the message/mutation ID so one report can be queried across ledgers; do not add a fifth event store or replace the projections in this PR.
## High-value missing tests
1. `stop_unknown` followed by positive PTY exit (F1) — required before merge.
2. Concurrent same-request local and federated `workerStart` coalescing (F2).
3. Crash after terminal close but before release settlement, then exact `exited` recovery (F3).
4. Table-driven coverage for every production transition edge used by `failDispatch`, stop, abandon, retry, start-unknown reconciliation, and worker report. The current lifecycle unit test has only two generic cases (`src/main/runtime/orchestration/db/lifecycle-transition.test.ts:4-49`).
5. Expand the direct-writer ratchet beyond its hard-coded eight-file list (`src/main/runtime/orchestration/db/lifecycle-transition-boundary.test.ts:5-23`) or replace it with a recursive production-source scan. Remote attachment lifecycle remains a parallel direct-SQL state machine in `remote-dispatch-attachment-authority.ts:116-160`, `remote-dispatch-attachment-stop.ts:8-66`, and `federation-relay-item.ts:23-57`; characterize stop/report/start races before centralizing it.
## Positive dispositions
- **Keep:** `BEGIN IMMEDIATE` ownership for read-before-write Task/Dispatch failures (`dispatch-completion.ts:109-218`) and nested-savepoint behavior; it addresses the WAL snapshot-upgrade race without weakening CAS.
- **Keep:** worker-done message, lifecycle projection, attempt fact, and mutation receipt in one transaction before notification (`orchestration.ts:905-914` plus the commit-notify characterization tests).
- **Keep:** conservative exact pane/process/host checks before terminal close (`orchestration-worker-release-completion.ts:95-129,194-245`) and `live/unverifiable/exited` separation.
- **Keep:** same-outcome duplicate worker reports preserving the first canonical Task result. The current contract explicitly accepts duplicate same-outcome federated reports even when bodies differ (`orchestration-federation-lifecycle-settlement.test.ts:614-656`); document that first-result-wins rule rather than adding content-based settlement branches.
## Verification note
A focused Vitest invocation covering lifecycle transitions, Task/Dispatch guards and races, mutation execution, and release recovery was started. The environment was concurrently running many other Vitest pools; the tool yielded after progress dots and the process later exited, but its final summary/exit status was not retained, so this audit does not claim that run as verification evidence. Static evidence for F1 is a direct production call-chain/transition mismatch and does not depend on a failing existing test.
-156
View File
@@ -1,156 +0,0 @@
# Orchestration-v3 complexity audit
Scope: read-only audit of activation/recovery, cold-park mailbox delivery,
federation/SSH authority, mixed-version behavior, and transcript/terminal-read
paths at `4a32d073fa`. Existing untracked audit files and five unrelated tracked
changes were preserved; no tracked source edits were made.
## Executive verdict
Most of the added state machines are justified by concrete failure modes: the
execution host remains authoritative, lifecycle recovery is transactional, and
mailbox/output reads are fenced by process/session identity. Two issues should
block promotion: the gated activation callback ignores `providesInitialSurface`,
and federation settlement mode is selected from persisted attachment protocol
even after a remote runtime can restart at an older capability level. Fleet and
relay time budgets are safe in outcome vocabulary but not budget-compliant under
retry/backlog; simplify or measure them before claiming latency guarantees.
## Findings
### A1 — Block: async activation can create an unwanted shell
The synchronous activation path honors the caller's non-terminal-surface opt-out
(`src/renderer/src/lib/worktree-activation.ts:274-287`). The gated callback instead
always invokes `ensureWorktreeHasInitialTerminal(..., { reseedEmptiedWorkspace: true })`
(`src/renderer/src/lib/worktree-activation.ts:257-269`), and the folder callback
does not forward the option at all (`:147-160`). When inventory hydration is in
flight, a file/diff navigation with `providesInitialSurface: true` can therefore
reseed a closed-last-terminal workspace after the gate resolves. Existing tests
prove the synchronous opt-out for worktrees and folders
(`src/renderer/src/lib/worktree-activation-emptied-workspace-reseed.test.ts:83-96,227-238`)
and gate hydration ordering (`src/renderer/src/lib/worktree-agent-activation-gate.test.ts:159-215`),
but not their combination. Add an async gated test for both workspace shapes that
holds inventory unresolved, activates with `providesInitialSurface: true`, resolves
to `empty`, and asserts the tab row remains empty; preserve the option through the
callback. This is a user-visible regression, not a speculative edge.
### A2 — Keep: lifecycle start-unknown and direct-write fixes are evidence-backed
Worker reports explicitly reconnect a blocked Task when its worker is
`start_unknown`, then settle Dispatch/Task/Worker through the lifecycle boundary
(`src/main/runtime/orchestration/db/dispatch-context/worker-report-settlement.ts:108-121,198-260`).
Start failures conservatively persist `start_unknown` (`src/main/runtime/orchestration/db/worker-dispatch/worker-dispatch-outcome.ts:112-154`).
Decision-gate and context-only release paths use `transitionLifecycleWithDb`, and
the boundary ratchet covers them (`src/main/runtime/orchestration/db/decision-gates/decision-gate-store.ts:84-90`,
`src/main/runtime/orchestration/context-only-dispatch-release.ts:39-75`,
`src/main/runtime/orchestration/db/lifecycle-transition-boundary.test.ts:5-23`).
Keep; earlier reviews that called start-unknown completion a blocker are stale at
this commit.
### A3 — Keep / simplify: cold-park mailbox mechanism vs fixed delay
Pointer delivery durably stages the pointer and delayed Enter, then revalidates
exact PTY, process incarnation, mailbox ownership, writability, and live idle state
before writing (`src/main/runtime/orchestration/mailbox-pointer-stage.ts:75-180`,
`src/main/runtime/orchestration/mailbox-pointer-submit.ts:62-153`). Restart recovery
settles conservatively without replaying an ambiguous Enter, and focused tests cover
delayed write, same-incarnation idle, explicit-check races, and duplicate suppression
(`src/main/runtime/orchestration/mailbox-pointer-submit.test.ts:82-153,325-381`,
`src/main/runtime/orchestration-mailbox-notification-consistency.test.ts:554-610`).
Keep this complexity: it protects real duplicate/lost-mail cases. The 500 ms
pointer-to-Enter constant (`mailbox-pointer-delivery.ts:17-23`) has no production
latency evidence; measure agent startup distributions or make the delay adaptive
before treating it as a contract (simplify/defer, not a correctness block).
### A4 — Keep: SSH/federation authority and liveness vocabulary
Remote attachment observation uses the execution host's verdict and maps absent
contact/verdict to `unverifiable`, never `exited`
(`src/main/runtime/rpc/methods/orchestration-federation-attachment-observation.ts:21-75`).
Federated operations pin peer fingerprint and pairing revision before every call
(`src/main/runtime/rpc/methods/orchestration-worker-observation.ts:172-194`,
`src/main/runtime/orchestration/federation-sync.ts:63-90,154-199`); release retains on
identity uncertainty. This matches `docs/reference/ssh-execution-boundary.md` and
is covered by liveness/transport safety suites. Keep; local terminal-connected
heuristics must not replace host authority.
### A5 — Block: settlement capability is stale across a remote downgrade/restart
`syncFederatedDispatchPages` decides `supportsLifecycleSettlement` solely from the
persisted attachment protocol (`src/main/runtime/orchestration/federation-sync.ts:75-86`),
and host ACK validation does the same (`src/main/runtime/orchestration/db/federation/federation-relay-ack.ts:80-99`).
The peer capability cache observes runtime epochs, but federation sync does not use a
status capability probe for this decision. A v3 attachment that survives a remote
restart/downgrade can make the new home send `replayUnacknowledged` and settlement
payloads to an old parser; zod strips those unknown fields
(`src/main/runtime/rpc/methods/orchestration-federation-relay.ts:13-18,20-42`),
so the remote may ACK relay rows without applying the lifecycle settlement or may
return a non-replayed sequence that the v3 home interprets under the wrong mode.
This violates the mixed-version rule that behavior changes require capability
negotiation (`docs/reference/remote-wire-compatibility.md:25-29,61-76`). Existing
tests cover protocol at attach time and epoch cache invalidation, but not one
attachment whose runtime capability changes. Add a two-runtime regression that
persists protocol 3, reports a new epoch with no settlement capability, and asserts
safe legacy mode (or `unverifiable`/retry without settlement mismatch). Block until
that scenario is handled or explicitly fenced.
### A6 — Simplify/defer: fleet and relay timing are not justified guarantees
Fleet advertises a 5 s total / 3 s host budget with concurrency four
(`src/main/runtime/rpc/methods/orchestration-federated-fleet-snapshot.ts:16-18,49-99`),
but capability resolution may retry once after an epoch race using the original
3 s probe timeout. Two probes can consume about 6 s before the snapshot call; pass
the remaining deadline into retries and add a fake-timer budget test. Relay sync
recurses through six 50-item pages with independent 15 s pull/ACK/import timeouts
(`src/main/runtime/orchestration/federation-sync.ts:23-24,79-90,154-165,194-207`),
allowing roughly 90 s and delaying fresh lifecycle reports. Carry one deadline
through pages or lower per-page timeout. These paths fail closed as host-unavailable
or unverifiable, so classify as simplify/defer timing work rather than authority
blocks.
### A7 — Keep: transcript/terminal reads are source- and host-fenced
Exact reads require a hook-attested transcript path and SSH filesystem provider;
desktop filesystem lookup is not attempted for remote sessions
(`src/main/runtime/orchestration/worker-transcript-read.ts:62-78`,
`src/main/runtime/rpc/methods/orchestration-worker-output.ts:39-71`). Local and remote
readers revalidate inode/size, boundary checkpoints, and provider identity
(`worker-transcript-local-read.ts:46-75,87-145`,
`worker-transcript-remote-read.ts:82-111`). Cursors pin dispatch/source identity,
and legacy federated peers reject transcript cursors rather than silently switching
to terminal output (`src/main/runtime/rpc/methods/orchestration-worker-legacy-federated-read.ts:17-47`).
Terminal fallback is explicitly `sourceExact:false/contentComplete:false`
(`src/main/runtime/rpc/methods/orchestration-worker-output.ts:211-220`). Keep these
guards; they address wrong-host and stale-process user failures.
### A8 — Simplify/defer: clipping metadata and pathological scans
Local initial transcript reads set `limited: page.hasMore` but return no explicit
non-pageable clipping warning (`src/main/runtime/orchestration/worker-transcript-local-read.ts:53-75`),
unlike the remote bounded path (`worker-transcript-remote-read.ts:205-213`). Add a
regression asserting `contentComplete`, `clipping`, and warning parity, plus CLI
`terminal read --screen` old-host formatting coverage. Defer stronger source hashing
for same-size rewrites with coarse mtime and a cap for files containing only malformed
records; those are theoretical and lack user-impact evidence.
## Regression suites and evidence
The focused audit run passed 8 files / 55 tests, covering mailbox cold-park and
restart recovery, lifecycle guards, peer capability cache, federation output and
transport safety, local/remote transcript reads, archive reads, and terminal-read
cursor behavior. Federation/SSH sub-suite independently passed 4 files / 35 tests;
transcript/terminal sub-suite passed 4 files / 34 tests. No tracked source edits were
made. Before release, add the A1 gated-surface test and A5 persisted-protocol/runtime-
capability downgrade test; use fake timers for A6 budget assertions and a real
two-build federation/relay harness (deferred infrastructure) before removing fallbacks.
## Classification summary
| Area | Keep | Simplify / defer | Block |
| --- | --- | --- | --- |
| Activation / recovery | lifecycle correction and recovery fencing | 500 ms delay measurement | A1 async `providesInitialSurface` regression |
| Cold-park mailbox | durable pointer + incarnation revalidation | adaptive delay / remote repark test | — |
| Federation / SSH | host authority, peer pinning, `unverifiable` | fleet/relay deadlines; two-build harness | — |
| Mixed-version settlement | explicit fallbacks elsewhere | — | A5 capability stale after downgrade |
| Transcript / terminal read | provider/path/source fences, cursor identity | clipping metadata; pathological scan tests | — |
@@ -1,45 +0,0 @@
# Dispatch failure race repair
## Outcome
Fixed the deterministic `database is locked` failure in `db-task-dispatch-races.test.ts` without weakening the stale-failure assertion, lifecycle compare-and-swap, receipt atomicity, or SQLite locking.
## Root cause
The lifecycle conversion changed `failDispatch` from an UPDATE-first savepoint to a read-then-update sequence. A top-level SQLite savepoint starts a deferred transaction, so the first connection acquired a read snapshot; when the fixture committed `settleWorkerReport` through a second WAL connection before the first lifecycle UPDATE, SQLite rejected the snapshot-to-writer upgrade as `SQLITE_BUSY_SNAPSHOT` (surfaced by `node:sqlite` as `database is locked`). `busy_timeout` does not resolve this class of snapshot upgrade, and retrying only the UPDATE would break the transition receipt/projection transaction.
## Fix
- `failDispatch` now uses `BEGIN IMMEDIATE` when it owns the outer transaction, reserving the WAL writer before its first lifecycle read. Competing writers wait at the transaction boundary and the winner's committed state is re-read before any CAS.
- When a caller already owns a transaction, `failDispatch` retains savepoint nesting and never commits or rolls back the caller's work. The sync SQLite adapter exposes `DatabaseSync.isTransaction` for this distinction.
- Projection updates and lifecycle receipts remain in the same transaction/savepoint. The existing rollback-on-task-transition-failure test remains unchanged and passing.
- The report-wins race fixture now commits the second connection immediately before the first connection executes `BEGIN IMMEDIATE`, a serializable boundary interleaving. It still asserts the completed task/dispatch/worker outcome and revoked capability, and now also asserts that no `dispatch_failed` receipt was written.
- A nested-transaction regression confirms `failDispatch` leaves the caller transaction open and that caller rollback removes the dispatch/task projections and lifecycle receipt together.
## Verification
- Reproduction before fix: `pnpm test src/main/runtime/orchestration/db-task-dispatch-races.test.ts` failed 1/3 at `transitionLifecycleWithDb(...).run(...)` with `Error: database is locked`.
- Stress after fix: race file passed 30/30 isolated iterations (4 tests per iteration).
- Final focused run: race file plus sync-database adapter tests passed 15/15 tests across 2 files.
- Related lifecycle/invariant run passed 88/88 tests across 8 files:
- `db/lifecycle-transition.test.ts`
- `db/lifecycle-transition-boundary.test.ts`
- `db-task-dispatch-invariant.test.ts`
- `db-task-dispatch-lifecycle-guards.test.ts`
- `lifecycle-reconciliation.test.ts`
- `dispatch-failure-idempotency.test.ts`
- `db-heartbeat-straggler-guard.test.ts`
- `orchestration-worker-dispatch-db.test.ts`
- `pnpm tc`: passed.
- `pnpm run check:code-quality:changed`: passed with 0 new findings across 127 changed files, including type-aware checks and React Doctor.
- `git diff --check`: passed with no output.
## Files modified
- `src/main/runtime/orchestration/db/dispatch-context/dispatch-completion.ts`
- `src/main/runtime/orchestration/db-task-dispatch-races.test.ts`
- `src/main/sqlite/sync-database.ts`
- `src/main/sqlite/sync-database.test.ts`
- `orchestration-db-task-dispatch-race-fix-report.md`
No work remains for this task.
@@ -1,407 +0,0 @@
# Orchestration CLI ergonomics evidence gate
## Decision status
**Closed: insufficient independent-trial evidence; no aliases or renames.** This
artifact froze the protocol, variants, answer key, and compatibility gate before
results were available. No fresh participant trial could be scheduled within
this gate, so behavioral thresholds are unmet by definition; repo-native static
evidence is reported separately and is not presented as a user study.
The benchmark targets the five workflows required by the vNext audit: one
worker, fan-out/join, ask/resume, a prompt queued during a busy turn, and
uncertain remote recovery. It compares discoverability without weakening the
current authority, idempotency, execution-host, or mixed-version contracts.
## Variants
Each participant sees one arm only. Do not describe another arm or reveal the
answer key before the trial.
| Arm | Surface | Purpose |
| ------------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- |
| A — current | Current source grammar and current terse output/help. | Baseline. |
| B — wording | Identical grammar, with the evidence wording below and exact next commands. | Isolate wording from command-shape effects. |
| C — candidate | B wording plus the proposed `work start`, `wait --attention`, `status`, `worker-recover`, and `worker-retry` surface. | Test nouns/aliases before implementation. |
Arm A means the current worktree source, not the installed production binary.
The installed app used during protocol preparation predates
`worker-start --spec` and terminal prompt receipts, while this worktree already
contains those additive changes. Participant cards are therefore authoritative;
do not let an older installed `orca --help` silently redefine the arm.
### Arm B wording
- Busy prompt accepted: `Queued for the agent's next turn. No resend is needed.`
- Ask timeout: `Question is still pending. Resume waiting; do not ask again:`
followed by the exact `ask --resume <message_id>` command.
- Remote contact loss: `Connection lost; process status is unverifiable. Do not
retry until the old worker is proven stopped or explicitly abandoned.`
- Delivery wait: `--types chooses which new message wakes this wait. Orca still
returns the oldest complete replayable delivery; process all rows before
acknowledging it.`
### Arm C candidate definitions
The candidate card must define behavior rather than rely on suggestive names:
- `orchestration work start --spec ...` means atomic Task creation plus
`worker-start`; it does not replace standalone Task creation for DAGs.
- `orchestration wait --attention` means wait on the current Run for questions,
escalations, accepted completions, proven failures, or an uncertain outcome.
It returns the oldest complete replayable delivery and does not hide older
rows. It does not enable desktop notifications, focus a terminal, or inspect
all Runs.
- `orchestration status` means a read-only, current-Run, attention-first fleet
projection; `--tree` adds dependency/parent edges.
- `worker-recover` is read-only diagnosis. It never reconnects, stops, abandons,
or retries by itself.
- `worker-retry --dispatch <id>` creates a linked new attempt only after the old
outcome is settled and still requires explicit placement.
These are semantic proposals, not assumptions that the existing alias resolver
can implement them.
## Trial setup
- Use fresh independent agent sessions with no prior orchestration-vNext
transcript. Minimum useful sample: six participants per arm; nine per arm is
preferred. If fewer than six complete an arm, report descriptive results only
and do not approve aliases.
- Balance provider/model across arms. Randomize scenario order with a Latin
square. Keep IDs, prose lengths, topology, and simulated response delays
identical between arms.
- Give the participant its arm card and one scenario at a time. It may request
command help. Record each help request before returning only that command's
arm-specific help.
- The trial is a command-planning simulation. Do not execute mutations against
the active development Run. The participant emits commands and a one-sentence
prediction of each command's effect; the harness returns the scripted receipt.
- Start the timer when the complete scenario becomes visible. Stop
time-to-first-correct-command when the first command that is both syntactically
valid for that arm and semantically safe for the fixture is emitted.
- Continue until the scenario reaches its terminal condition or the participant
requests coordinator intervention. Cap each scenario at 10 minutes and 12
participant commands.
## Scenario cards
Use the text in this section verbatim except for arm-specific command-card
references.
### S1 — one supervised worker
> You are the coordinator of bound Run `run_bench`. Start one fresh Codex worker
> in the current folder workspace to “review authentication fallback handling”.
> The current workspace must be reused; do not create a worktree or a second Run.
> After the scripted accepted completion arrives, preserve its readable output
> and account for the worker terminal. Show every command you would run and state
> what it will do before running it.
Scripted facts: start succeeds with Task `task_auth` / Dispatch `dsp_auth`;
`check` later returns delivery `dlv_auth` containing accepted `worker_done`;
release archives then closes only the fresh owned agent terminal.
### S2 — fan-out and join
> In bound Run `run_bench`, create two independent Tasks, `api` and `ui`, then an
> `integrate` Task that depends on both. Start `api` with Codex and `ui` with
> Claude in the current git worktree before waiting. Start `integrate` only after
> both predecessors complete. Do not create another Run, another worktree, or
> infer success from terminal silence. Show every command and predicted effect.
Scripted facts: Task creation returns `task_api`, `task_ui`, and `task_join`;
the first delivery completes `api` and asks a question from `ui`; replying makes
`ui` complete in the next delivery; `task-list --ready` then exposes the join.
### S3 — ask timeout and resume
> You are dispatched worker `dsp_ask`. Ask your coordinator whether the legacy
> flag should be preserved, offering `preserve,remove`, and wait up to 30 seconds.
> The first wait times out with pending message `msg_ask`; the answer arrives
> later. Continue waiting without creating a second question. Show every command
> and predicted effect.
### S4 — busy-turn queued prompt
> Agent terminal `term_busy` is proven to be in the middle of a turn. Send
> “Keep the Git 2.25 fallback.” with Enter and observe for up to 10 seconds
> whether the provider submission occurs. The first receipt says the prompt is
> accepted and queued, but submission is not yet observed. A simulated transport
> disconnect then makes the next result ambiguous. Continue safely without
> adding the text or Enter twice. Show every command and predicted effect.
Scripted facts: the first receipt issues request `req_busy`; replay/observation
of that request later proves one submission and one new turn.
### S5 — uncertain remote recovery
> Dispatch `dsp_remote` was working on saved environment `linux-ci` when contact
> was lost. Its last process verdict is `unverifiable`; the remote process may
> still be live. Determine the state and present safe actions. The operator then
> chooses “retry only after the exact old worker is proven stopped”. Do not run a
> local fallback, kill by pane/title, or start a duplicate attempt while the old
> outcome is unknown. Show every command and predicted effect.
Scripted facts: diagnosis remains `unverifiable`; exact `worker-stop` later
returns proven `stopped`; the linked retry must explicitly name `--on linux-ci`,
an exact remote worktree selector, and an agent.
## Current-command answer key
Equivalent quoting and optional `--json` are accepted. Extra read-only
inspection is counted but not failed unless it changes the target or treats
advisory output as proof.
| Scenario | Required current commands / decisions |
| -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| S1 | `orca orchestration worker-start --spec "review authentication fallback handling" --worktree current --agent codex`; `check --wait --types worker_done,escalation,question`; after processing accepted completion, `worker-release --dispatch dsp_auth`; acknowledge only after the release decision. |
| S2 | Three `task-create` calls, with `--deps '["task_api","task_ui"]'` on the join; start both ready branches before waiting; process every delivery, use `reply --id <question_id>` for the question, acknowledge whole deliveries, query `task-list --ready`, then `worker-start --task task_join ...`. No name/`--after` syntax exists in the baseline. |
| S3 | `ask --question ... --options preserve,remove --timeout-ms 30000`; after timeout, `ask --resume msg_ask --timeout-ms ...`. Repeating `--question` is a duplicate. |
| S4 | `orca terminal send --terminal term_busy --text "Keep the Git 2.25 fallback." --enter --wait-submit 10`; after an ambiguous result, repeat the exact payload with `--retry-request req_busy` (and optionally `--wait-submit 10`). A fresh send without the request ID is a duplicate risk. |
| S5 | `worker-show --dispatch dsp_remote`; do not infer exit or retry from contact loss; after the operator choice, `worker-stop --dispatch dsp_remote`, inspect the stop receipt, then `worker-start --task <original_task> --retry-of dsp_remote --on linux-ci --worktree <exact_remote_selector> --agent <agent>`. `worker-abandon` is a safe alternative only if the operator chooses to stop tracking without process action. |
Arm C accepts its defined composed commands when they preserve the same facts
and targeting. It does not accept `wait --attention` as a substitute for Task
creation, delivery acknowledgement, terminal release, or remote diagnosis.
## Capture template
Record one row per participant/scenario/arm. Preserve the timestamped transcript
or attach its path.
| Field | Definition |
| -------------------------- | ------------------------------------------------------------------------------------------------------------ |
| `ttfc_ms` | Prompt-visible timestamp to first correct command. |
| `commands_total` | Every emitted CLI command, including read-only and help. |
| `mutations_total` | Commands that would change orchestration/terminal state. |
| `retries` | Explicit recovery attempts, valid or invalid. |
| `duplicate_risks` | New question/message/prompt/worker attempts where an existing identity should have been resumed or observed. |
| `wrong_target_actions` | Any mutation aimed at another Run/Dispatch/host/workspace/terminal or an unsafe local fallback. |
| `help_lookups` | Number and command path of help requests. |
| `interventions` | Coordinator corrections needed to proceed. |
| `completed` | Scenario reached the answer-key terminal condition within limits. |
| `prediction_correct` | Participant correctly described each mutation before execution. |
| `attention_interpretation` | Verbatim answer plus coded category below. |
For all arms, finish with this comprehension probe before revealing results:
> Without running it, explain exactly what `orca orchestration check --wait
--attention` would watch, which Run(s) it covers, whether it filters returned
> rows, whether it acknowledges anything, and whether it changes desktop
> notifications or terminal focus.
Code the answer as:
- `U` — recognizes the flag is unsupported in the current arm.
- `W` — wake predicate on one bound Run, full oldest delivery still returned,
no acknowledgement/notification/focus side effects.
- `F` — believes attention rows are silently filtered from the returned batch.
- `A` — believes it acknowledges/clears attention.
- `N` — believes it controls native notifications, unread badges, or focus.
- `G` — believes it watches all Runs/global fleet state.
- `O` — other or incomplete; preserve verbatim text.
Arm A's only correct answer is `U`. Arm B should also be `U` unless its card
explicitly introduces the flag. Arm C's correct answer is `W` only if the
candidate is narrowed to the existing mailbox contract. The broader audit
definition that also wakes on synthesized uncertain worker state is not an
alias: score that answer separately as `W+state` and require an implemented,
capability-safe event source before promotion.
The Arm C definition in this protocol is the broader form, so `W+state` is its
expected answer. `W` is recorded as a narrower interpretation, not silently
accepted as equivalent.
## Analysis and decision rule
Report per arm and scenario: completion proportion; median and range for
`ttfc_ms`, commands, and help; totals for duplicate risks, wrong targets, and
interventions; and the attention interpretation distribution. Keep raw counts
beside percentages because the intended sample is small.
Wording (B) may be adopted when it causes no new wrong-target or duplicate-risk
actions and improves at least one error/comprehension measure. It does not need
to reduce command count because its purpose is truthful interpretation.
No alias/rename from C is approved unless all of these hold:
1. Completion is non-inferior to both A and B, with no wrong-target actions.
2. Median time to first correct command improves by at least 20% versus B in at
least three scenarios, including one of fan-out/join or remote recovery.
3. Total commands or help lookups improve in at least three scenarios without
increasing retries, duplicate risks, or interventions in any scenario.
4. At least 90% of participants interpret its scope and side effects correctly;
no participant mistakes `--attention` for acknowledgement or notification/
focus policy.
5. The compatibility review below classifies the exact implementation as safe.
If results are mixed, keep the canonical commands and adopt only the successful
wording/help changes. Do not combine B and C results or attribute wording gains
to aliases.
## Compatibility review before any implementation
| Candidate | Compatibility finding |
| ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Improved receipt/help wording | Additive and low risk if JSON fields, exit codes, stable status/verdict vocabulary, and exact recovery argv remain unchanged. Mixed-version output must state host capability honestly. |
| `work start` | Not a mechanical alias in the current parser: it changes command depth relative to `worker-start`, and existing alias-policy tests reject depth-changing aliases. Its proposed atomic semantics also do not cover DAG starts. Requires a real command handler and continued canonical support; benchmark gains must clear the full gate. |
| `wait --attention` | Not a safe alias to `check --wait --types ...` if it includes uncertain worker state. Current type filters only decide wakeup while returning the oldest full delivery. A broader predicate needs additive runtime/RPC capability negotiation and mixed-version fallback. |
| `status` | Cannot be a blind alias to `worker-list`: current worker-list includes terminal-resource accounting, while the proposal is current-Run, attention-first work/worker state with dependencies. Keep `worker-list` stable. |
| `worker-recover` | Can remain a new read-only composition over `worker-show` only if it performs no mutations and preserves `live` / `unverifiable` / `exited`. It must route reads to the execution host and degrade explicitly on old peers. |
| `worker-retry` | Not a synonym for `worker-start`: it must look up/link the prior Dispatch, reject unverifiable old liveness, require explicit host/worktree/agent placement, and preserve the existing `--retry-of` path for old clients. |
| Task names / `--after` | New Run-scoped identifiers and parsing, not aliases. IDs remain authoritative; old clients/hosts need ID-based fallback and duplicate/ambiguous names must fail before mutation. |
| Context-bound `done` | Authority-sensitive new behavior, not a rename. Zero or multiple active Dispatch roles must fail closed; the explicit `send --type worker_done --task-id ... --dispatch-id ...` contract remains canonical and mixed-version safe. |
## Repo-native discoverability and command-floor evidence
This section is objective implementation evidence, not a substitute for the
independent participant benchmark.
The worktree CLI was rebuilt successfully with `pnpm run build:cli`, then queried
through `node out/cli/index.js` so the measurements include the current source
rather than the older installed app.
### What the current registry exposes
- `agent-context` reports 233 total commands and 30 orchestration commands.
None of the scenario commands has an alias or an example. The only two
orchestration aliases are `run` and `run-stop`, both attached to retired
no-effect compatibility commands.
- `worker-start --help` exposes atomic `--spec`, but its usage has 27 flags and
eight notes. This already removes the Task-ID plumbing for S1, so `work start`
cannot reduce S1's command count.
- `task-create --help` exposes `--deps <json_array>` with no note or example.
Fan-out/join therefore has a specific documentation gap; changing the noun
does not teach dependency ordering.
- `ask --help` explicitly says a timeout leaves the question pending and to
resume with the original message ID. Its current non-JSON timeout output,
however, prints only the thread and elapsed time rather than the exact resume
command.
- `terminal send --help` explicitly states that `--wait-submit` observes without
resending and that ambiguous transport recovery reuses `--retry-request`.
Human output currently prints stage names, not the proposed plain-language
“No resend is needed” sentence.
- `worker-show --help` explains interactive-wait evidence but does not present
the `live` / `unverifiable` / `exited` recovery decision. Its human formatter
omits process liveness and ordered recovery actions, so S5 currently depends
on JSON or skill guidance.
- `worker-list` already prints projected attention categories, but an omitted
`--run` queries all Dispatches and the command is described as terminal
resource accounting. A proposed current-Run `status` view is therefore not an
alias-equivalent spelling.
- The global and command-specific help disagree about retired coordinator
commands: global help still describes `coordinator-start` as starting a legacy
loop, while orchestration help correctly says it is retired. Its intuitive
`orchestration run` alias performs no effects and returns skill recovery.
- `orchestration check --help` incorrectly borrows the screenshot flag label for
`--format`, displaying `--format <png|jpeg> Screenshot image format`; the same
help page's note correctly says it locally renders message rows. This is a
concrete wording defect that can be repaired without inventing a new noun.
### Unsupported-candidate behavior
Each proposed noun currently exits 1 as unknown. `work start`, `wait`,
`status`, `worker-recover`, and `worker-retry` print 365–366 lines of global
help with no nearest-command recovery. `check --attention --peek --json` fails
locally with `Unknown flag --attention`, lists 17 valid flags, and offers no
suggestion. This establishes a discoverability cost in the current CLI but does
not establish that these particular nouns are the remedy.
The current alias resolver canonicalizes an exact alias before validation and
dispatch, so same-semantics/same-depth aliases can safely share a handler. The
vocabulary-policy test deliberately rejects aliases at a different command
depth. `work start` is one token deeper than `worker-start`, and the candidate
recovery/status commands have different behavior, so none qualifies as a
mechanical alias under the existing compatibility mechanism.
### Static CLI command floor
| Scenario | Current A/B floor | Candidate C floor | Finding before participant timing |
| -------- | --------------------------------------------------: | ----------------------: | -------------------------------------------------------------------------------------------------------------------------------------- |
| S1 | 4: start, wait, release, acknowledge/continue | 4 | `work start` changes spelling only. |
| S2 setup | 5: create three Tasks, start two branches | 5 | Names/`--after` avoid ID extraction but do not remove a CLI command. A combined status view may later remove one read, not a mutation. |
| S3 | 2: ask, resume | 2 | Wording is the proposed intervention; no alias is needed. |
| S4 | 2: accepted send, identity-bound retry/observation | 2 | Wording is the proposed intervention; no alias is needed. |
| S5 | 3: show, exact stop, explicitly placed linked start | 3: recover, stop, retry | Candidate nouns do not reduce actions and must not hide placement or uncertainty. |
Thus the candidate arm has no pre-trial command-count advantage in four complete
scenarios or in S2 graph setup. Any case for it must come from measured time,
help, or targeting accuracy and still clear the compatibility gate.
The spelling-only savings are also uneven: `worker-start` is 12 characters and
`work start` is 10 but adds a command token; `worker-show --dispatch dsp_remote`
is three characters shorter than `worker-recover --dispatch dsp_remote`.
`wait --attention` is materially shorter than the 52-character explicit wait
filter, and a composed retry can be shorter while retaining explicit placement.
Those two cases remain semantic compositions, so timing/comprehension—not string
length—must justify them.
### Remote and mixed-version boundary
- Direct SSH loss disconnects Orca's client-resident control plane while the
execution-host PTY may remain live. A recovery noun cannot imply reconnect or
exit; only host evidence may change `unverifiable` to `live` or `exited`.
- Paired runtimes update independently. An optional response field is safe only
while every reader treats absence as unknown. New behavior that depends on a
field or event requires capability negotiation and a safe old-peer fallback.
- A broad `--attention` sent to an older host would otherwise be especially
hazardous: an old decoder may strip an unknown optional parameter, leaving a
new client waiting under ordinary message semantics while believing worker
outcome state is included. Promotion needs an advertised wait capability or
a client-side composition whose partial-host coverage is explicit.
- The current tree already has narrow capabilities for prompt delivery, worker
stop verdicts, federated fleet snapshots, structured remote reads, and remote
release. Candidate wording must expose their typed downgrade; it must not
silently run a remote operation on the desktop or reuse a local worker.
### Validation state observed during protocol preparation
The focused alias parser and vocabulary tests passed. The combined focused run
finished 77 of 81 tests green, with one live registry-parity failure because
`worker-cleanup` is exposed by the spec but absent from the handler registry,
plus three lifecycle-check failures where current shared-worktree settlement
could not find a worker row. These failures are outside the benchmark artifact,
but they mean the current merged tree is not a safe base for speculative command
surface changes until its existing registry/lifecycle integration is green.
The scenario-specific current grammar is independently green: the worker-start,
terminal-send, timeout/resume, worker-show wait, and durable prompt receipt suites
passed 49 of 49 tests. The failing integration checks above therefore do not
invalidate the S1/S3/S4/S5 answer-key syntax, but they still block promotion of a
new public surface in this shared tree.
## Results
No independent participant completed an arm (`n=0` for A, B, and C). The
coordinator explicitly chose to close this gate as insufficient evidence rather
than substitute a single evaluator's self-timed dry run.
| Required measure | Result | Decision use |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------- |
| Time to first correct command | Not measured; no valid participant timing sample. | Cannot support a new noun. |
| Command count | Static mutation/read floors measured above. Candidate C removes no command in S1, S3, S4, S5, or S2 setup. | No objective count benefit. |
| Retries / duplicate sends | Scenario-specific current receipt/resume suites pass; no participant error rate was measured. | Existing identities remain canonical. |
| Wrong-target actions | No participant rate measured. Remote boundary analysis shows a candidate retry/recover composition could be unsafe if it hides host or placement. | Requires behavioral and skew trials before promotion. |
| Help lookups | No participant rate measured. Current-source discovery exposes zero examples for the seven scenario commands; unsupported candidates dump 365–366 lines of global help. | Justifies testing focused help/wording, not aliases. |
| `--attention` interpretation | No participant comprehension sample. Current source rejects it; the narrow mailbox meaning and broad mailbox-plus-worker-state meaning require different implementations. | The flag remains unapproved. |
The evidence is sufficient to reject implementation now, not to claim that the
candidate labels are intrinsically worse. The follow-up protocol above remains
the promotion path: run balanced fresh sessions, preserve raw timestamps and
predictions, and apply the predeclared thresholds without combining wording and
alias effects.
## Final decision
Keep the current command grammar. Do not add `work start`, `wait --attention`,
`status`, `worker-recover`, `worker-retry`, Task-name/`--after`, or context-bound
`done` on this evidence. They either do not reduce the command floor or require
new semantics, authority checks, and mixed-version negotiation that cannot be
justified without behavioral gains.
No CLI noun, alias, flag, or command behavior was changed. A later wording-only
change may address the concrete help defects (`check --format`, retired
coordinator wording, exact ask-resume/no-resend guidance) after it is tested as
Arm B; it must preserve JSON, exit codes, stable verdicts, and emitted recovery
argv.
-215
View File
@@ -1,215 +0,0 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Orca orchestration vNext — PR decision brief</title>
<style>
:root { color-scheme: light; --ink:#17202a; --muted:#5d6873; --line:#dfe5ea; --bg:#f6f8fa; --card:#fff; --good:#087443; --warn:#9a5b00; --bad:#a12a2a; }
* { box-sizing:border-box; } body { margin:0; background:var(--bg); color:var(--ink); font:15px/1.55 -apple-system,BlinkMacSystemFont,"Segoe UI",sans-serif; }
main { max-width:1000px; margin:0 auto; padding:34px 22px 64px; } h1 { margin:0 0 5px; font-size:30px; letter-spacing:-.02em; }
h2 { margin:30px 0 9px; font-size:20px; } p { margin:8px 0; }
.sub,.small { color:var(--muted); } .small { font-size:13px; }
.decision { background:var(--card); border:1px solid #e7c98b; border-left:5px solid var(--warn); border-radius:10px; padding:15px 18px; margin:20px 0; }
.decision strong { color:var(--warn); } .grid { display:grid; grid-template-columns:repeat(auto-fit,minmax(210px,1fr)); gap:10px; }
.card { background:var(--card); border:1px solid var(--line); border-radius:9px; padding:12px 14px; } .label { color:var(--muted); font-size:12px; text-transform:uppercase; letter-spacing:.06em; }
.value { font-size:20px; font-weight:650; margin-top:3px; } .good { color:var(--good); } .warn { color:var(--warn); } .bad { color:var(--bad); }
table { width:100%; border-collapse:collapse; background:var(--card); border:1px solid var(--line); border-radius:9px; overflow:hidden; }
th,td { padding:9px 11px; border-bottom:1px solid var(--line); text-align:left; vertical-align:top; } th { background:#eef2f5; font-size:13px; } tr:last-child td { border-bottom:0; }
code { background:#edf1f4; padding:1px 4px; border-radius:4px; font-size:.91em; } ul { margin:7px 0 7px 20px; padding:0; }
.flow { display:flex; flex-wrap:wrap; gap:7px; align-items:center; margin:12px 0; } .node { background:var(--card); border:1px solid #cbd5dc; border-radius:7px; padding:8px 11px; } .arrow { color:var(--muted); }
.evidence { border-left:3px solid var(--line); padding-left:13px; } footer { margin-top:30px; color:var(--muted); font-size:13px; }
</style>
</head>
<body>
<main>
<h1>Orchestration vNext: whole-PR decision brief</h1>
<p class="sub">PR <a href="https://github.com/stablyai/orca/pull/16904">#16904</a> · report snapshot <code>4a32d073fa</code> · current workspace head <code>548ca8529f</code> · 30 Aug 2026</p>
<div class="decision">
<strong>Recommendation: do not treat this brief as merge approval yet.</strong>
The implementation has strong focused evidence, but this document's GitHub/CI snapshot is stale
(it was generated for <code>4a32d073fa</code>, while the workspace is now at <code>548ca8529f</code>).
Re-run repository CI and review the current diff before deciding. The local tests prove specific
behaviors; they do not prove the absence of bugs across every provider, OS, remote topology, or failure timing.
</div>
<div class="grid">
<div class="card"><div class="label">GitHub state</div><div class="value warn">STALE SNAPSHOT</div><div class="small">Must refresh against current PR head</div></div>
<div class="card"><div class="label">CI</div><div class="value warn">NOT CURRENT</div><div class="small">36 checks were recorded for the older snapshot</div></div>
<div class="card"><div class="label">PR size</div><div class="value">289 files</div><div class="small">23,272 additions / 4,162 deletions</div></div>
<div class="card"><div class="label">Review coverage</div><div class="value">7 areas</div><div class="small">lifecycle, mailbox, terminal, transcript, federation, cleanup, ergonomics</div></div>
</div>
<h2>What “technically ready” did—and did not—mean</h2>
<p>The earlier wording meant that the reviewed code paths had focused tests and no known blocker in that
snapshot. It was not a claim that the PR was bug-free, that current CI was green, or that every provider and
execution topology had been physically exercised. Because the snapshot predates the current workspace head,
that wording was too strong and has been withdrawn here.</p>
<table>
<thead><tr><th>Evidence level</th><th>What it establishes</th><th>What remains unproven</th></tr></thead>
<tbody>
<tr><td><span class="good"><b>Strong local evidence</b></span></td><td>Focused lifecycle, mailbox, federation, archive, and provider-transcript paths pass their targeted tests; the orchestration Electron matrix is green.</td><td>It cannot cover all OS/provider/remote combinations or arbitrary crash timing.</td></tr>
<tr><td><span class="warn"><b>Needs refresh</b></span></td><td>GitHub status and the CI counts shown in this document came from the older commit.</td><td>Current PR CI, current merge base, and any checks affected by later commits.</td></tr>
<tr><td><span class="warn"><b>Explicit product tradeoffs</b></span></td><td>Ambiguous mailbox writes fail closed; hook-only providers may not expose historical turn proof; composer auto-submit is not guaranteed.</td><td>Whether those boundaries are acceptable to the product owner.</td></tr>
</tbody>
</table>
<h2>Why this PR exists</h2>
<p>Orca’s original orchestration made it possible to coordinate agents, but delivery and lifecycle evidence
were too easy to lose, duplicate, or misinterpret. A coordinator could see a transport failure even when
text was queued, wake a busy agent twice, read the wrong transcript, or infer that a remote process had
exited when the host was simply unreachable.</p>
<p>This PR makes those outcomes durable and explicit while keeping agents responsible for decomposition,
placement, parallelism, and recovery choices.</p>
<h2>The model an agent actually uses</h2>
<div class="flow">
<span class="node"><b>Run</b><br><span class="small">coordination namespace</span></span><span class="arrow">→</span>
<span class="node"><b>Task</b><br><span class="small">piece of work + dependencies</span></span><span class="arrow">→</span>
<span class="node"><b>Dispatch</b><br><span class="small">one owned attempt</span></span><span class="arrow">→</span>
<span class="node"><b>Worker / terminal</b><br><span class="small">execution surface</span></span><span class="arrow">→</span>
<span class="node"><b>Delivery</b><br><span class="small">durable coordinator inbox</span></span>
</div>
<p class="small">These are storage and responsibility boundaries, not a new scheduler or workflow language. Existing low-level terminal/worktree commands remain available.</p>
<h2>What changed, grouped by responsibility</h2>
<table>
<thead><tr><th>Area</th><th>What the PR does</th><th>Benefit / boundary</th></tr></thead>
<tbody>
<tr><td><b>Mailbox durability</b></td><td>Persists message state before a disposable PTY pointer; replays the same Delivery until acknowledged; acknowledgments are generation-bound and idempotent.</td><td>No lost coordinator mail or duplicate processing after restart.</td></tr>
<tr><td><b>Prompt receipts</b></td><td>Reports distinct evidence for <code>input_accepted</code>, queued, submitted, and turn-started. FIFO claims use PTY/process generation and request baselines.</td><td>Agents can stop resending based on an ambiguous transport error.</td></tr>
<tr><td><b>Lifecycle ledger</b></td><td>Central guarded transitions, compare-and-set semantics, mutation request IDs, and durable lifecycle receipts.</td><td>Concurrent writers and retries cannot silently overwrite state.</td></tr>
<tr><td><b>Worker composition</b></td><td><code>worker-start/show/read/stop/abandon/release</code> compose existing worktree, setup, terminal, and dispatch primitives with explicit effects and residual resources.</td><td>One understandable supervised path without a hidden scheduler.</td></tr>
<tr><td><b>Crash/recovery</b></td><td>Unknown outcomes stay unknown; restart recovery is identity- and incarnation-fenced; completed output is archived before release.</td><td>Orca never “fixes” uncertainty by killing or duplicating work.</td></tr>
<tr><td><b>Transcripts</b></td><td>Reads are routed to the execution host and carry source identity, cursor, and completeness metadata; archived output remains readable after release.</td><td>Coordinator can tell “missing” from “not visible on this provider.”</td></tr>
<tr><td><b>Federation</b></td><td>Remote attachment authority, stale-generation fencing, capability caching, and host-owned liveness/release decisions.</td><td>Disconnect is <code>unverifiable</code>, not falsely <code>exited</code>.</td></tr>
<tr><td><b>Skill / CLI</b></td><td>Compact generated guidance teaches batch Delivery handling, exact handles, setup policy, wait semantics, recovery, and cleanup.</td><td>Less agent guesswork without introducing commands that are not implemented.</td></tr>
</tbody>
</table>
<h2>Review findings and decisions</h2>
<table>
<thead><tr><th>Finding</th><th>Disposition</th><th>Evidence</th></tr></thead>
<tbody>
<tr><td>Duplicate pointer text / naked Enter after a crash</td><td><span class="good"><b>Fixed.</b></span> Attempt phases are persisted before each external write; ambiguous restart recovery never emits pointer text or standalone <code>\r</code>.</td><td><code>mailbox-pointer-stage.ts</code>, <code>mailbox-pointer-submit.ts</code>, <code>mailbox-pointer-resume.ts</code>; focused pointer/DB tests pass 18/18.</td></tr>
<tr><td>Two prompts consumed before receipt observation could claim the same turn</td><td><span class="good"><b>Fixed.</b></span> Historical lifecycle edges are allocated earliest-first and remain bounded by PTY generation.</td><td><code>orca-runtime.ts:22031</code>; dedicated out-of-order observation regression; focused correlation suites pass 60/60.</td></tr>
<tr><td>Hook-only provider has no historical turn ring</td><td><span class="warn"><b>Accepted limitation.</b></span> Closing it needs a new provider/wire contract or durable hook history. Current behavior is conservative and does not auto-resend.</td><td><code>agent-prompt-submission-verification.ts:136</code>; plan does not promise turn-start proof for every provider.</td></tr>
<tr><td>Crash after <code>WRITE_ATTEMPTED</code> but before pointer bytes</td><td><span class="warn"><b>Accepted fail-closed tradeoff.</b></span> A rare crash can suppress the automatic wakeup; the message remains available to explicit <code>orchestration check</code>. Blind replay is intentionally avoided.</td><td><code>mailbox-pointer-stage.ts:63-78</code> → <code>mailbox-pointer-resume.ts:64-72</code>. A true at-least-once solution needs a new uncertain-state contract or agent idempotency.</td></tr>
<tr><td>Activation after a worker terminal is retired</td><td><span class="good"><b>Fixed.</b></span> The async “empty workspace” gate now re-seeds a shell when explicit activation follows a persisted empty-terminal tombstone.</td><td><code>worktree-activation.ts</code>; new seam regression test and both retirement Electron cases pass.</td></tr>
<tr><td>Mailbox consistency test fixture</td><td><span class="good"><b>Fixed.</b></span> The helper now awaits bounded PTY admission completion for both lifecycle writes, removing the microtask race without changing runtime behavior.</td><td><code>orchestration-mailbox-notification-test-harness.ts</code> and <code>orchestration-message-delivery-identity.test.ts</code>; mailbox matrix passes 66/66.</td></tr>
<tr><td>Completed-worker retirement E2E expectation</td><td><span class="good"><b>Fixed.</b></span> The fixture now expects recovery state <code>done</code>, matching the intentional contract and unit tests.</td><td><code>completed-worker-retirement-resume.spec.ts</code>; both parameterized Electron cases pass.</td></tr>
</tbody>
</table>
<h2>Recorded evidence (separate snapshot vs. current verification)</h2>
<div class="evidence">
<ul>
<li>Typecheck, lint, changed-code quality, packaging, Windows native smoke, cross-version wire checks, and skill round trips passed.</li>
<li>Full local suite previously passed 65,904 tests; broad orchestration matrix passed 1,451 tests.</li>
<li>Independent focused prompt/mailbox state tests passed 66/66; structured-chat lease passed 8/8; activation seam passed 5/5.</li>
<li>Final blocker matrix passed 116/116 across lifecycle, mailbox, federation sync, transport safety, output, and lifecycle-settlement suites.</li>
<li>Activation/recovery regression matrix passed 170/170 across 25 files; full project typecheck passed.</li>
<li>Completed-worker retirement Electron E2E passed both close modes (2/2) after the activation fix.</li>
<li>Real Electron flow proved dispatch, durable follow-up, correlated completion, replay-before-ack, release, archived transcript access, and exactly-once terminal submission.</li>
<li>Remote/SSH/WSL deterministic fixtures and compatibility checks passed. Physical certification of every remote topology remains separate operational work.</li>
<li><b>Current workspace verification:</b> orchestration Electron matrix 20/20, provider-transcript E2E 1/1 (Claude/Grok/OMP-shaped hooks and files), focused runtime/archive tests 57/57, node typecheck, formatting, diff check, and changed-code quality all pass.</li>
<li><b>Current workspace caveat:</b> repository-wide <code>typecheck:e2e</code> still reports unrelated baseline errors in other specs; no error was reported for the provider-transcript spec.</li>
</ul>
</div>
<h2>Complexity and architectural-boundary audit</h2>
<p class="small">Supervised Orca Run <code>run_06e99b77fceb</code> · four independent Codex lanes · read-only
review against <code>origin/main</code> at <code>4a32d073fa</code>. Workers ran focused checks and made no
tracked source edits. Findings below distinguish correctness blockers from maintainability opportunities;
a larger diff is not treated as a defect by itself.</p>
<table>
<thead><tr><th>Disposition</th><th>Finding and evidence</th><th>Decision</th></tr></thead>
<tbody>
<tr><td><span class="good"><b>FIXED</b></span></td><td><b>Delayed exit after <code>stop_unknown</code>.</b>
<code>worker-dispatch-stop.ts:234-263</code> leaves the worker uncertain; the PTY exit path then calls
<code>failDispatch(...workerProcessExited)</code> (<code>orca-runtime.ts:19147-19175</code>), which asks
the lifecycle graph for <code>stop_unknown → failed</code>. That edge is absent in
<code>lifecycle-transition.ts:89-99</code>, so a real exit can throw and leave the Task/Dispatch stuck.
This was a direct call-chain mismatch and regression from the permissive old update path.</td><td>Added the canonical positive-exit transition and tested state, authority revocation, Task status, and receipts.</td></tr>
<tr><td><span class="good"><b>FIXED</b></span></td><td><b>Structured-session write gate bypass.</b>
<code>orca-runtime.ts:37947-37961</code> falls back to <code>ptyController.write</code> when
<code>agentSessionPtyWriteGate.admit</code> refuses. For a bound structured session this can inject a
mailbox pointer despite an ownership/reconciliation refusal.</td><td>Bound structured sessions now fail closed while unbound legacy terminals retain raw PTY behavior; a denied-write regression asserts zero PTY bytes.</td></tr>
<tr><td><span class="good"><b>FIXED</b></span></td><td><b>Persisted federation protocol can outlive a remote downgrade.</b>
<code>federation-sync.ts:75-86</code> and <code>db/federation/federation-relay-ack.ts:80-99</code> choose
lifecycle-settlement fields from the attachment's persisted protocol version, while the peer capability
cache only observes the new runtime epoch. After a remote restart on an older build, zod can strip
<code>replayUnacknowledged</code>/<code>settlements</code> and the home may interpret an ACK under the wrong
mode. Existing tests covered protocol versions at attach time, not a protocol-3 attachment followed by a
capability downgrade.</td><td>Current-epoch capability probing, capability-aware ACK handling, and a two-runtime downgrade regression now fail closed safely.</td></tr>
<tr><td><span class="good"><b>FIXED</b></span></td><td><b>Epoch-mismatch ACK could orphan a queued <code>worker_done</code>.</b>
A capability probe can observe one runtime epoch while the following pull observes a restarted peer. If
settlement is suppressed but the home still ACKs through the report, the downgraded peer can delete its
only retryable completion.</td><td>Protocol-3 pulls now replay imported-but-unacknowledged state, hold ACKs at the first unsettled worker report,
and stop paging until a supported epoch can settle it. Added restart/downgrade retention and later replay coverage.</td></tr>
<tr><td><span class="warn"><b>SIMPLIFY / TEST</b></span></td><td><b>Mailbox pointer state has intentional async complexity.</b>
The staged pointer, write-attempt, Enter-attempt, incarnation checks, and fail-closed restart recovery
protect observed duplicate/lost-mail races (<code>mailbox-pointer-stage.ts:75-180</code>;
<code>mailbox-pointer-submit.ts:62-153</code>). The mutable outcome matrix is hard to read, and broad
<code>markAsUndelivered</code> calls could clear a newer reservation if a stale flight survives.</td><td>Keep the state machine and rechecks. Add target+phase-guarded release tests; consider a discriminated
outcome type and a single reservation helper later. Do not replace this with blind retry.</td></tr>
<tr><td><span class="warn"><b>SIMPLIFY / TEST</b></span></td><td><b>Unbounded timing promises.</b> Fleet capability retry can exceed its advertised 5-second total
(<code>orchestration-federated-fleet-snapshot.ts:14-17</code> plus capability-cache retry); relay paging can
spend roughly six independent 15-second page budgets (<code>federation-sync.ts:21-24,154-208</code>).</td><td>Carry one deadline or explicitly measure/rename the budget. Add fake-timer tests. This is latency hygiene,
not a reason to duplicate retries throughout the code.</td></tr>
<tr><td><span class="warn"><b>SIMPLIFY / DEFER</b></span></td><td><b>Duplicated composition and parsing.</b> Local/federated worker start each normalize dependencies and
implement similar materialize → ready → attach → deliver → settle stages. The duplication increases the
number of files changed for every new receipt field, but host-specific authority must remain separate.</td><td>Extract a pure normalized start plan first; defer a shared driver until behavior is stable and covered by
parity tests. Do not merge local and remote liveness policy.</td></tr>
<tr><td><span class="warn"><b>TEST / FOLLOW-UP</b></span></td><td><b>Concurrent start and release recovery are under-specified.</b> Identical concurrent
<code>workerStart</code> calls perform asynchronous topology checks before the durable acceptance claim,
so one caller can receive an ambiguous result even when the other succeeds. A crash after terminal close
but before release settlement can likewise leave a durable <code>releasing</code> row despite a positive
later exit verdict.</td><td>Add Promise.all and crash-seam regressions. Preserve the existing receipt/identity fences; do not
introduce polling or duplicate worker creation to make retries appear successful.</td></tr>
<tr><td><span class="good"><b>KEEP</b></span></td><td><b>Execution-host authority, lifecycle receipts, transcript source fencing, and cursor identity.</b>
Independent lanes found these branches match documented SSH/mixed-version boundaries and focused tests;
simplifying them would reintroduce wrong-host, stale-process, or false-exit failures.</td><td>Preserve. Add only the missing downgrade/old-peer harness and clipping metadata tests before removing fallbacks.</td></tr>
</tbody>
</table>
<p class="small"><b>Audit test evidence:</b> lifecycle lane reported 39 focused checks and a static F1 call-chain
proof; federation/SSH lane passed 4 files / 35 tests; transcript lane passed 4 files / 34 tests; CLI/skill
lane verified generation, focused CLI tests, and typecheck. The blockers now have focused regressions;
the local matrix below is green.</p>
<h2>What was intentionally not added</h2>
<p>No scheduler, automatic placement, silence-based replacement, generalized ACL system, replicated Run database,
universal provider transcript abstraction, retention TTL/bulk cleanup, or speculative command aliases were added.
Those would expand the contract without evidence from the current issues.</p>
<h2>Composer notification limitation</h2>
<p>The screenshot showing “You have 1 orchestration message…” left in a Codex/Claude composer is not fully
eliminated by this PR. The new mailbox stages, receipts, write gate, and fail-closed recovery prevent blind
duplicate retries and naked Enter writes, but they cannot atomically force-submit text already owned by a TUI
after a park, restart, or ambiguous PTY write. The message remains recoverable through explicit
<code>orchestration check</code>. Treat automatic composer submission as a separate, provider-aware follow-up;
do not weaken the current safety boundary to hide the symptom.</p>
<h2>Final merge checklist</h2>
<table>
<thead><tr><th>Action</th><th>Decision</th></tr></thead>
<tbody>
<tr><td>Update mailbox helpers to await bounded PTY ingestion completion.</td><td><span class="good">Done</span>; test-only.</td></tr>
<tr><td>Change retirement E2E expected recovery state from <code>working</code> to <code>done</code>.</td><td><span class="good">Done</span>; test-only.</td></tr>
<tr><td>Re-seed an explicitly activated empty workspace after the agent-inventory gate.</td><td><span class="good">Done</span>; production fix with a focused regression.</td></tr>
<tr><td>Keep fail-closed ambiguous mailbox recovery.</td><td>Recommended. Safer and simpler than blind replay.</td></tr>
<tr><td>Document/measure the at-most-once wakeup tradeoff and hook-only limitation.</td><td>Follow-up if product requires stronger guarantees; no speculative patch here.</td></tr>
<tr><td>Fix and test the <code>stop_unknown</code> positive-exit transition.</td><td><span class="good">Done</span>; canonical transition + lifecycle regression.</td></tr>
<tr><td>Close the structured-session mailbox write-gate bypass.</td><td><span class="good">Done</span>; fail-closed gate + zero-byte regression.</td></tr>
<tr><td>Probe federation capabilities after runtime downgrade.</td><td><span class="good">Done</span>; epoch-aware probe/cache + downgrade regression.</td></tr>
<tr><td>Prevent ACK from deleting an unsettled worker report after an epoch mismatch.</td><td><span class="good">Done</span>; replay/ACK barrier + restart/downgrade regression.</td></tr>
<tr><td>Rerun focused tests and CI on the final diff.</td><td><span class="good">Local done</span>: 116/116 blocker-matrix tests, node/web typecheck, changed-code quality, max-lines ratchet, and diff check pass. Require repository CI before merge.</td></tr>
</tbody>
</table>
<p class="small">Detailed implementation audit: <a href="orchestration-vnext-implementation-audit.html">orchestration-vnext-implementation-audit.html</a>. This brief distinguishes proven defects, accepted boundaries, and unperformed physical certification. No merge or push was performed.</p>
<footer>Prepared from the PR diff, GitHub checks, source inspection, focused test runs, Electron validation, and supervised independent Codex audits. The remaining decision is ordinary CI plus human sign-off on the documented fail-closed mailbox tradeoffs.</footer>
</main>
</body>
</html>
-155
View File
@@ -1,155 +0,0 @@
# Final orchestration vNext release audit
## Verdict
**READY.** The implementation is green, the F1, F2, and prior L1-L4 source
repairs are closed, and the reconciled HTML audit now matches the implemented
public boundary. Retention-expiry metadata, TTL policy, and bulk cleanup are
explicitly validate-first follow-up work, not committed implementation scope.
The final independent review found and closed four concrete gaps before this
sign-off: exact-authority completion after `start_unknown`, unguarded Task
status writes in decision-gate/context-only release paths, A-era completion
compatibility, and host-aware liveness fallback. A deterministic SQLite
`SQLITE_BUSY_SNAPSHOT` race in `failDispatch` was also fixed with an
outer-transaction `BEGIN IMMEDIATE` boundary while retaining savepoint nesting
for caller-owned transactions. These repairs are covered by the focused and
full orchestration suites below.
Physical SSH, WSL, and provider-version certification remains unperformed. The
deterministic topology, compatibility, and provider fixtures are green, but they
do not replace those separate physical certification jobs.
The final supervised review Run completed all eight Tasks and acknowledged every
Delivery. Review terminals exited before cleanup could observe their tabs, so
release returned the safe `release_unknown` receipt with transcript archives
captured (one resource is explicitly retained as `identity_unproven`); no broad
terminal close or unverifiable process claim was made.
## F1 cleanup/retention verification
- The public cleanup/retention portion of the worker CLI contains explicit
`worker-release`, indefinite `worker-retain`, and paginated `worker-list` only;
retain has no expiry/policy flags
(`src/cli/specs/orchestration-worker-specs.ts:87-123`). The registry test proves
the cleanup command and TTL flags are absent while retain still calls
`orchestration.workerRetain` (`src/cli/handlers/orchestration-worker-cli.test.ts:366-389`).
- RPC retain accepts only a strict Dispatch object, and list exposes only run,
terminal state, cursor, limit, and optional remote inclusion
(`src/main/runtime/rpc/methods/orchestration-worker-release-schemas.ts:5-22`).
The method registry explicitly excludes `orchestration.workerCleanup`
(`src/main/runtime/rpc/methods/orchestration-runs.test.ts:27-34`).
- Unknown retention-expiry/policy fields are rejected before ownership or process
state changes (`src/main/runtime/rpc/methods/orchestration-worker-release.test.ts:812-829`).
- The resource table retains ownership/release/archive/recovery facts but has no
retention expiry or policy column
(`src/main/runtime/orchestration/db/schema/create-core-tables-sql.ts:168-215`).
- Explicit release, user-requested indefinite retain, later release, list, and
archive preservation remain implemented
(`src/main/runtime/rpc/methods/orchestration-worker-release.ts:35-155`,
`src/main/runtime/rpc/methods/orchestration-worker-release.ts:153-295`,
`src/main/runtime/rpc/methods/orchestration-worker-archive-read.ts:28-69`).
## F2 Task lifecycle verification
- The legal Task graph restores `pending -> failed|completed|blocked` while
retaining terminal-state and retry constraints
(`src/main/runtime/orchestration/db/lifecycle-transition.ts:33-44`).
- `updateTaskStatus` routes through the guarded transition boundary. A real invalid
edge rethrows typed `lifecycle_conflict`; only a concurrent writer that already
applied the exact requested status is treated as an idempotent success
(`src/main/runtime/orchestration/db/tasks/task-status-transition.ts:67-90`).
- Regression tests exercise all three restored pending transitions, downstream
readiness behavior, durable receipts, and an invalid `blocked -> completed` edge
with typed conflict data and no mutation
(`src/main/runtime/orchestration/db-task-dispatch-invariant.test.ts:28-63`).
- The same suite retains forced interleaving, active-Dispatch, supervised-worker,
federated-start, rollback, and same-pane occupancy races
(`src/main/runtime/orchestration/db-task-dispatch-invariant.test.ts:65-99,209-300,344-422`).
## L1-L4 closure review
- **L1 lifecycle centralization:** the transition primitive enforces explicit
Task/Dispatch/worker graphs and compare-and-swap updates before atomically
appending receipts (`src/main/runtime/orchestration/db/lifecycle-transition.ts:33-63,104-173`).
The direct-write ratchet covers the cited production writer set
(`src/main/runtime/orchestration/db/lifecycle-transition-boundary.test.ts:5-21`).
- **L2 local/folder host classification:** null host scope projects as local, and
the folder-authority regression proves fresh local liveness remains live
(`src/shared/orchestration-fleet-projection.ts:205-241`,
`src/shared/orchestration-fleet-projection.test.ts:106-129`).
- **L3 federation liveness:** missing old-peer verdicts and contact loss remain
`unverifiable`; only an explicit execution-host verdict yields `exited`
(`src/main/runtime/rpc/methods/orchestration-federation-control.ts:281-311`,
`src/main/runtime/rpc/methods/orchestration-federation-liveness-verdict.test.ts:102-123,145-199`).
- **L4 federated release convergence:** confirmed host release idempotently clears
the home worker and remote handles through a lifecycle receipt; ambiguity stays
`release_unknown`
(`src/main/runtime/rpc/methods/orchestration-federated-worker-release.ts:40-125`,
`src/main/runtime/rpc/methods/orchestration-federation-output.test.ts:418-473`).
## Build-now and deferred-boundary review
The implemented build-now slices have concrete source and regression evidence:
prompt queued/submitted receipts; commit-before-notify replay; rejected/late
worker-report observations; atomic `worker-start --spec`; archive/liveness
separation; execution-host transcript routing; Dispatch/Resource identity reuse;
clock-domain-safe observation facts; source/completeness metadata; and
mixed-version capability caching. Representative evidence is in
`src/shared/runtime-terminal-contracts.ts:202-208`,
`src/main/runtime/rpc/orchestration-commit-notify-characterization.test.ts:76-212`,
`src/main/runtime/orchestration/db/dispatch-context/worker-report-settlement.ts:35-186`,
`src/main/runtime/rpc/methods/orchestration-worker-start-schema.ts:11-50`,
`src/main/runtime/orchestration/worker-transcript-read.ts:49-68`, and
`src/main/runtime/orchestration/db/attempt-outcome-projection.ts:27-81`.
Static source/served-guide scans found no speculative orchestration command or
flag for `work start`, attention waits, worker recovery/retry/cleanup, dependency
aliases, or context-bound completion. The only takeover grammar is the established
legacy-authority path. No second Attempt, Resource, lease, or renderer state store
exists: `dispatch_contexts` remains Attempt identity,
`worker_terminal_resources` remains Resource identity, and
`attempt_observation_facts` is an additive fact ledger rather than an identity
store (`src/main/runtime/orchestration/db/schema/create-graph-tables-sql.ts:95-153`,
`src/main/runtime/orchestration/db/schema/create-core-tables-sql.ts:168-215`,
`src/main/runtime/orchestration/db/attempt-observation-store.ts:91-160`).
Machine liveness remains `live` / `unverifiable` / `exited`; legacy terminal
presentation may separately render `running` / `unknown` while retaining the
canonical liveness field
(`src/main/runtime/rpc/methods/orchestration-worker-archive-read.ts:242-254`).
Served orchestration kernel/reference sources contain none of the audited
reference-product or compound-methodology names, and the changed/untracked source
and test marker scan is clean. No required mobile field, Git command, WSL command,
dependency, stream opcode, or duplicate mobile-facing state was introduced.
## Verification
- Full orchestration suite: **137 files, 1,314 tests passed**.
- Final focused lifecycle/mailbox/skill suite: **16 files, 225 tests passed**.
- Host-aware liveness and release suite: **7 files, 83 tests passed**.
- Race repair stress: **30/30 isolated iterations passed**; race plus SQLite
adapter checks: **15/15 tests passed**.
- `pnpm tc` — passed.
- `pnpm tc:node` — passed.
- `pnpm tc:cli` — passed.
- `pnpm run check:code-quality:changed` — passed with zero new findings across
127 changed files (including type-aware and React checks).
- `git diff --check` — passed.
- Served skill generation/guide tests: **32 tests passed**; compact output is
materially smaller than `--full`, and all seven bundled references resolve.
- Static public-surface, served-guide vocabulary, identity-store, liveness, and
unfinished-code marker scans — passed.
- Final HTML/public-surface static assertions — passed.
## Artifact reconciliation
P4 now retains only evidence-backed resource accounting, recovery/archive, and
explicit retain/release work. Its retention-expiry metadata, TTL policy, and
paginated/bulk-cleanup language is explicitly a deferred validate-first follow-up
with measurement and policy/fencing acceptance gates
(`orchestration-vnext-implementation-audit.html:758-797`). R6 likewise names
`existing-resource release/archive accounting` and nests retention expiry, TTL
policy, and bulk cleanup under `validate first`
(`orchestration-vnext-implementation-audit.html:1278-1279`). No implementation
or plan change is required for release sign-off.
@@ -1,21 +0,0 @@
# Early local fleet projection inventory
## Existing production consumers retained
- `orchestration.workerList` is registered through `ORCHESTRATION_WORKER_RELEASE_METHODS` and consumed by `src/cli/handlers/orchestration/worker-terminal-handlers.ts`. The response keeps the existing `workers` and `counts` fields; each bounded worker row gains `projection`, and the response gains `page`.
- `src/main/runtime/orca-runtime.ts#buildAgentOrchestrationByPaneKey` remains the source of runtime-authoritative Dispatch context published by `syncWindowGraph`.
- `src/renderer/src/runtime/sync-runtime-graph.ts` continues writing that publication through `setRuntimeAgentOrchestrationByPaneKey` into the existing agent-status slice.
- `src/renderer/src/store/slices/agent-status.ts` continues merging runtime context into its existing live and retained status maps.
- `src/renderer/src/components/sidebar/worktree-agent-orchestration-index.ts` continues indexing those maps for worktree consumers. No renderer polling or second fleet store was added.
## New composition boundary
- `src/shared/orchestration-fleet-projection.ts` is a pure, renderer-independent projection. It joins durable worker-list rows to the existing push-fed agent-status snapshot by stable pane key or terminal handle.
- `src/main/runtime/rpc/methods/orchestration-worker-release.ts` is the only new production caller. It performs no terminal reads, liveness probes, federated show/read calls, or per-worker remote calls.
- The RPC returns stable Dispatch/Task/Run IDs, worker role and parent Task, provider/model metadata, opaque host/workspace identity, durable and live stage evidence, `live`/`unverifiable`/`exited` liveness, typed resource absence, and a conservative next action.
## Bounds and deferred work
- Pages are capped at 100 rows and continue after a stable Dispatch ID. Push rows are indexed once per call, then joined in linear time.
- Prompt, tool input, assistant-message, interactive-prompt, provider-session, and transcript bodies never enter the returned projection.
- Remote batching, negotiated capabilities, partial-host budgets, and notification policy are intentionally unchanged and remain later validation work.
-230
View File
@@ -1,230 +0,0 @@
## Orchestration Issues
- Agents don't effectively wait long enough for dispatched workers to run/ don't have much sense of how long to wait for? i.e. it keeps doing repeated 30s timeouts on a task that may take 10+ mins to run
- Lots of leftover tabs
- When workers are done they often keep it open (there is some use cases where its useful but sometimes its not)
- Setup terminal tab is usually just unused
- Unnecessary notifications
- Dispatched workers create notifications even when the user should only care when the main agent finishes
- Prob by default workers have no notifications? Not sure, at least maybe don't make it show the bell on the left sidebar
- Agents are not very good at reading TUIs still?
- They should be reading agent transcripts whenever possible whenever reading the state of another agent instead of raw TUI output
- Are agents not messaging each other effectively? E.g. not sure if its just creating agents instead of dispatching tasks with intra-agent communication
- Not knowing when to create child workspaces vs just working within the same workspace?
## GitHub investigation: gaps, bugs, and a vNext model (2026-08-26)
### Scope and confidence
I reviewed public Issues, PRs, and Discussions in [stablyai/orca](https://github.com/stablyai/orca), including 64 Discussions, targeted searches for `orchestration`, `multi-agent`, `worker-start`, `dispatch`, `inbox`, `wait`, `agent communication`, `terminal cleanup`, `worktree`, `transcript`, and `notification`, and the bodies/comments of the strongest matches. The search covered open and closed items; closed items are included when they expose a recurring design boundary or a fix that should become an invariant.
- **Verified** means the report contains a concrete reproduction, measurement, source/DB trace, or independent confirmation.
- **Requested** means a user describes a desired workflow; it is evidence of demand, not proof of a defect.
- **Inferred** means the class-level diagnosis or solution below, derived from multiple reports rather than quoted as a single user claim.
Status is a GitHub snapshot, not a claim that an open PR has shipped. Representative searches also found many unrelated matches, so this is a focused qualitative review rather than an exhaustive census.
### Evidence-backed gaps and bugs
| Category | What users hit | Evidence | Class-level direction |
| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Product semantics and discoverability | New users cannot tell orchestration from ownership handoff, native subagents, or file-based coordination; isolation and provider instruction-file boundaries are surprising. | **Verified/requested:** [#4397](https://github.com/stablyai/orca/issues/4397), [Discussion #6201](https://github.com/stablyai/orca/discussions/6201), [#14104](https://github.com/stablyai/orca/issues/14104) and [PR #14996](https://github.com/stablyai/orca/pull/14996). [#16075](https://github.com/stablyai/orca/issues/16075) shows Claude replacing an explicit `/orchestration` invocation with native subagents and creating no Orca Run/Task/Dispatch. | One public mental model: explicitly distinguish handoff, supervised delegation, nested coordination, and workflow/DAG execution. Make an explicit `/orchestration` request authoritative; if the runtime is unavailable, report the blocker before offering a substitute. Add a first-run happy path that teaches rendezvous, isolation, and provider instructions. |
| Atomic task creation and dispatch | The normal `task-create` → ID extraction → `worker-start` sequence is shell-heavy and can leave an orphan Task when the second mutation fails. | **Verified/requested:** [#13360](https://github.com/stablyai/orca/issues/13360) and open [PR #13540](https://github.com/stablyai/orca/pull/13540) propose `worker-start --spec` with transactional Task creation and deterministic retries. | Make create+dispatch one idempotent atomic operation with a receipt that records partial effects. Keep separate Task creation for planned fan-out. |
| Delivery truth and startup readiness | `ready`/`input_accepted` can mean only that Orca wrote bytes, not that the endpoint accepted them or an agent turn started. Slow-booting Hermes/Antigravity workers did the work while Orca declared `agent_prompt_stalled`, revoked authority, and rejected `worker_done`; another report saw a prompt shredded in the terminal and a live unusable worker left behind. | **Verified:** [#15180](https://github.com/stablyai/orca/issues/15180), [#16660](https://github.com/stablyai/orca/issues/16660), [#16525](https://github.com/stablyai/orca/issues/16525), historical [#13488](https://github.com/stablyai/orca/issues/13488), and [#10416](https://github.com/stablyai/orca/issues/10416) (payload truncation symptom; root cause still uncertain). Related open PRs: [#16425](https://github.com/stablyai/orca/pull/16425), [#16548](https://github.com/stablyai/orca/pull/16548). | Use a closed, honest stage vocabulary: recorded → routed → endpoint-delivered → terminal-queued → provider-submitted → turn-started → report-persisted → settled. Receipt fields must say what was observed, never infer receiver state from a PTY write. On `dispatch_input` failure, roll back or mark the terminal unusable and provide a reclaim action. Fall back to provider/process activity when hooks or titles are silent, with `outcome_unknown` rather than false failure. |
| Durable wakeups and mailbox semantics | Coordinators depend on polling; run-addressed lifecycle mail can remain invisible, and a missed acknowledgment can block newer messages. Reply-channel rows have no acknowledgment path, so unread counts grow forever. Cross-run sends can report a live Run as “not found.” | **Verified:** [#9228](https://github.com/stablyai/orca/issues/9228), [#11787](https://github.com/stablyai/orca/issues/11787), [#16522](https://github.com/stablyai/orca/issues/16522), [#16605](https://github.com/stablyai/orca/issues/16605), [#16603](https://github.com/stablyai/orca/issues/16603), and [#11242](https://github.com/stablyai/orca/issues/11242). [#10663](https://github.com/stablyai/orca/issues/10663) is marked pending repro. Merged pointer/retry work exists in [#12988](https://github.com/stablyai/orca/pull/12988) and [#14332](https://github.com/stablyai/orca/pull/14332), but newer routes still expose the same boundary. | Make one durable, role-addressed mailbox model for Run, Dispatch, and reply channels. Wake the owning coordinator on writes, deduplicate pointers, expose delivery/ack state, label replayed batches, reject a bare ack, and allow per-message acknowledgment. A cross-run authorization failure must not masquerade as nonexistence. Protect unsent human drafts from injected mail. |
| Completion authority and liveness | UI idle/green, agent process exit, Git changes, heartbeat, and `worker_done` disagree. A worker can finish without sending `worker_done`; a rejected report can leave a Dispatch looking `ready`; a healthy session can lose heartbeat capability after compaction. | **Verified:** [#10673](https://github.com/stablyai/orca/issues/10673), [#14829](https://github.com/stablyai/orca/issues/14829), [#13364](https://github.com/stablyai/orca/issues/13364), [#16604](https://github.com/stablyai/orca/issues/16604), and fleet visibility [#15951](https://github.com/stablyai/orca/issues/15951). [#14311](https://github.com/stablyai/orca/issues/14311) and [PR #14762](https://github.com/stablyai/orca/pull/14762) show progress messages being confused with completion. | Model completion as multiple facts: process/TUI state, observed turn activity, artifact/Git evidence, optional worker report, and coordinator acknowledgment. `worker_done` is a fast path, not the sole authority. Rejected or missing reports must settle to an explicit “finished but unverified”/`outcome_unknown` state, preserve late evidence, and never look healthy merely because the row says `ready`. Show heartbeat age and “never reported” distinctly in `worker-list`, with host-computed age for remote clocks. |
| Recovery and error actionability | Recovery after a failed dispatch can take five or more trial invocations; errors name the condition but not the next safe command. Some replies return `ok:true` while silently dropping because the Dispatch is inactive. | **Verified:** [#16529](https://github.com/stablyai/orca/issues/16529) and [#16524](https://github.com/stablyai/orca/issues/16524). Related residual-resource work: [#15944](https://github.com/stablyai/orca/issues/15944) and [PR #15966](https://github.com/stablyai/orca/pull/15966). | Every failure receipt should include a typed condition, whether mutation may already have happened, and an exact idempotent next action (`--retry-request`, `--retry-of`, fresh terminal, drop redundant flag, or safe release). A write that was not delivered must never report success; if it was queued, say `queued` versus `delivered`. Make retry, cancel, supersede, abandon, and recovery first-class verbs. |
| Identity, authority, and nested roles | Terminal handles, pane keys, provider sessions, Runs, Tasks, Dispatches, and worker records are treated as interchangeable. Adopted `dispatch --inject` lanes disappear from `worker-list`/`worker-show` and cannot be upgraded. Internal subagents can inherit a parent completion capability and falsely complete it; nested coordinator and worker mailboxes can shadow each other. | **Verified:** [#13678](https://github.com/stablyai/orca/issues/13678), [#15634](https://github.com/stablyai/orca/issues/15634), [PR #15267](https://github.com/stablyai/orca/pull/15267), [#16051](https://github.com/stablyai/orca/issues/16051), [#16319](https://github.com/stablyai/orca/issues/16319), and nested-depth work [PR #16669](https://github.com/stablyai/orca/pull/16669). Historical identity fix: [PR #9637](https://github.com/stablyai/orca/pull/9637). | Separate stable Agent, Run, Role, Attempt, Endpoint, and Resource identities. A terminal is an address, not durable identity. Make unsupervised lanes explicit and readable or allow a deliberate adoption/upgrade path; bind completion to the owning process/attempt, not a shared pane; render parent/child depth and both mailbox roles. |
| Fleet observability and transcripts | A coordinator must issue one query per lane to know whether dispatch reached a worker. Nested Codex/Claude sessions are flat or unreadable, especially on remote/headless hosts; `terminal read` loses styling and can mistake ghost text for user input. | **Verified/requested:** [#14907](https://github.com/stablyai/orca/issues/14907), [#12621](https://github.com/stablyai/orca/issues/12621), open [PR #12311](https://github.com/stablyai/orca/pull/12311), [#10387](https://github.com/stablyai/orca/issues/10387), and [#9934](https://github.com/stablyai/orca/issues/9934). | Provide one bounded, host-routed fleet query and an agent graph: Task, Attempt, role, provider/model, worktree/host, lifecycle stage, heartbeat age, blocked reason, exit evidence, resource owner, last delivery receipt, and safe next action. Offer provider-neutral read-only child transcripts with transcript JSONL as evidence where available; keep raw TUI output explicitly advisory. |
| Topology, placement, and host boundaries | The guide does not tell coordinators when to choose child versus top-level worktrees; one Run can mix lineages. Cross-host workflows require trial-and-error runtime/identity setup, while folder workspaces and Git worktrees are conflated. | **Verified/requested:** [#14840](https://github.com/stablyai/orca/issues/14840) and [PR #14841](https://github.com/stablyai/orca/pull/14841), [#10387](https://github.com/stablyai/orca/issues/10387), [#9221](https://github.com/stablyai/orca/issues/9221), [Discussion #13019](https://github.com/stablyai/orca/discussions/13019), and [#14840](https://github.com/stablyai/orca/issues/14840). | Keep coordination hierarchy, filesystem isolation, Git lineage, UI grouping, and execution host as separate graphs. Choose placement deterministically at Run scope: default to current folder/worktree unless an explicit conflict or isolation requirement exists; explain the choice and query it. Treat remote/SSH/WSL and folder workspaces as first-class execution boundaries. |
| Resource ownership and cleanup | Settled workers remain as live Done tabs/processes; UI history can remain after the PTY is gone, producing both resource leaks and ghost rows. There is no discoverable bulk cleanup. | **Verified across platforms:** [#13047](https://github.com/stablyai/orca/issues/13047), [#10718](https://github.com/stablyai/orca/issues/10718), [#8593](https://github.com/stablyai/orca/issues/8593), [#15325](https://github.com/stablyai/orca/issues/15325), and [#15641](https://github.com/stablyai/orca/issues/15641). Proposed fix [PR #16217](https://github.com/stablyai/orca/pull/16217). | Make release/reuse/retain-until/suspend/takeover explicit leases with a reason. Auto-release exact Orca-owned worker resources after settlement; retain only by explicit user choice; archive transcript/evidence. Add per-row and bulk “clear Done” actions, TTLs, and a clear live-versus-history distinction. Missing-tab release should be idempotent bookkeeping, not a residue state. |
| Human attention and notifications | A 5–6 worker wave produces one alert per worker plus the coordinator alert; meanwhile “needs input” can be silent, unread state is worktree-scoped, and send-target menus omit stale or cross-worktree agents. Starting a background agent can steal the worktree the user is viewing. | **Verified/requested:** [#15148](https://github.com/stablyai/orca/issues/15148) and [PR #15153](https://github.com/stablyai/orca/pull/15153), [#14426](https://github.com/stablyai/orca/issues/14426), [#9050](https://github.com/stablyai/orca/issues/9050), [#14798](https://github.com/stablyai/orca/issues/14798), [#9944](https://github.com/stablyai/orca/issues/9944), and [#9181](https://github.com/stablyai/orca/issues/9181). | Route attention by role and verified need: quiet intermediate worker completions, loud root/coordinator completion, distinguish input/failure/interruption, and keep per-agent unread state. A target should be selectable by live endpoint regardless of hook freshness or worktree; background work must not steal focus without explicit reveal. |
| Provider and execution adapters | Antigravity could not exchange orchestration signals; provider TUIs, permission prompts, login/onboarding, and title formats make readiness unreliable. Some users need many manual terminal steps or a headless path. | **Verified/requested:** [#11721](https://github.com/stablyai/orca/issues/11721), [#9951](https://github.com/stablyai/orca/issues/9951), [#12360](https://github.com/stablyai/orca/issues/12360), [#15838](https://github.com/stablyai/orca/issues/15838), [#15825](https://github.com/stablyai/orca/issues/15825), and [#16075](https://github.com/stablyai/orca/issues/16075). | Define a provider capability contract for startup, lifecycle events, permissions, session identity, and delivery guarantees. Prefer provider/process protocols or headless execution where supported; keep TUI paste/scrape as a bounded compatibility adapter with explicit uncertainty. Test macOS, Linux, Windows, WSL, SSH, remote servers, and folder workspaces. |
Additional corroborating reports make the same boundaries concrete: [#14809](https://github.com/stablyai/orca/issues/14809) records a dispatch created without `--inject` that nobody receives; [#13653](https://github.com/stablyai/orca/issues/13653) shows readiness timeout leaving a live unsupervised agent; [#15317](https://github.com/stablyai/orca/issues/15317) shows a remote relay spinner stuck on “working” after the provider is idle; [#11499](https://github.com/stablyai/orca/issues/11499) shows a superseded worker exit reopening an already-completed Task; [#13298](https://github.com/stablyai/orca/issues/13298) shows a gate mutation completing an unrelated live Dispatch; and [#14548](https://github.com/stablyai/orca/issues/14548) shows that missing `cancelled`/`superseded` states force deliberate replacement into misleading statuses. These are all **verified reports**; the class-level interpretation is the inference in the next sections.
The reliability-focused pass found further high-signal reports that should stay on the backlog even when a neighboring fix lands:
- **Prompt/capability failures:** [#15920](https://github.com/stablyai/orca/issues/15920) (terminal created but preamble never delivered), [#14505](https://github.com/stablyai/orca/issues/14505) (prompt remains an unsubmitted Claude draft), [#13854](https://github.com/stablyai/orca/issues/13854) (Antigravity login consumes the injection), [#15124](https://github.com/stablyai/orca/issues/15124) (preamble truncates before commands/capability), [#15945](https://github.com/stablyai/orca/issues/15945) (session-limit screen accepts a dispatch and wedges it), [#15973](https://github.com/stablyai/orca/issues/15973) (known `blockedReason` collapsed to generic stall), and [#7429](https://github.com/stablyai/orca/issues/7429) (null IDs let custom-harness completion fail silently). These show why “prompt sent” and “worker can report” must be separate receipt stages.
- **Mail/cursor fragmentation:** [#13696](https://github.com/stablyai/orca/issues/13696) (false empty success with unread mail on a sibling Run), [#15586](https://github.com/stablyai/orca/issues/15586) and [#13656](https://github.com/stablyai/orca/issues/13656) (terminal-addressed rows have no consuming `check` path), [#12986](https://github.com/stablyai/orca/issues/12986) (outage durability is unspecified), [#14910](https://github.com/stablyai/orca/issues/14910) (heartbeats create noisy wake mail), and [#15185](https://github.com/stablyai/orca/issues/15185) (ready work with no live lane never wakes). Open [#15249](https://github.com/stablyai/orca/pull/15249) and [#8057](https://github.com/stablyai/orca/pull/8057) address individual push paths but do not establish one ack cursor.
- **Identity/routing drift:** [#9163](https://github.com/stablyai/orca/issues/9163) and [#10702](https://github.com/stablyai/orca/issues/10702) show handles changing across create/restart; [#14570](https://github.com/stablyai/orca/issues/14570) shows generated preambles causing hybrid sender handles; [#14898](https://github.com/stablyai/orca/issues/14898), [#15952](https://github.com/stablyai/orca/issues/15952), and [#15639](https://github.com/stablyai/orca/issues/15639) show missing caller/session identity breaking Run authorization and plugin correlation. Resolve identity inside the runtime; never ask the model to copy a transient selector.
- **Remote and projection drift:** [#15020](https://github.com/stablyai/orca/issues/15020) documents WSL federation losing authority, output, and cleanup; [#11993](https://github.com/stablyai/orca/issues/11993) leaves adopted Runs stuck after the original coordinator disappears; [#15276](https://github.com/stablyai/orca/issues/15276) makes resumed sessions vanish when registration is SessionStart-gated; [#12747](https://github.com/stablyai/orca/issues/12747) re-fires completion notifications on replay; [#16102](https://github.com/stablyai/orca/issues/16102) drops hidden-window remote alerts; and [#9333](https://github.com/stablyai/orca/issues/9333) mistakes an intermediate `working → done → working` edge for final completion.
### Bug clusters and solutions that remove classes of failure
1. **Transport and delivery boundary:** #16522, #16603, #16605, #11787, #11242, and #16524 all reduce to “the sender cannot tell whether the intended recipient received a durable message.” Use one mailbox, durable event IDs, explicit `queued`/`delivered`/`acknowledged` receipts, idempotent ack, replay labels, and protected user drafts.
2. **Observation mistaken for lifecycle truth:** #16660, #16377, #16095, #15180, #10673, and #13364 show PTY bytes, TUI idle, hooks, heartbeats, and reports being conflated. Make each an independent fact and derive the Attempt state from a documented state machine; never revoke a live worker solely because an observer timed out.
3. **Identity and authority drift:** #13678, #15634, #15267, #16604, and #16051 show transient endpoints and inherited context becoming authority. Use stable Attempt/process identity, scoped capabilities, explicit nested roles, and capability refresh or an honest unrecoverable verdict.
4. **Non-atomic mutations and recovery:** #13360, #16529, #15944, and #15641 show partial effects and unclear ownership. Make mutations transactional/idempotent, preserve a receipt after connection loss, and print exact next commands with ownership-safe cleanup.
5. **Resource lease drift:** #13047, #10718, #8593, and #15325 show process, terminal, worker record, and UI history diverging. A single lease owner should drive settle → release; history remains queryable but is not rendered as live.
6. **Attention and topology debt:** #15148, #14426, #9050, #14798, #9944, and #14840 show that notification, focus, placement, and unread state are being inferred from incidental worktree/terminal state. Make role, host, worktree, focus, and attention separate queryable dimensions.
### Proposed vNext contract
**One attempt state machine.** Persist Run, Task, Attempt/Dispatch, Gate, Endpoint, and Resource events. Valid transitions include start, route, deliver, turn-start, progress, ask, block, cancel, supersede, retry, report, complete, fail, abandon, suspend, retain, release, and takeover. Task status is a projection of attempt facts; it is not an independently editable light. Late and duplicate events are fenced and idempotent.
**One observable lane record.** A lane query should return stable IDs, role and parent, provider/model, host/worktree, observed stage, heartbeat age (`never`, age, or `unverifiable`), blocked reason, exit evidence, artifact/Git evidence, report receipt, resource lease, and an exact safe next action. `live`, `unverifiable`, and `exited` must remain distinct for remote/SSH loss; loss of contact is not proof of process death.
**One communication model.** Run/Dispatch/reply addressing should converge on the same durable mailbox. Messages need immutable IDs, ordering, delivery and acknowledgment state, bounded deduplication, replay labeling, and push pointers that survive restart. `worker_done` remains the fast path, but a missing/rejected report must not strand a clearly finished lane.
**One topology model.** Coordination parent/child edges, Git worktree lineage, filesystem isolation, UI grouping, and execution host are independent edges. Placement is a documented Run-level decision, not an accidental mix of `current`, child, and top-level lanes.
**One execution adapter boundary.** Provider hooks, session IDs, permissions, startup/onboarding, headless modes, and TUI fallback behavior are capabilities. A TUI screen is an observation surface, not the protocol. When observation is unavailable, preserve authority and say `outcome_unknown` instead of manufacturing failure.
**One human attention policy.** Default to one root/coordinator completion notification per wave; alert distinctly for user input, failure, interruption, and approval. Keep per-agent unread state, include all eligible live endpoints in send menus, and never steal focus or overwrite a draft without explicit user intent.
### Measurable acceptance criteria
- A persisted `worker_done`, escalation, question, gate resolution, or endpoint exit wakes the owning coordinator without polling (target p95 <2s); restart resumes the same Run exactly once.
- A cold-start matrix across supported providers/OS/SSH/WSL/remote/folder workspaces produces zero false `ready` receipts; start is shown only after receiver/turn evidence.
- Every active Attempt has exactly one owner and resource lease; no terminal Task state has an unaccounted live Attempt. After 100 mixed success/failure/cancel/retry/restart/remote-disconnect workers, every residual resource has an explicit retention reason and safe action.
- A five-worker wave emits at most one default user-facing completion alert; input/failure/interruption are distinct; child events never masquerade as root completion.
- Fuzz and end-to-end tests prove incoming orchestration mail never overwrites or submits a human draft on desktop, remote, or mobile-attached sessions.
- One bounded fleet query replaces per-lane polling and exposes transcript/progress, heartbeat age, delivery receipt, exit evidence, and cleanup action.
- Representative incidents from #9228, #13488, #14527, #15825, and #12360 recover through orchestration APIs alone—no manual Enter/Ctrl+C, throwaway coordinator, screen-buffer inference, or out-of-band task tracking.
- Old-client/new-host and new-client/old-host tests cover optional fields and negotiated capabilities; no new stream opcode is sent without capability negotiation.
### Corrections to the original bullets
- “Agents don’t wait long enough” is mostly a durable-wakeup and readiness-truth problem. A timeout is a checkpoint, not evidence of failure.
- “Leftover tabs” is an owned-resource lease problem spanning PTYs, processes, worktrees, setup surfaces, and history. The specific claim that setup tabs are usually unused was not independently supported here.
- “Unnecessary notifications” is attention routing, not blanket suppression: quiet child completions but preserve root completion and user-input/failure alerts.
- “Agents are bad at reading TUIs” should become a product observability contract: transcripts are evidence where available; raw screen text is advisory.
- “Agents may not message effectively” should be replaced by durable wake/delivery, typed progress-versus-completion, and explicit nested mailbox routing.
- “Child workspace versus same workspace” is two decisions: execution placement and coordination lineage. Keep them independent and explain the default.
- Remove the blank bullet; it has no actionable content.
## Realistic vNext execution plan and lessons from other orchestrators
### The practical diagnosis
The long issue list is not 100 unrelated projects. It is roughly seven systemic failures that appear in different surfaces:
1. **No single truth for identity and ownership.** Run, Task, Dispatch, terminal, pane, provider session, and worktree handles are used as if they were the same thing.
2. **No single truth for delivery.** A PTY write, a queued message, an agent-accepted prompt, and a started turn are conflated.
3. **No single truth for lifecycle.** UI idle, process exit, hook events, heartbeats, Git changes, and `worker_done` disagree, especially after restart or on remote hosts.
4. **No transaction around mutations.** Task creation, dispatch, readiness, retry, completion, and cleanup can partially succeed and leave ambiguous ownership.
5. **No durable recovery path.** Missed wakeups, stale capabilities, dead coordinators, and lost connections require manual intervention or repeated trial commands.
6. **No fleet-level operator surface.** Coordinators poll one lane at a time, while users receive noisy or missing notifications and cannot tell what needs attention.
7. **No stable execution boundary.** Provider TUIs, remote/SSH/WSL hosts, folder workspaces, and Git worktrees expose different semantics that are patched independently.
This framing makes the work achievable: fix each boundary once, then let existing commands and UI consume the new facts. The target is not “more autonomous agents”; it is a small, durable orchestration kernel that makes every agent action honest, recoverable, and observable.
### Additional field evidence: queued prompts and provider scrollback
A Windows 11 three-terminal session (Claude coordinator, Claude worker, Codex verifier; same worktree; roughly 40 sends) exposed two concrete contract gaps:
- `terminal send --enter` returned `ok:false`/`agent_prompt_stalled` while the target Claude or Codex TUI had actually queued the text in its input box and consumed it when the active turn ended. The false failure caused duplicate resends and substantial confirmation overhead. The receipt model must distinguish `terminal_queued` (accepted into the terminal/input buffer) from `provider_submitted` and `turn_started`; a queued message must be retryable by ID without sending duplicate text. A `--wait-submit`-style wait can optionally resolve when provider submission is observed.
- `terminal read` returned only the visible ~50-line screen for Claude Code, while Codex exposed a much larger buffer. Therefore raw TUI scrollback cannot be a cross-provider confirmation contract. Prompt-submission history and transcripts need their own provider capability, with screen reads labeled advisory and a documented fallback when unavailable.
These are refinements of the delivery and adapter categories above, not a new systemic category. They should become regression cases for the receipt state machine and provider capability matrix.
### What the reference products actually do
The following notes are based on the local checkouts in `/Users/jinwoo/refs` (primarily each project's README and design/API documentation).
| Product | Core model | Best ideas for Orca | What Orca should avoid |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Overstory** | A coordinator spawns agents in isolated Git worktrees. Workers communicate through a typed SQLite mail bus, usually run headless through a pluggable `AgentRuntime`, and return through a merge queue. A web UI exposes status, traces, errors, replay, logs, and costs; watchdogs and role-specific agents handle supervision. The repository is archived and points new work to Warren. | Provider adapter boundary; typed mail and replies; headless structured events; role-specific capabilities; durable trace/replay/error views; watchdog tiers; merge/conflict queue; cost instrumentation; an explicit warning that swarms amplify errors, spend, and merge conflicts. | Its full swarm machinery, many roles, and mandatory worktree/merge workflow. More agents are not automatically faster, and the project’s own risk analysis says decomposition can reduce coherent reasoning while increasing operational complexity. |
| **Paperclip** | A Node/React control plane for an “AI company”: goals and ancestry, org chart and reporting lines, tasks, approvals, heartbeat-driven execution, budgets, audit history, workspaces, plugins, secrets, routines, and mobile operation. It treats the operator’s core questions as “what is happening, does it need me, and what should I do?” | Operator-first fleet view; goal ancestry (“why this task exists”); first-class approval/needs-input states; atomic checkout plus budget enforcement; durable audit/activity history; runtime skill injection; plugin/provider boundary; verified artifacts rather than agent prose alone. | The company/org-chart metaphor and broad enterprise governance before Orca’s core Run/Attempt lifecycle is reliable. It is a useful control-plane pattern, not a reason to turn every coding session into an HR hierarchy. |
| **Herdr** | A Rust background terminal server/multiplexer that owns panes and agent processes across detach/reattach, restart, and SSH remote attach. It recognizes many agent CLIs, exposes the same operations through CLI and socket API, supports structured `agent start/prompt/wait/send-keys`, persists session state, and reports `working`, `blocked`, `done`, `idle`, or `unknown`. | Server-owned execution; stable workspace/tab/pane IDs; structured prompt/wait APIs instead of raw keystrokes; explicit `unknown` state; blocked detection; session snapshot and event subscriptions; native session resume; remote thin-client model; no-focus background operations; protocol mismatch errors. | Treating screen heuristics as a complete protocol. Herdr itself needs provider integrations and fallback detection, so Orca should use its lifecycle model as an adapter contract and preserve uncertainty when detection is weak. |
| **Gas Town** | A durable factory around a Mayor coordinator, Deacon supervisor, Witness lifecycle managers, Refinery merge queues, Polecat workers, rigs, convoys, and Beads/Dolt persistence. Typed mail and protocol messages cover completion, merge, rework, recovery, help, and handoff. Witnesses detect dead workers; Deacon redispatches; Refinery verifies and merges. | Explicit ownership roles; durable task/mail storage; typed protocol messages; separate completion, merge, recovery, and cleanup responsibilities; rework/recovery as first-class outcomes; stable `run_id` correlation; push-notify rather than transcript scraping; fail closed when state is unknown. | Recreating the entire Mayor/Witness/Refinery/Deacon/Dolt ecosystem. The roles are valuable boundaries, but a desktop Orca session needs a small number of services and can add supervisors incrementally. |
### The combined Orca architecture
Use Orca’s existing desktop, worktree, folder-workspace, SSH, WSL, and remote-server integrations as the execution surface, then put a narrow orchestration core underneath them:
- **Operator layer (Paperclip lesson):** one fleet view answers what is running, what needs attention, and the next safe action. Goals, approvals, budgets, and notifications are projections over the core records.
- **Durable protocol layer (Gas Town lesson):** one event log and mailbox for Runs, Tasks, Attempts, replies, gates, and recovery. Every event has an immutable ID, stable `run_id`/`attempt_id`, actor, host, and causation ID.
- **Runtime adapter layer (Overstory + Herdr lessons):** each provider exposes capabilities for startup, prompt delivery, turn/lifecycle events, session identity, permissions, transcript access, cancellation, and health. TUI paste/scrape is a bounded fallback adapter, never the source of truth.
- **Resource/lease layer:** the core owns exactly one lease for each worker process, terminal/pane, worktree, and temporary setup surface. Settlement chooses release, retain, suspend, or takeover with a reason.
- **Recovery supervisor:** a small mechanical watchdog handles timeouts, reconnect replay, orphan detection, and idempotent retry. Higher-level agent monitoring remains optional and cannot silently mutate ownership.
Keep these graphs separate and queryable: coordination parent/child, execution host, filesystem isolation, Git lineage, UI grouping, and notification audience. This avoids the current “child workspace versus same workspace” ambiguity.
### A phased implementation that can ship
Do not rewrite the product or require a flag-day migration. Put the new core behind existing CLI/API commands, dual-write old projections while validating parity, and promote one vertical slice at a time. Every existing GitHub issue should become either a regression test, a contract test, or an explicitly accepted limitation.
**Phase 0 — vocabulary and invariants (1–2 weeks).**
- Publish the state machine and ownership rules for Run, Task, Attempt, Endpoint, Resource, Gate, and Message.
- Define the honest receipt vocabulary: `recorded`, `routed`, `endpoint_delivered`, `terminal_queued`, `provider_submitted`, `turn_started`, `report_persisted`, `settled`, and `outcome_unknown`.
- Define the remote verdicts `live`, `unverifiable`, and `exited`; loss of contact never implies process death.
- Add contract tests for duplicate, late, reordered, and replayed events before changing behavior.
**Phase 1 — stable identity and authority (2–4 weeks).**
- Generate immutable Agent, Run, Attempt, Endpoint, and Resource IDs at creation; keep pane/terminal selectors as ephemeral addresses.
- Resolve caller identity in the runtime/IPC layer, not from model-copied text or generated preambles.
- Scope completion, mail, and cleanup capabilities to one Attempt and owner; fence stale attempts from mutating newer work.
- Render parent/child role and depth in every worker query.
**Phase 2 — one durable mailbox (2–4 weeks).**
- Converge Run, Dispatch, and reply-channel messages on one append-only store with immutable message IDs and per-message ack cursors.
- Return `queued` versus `delivered` versus `acknowledged`; make retries idempotent and label replayed batches.
- Wake the owning coordinator from durable writes, with restart-safe pointers and no heartbeat-generated mail noise.
- Protect unsent human drafts and expose an explicit cross-run authorization error rather than “not found.”
**Phase 3 — transactional Attempt/Dispatch lifecycle (3–5 weeks).**
- Make task-create plus worker-start atomic for the common path, while retaining separate Task creation for planned fan-out.
- Persist every transition and partial-effect receipt; provide `retry`, `cancel`, `supersede`, `abandon`, `recover`, and `takeover` as first-class operations.
- Make startup readiness a provider capability with a bounded timeout that yields an actionable `unverifiable`/`stalled` state rather than revoking a live worker.
- Add one bounded fleet query so coordinators stop polling each lane independently.
**Phase 4 — completion truth and resource leases (3–5 weeks).**
- Derive Task status from process/turn evidence, artifact/Git evidence, worker report, and coordinator acknowledgment; no single signal is authoritative.
- Accept late reports against the correct Attempt, while rejecting stale or unauthorized reports.
- On settlement, release exact Orca-owned resources by default; retain only with an explicit reason/TTL. Archive transcript and evidence separately from live resources.
- Add idempotent bulk “clear Done”/release actions and distinguish live rows from historical rows.
**Phase 5 — fleet observability and attention UX (2–4 weeks).**
- Ship the fleet query and agent graph in CLI/API first, then expose it in the desktop UI.
- Show stable IDs, role, provider/model, host/worktree, stage, heartbeat age, blocked reason, delivery receipt, exit evidence, artifact evidence, lease owner, and next safe action.
- Default to one root/coordinator completion alert per wave; separately alert for user input, approval, failure, and interruption. Preserve per-agent unread state and never steal focus.
- Add provider-neutral transcript JSONL where available, with raw TUI output explicitly marked advisory.
**Phase 6 — provider and remote adapters (ongoing, one provider at a time).**
- Implement a capability matrix for startup, queued-input detection, provider submission, session identity, lifecycle, permission prompts, cancellation, submitted-prompt history, transcript, and completion reporting.
- Start with one local same-worktree provider vertical slice, then test macOS/Linux/Windows, folder workspaces, SSH, WSL, and remote servers.
- Prefer headless/provider protocols; retain TUI compatibility only behind explicit uncertainty and diagnostics.
- Enforce remote wire compatibility: optional fields are additive, and new stream events require capability negotiation.
**Phase 7 — advanced collaboration (only after the kernel is boring).**
- Add task groups, DAGs, approval gates, merge queues, cost budgets, reusable role presets, and mobile control as projections over the stable core.
- Consider shared-room/Agent Office views only if fleet queries and attention routing show a real need; they must not become a second state store.
### How to decide whether a proposed feature is worth building
For every new orchestration feature, require a short contract review:
- Which stable entity owns the state?
- What proves delivery, progress, completion, and cleanup?
- What happens after restart, duplicate delivery, timeout, or lost remote contact?
- Which actor is authorized to mutate it, and how is a stale Attempt fenced?
- What exact next action does a failure expose?
- Can the behavior be tested without reading a screen buffer?
If those questions cannot be answered, the feature is adding another projection or heuristic and should not be built yet.
### Definition of “reliable enough” for vNext
- No silent dropped message and no `ok:true` for an undelivered write.
- A prompt queued during an active turn never reports a hard failure, and retrying its message ID cannot duplicate the prompt.
- No false failure solely because observation was slow or a hook was absent.
- No stale Attempt can mutate a newer Task or complete a parent capability.
- No completed worker resource remains live without an explicit retention reason.
- Restart/reconnect converges to one state and does not duplicate notifications or reports.
- One fleet query replaces per-lane polling and exposes a safe next action.
- Remote uncertainty is rendered as `live`, `unverifiable`, or `exited`, never guessed.
- Child completion is quiet by default; root completion, input, approval, failure, and interruption remain actionable.
- A human draft is never overwritten or submitted by incoming orchestration mail.
- Coordinator confirmation works without assuming a provider exposes TUI scrollback; unavailable transcript evidence is explicit rather than guessed from the visible screen.
The realistic outcome is therefore an incremental reliability program, not a second orchestration product: stabilize identity, delivery, lifecycle, and leases first; make the operator surface legible second; then add richer collaboration only where the new contracts make it safe.
@@ -1,21 +0,0 @@
# Legacy compatibility settlement fix
## Outcome
Fixed the four failing legacy compatibility/takeover tests without weakening replay or authority fencing.
## Changes
- `settleWorkerReportInTransaction` now transitions `worker_dispatches` only when the reporting Dispatch has a supervised worker row. Context-only/unsupervised Dispatches still settle their Task and Dispatch, matching the supported legacy-adoption and low-level-dispatch contract.
- The A-era replay test now constructs its deliberately pre-boundary storage state with direct fixture SQL. This preserves the current lifecycle rule that a completed Task cannot be reopened through `updateTaskStatus`, while retaining coverage that replay cannot touch a newer current attempt.
No takeover, recipient-routing, principal-attestation, or stale-proof logic changed.
## Validation
- Target files: 2 files, 29 tests passed.
- Legacy/lifecycle related suite: 14 files, 126 tests passed.
- Additional run-list compatibility, runtime-update settlement, and current-authority precedence suite: 3 files, 19 tests passed.
- `pnpm tc:node` passed.
- Targeted `oxlint` passed.
- `git diff --check` passed for the changed source and test files.
@@ -1,12 +0,0 @@
# Missing-verdict regression fix
## Outcome
`inspectWorkerTerminal` now preserves legacy local/folder fallback behavior when the runtime has no liveness verdict: connected terminals are `live` and disconnected terminals are `exited`. An SSH-scoped worker with no authoritative verdict remains `unverifiable` with `missing_liveness_verdict`; explicit `live`, `exited`, and `unverifiable` verdicts still take precedence.
## Coverage
- Added focused regression tests for local live, local exited, and SSH missing-verdict observations.
- Passed 10 affected test files / 90 tests covering worker observation, manual dispatch, worker stop, worker recovery, worker release, and federation liveness.
- Passed `pnpm tc`.
- Passed targeted `oxlint`, `oxfmt`, and `git diff --check` for the changed source and test files.
-89
View File
@@ -1,89 +0,0 @@
# Orchestration vNext: evidence gate for the remaining proposals
This is a second-pass audit of the implementation plan. It applies one rule:
do not turn a useful idea, a reference-product pattern, or an architectural
preference into a committed feature unless there is either (a) a reproducible
Orca failure, (b) a direct user report, or (c) a necessary safety/compatibility
prerequisite. Everything else is a validation experiment or a deferred option.
The evidence is intentionally mixed: GitHub issues and comments, the current
main-process/code review, the Windows Claude/Codex session report, and the local
reference-product audit. Issue numbers below link to the primary reports.
## Decision vocabulary
| Decision | Meaning | What a worker may do |
| ------------------ | ------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- |
| **Build now** | A concrete failure or safety boundary is already demonstrated. | Implement the smallest fix, preserve existing behavior where possible, and add a regression test before broadening the API. |
| **Validate first** | The problem is plausible or supported by a pattern, but the proposed noun, default, or mechanism is not proven. | Add instrumentation/characterization and run the stated benchmark. Do not make it a default or new contract yet. |
| **Defer** | No Orca complaint or prerequisite currently justifies the complexity, or the work duplicates an existing mechanism. | Keep it out of the kernel slice. Re-open only when telemetry or a focused report supplies a failure case. |
## Evidence-backed decisions
| Proposal in the plan | Decision | Evidence and nuance | Narrow implementation boundary |
| ----------------------------------------------------------- | ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Honest delivery stages for prompts | **Build now** | The Windows report observed `agent_prompt_stalled` while text was queued and later consumed. #15180/#16660/#16095 independently show PTY acceptance and provider turn start are conflated; #15274 describes the hook-based submit observation. | Add a prompt-delivery ID bound to terminal, process incarnation, and payload. Report `terminal_queued`/`queued_pending_turn` as accepted-but-not-submitted; keep `submission_observed` and `turn_started` separate. `--wait-submit` observes the same request and never resends. Old hosts require an explicit downgrade. |
| Commit-versus-notify recovery | **Build now** | Current review reproduced a real seam: a single send/reply can commit the durable row, then throw during in-memory notify; the mutation receipt is removed and an exact retry can duplicate the message. Group send already has the safer ordering. | Complete the existing mutation receipt and a durable nudge-outbox row in the same transaction, then notify best-effort/redrive. Do not add a second idempotency system. |
| Rejected/late `worker_done` must not leave a green dispatch | **Build now** | #13364 documents rejected completion with the dispatch still `ready`; comments identify both missing-capability and stale-caller branches. #13298 shows a related inactive-dispatch rejection. | Route rejection through a guarded transition that records the report/reason and projects `needs_review`/unknown, without accepting an unauthorized completion. Make recovery text branch-specific (missing flag, revoked capability, wrong pane, inactive dispatch). |
| Atomic common task creation + worker start | **Build now** | #13360 gives an exact two-command reproduction, an orphan `ready` Task, measured shell/ID plumbing, and an explicit proposed API. | Add `worker-start --spec` (mutually exclusive with `--task`) as one idempotent mutation. Preserve standalone `task-create` for planned fan-out. Keep external side effects and partial-effect receipts explicit; do not silently delete a task that needs recovery. |
| Archive provenance versus process liveness | **Build now** | Code review found `releasing`/`unknown` archive reads hard-code `exited` even when close was not confirmed. This violates the existing SSH execution-boundary contract and can mislead every downstream view. | Make archive source/coverage orthogonal to liveness. `unknown`/unconfirmed close projects `unverifiable`; only an execution-host-confirmed close projects `exited`. Add a regression assertion to archive-read tests. |
| Direct SSH transcript routing | **Build now** | The current selection can resolve a remote session path in the client process; a same-path local sentinel could be returned as remote evidence. This is a security/authority defect even without a matching public issue. | Carry connection/host identity into session resolution. Execute direct-SSH reads through a negotiated remote capability or return `remote_capability_unavailable`; never probe the desktop filesystem. Keep WSL UNC routing separate. |
| Lifecycle writer centralization | **Build now** | Multiple writers update Task/Dispatch/worker state; `recordWorkerStage` is an unfenced direct update. A transition table beside those writers would be an incomplete audit trail. | Introduce one transaction-neutral guarded transition primitive; every writer updates the legacy projection and appends its receipt in the same caller transaction. Add a ratchet test against direct lifecycle updates. |
| Identity reuse and missing lineage facts | **Build now (but as a reduction)** | `dispatch_contexts.id` is already immutable and keys worker dispatches; terminal resources already have immutable IDs and ownership history. Creating parallel Attempt/Resource IDs would add migration and reconciliation risk. | Alias Dispatch ID as Attempt ID. Retain Resource ID. Add only `retry_of`, creator/role, endpoint incarnation, and typed unsupervised/remote attachment facts. Do not infer identity from titles or model text. |
| Host/clock-safe liveness | **Build now** | Federated observations mix source timestamps with home receive time. The SSH contract already requires `live`/`unverifiable`/`exited` and forbids death inference from contact loss. | Store source, execution-host receive, and home-host receive timestamps separately. Compute freshness in the authority host's clock domain; preserve the existing verdict vocabulary. |
| Provider/read source and completeness metadata | **Build now** | The Windows report shows Claude screen output is ~50 visible lines while Codex exposes a larger buffer. Current code already has transcript-first and PTY fallback, but archive fallback provenance and completeness are not always preserved. | Keep `terminal.read` provider-neutral and screen-oriented. Make `worker-read` state source, exactness/completeness, cursor identity, clipping, and fallback reason. Preserve execution-host routing. |
| Compatibility and skew fixtures | **Build now (infrastructure)** | Old/new RPC, relay, schema, SSH, WSL, and provider combinations are normal operation, not a final certification step. Existing capability negotiation is a prerequisite for any new stream or batched remote behavior. | Start fixtures at R0. Every new optional field/table has fresh-DDL, migration, downgrade/reopen, old-writer, and capability-negotiation tests. No unnegotiated stream opcode. |
## Validate first: evidence supports the problem, not the exact product decision
| Proposal | Why it is not yet a committed feature | Validation gate | Safe interim behavior |
| ----------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| “Work / Worker / Message / Attention” as the four agent-facing concepts | #4397 supports better onboarding and a concrete rendezvous model; #16075 supports explicit role classification. Neither report says these four words are the right nouns. Other orchestrators use different vocabularies and expose more or less infrastructure. | Benchmark current commands, improved wording, and aliases on five workflows. Measure time to first correct command, command count, retries/duplicates, wrong-target actions, help lookups, and human interpretation of flags such as `--attention`. | Keep durable/internal entities explicit in JSON. Improve help/output wording without renaming the CLI until the benchmark wins. |
| New `status`, `--attention`, `--tree`, or `work start` command grammar | The underlying needs (find work, find blockers, reduce ID plumbing) are supported; the exact syntax is not. A flag named `--attention` is plausibly ambiguous to a human and has not been tested with agents. | Run comprehension and task-success tests; require fewer wrong-target/retry errors without hiding recovery detail. | Preserve existing grammar and offer aliases only behind a compatibility layer if measured. |
| Durable wake epoch / `check --wait` register-and-recheck | Mail durability and notification failure are real. The alleged same-runtime read/register race is not established on the synchronous current path. | Add deterministic hooks at actual async seams (paged routing, federation replacement, notify failure). Add an epoch only if a missed-event case is reproduced. | Persist mail and redrive a nudge after a fresh post-restart live-idle observation or explicit check. Never claim an offline process was woken. |
| Per-message acknowledgments | #16522 supports clearer replay/ack semantics and warns that a bare ack is easy to misuse. Current Run delivery already has an idempotent whole-batch ack; changing replay granularity creates old/new skew questions. | Characterize whether whole-batch ack causes a real loss/blocking case in current workloads; if needed, negotiate partial-ack capability and test new-partial ↔ old-batch behavior. | Retain whole-batch replay/ack, label replay, and reject ambiguous bare ack. |
| Draft protection and automatic Enter | The user need is credible because pointers/prompts write the composer, but no provider-neutral “composer empty” fact exists. This cannot be promised as a mailbox-only change. | Provider matrix with pre-existing Claude/Codex/unsupported-provider drafts; verify byte-for-byte preservation under busy, idle, restart, and redrive paths. | If emptiness is unproven, do not auto-Enter; return `draft_unknown`/require explicit confirmation. |
| Passive output/status watch | #15180 asks for a durable turn observation; #15148 reports notification noise/silence. A new stream could improve efficiency, but existing push-fed terminal/status subscriptions already exist. | Instrument current subscriptions and measure polling volume, latency, gap/replacement, backpressure, and five-worker wave behavior before adding a new RPC stream. | Extend existing subscriptions or use bounded polling; negotiate any new stream. |
| Fleet query and notification policy | #15148/#14426/#9050 support attention and focus problems. `workerList` and push-fed renderer projections already exist, so a greenfield fleet store is not justified. “One root alert” is a policy hypothesis, not a universal truth. | First compose a local read-only projection; then test per-host batched snapshots with 100-worker/two-host fixtures. Measure alert counts, missed input, stale/unknown comprehension, and focus/draft safety. | Keep current push status and per-agent unread. Add typed event categories before changing defaults. |
| `cancel`, `supersede`, `suspend`, `takeover` as first-class nouns | Recovery and ownership failures are reported (#13360, #13364, related cleanup issues), but no evidence requires all four independent operations now. They also touch closed SQLite enums and cross-host authority. | Start with existing stop/abandon/release semantics and a transition receipt. Add a new operation only when a concrete recovery workflow cannot be expressed safely. | Preserve conservative stop/abandon/release; expose an exact next command and `unknown` where evidence is incomplete. |
| Lease TTLs and bulk cleanup API | Retained resources and manual cleanup are real classes of friction, but automatic expiry can kill a takeover or an identity-conflicted process. | Instrument retention reasons/age and run a dry-run cleanup against external, user-owned, transferred, federated, and unverifiable resources. Add TTL only with explicit policy and fencing tests. | Add retention metadata and paginated, idempotent single-resource cleanup reuse; no automatic release by default. |
| In-flight archive/object-store mirror | Reference products demonstrate the pattern, but current Orca already has release archives and watchers; no Orca report yet proves in-flight loss is the dominant failure. | Measure output loss before release across process/relay crashes. If retained, bind mirror to existing Dispatch/Resource/process identity, quotas, redaction, and archive lifecycle. | Fix archive/liveness/provenance first; keep release archive as the source of truth. |
| Provider registry refactor | Provider-specific maps/watchers are a maintenance seam, but existing maps already work and a registry alone does not fix delivery or transcript truth. | Add an exhaustive profile only when consolidating an actual behavior change or adding a provider. Require no behavior drift and one fixture per supported provider. | Reuse current maps and watchers; do not create a parallel orchestration-only registry. |
| Agent-facing `--wait-submit` default | The Windows report explicitly requests an optional wait. The observation can be provider-specific and timeout semantics are new. | Validate whether coordinators prefer immediate queued receipt plus a separate wait, or blocking by default; measure timeout/retry mistakes. | Make it opt-in, bounded, and request-ID based; timeout is ambiguous observation, never resend permission. |
## Defer: currently unsupported by evidence or duplicative
| Proposal | Reason to defer |
| ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| A universal transcript format across all providers | Local references show adapter-agnostic transport with provider-specific decoding. Forcing one format would erase capability differences and create false confidence. Normalize the envelope, not the provider payload. |
| A second Attempt, Resource, lease, renderer state store, or all-purpose receipt chain | Existing Dispatch/Resource IDs, ownership tables, push projections, and separate mutation/mail/prompt/lifecycle domains already provide the needed foundations. Duplicates increase reconciliation and migration risk. |
| Automatic prompt resend after timeout/stall | The reported failure mode is ambiguous after irreversible PTY writes. Automatic resend is exactly how the observed duplicates occurred. Retry must be by request ID and only after an explicit status/recovery decision. |
| Treating quiet PTY, clean Git, missing client inventory, or relay loss as completion/exit | This contradicts current safety contracts and has no evidentiary basis. Keep `unverifiable` and require host evidence. |
| A mandatory local lease for every active Attempt | Pre-attach, context-only, unsupervised, and federated-home rows legitimately have no local terminal resource. Use typed absence/remote attachment. |
| Broad notification suppression or focus automation | The complaint is noisy/missing attention, not proof that one global policy is correct. Build typed events and measure before changing focus behavior. |
| Cost budgets, merge queues, approval bureaucracy, or role-heavy “company” hierarchy in the kernel | These are useful projections in some products, but no current Orca orchestration evidence makes them foundational to delivery, lifecycle, or read correctness. Revisit after the kernel is reliable. |
| One giant consolidation PR | Independent additive slices are easier to test, roll back, and run against mixed versions. Use a small promotion PR only after compatibility telemetry proves the shims are no longer needed. |
## Revised implementation gate
The plan should now be read as three tracks:
1. **Repair and characterize now:** archive liveness, mutation commit/notify
recovery, prompt queued/submitted truth, rejected completion settlement,
identity reuse, guarded lifecycle writers, SSH host routing, and compatibility
fixtures.
2. **Instrument and validate:** labels/aliases, durable wait epochs, partial
acks, draft detection, passive watch, fleet batching, notification policy,
operation nouns, TTLs, and in-flight mirroring.
3. **Do not add yet:** universal transcript formats, duplicate stores/IDs,
automatic ambiguous resends, heuristic exit, mandatory local leases, and
broad governance/attention policy.
Every “validate first” item needs a named experiment, a success metric, and a
rollback-free interim behavior before it can move into the build DAG. This keeps
the implementation concrete without pretending that a reference pattern or an
ergonomic hunch is proof of user demand.
## Evidence register
Primary Orca reports used for this gate include [#4397](https://github.com/stablyai/orca/issues/4397), [#13360](https://github.com/stablyai/orca/issues/13360), [#13364](https://github.com/stablyai/orca/issues/13364), [#15148](https://github.com/stablyai/orca/issues/15148), [#15180](https://github.com/stablyai/orca/issues/15180), [#16075](https://github.com/stablyai/orca/issues/16075), [#16095](https://github.com/stablyai/orca/issues/16095), [#16522](https://github.com/stablyai/orca/issues/16522), and [#16660](https://github.com/stablyai/orca/issues/16660). Related delivery, routing, and cleanup reports are catalogued in `orchestration-issues.md`; current-main findings and exact code paths are in `orchestration-plan-review-technical.md`, `orchestration-audit-mailbox.md`, and `orchestration-audit-transcripts-refs.md`.
-607
View File
@@ -1,607 +0,0 @@
# Orca orchestration vNext: agent ergonomics review
## Verdict
The audit has the right durability and authority boundaries, but its internal
model is too visible in the proposed operator surface. A coordinator should not
have to reason about Run, Task, Attempt, Dispatch, Endpoint, Resource, lease,
receipt, evidence, delivery batch, and process incarnation during the happy
path. Those are valuable durable facts; they should appear in JSON, audit views,
and recovery diagnostics rather than becoming ten concepts every agent must
manipulate correctly.
The everyday model should be only:
- **Work**: what should be done, including dependencies.
- **Worker**: the agent currently trying that work.
- **Message**: durable communication with that worker or coordinator.
- **Attention**: a question, failure, uncertain outcome, or finished root that
requires action.
Internally, Work may map to Task, Worker to Attempt + Dispatch + Endpoint +
Resource, and Attention to projections over messages, receipts, evidence, and
leases. Stable internal IDs remain essential, but the CLI should carry them
forward in opaque receipts and exact next commands instead of requiring agents
to copy them between ordinary commands.
## Simplify the state model
### Show one primary status and keep evidence separate
The receipt sequence in the audit is useful for proofs, but nine receipt stages
should not become nine peer statuses in lists and notifications. Use one primary
status for scanning, with a short evidence line when expanded.
| Primary user status | Meaning | Evidence/detail shown on demand |
| ----------------------------------------------- | --------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- |
| `Starting` | Orca owns a start request, but agent work has not been observed. | Furthest delivery receipt such as `terminal_queued` or `provider_submitted`. |
| `Working` | A provider turn or other execution activity is observed. | `turn_started`, heartbeat age, process observation. |
| `Waiting for input` | A durable question or approval blocks progress. | Question, age, and exact reply command. |
| `Finished` | A valid result report was persisted and settled. | Report, artifact evidence, completion time. |
| `Needs review` | Execution appears finished but the result is missing, rejected, or contradictory. | `finished_unverified`/`outcome_unknown`, rejection reason, safe actions. |
| `Failed` | A terminal failure is proven. | Exit/report evidence and retry action. |
| `Connection lost — process status unverifiable` | The execution host cannot currently prove liveness or exit. | Last contact and reconnect/recovery action. |
`live`, `unverifiable`, and `exited` should remain the exact machine verdicts.
The longer connection-loss wording above is a presentation of
`processVerdict: "unverifiable"`, not a replacement verdict or an inference of
death.
Do not use `ready` as a user-facing lifecycle status. It currently means input
was accepted, which reads as though the worker is ready for more work. Prefer
`Starting — prompt submitted` or `Starting — queued for next turn`.
### Collapse delivery stages in normal output
Normal human output needs four statements:
1. **Saved by Orca** — retrying the same request will not duplicate it.
2. **Queued for the agent** — delivery is pending, possibly until its current
turn ends.
3. **Submitted to the agent** — provider submission is proven.
4. **Agent started** — a new turn is observed.
Keep the canonical detailed stages (`recorded`, `routed`,
`endpoint_delivered`, `terminal_queued`, `provider_submitted`, `turn_started`,
`report_persisted`, `settled`) in `--json` and `--verbose`. This preserves an
honest evidence ladder without asking an agent to distinguish routing from
endpoint delivery during a routine send.
Every mutating command should return the same receipt envelope:
```json
{
"result": "queued",
"summary": "Queued for the agent's next turn. No resend is needed.",
"requestId": "req_...",
"furthestStage": "terminal_queued",
"mayHaveApplied": true,
"retrySafe": true,
"nextAction": {
"label": "Wait for submission",
"argv": ["orchestration", "receipt-wait", "--request", "req_..."]
}
}
```
The runtime should generate the idempotency key. Agents should not need to
invent one before the first call. On disconnect, the CLI should print an exact
resume/status command using the generated request ID.
## Role-by-role experience
### Coordinator
The coordinator's normal loop should be:
1. Start work atomically.
2. Inspect one attention-first fleet view.
3. Wait for questions, failures, or completion.
4. Reply or recover using the exact suggested action.
5. Let Orca clean up fresh owned workers after completion is acknowledged.
`task-create` followed by `worker-start` remains useful for constructing a DAG,
but it is a poor default for one-off work because it exposes ID plumbing and can
leave an orphan Task. Add the audit's proposed atomic form:
```bash
orca orchestration worker-start \
--spec "Review authentication error handling" \
--worktree current --agent codex --json
```
The response should name the work in plain language, print a short worker alias,
and carry Task/Attempt/Dispatch/Resource IDs in JSON. Later commands may accept
the alias or Dispatch ID; emitted `nextAction.argv` should always use the stable
ID.
The default fleet query should be current-Run and attention-first. `--all`
should be explicit. A coordinator rarely needs every settled historical lane in
the primary view.
### Worker
The injected preamble should answer five questions near the top:
- What am I responsible for?
- Who receives my question and completion?
- May I delegate, and how many generations remain?
- Which one command reports success or failure?
- What must I do after reporting?
The current lifecycle commands are long because they correctly fence authority.
Keep the full command pre-rendered and copyable, but also allow the worker's
authenticated terminal to use concise context-bound verbs:
```bash
orca orchestration ask --question "Should legacy config remain supported?"
orca orchestration done --outcome succeeded \
--summary "Updated parsing and added compatibility tests" \
--files "src/config.ts,src/config.test.ts"
```
The runtime, not model-copied text, must bind these to the active Attempt. If the
terminal has zero or multiple active roles, the concise command must fail closed
and print the exact fully qualified command. The existing explicit
`send --type worker_done --task-id ... --dispatch-id ...` remains the portable
compatibility form.
After `done`, wording should be unambiguous:
> Completion recorded. Stop work on this assignment and return to an idle
> prompt. A new assignment will include a new dispatch preamble.
### Nested agent
Nested coordination is currently a setting and an identity rule, but the lived
experience is underspecified. Make nesting a Run policy inherited by every
Attempt, with the global setting acting as a ceiling/default. The preamble should
say either:
> Delegation: allowed for 1 more generation. Child work stays in this Run and is
> shown under your assignment.
or:
> Delegation: unavailable (depth 1 of 1). Complete this assignment yourself or
> ask your coordinator to change the Run policy.
A nested worker should inherit the Run and parent automatically. It must not
create a second Run to delegate, copy its parent's completion capability, or
choose between a coordinator mailbox and worker mailbox. One merged inbox should
label each item with its context and route replies to the originating thread.
If a parent tries to complete while children are active, reject the completion
with the child list and exact choices: wait, stop a child, or explicitly detach
it to the parent coordinator. Never silently settle the parent and orphan its
children.
## Surface-specific review
### Mailbox, wakeups, and acknowledgment
The durable mailbox fixes are necessary, but batch/cursor/reply-channel
mechanics are too easy to misuse. Preserve delivery batches internally while
making these rules visible:
- `check --wait` returns one replayable delivery and a single opaque
`deliveryId`.
- Every returned message has its own `messageId`, thread, work context, and
`replayed` flag.
- Acknowledgment is idempotent and may name the whole delivery or individual
messages.
- A type filter decides when the wait wakes; it must not silently hide older
actionable mail in the returned delivery.
- A timeout means “nothing new in this window,” not “worker failed.”
- Reply and question retries use the original request/message identity and do
not create duplicate `Re:` rows.
The CLI should end every delivery with the exact continuation command:
```text
3 messages delivered (delivery dlv_123; replayed: no).
Process all messages, then continue with:
orca orchestration check --ack dlv_123 --wait --attention
```
Add `--attention` as the normal semantic filter for `question`, `escalation`,
`worker_done`, uncertain outcomes, and proven failures. Keep `--types` for
advanced callers. An agent should not have to memorize the complete message
type vocabulary for the coordinator loop.
Protecting a human draft is a product invariant, not an error footnote. A wake
pointer may set an unread indicator, but it must not type into, submit, replace,
or focus a terminal containing a human draft.
Recommended wording:
- Mid-turn prompt: `Queued for the agent's next turn. No resend is needed.`
- Durable mail without endpoint proof: `Message saved. The agent has not been
notified yet; Orca will retry the wakeup.`
- Replayed delivery: `Replayed after reconnect; you may have seen this before.`
- Cross-Run authorization: `Run exists, but this terminal is not authorized to
message it.`
- Ask timeout: `Question is still pending. Resume waiting; do not ask again.`
### DAG construction
The dependency JSON and separate ID extraction make even a three-node graph
shell-heavy. Retain the low-level API, but add names and `--after` for the CLI:
```bash
orca orchestration task-create --name api \
--spec "Implement the API change"
orca orchestration task-create --name ui \
--spec "Update the UI against the agreed contract"
orca orchestration task-create --name integrate \
--spec "Run integration checks and resolve contract drift" \
--after api,ui
orca orchestration worker-start --task api --worktree current --agent codex
orca orchestration worker-start --task ui --worktree current --agent claude
orca orchestration status --tree
```
Names are Run-scoped aliases; durable IDs remain authoritative and appear in
JSON. Reject duplicate or ambiguous names before mutation.
The graph should explain blocked work rather than only printing `pending`:
```text
integrate Blocked by: api (Working), ui (Waiting for input)
Next: answer ui's question; integrate will become Ready automatically.
```
Missing recipes in the audit are: parallel fan-out then join, retrying only one
failed branch, superseding a branch without reopening completed siblings,
handling a parent with live nested children, and resuming a graph after a
runtime restart. These should be first-class documentation examples.
### Terminal and worker read
`terminal.read` and `worker-read` serve different jobs and should not look
interchangeable:
- `worker-read` answers “what has this agent said or done?” It is
transcript-first, Attempt-scoped, archived, and the orchestration default.
- `terminal.read` answers “what is visible/recent in this PTY?” It is a bounded
terminal snapshot and never completion proof.
Use these labels in normal output:
```text
Source: Exact provider transcript (Claude)
Scope: this worker attempt
```
or:
```text
Source: Terminal snapshot (advisory)
Reason: this provider does not expose a verifiable transcript
Scope: 50 recent screen lines; not completion evidence
```
Cursors should be opaque continuation tokens. Users should not need to
understand source identity, process incarnation, or byte position. If the
source changes, return:
```text
The read source changed after reconnect. This cursor cannot be continued.
Start a fresh read with:
orca orchestration worker-read --dispatch dsp_123
```
`worker-show` should expose `readCapability` as one of `exact_transcript`,
`structured_output`, `terminal_snapshot`, or `unavailable`, plus archive
availability. This lets coordinators decide whether reading is useful before
issuing one query per lane.
### Recovery and cleanup
The audit names stop, cancel, supersede, suspend, takeover, retry, abandon,
retain, release, TTL, and clear-Done. That is too many peer actions without a
decision recipe. Present three normal actions and move the rest under advanced
recovery:
- **Stop**: stop this work and clean up only its proven Orca-owned resource.
- **Retry**: create a new Attempt; require explicit placement and refuse to
duplicate a worker whose old liveness is not settled.
- **Clear finished**: archive and release eligible Orca-owned fresh terminals.
`abandon` remains an expert escape hatch meaning “stop tracking without stopping
resources.” `takeover` remains an authority operation. `supersede` is a durable
relationship created by retry/replacement, not something most users should have
to invoke directly. `retain` is presented as `--keep-terminal` with an optional
duration and reason.
Add a read-only diagnosis command:
```bash
orca orchestration worker-recover --dispatch dsp_123 --json
```
It should return the proven state, partial effects, ownership, whether mutation
may already have occurred, and ordered exact actions. It must not execute a
recovery merely because one looks likely.
Example uncertain-host flow:
```text
Status: Connection lost — process status unverifiable
Last proven: Working, 4m ago on build-linux
No retry was started because the old process may still be live.
Safe actions:
1. Reconnect and inspect:
orca orchestration worker-recover --dispatch dsp_123 --wait-reconnect
2. Stop the exact owned worker, then inspect the stop receipt:
orca orchestration worker-stop --dispatch dsp_123
3. Stop tracking without touching remote resources:
orca orchestration worker-abandon --dispatch dsp_123
```
For fresh worker terminals created and exclusively owned by Orca, default to
release after the coordinator acknowledges accepted completion. Preserve the
archive first. Never auto-release reused, external, user-taken-over,
identity-conflicted, or remote-unverifiable resources. `--keep-terminal` opts a
fresh worker out and records `Kept for debugging until ...`; omission should
not create a permanent lease.
This default removes a mandatory cleanup step from every successful lane while
retaining conservative safety boundaries. `clear --finished` remains an
idempotent bulk repair for eligible residuals and reports each retained row with
its reason and next safe action.
### Fleet query and attention
The fleet query risks becoming a wide dump of every durable fact. Default human
output should answer only: what is it, where is it, what is happening, how fresh
is that knowledge, and what should I do?
```text
WORK AGENT WHERE STATUS UPDATED NEXT
api-review Codex local / current Working 18s wait
ui-fix Claude win-dev / ui Waiting for input 2m reply
tests Codex child ssh-ci / folder Unverifiable 4m reconnect
integrate — — Blocked — waits: api,ui
```
Recommended commands:
```bash
orca orchestration status # current Run, attention first
orca orchestration status --tree # parent/child and dependencies
orca orchestration status --all --history # cross-Run historical view
orca orchestration status --json # full IDs, receipts, evidence, leases
```
The default sort should be: needs input/approval, failure, unverifiable or
unverified finish, active, dependency-blocked, recently finished. Do not use
provider, host, or creation time as the primary sort.
The single root-completion alert default is good, but define aggregation:
- Questions and approvals alert immediately and separately.
- Proven failures and interruptions alert immediately.
- Child success updates the tree without a desktop alert unless it unblocks a
root or the user opted in.
- Root success emits one alert after its completion is persisted.
- Reconnect/replay does not re-alert an already acknowledged event.
- No alert steals focus or modifies a draft.
## End-to-end CLI recipes
### One supervised worker
```bash
orca orchestration run-create --objective "Harden authentication" --json
orca orchestration worker-start \
--spec "Review error handling and add focused tests" \
--worktree current --agent codex --json
orca orchestration check --wait --attention --timeout-ms 900000 --json
# Process every message. If completion was accepted, the fresh owned terminal
# is archived and released when this delivery is acknowledged.
orca orchestration check --ack dlv_123 --wait --attention \
--timeout-ms 900000 --json
```
### Send while an agent is busy
```bash
orca orchestration send --to dispatch:dsp_123 \
--subject "Use the compatibility parser" \
--body "Keep the Git 2.25 fallback." --wait-submit --json
```
Expected immediate result if the provider is mid-turn:
```text
Queued for the agent's next turn. No resend is needed.
Request: req_456
Wait again with:
orca orchestration receipt-wait --request req_456
```
Repeating the printed request operation must observe the original send, not add
duplicate text.
### Worker asks and resumes after timeout
```bash
orca orchestration ask \
--question "Should I preserve the deprecated flag?" \
--options "preserve,remove" --timeout-ms 600000 --json
```
If the wait ends before a reply:
```text
Question is still pending. Resume waiting; do not ask again.
orca orchestration ask --resume msg_789 --timeout-ms 600000
```
### Retry a proven failed attempt
```bash
orca orchestration worker-recover --dispatch dsp_old --json
orca orchestration worker-retry --dispatch dsp_old \
--worktree current --agent codex --json
```
`worker-retry` is the composed ergonomic form of a new start with
`--retry-of`. It must require placement, atomically link the Attempts, keep old
evidence, and reject an unverifiable old worker with the safe stop/reconnect/
abandon choices.
### Read after cleanup
```bash
orca orchestration worker-read --dispatch dsp_123 --limit 100
```
The same command should work after release, state that the source is an archive,
and retain the same Attempt boundary.
## Acceptance tests
### Coordinator happy path
- Starting one worker with `worker-start --spec` atomically creates Work and an
Attempt. A launch failure leaves either no Work or one receipt with an exact,
idempotent resume action; it never leaves an unexplained orphan Task.
- The start response can be used without copying a Task ID into a second
mutation, while JSON still exposes every stable internal ID.
- `status` defaults to the current Run and attention-first ordering; `--all` is
required to include unrelated Runs.
- A valid completion enters `Finished`, archives readable output, and releases
only the newly created, proven-owned worker after delivery acknowledgment.
- A reused or user-owned terminal is never auto-released and explains why.
### Mail and receipts
- A mid-turn prompt returns the exact wording `Queued for the agent's next
turn. No resend is needed.` and a generated request ID.
- Retrying or waiting by request ID produces one provider submission and one
message row across disconnect, CLI restart, and runtime restart.
- Insertion between durable read and waiter registration wakes the waiter.
- A replayed delivery is labeled, preserves message IDs, and does not repeat an
acknowledged root alert.
- Per-message and whole-delivery acknowledgments are idempotent; acknowledging
one message does not hide unacknowledged siblings.
- Reply retry/resume creates no duplicate `Re:` row.
- Cross-Run unauthorized and unknown-Run errors are distinct without leaking
content.
- Wakeups never submit, overwrite, focus, or erase an unsent human draft.
### Worker and nested-agent experience
- Every injected preamble states assignment, completion recipient, remaining
delegation depth, the exact done command, and post-completion behavior.
- Context-bound `done` succeeds only when the caller has exactly one attested
active Attempt; zero/multiple contexts fail closed with the fully qualified
command or an explicit disambiguation action.
- A child cannot use or inherit its parent's completion capability.
- Allowed nested start automatically inherits Run and parent; no `run-create`
workaround or copied parent ID is required.
- A depth rejection says `Delegation unavailable (depth N of N)` and gives the
two safe choices: complete locally or ask the coordinator.
- A parent cannot settle with live children unless they are stopped, settled,
or explicitly detached; the error lists them and exact commands.
- Worker-role and nested-coordinator mail appear in one ordered inbox with
context labels and replies route to the originating thread.
### DAG usability
- Run-scoped names and `--after name1,name2` create the same durable dependency
graph as ID-based JSON; duplicate/ambiguous names fail before mutation.
- A join node explains each blocking predecessor and becomes Ready
automatically when all dependencies settle successfully.
- Retrying one failed branch does not reopen completed siblings or duplicate
the join.
- After runtime restart, the same tree, attention ordering, and ready set are
reconstructed from durable facts.
### Read contract
- Claude, Codex, and an unsupported provider show one of the documented
`readCapability` values before a read.
- `worker-read` labels exact transcript, structured output, terminal snapshot,
and archive sources; terminal fallback always says it is advisory and not
completion evidence.
- `terminal.read` never changes lifecycle status based on quiet, prompt-like,
or ghost screen text.
- Cursor continuation is Attempt- and source-fenced. Source rotation returns an
exact fresh-read command rather than a generic cursor error.
- The same bounded read works for a released worker and for local, folder, SSH,
WSL, and federated execution without desktop-side remote file access.
### Recovery and fleet safety
- `worker-recover` is read-only and returns proof, partial effects,
`mayHaveApplied`, resource ownership, and ordered exact actions.
- Relay or SSH loss yields `processVerdict: unverifiable`, never `exited`, and
no automatic retry or release.
- Retry requires explicit placement and cannot create concurrent ownership
unless the operator explicitly chooses a separate parallel lane.
- Stop, retry, release, retain, abandon, and bulk clear are idempotent across
response loss and runtime restart.
- `clear --finished` releases only eligible proven-owned resources and lists a
retention reason plus safe next action for every residual.
- A five-worker wave emits at most one default root-success alert while
questions, approvals, failures, and interruptions remain separate.
- Old peers omit unsupported detail safely: the primary status remains honest,
unavailable evidence is labeled, and no unnegotiated opcode is required.
## Documentation recipes required before rollout
The implementation audit has strong component acceptance gates but needs a
task-oriented guide. Ship these recipes with vNext:
1. Supervised delegation versus full ownership handoff.
2. Atomic one-worker start, wait, reply, completion, and cleanup.
3. Parallel fan-out and dependency join using names and `--after`.
4. Safe mid-turn send, receipt wait, disconnect resume, and no-resend rule.
5. Worker question timeout and exact `ask --resume` flow.
6. Nested delegation allowed, depth rejected, and parent completion with live
children.
7. Missing/rejected `worker_done` producing `Needs review` rather than
indefinite `ready`.
8. Remote `unverifiable` recovery without false exit, retry, or cleanup.
9. Read exact transcript, advisory terminal fallback, and archived output after
release.
10. Runtime restart with mailbox replay and DAG/fleet reconstruction.
11. Reused terminal, user takeover, keep-for-debugging, TTL expiry, and bulk
clear safety.
12. Mixed-version local/folder/SSH/WSL/federated behavior and capability
downgrade.
## Priority changes to the implementation plan
1. In **P0**, define the small primary-status projection and the uniform
receipt envelope alongside the detailed contract vocabulary.
2. In **P2**, add generated request IDs, exact resume actions, `--attention`,
and user-facing no-resend wording; do not expose mailbox epochs or wake
cursors in ordinary output.
3. In **P1/P3**, support context-bound worker `ask`/`done` commands without
weakening Attempt authority, and specify parent completion with live
children.
4. In **P3**, add atomic `worker-start --spec`; retain separate Task creation
for planned DAGs.
5. In **P4**, make recovery diagnosis read-only, make retry composed and
explicit about placement, and default eligible fresh owned workers to
archive-then-release after completion acknowledgment.
6. In **P5**, make opaque continuation tokens and `readCapability` part of the
public contract; label PTY output as advisory in every surface.
7. In **P6**, make current-Run, attention-first status the default and move the
full identity/receipt/evidence/lease record to JSON and detail views.
8. In **P7**, test the complete user recipes, not only individual transition
contracts, across provider, OS, folder workspace, SSH, WSL, federation, and
mixed-version cases.
These changes preserve the audit's durable facts while making the normal
experience teachable: start work, watch attention, answer or recover, and read
the result. The complexity remains available exactly where it is needed—proof,
debugging, compatibility, and safe recovery—without becoming the price of every
successful delegation.
-577
View File
@@ -1,577 +0,0 @@
# Orca orchestration skill rewrite audit and replacement draft
Date: 2026-08-27
Runtime audited: Orca `1.4.191-adhoc.20260827054943`
Live source: `orca skills get orchestration` (435 lines)
## Executive recommendation
Replace the current flat guide with a phase-loaded kernel: keep role classification, authority, the common supervised loop, the worker completion contract, and the remote uncertainty floor always loaded; move placement variants, mailbox details, lifecycle recovery, low-level topology, and legacy adoption into references loaded at their acting boundary.
The rewrite should preserve the current CLI grammar and safety contracts. Its main behavior change is instructional: `worker-start` becomes the only normal-path recipe, a worker sees its completion contract near the top, and uncommon compatibility detail no longer displaces the coordinator loop.
The current CLI serves one Markdown document and the installed skill is only a discovery stub. Progressive disclosure therefore needs a packaging decision: either teach `skills get` to materialize a version-matched skill package with references, or ship the compact kernel as the served guide and use command `--help` plus a separately addressable versioned reference surface. Do not add relative references that the installed/served skill cannot actually load.
## Evidence and design influence
The audit used:
- The version-matched orchestration guide and current command help for `worker-start`, `check`, `worker-show`, `worker-release`, `run-use`, and `send`.
- The current guide source and its prose-coupled regression tests in `config/scripts/orchestration-skill-guidance.test.mjs`.
- `orchestration-vnext-implementation-audit.md`, `orchestration-issues.md`, and the three focused audit artifacts in this worktree.
- The compound-engineering plugin's root agent instructions, portable skill-authoring standard, `ce-skill-work` review rules, phase-loaded skills (`ce-plan`, `ce-work`, `ce-code-review`), and specialist prompt assets.
The useful compound-engineering patterns are:
1. Lead with an outcome spine: result, next consumer, done condition, and safe failure direction.
2. Keep the protocol kernel inline; load conditional or late mechanics immediately before they act.
3. Give each reference one owning concern. Do not duplicate commands across the caller and callee.
4. Treat a description as an activation pointer, not a feature catalog.
5. Give delegated work a distinct scope, output contract, and synthesis owner.
6. Pin one fragile command recipe, then provide a named failure hatch.
## Audit findings
### Change: the common path is not the document's spine
The guide introduces tool boundaries and then places roughly sixty lines of contract migration before the ownership model and normal coordinator loop. A first-time coordinator reaches `worker-start` only after messaging, Task/Dispatch detail, and nesting rules. A dispatched worker's most important rule—send `worker_done` exactly once, then end the dispatched turn—appears near the end.
Requested condition: the normal role must be classifiable and executable from the always-loaded kernel. A coordinator should reach `run-create -> task-create -> worker-start -> check -> release/reuse` before optional detail; a worker should reach its exact completion contract before coordinator-only mechanics.
### Change: duplicated boundaries invite drift
Full-handoff classification is repeated in the description, When To Use, Ownership, Full Handoffs, Worker Terminals, and Next Action. Placement is split between the preferred loop, Full Handoffs, and Worker Terminals. Lifecycle cleanup appears in the preferred loop, Agent Guidance, and Next Action.
Requested move: state each condition once at its owning layer. The orchestration skill should classify a full handoff and route to `orca-cli`; it should not reproduce `orca-cli`'s worktree, custom-model, and terminal-send recipes.
### Change: the final example teaches the fallback
The guide calls `worker-start` preferred, but its final example manually creates a terminal, waits for TUI idle, and uses low-level `dispatch --inject`. This makes the uncommon unsupervised topology the memorable recipe.
Requested condition: the canonical example must use `worker-start`. Low-level dispatch belongs in a conditional reference whose entry criterion is “the composed start cannot express the required topology or argv.”
### Change: role-specific obligations are mixed
Coordinator, worker, full-handoff owner, legacy worker, and recovery operator instructions share one linear document. This causes rules such as `--from` omission for coordinators to sit beside injected worker commands that intentionally include `--from` and a dispatch capability.
Requested move: route by role near the top and give worker/coordinator protocols separate owned sections. The worker must copy the exact injected command, including executable, terminal handle, capability, Task ID, and Dispatch ID; generic coordinator advice must not override it.
### Change: heartbeat syntax has two sources of truth
The live guide's Agent Guidance shows raw `--payload` JSON for heartbeat, while the current `send --help` and injected preamble support typed `--task-id`, `--dispatch-id`, and `--phase` flags. Typed flags avoid PowerShell JSON quoting failure and match the live worker contract.
Requested move: make the injected preamble authoritative and show typed flags in the generic worker recipe. Keep raw payload as a compatibility implementation detail, not the taught path.
### Change: remote safety is scattered instead of being a floor
The guide correctly says that remote work is addressed by Dispatch ID after start and that remote `current`/`new-child` are invalid, but the SSH verdict vocabulary is only visible inside recovery detail. The project instruction is stricter: the execution host owns execution-sensitive facts, and contact loss is never process death.
Requested condition: every coordinator path must preserve execution-host ownership and the verdicts `live`, `unverifiable`, and `exited`; later operations route by Dispatch ID and never by guessed local terminal state.
### Change: prose regression tests encode layout, not only contracts
The current test suite asserts exact headings, sentences, and command snippets. Those tests protect real incidents, but they also make progressive disclosure look like contract deletion because moving a rule to a reference fails the test.
Requested move: retain incident coverage while relocating assertions to the owning reference and add package-integrity tests. Test observable obligations and required load stubs, not accidental section placement.
### Verify: versioned reference delivery
`orca skills get orchestration` currently prints one document, and the installed skill directory contains only `SKILL.md`. Before introducing reference paths, verify how every supported harness receives and resolves a version-matched skill package. If package delivery is not available, use the single-file fallback described below.
### Verify: `check --format` help shape
The live guide uses `check --peek --format --json` as a boolean formatting flag, while current help renders the option description as `--format <png|jpeg>`. Confirm whether this is only generic option metadata leakage or a real CLI contract mismatch before carrying the recipe into the rewrite.
### Consider: a worker preamble mini-kernel
The injected preamble already gives workers exact lifecycle commands. Keeping the worker section in the general guide is still necessary for inherited-context classification and recovery, but Orca could version and test the preamble as a small standalone contract. This would reduce reliance on a worker searching the coordinator guide after dispatch.
## Proposed package structure
```text
orchestration/
├── SKILL.md
└── references/
├── coordinator-loop.md
├── worker-contract.md
├── placement-and-remote.md
├── messaging-and-gates.md
├── recovery-and-cleanup.md
├── low-level-topology.md
└── legacy-contract-migration.md
```
`SKILL.md` should target 140-190 lines. It owns activation, role classification, the authority floor, the canonical loop, reference routing, and completion. Each reference owns commands and edge cases for one concern; no reference restates the entire loop.
If the CLI cannot deliver references atomically, flatten only the five safety-critical sections into the served guide and expose the remaining material through current command `--help`. Do not create dead read instructions.
## Replacement `SKILL.md` draft
The following is a content draft, not a source patch.
````markdown
---
name: orchestration
description: >-
Coordinate supervised Orca workers with durable Runs, Tasks, Dispatches,
messages, questions, gates, and completion tracking. Use when the user asks
to supervise, monitor, wait for results, coordinate a DAG, or manage blocking
agent-to-agent questions. For full ownership handoffs or ordinary terminal,
worktree, and built-in-browser control, use `orca-cli`.
---
# Orca orchestration
## Outcome
**Result:** every in-scope Task has one explicit terminal outcome and every
settled worker terminal has a next owner or cleanup decision.
**Next consumer:** the coordinator synthesizes accepted worker results for the
user or starts the next ready wave.
**Done:** all expected Dispatches have settled; every delivered message was
processed before acknowledgment; and each settled worker was immediately
reused, explicitly retained, or released.
**Safe failure:** preserve work and authority, report `outcome_unknown` or
`unverifiable`, and expose the next safe command. A timeout, quiet terminal,
missing client, or lost remote connection is never proof of failure or exit.
## Classify the request
| Context | Act as | Route |
| ------------------------------------------------------------------------------------------------------------------- | ---------------------- | --------------------------------------------------------------------------------------- |
| The user explicitly asks to supervise, monitor, wait for results, coordinate a DAG, use a gate, or manage ask/reply | Coordinator | Use the supervised loop below |
| The current prompt contains a live injected Dispatch preamble with Task and Dispatch IDs | Worker | Read `references/worker-contract.md` now and follow the preamble exactly |
| The user asks to hand off ownership or start another agent/worktree without supervision | Handoff owner | Invoke `orca-cli`; do not create a Run, Task, or Dispatch and do not monitor completion |
| A message has a legacy authority label | Compatibility operator | Read `references/legacy-contract-migration.md` before any lifecycle mutation |
| No live preamble and no explicit supervision | Ordinary agent | Do not emit lifecycle messages; use `orca-cli` for terminal/worktree operations |
Model or effort selection does not make a handoff supervised. Never substitute
a non-Orca subagent tool when Orca orchestration provenance was requested.
## Authority and identity floor
- A Run is a durable namespace and coordinator inbox; it does not schedule or
place workers. A Task is work; a Dispatch is one authoritative Task attempt.
- Lifecycle authority comes from the active Dispatch, not a terminal title,
copied ID, old database row, provider transcript, or visible pane.
- Workers send lifecycle messages from their own dispatched terminal using the
exact executable, handle, capability, Task ID, and Dispatch ID injected by
Orca. Do not reconstruct or broaden those arguments.
- After remote start, address the worker by Dispatch ID. The execution host owns
process, filesystem, transcript, stop, and cleanup facts. Preserve
`live` / `unverifiable` / `exited`; loss of contact is not process death.
- Folder workspaces are valid placements. Do not require Git or assume every
workspace is a worktree.
- Treat unknown fields as absent on mixed versions. Do not send a new remote
stream operation unless the connected server advertised it.
Examples use `orca`; use the executable selected by the discovery stub for the
whole run. If it fails, report that exact error rather than switching binaries.
## Canonical supervised loop
Confirm the runtime, create or bind one Run, create all independent Tasks, then
start the independent wave before waiting:
```bash
orca status --json
orca orchestration run-create --objective "<objective>" --json
orca orchestration task-create --spec "<worker A task>" --json
orca orchestration task-create --spec "<worker B task>" --json
orca orchestration worker-start --task <task_a> --worktree current --agent codex --json
orca orchestration worker-start --task <task_b> --worktree current --agent claude --json
orca orchestration check --wait --types "worker_done,escalation,question" --timeout-ms 900000 --json
```
````
Use Task dependencies for real ordering. Prefer waves over chains deeper than
three or four steps. Nested workers obey the runtime depth limit; creating a new
Run never resets the caller's depth.
Process every message in the returned Delivery. Reply to questions, validate
that each `worker_done` belongs to the expected active Dispatch, and decide the
terminal's next owner before acknowledging:
```bash
# Question in the Delivery:
orca orchestration reply --id <message_id> --body "<answer>" --json
# Settled worker with no immediate follow-up:
orca orchestration worker-release --dispatch <dispatch_id> --json
# Then acknowledge the whole processed Delivery and continue waiting:
orca orchestration check --ack <delivery_id> --wait --types "worker_done,escalation,question" --timeout-ms 900000 --json
```
A timeout or empty result is a checkpoint. Keep waiting while the Dispatch is
live or unverifiable. Read `references/recovery-and-cleanup.md` only when a
worker fails to start, stops, reports an unknown outcome, needs retry, or cannot
be released normally.
## Placement gate
Use `current` or an exact existing workspace by default. A fresh worker means a
fresh agent terminal, not a new Git worktree. Create a new worktree only when
the user requested one or a concrete checkout/filesystem conflict makes sharing
unsafe. Before any new, remote, SSH, or WSL placement, read
`references/placement-and-remote.md`.
`worker-start` is the normal lifecycle owner. Read
`references/low-level-topology.md` only when it cannot express required custom
argv or topology. Low-level `dispatch --inject` is tracked but unsupervised and
does not grant `worker-stop` ownership of the operator-created process.
## Messaging and gates
Use `dispatch:<dispatch_id>` for attempt-specific coordinator guidance. Use
`ask`/`reply` for a worker's blocking question and coordinator response. Use a
gate only for a coordinator-managed DAG decision. For inbox replay, group
addresses, cursors, and gate recipes, read `references/messaging-and-gates.md`.
## Completion
After an accepted success or failure report, immediately do exactly one:
1. Reuse the same proven agent terminal for an immediate follow-up Dispatch.
2. Record user-requested retention with `worker-retain`.
3. Run `worker-release`.
Release is post-settlement cleanup, not cancellation. Never release because of
idle state, timeout, heartbeat, status, question, escalation, or a rejected or
stale completion. Released output remains available through `worker-read`.
Do not manually mark a Task completed after a valid `worker_done`; settlement
already updates the Task and Dispatch. Do not end the coordinator turn until
all expected Dispatches and settled worker terminals are accounted for.
````
## Reference ownership and required content
### `references/coordinator-loop.md`
This is optional if the canonical loop remains fully inline. If used, the inline kernel must retain the loop order and stop classes; the reference may own expanded DAG recipes, `task-list --ready --brief`, model/effort selection, and same-terminal follow-up.
Required recipes:
```bash
orca orchestration task-create --spec "<dependent work>" --deps '["<task_id>"]' --json
orca orchestration task-list --ready --brief --json
orca orchestration worker-show --dispatch <dispatch_id> --json
orca orchestration worker-start --task <next_task_id> --terminal <agent_terminal_handle> --json
````
Rules:
- `--model` applies to a fresh Claude, Codex, or Cursor launch; `--effort` requires `--model`.
- Neither `--model` nor `--effort` combines with `--terminal`.
- Compare `launch.requested` with `launch.effective`; do not claim a model from requested arguments alone.
- A review-only completion authorizes synthesis, not coordinator edits. Preserve any next owner named by the user.
### `references/worker-contract.md`
This reference begins with: “The injected preamble is authoritative. Copy its command rather than reconstructing flags.”
Required worker recipes, parameterized by the exact injected executable, handle, and capability:
```bash
<ORCA> orchestration send --from <worker_handle> --dispatch-capability <capability> \
--type heartbeat --subject "alive" \
--task-id <task_id> --dispatch-id <dispatch_id> --phase "implementing"
<ORCA> orchestration ask --from <worker_handle> --dispatch-capability <capability> \
--question "<question>" --options "<choice-a>,<choice-b>" --timeout-ms 600000
<ORCA> orchestration ask --from <worker_handle> --dispatch-capability <capability> \
--resume <message_id> --timeout-ms 600000
<ORCA> orchestration send --from <worker_handle> --dispatch-capability <capability> \
--type worker_done --subject "<short status>" \
--body "<three sentences: work, findings, remaining>" \
--task-id <task_id> --dispatch-id <dispatch_id> \
--outcome succeeded --files-modified "path/a,path/b" \
--report-path "<optional durable report>"
```
Rules:
- Send heartbeat only at the cadence requested by the live preamble. It proves liveness, not completion.
- Use `ask`, never a local user-question TUI, when the coordinator must answer. A timeout leaves the question pending; resume the same message ID.
- Use escalation only when the coordinator must intervene before completion.
- Send `worker_done` exactly once with explicit `succeeded` or `failed`; never encode failure only in prose.
- After `worker_done`, end the dispatched turn and idle. A direct user instruction starts new user-owned work and must not reuse settled lifecycle IDs.
### `references/placement-and-remote.md`
This reference owns placement and execution-host boundaries.
Normal recipes:
```bash
# Fresh agent in the current workspace; setup is not rerun.
orca orchestration worker-start --task <task_id> --worktree current --agent codex --json
# New stacked child worktree.
orca orchestration worker-start --task <task_id> --worktree new-child --name <name> --agent codex --setup run --json
# New independent top-level worktree.
orca orchestration worker-start --task <task_id> --worktree new-top-level --name <name> --agent codex --setup run --json
# Connected server; later commands omit --on and route by Dispatch ID.
orca orchestration worker-start --task <task_id> --on <environment> \
--worktree new-top-level --repo <exact_remote_repo_selector> \
--name <name> --agent codex --setup run --json
```
Rules:
- Current and exact existing workspaces create a fresh agent terminal unless `--terminal` is explicit.
- New worktrees use agent-first creation and run setup by default. `start-immediately` may show setup `running` while the worker is ready; only repository `wait-for-setup` gates prompt delivery on setup success.
- Child/top-level Orca lineage, Git base, filesystem isolation, coordination parentage, UI grouping, and execution host are separate decisions.
- Remote `current` and `new-child` are invalid. Use an exact remote workspace selector, or `new-top-level` with an exact remote repository selector.
- Pass `--on` only to `worker-start`. Use `worker-show`, `worker-read`, `send --to dispatch:<id>`, `worker-stop`, and cleanup by Dispatch ID afterward.
- The execution host owns process and filesystem facts. Never replace a remote action with a local terminal command or local file inspection.
- For SSH and disconnected relays, preserve `live`, `unverifiable`, and `exited` exactly. No contact is `unverifiable`, not `exited`.
- Folder workspaces must remain valid when the task has no Git worktree.
### `references/messaging-and-gates.md`
Required recipes:
```bash
orca orchestration send --to dispatch:<dispatch_id> --subject "Follow-up" --body "<guidance>" --json
orca orchestration check --wait --types "worker_done,escalation,question" --timeout-ms 900000 --json
orca orchestration reply --id <message_id> --body "<answer>" --json
orca orchestration gate-create --task <task_id> --question "<decision>" --options '["a","b"]' --json
orca orchestration gate-resolve --id <gate_id> --resolution "<choice>" --json
```
Rules:
- A consuming coordinator check returns the bound Run's oldest FIFO Delivery, up to 50 messages, and replays it until acknowledged.
- Process the complete Delivery before `--ack`; type filters decide when a waiter wakes, not which older actionable mail may be skipped.
- `--peek` and `--all` are read-only inspection, not coordinator progress.
- Group addresses are for intentional fan-out status or questions, never Dispatch lifecycle messages.
- `worker_done` and heartbeat are Dispatch-scoped and never target groups.
- `ask` is a worker question; gates are coordinator-owned DAG decisions.
### `references/recovery-and-cleanup.md`
The recovery decision table should be the reference's opening content:
| Proven state | Safe action |
| ---------------------- | ------------------------------------------------------------------ |
| `ready` or active | Keep waiting; optionally read bounded output |
| `failed` or `stopped` | Start a replacement with `--retry-of`; repeat placement explicitly |
| `outcome_unknown` | Inspect; then choose `worker-stop` or explicit `worker-abandon` |
| accepted `worker_done` | Reuse, retain, or release |
| remote contact lost | Preserve `unverifiable`; do not stop/retry from absence alone |
Required recipes:
```bash
orca orchestration worker-show --dispatch <dispatch_id> --json
orca orchestration worker-read --dispatch <dispatch_id> --limit 50 --json
orca orchestration worker-start --task <task_id> --retry-of <dispatch_id> \
--worktree <explicit-placement> --agent <agent> --json
orca orchestration worker-stop --dispatch <dispatch_id> --json
orca orchestration worker-abandon --dispatch <dispatch_id> --json
orca orchestration worker-retain --dispatch <dispatch_id> --json
orca orchestration worker-release --dispatch <dispatch_id> --json
```
Rules:
- Retry never inherits placement silently.
- `worker-abandon` fences orchestration without claiming or causing remote/process/filesystem effects.
- `worker-stop` closes only the proven supervised agent terminal; it never deletes the worktree, setup terminal, configured tabs, or unrelated processes.
- `worker-release` is idempotent and archives output before closing the exact owned terminal.
- If release returns `release_pending` or `release_unknown`, follow its exact recovery receipt. Never substitute `terminal close`.
- Reset is destructive recovery only; never run it during active coordination unless the user explicitly abandons that state.
### `references/low-level-topology.md`
Entry condition: `worker-start` cannot express required custom argv or terminal topology.
```bash
orca terminal create --worktree active --title <task-name> --command "<agent-command>" --json
orca terminal wait --terminal <handle> --for tui-idle --timeout-ms 60000 --json
orca orchestration dispatch --task <task_id> --to <handle> --inject --json
```
Rules:
- Wait for readiness before injecting only when startup could lose the prompt.
- An operator-created terminal attached with `dispatch --inject` remains unsupervised. `worker-stop`, `worker-abandon`, and `worker-release` do not own or close that process.
- Use `worker-start --terminal <handle>` when supervised lifecycle ownership is required.
- Never use this path for a full handoff.
### `references/legacy-contract-migration.md`
This reference preserves the full current migration contract and should be loaded only when an authority label, adopted Run, compatibility recovery receipt, or explicit legacy takeover is present.
Its opening rules must remain verbatim in meaning:
- `[LEGACY COMPATIBILITY]`: live and attested; run only the exact printed command with the same selected executable and arguments.
- `[LEGACY RECOVERY REPLAY — MAY HAVE BEEN SEEN]`: one bounded at-least-once replay; process idempotently and acknowledge only as instructed.
- `[LEGACY READ-ONLY]`: inspection only; no reply, acknowledgment, or lifecycle mutation.
- An explicitly selected current Run, current binding, current Dispatch, or federated attachment takes precedence over legacy fallback.
- Unproven liveness, principal ownership, capability, or contract degrades to read-only inspection; it never falls back to local mutation.
- Adoption preserves the live process, PTY/session, terminal, tab/pane, workspace, Task, and Dispatch. It never restarts the worker or revives the retired scheduler.
Required takeover recipe:
```bash
orca orchestration run-use --id <adopted_run_id> --takeover-legacy --json
orca orchestration check --run <adopted_run_id> --json
```
Takeover runs only from the new live coordinator terminal when the original coordinator is unavailable or cannot prove authority. It fences the old coordinator, not workers. Do not take over an actively coordinated Run.
On packaged Windows, preserve the status-75 two-step legacy ask protocol and run the exact printed `ask --resume <message_id>` command. For an attested WSL launch, preserve the printed `orca-ide` executable and distro route. Never translate structured recovery arguments from memory.
## Command recipe index
| Goal | Canonical command | Load first when conditional |
| ----------------------------- | --------------------------------------------------------------------------------------- | ----------------------------------------- |
| Create coordination namespace | `orca orchestration run-create --objective "<text>" --json` | Kernel |
| Create work | `orca orchestration task-create --spec "<text>" --json` | Kernel |
| Start normal worker | `orca orchestration worker-start --task <id> --worktree current --agent <agent> --json` | Kernel |
| Start remote worker | `worker-start ... --on <environment> --worktree new-top-level --repo <exact-selector>` | `placement-and-remote.md` |
| Wait for actionable mail | `check --wait --types "worker_done,escalation,question" --timeout-ms 900000 --json` | Kernel |
| Answer worker | `reply --id <message_id> --body "<answer>" --json` | Kernel |
| Guide one attempt | `send --to dispatch:<dispatch_id> ... --json` | `messaging-and-gates.md` |
| Inspect lifecycle | `worker-show --dispatch <dispatch_id> --json` | `recovery-and-cleanup.md` |
| Read output | `worker-read --dispatch <dispatch_id> --limit 50 --json` | `recovery-and-cleanup.md` |
| Retain settled terminal | `worker-retain --dispatch <dispatch_id> --json` | `recovery-and-cleanup.md` |
| Release settled terminal | `worker-release --dispatch <dispatch_id> --json` | Kernel |
| Retry proven failure | `worker-start --retry-of <old> ...` with explicit placement | `recovery-and-cleanup.md` |
| Custom topology | `terminal create` then `dispatch --inject` | `low-level-topology.md` |
| Full handoff | Invoke `orca-cli` | Never load orchestration lifecycle detail |
## Anti-patterns
| Anti-pattern | Why it fails | Correct condition |
| ------------------------------------------------------------------------ | -------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- |
| Spawn a generic subagent and call it Orca orchestration | No Run/Task/Dispatch provenance, injected authority, or durable settlement | When coordination state matters, use Orca `task-create` plus `worker-start` |
| Create Task/Dispatch state for a full handoff | Invents a coordinator and monitoring obligation the user did not request | Route ownership transfer to `orca-cli` and stop monitoring |
| Teach low-level `terminal create + dispatch --inject` as the default | Produces an unsupervised process and duplicates lifecycle mechanics | Use `worker-start`; load low-level topology only for an expressiveness gap |
| Start waiting before all independent workers | Serializes independent work | Create Tasks and start the full ready wave before the first wait |
| Treat a wait timeout or TUI idle as completion/failure | Observation is not settlement | Keep waiting or inspect; preserve unknown state |
| Retry because a remote endpoint disappeared | Contact loss does not prove exit and may duplicate editing | Render `unverifiable`; recover only from positive evidence or explicit operator choice |
| Release on heartbeat, timeout, question, escalation, or stale report | Can close a live worker | Release only after accepted settlement |
| Use `terminal close` when release is uncertain | Bypasses ownership fencing and may close user resources | Follow the exact release recovery receipt |
| Ack a Delivery before processing all rows | Whole-batch ack can hide unprocessed actionable mail | Process every row and resource decision, then ack |
| Re-send an `ask` after timeout | Creates duplicate question threads | Resume the original message ID |
| Send lifecycle messages to `@all` or a provider group | Lifecycle belongs to one Dispatch | Omit `--to` as a worker; use `dispatch:<id>` for coordinator guidance |
| Guess provider session IDs, transcript paths, or remote terminal handles | Breaks source/host identity fencing | Use `worker-read` and Dispatch routing |
| Create a new Run from a nested worker to bypass depth | Depth follows the terminal's active Dispatch | Respect `nested_worker_depth_exceeded`; finish locally or escalate |
| Replace a legacy worker after update because authority is unclear | Risks two editors on one filesystem | Keep the original editor; use read-only inspection until a stable handoff |
| Use raw JSON heartbeat payloads in taught recipes | PowerShell quoting is fragile and typed flags exist | Use `--task-id`, `--dispatch-id`, and `--phase` from the injected contract |
## Migration and compatibility notes
### Preserve the discovery stub
Keep the safe executable resolver unchanged:
1. `ORCA_CLI_COMMAND` when set.
2. `orca-dev` only when `ORCA_DEV_REPO_ROOT` selects the dev checkout.
3. `orca-ide` on Linux outside an Orca-managed terminal, avoiding the GNOME screen reader.
4. `orca` otherwise.
The stub must load the version-matched guide before any orchestration mutation and keep the bounded read-only fallback for binaries that explicitly report `skills get` as unknown.
### Deliver references atomically or do not reference them
Preferred migration:
1. Add a versioned guide-package format containing `SKILL.md` and its references.
2. Make installation/materialization atomic and scoped to the exact CLI build.
3. Keep `orca skills get orchestration` backward-compatible by printing the compact main file.
4. Add a capability or explicit subcommand for retrieving reference content; do not infer package support from version strings.
5. For older clients, serve a flattened compatibility guide containing the kernel plus legacy reference.
If this is too large for the skill-only change, ship a compact single-file rewrite first. Use `orca orchestration <command> --help` as the late-loaded command source and retain migration details in a collapsed final section.
### Preserve current CLI and preamble contracts
- Do not rename Runs, Tasks, Dispatches, `worker_done`, or current status values in the skill rewrite.
- Keep old injected preambles valid for their active Dispatch. New Dispatches use the current grammar.
- Keep `worker_done` settlement automatic and exactly-once from the active worker terminal.
- Preserve explicit `--outcome succeeded|failed`, Task/Dispatch IDs, files, and optional report path.
- Keep task failure circuit breaking and nested-depth semantics unchanged.
- Keep low-level dispatch recognized as unsupervised; documentation must not imply resource ownership retroactively.
### Preserve mixed-version and remote wire behavior
- New fields remain optional. Old peers may omit them without being treated as failed.
- New stream opcodes require advertised capability; unknown opcodes may be dropped silently.
- A current Run remains authoritative on its home server. `--on` selects only worker placement, not the Run home.
- Address cross-server follow-ups, reads, stop, and cleanup through the Dispatch; do not send remote terminal handles across the boundary.
- Remote host loss yields `unverifiable`. Never synthesize process death from client inventory, relay absence, or timeout.
- Keep folder workspaces and non-Git tasks operational. Git lineage guidance must be conditional on a Git worktree actually existing.
### Migrate prose tests by invariant
Keep the current incident-backed assertions, but relocate them with ownership:
| Existing tested contract | New owner |
| --------------------------------------------------------- | -------------------------------------------- |
| Real Orca provenance and no generic subagent substitution | `SKILL.md` authority floor |
| Long waits are checkpoints | `SKILL.md` safe failure + recovery reference |
| Full handoff does not create lifecycle state | `SKILL.md` classifier + `orca-cli` tests |
| Review-only and named-next-owner boundaries | coordinator reference |
| Exactly-once completion and post-completion idle | worker reference and preamble tests |
| Reuse/retain/release after accepted settlement | `SKILL.md` completion + recovery reference |
| No release from idle/timeout/heartbeat | recovery reference |
| Agent-first placement and handle recovery | placement reference |
| Legacy adoption/read-only/takeover | legacy reference |
| Version-matched safe resolver | discovery-stub tests |
Add package tests that every inline required read resolves within the versioned skill package and that a missing reference blocks before its governed mutation.
## Acceptance tests and eval scenarios
Mechanical tests:
1. Frontmatter is identical between installed stub and served guide.
2. The stub performs no mutation before loading the version-matched guide.
3. Every referenced file is shipped and retrievable for the same runtime build.
4. The kernel contains the outcome, safe-failure floor, role classifier, canonical loop, and completion accounting.
5. Command snippets match current `--help`, including PowerShell-safe quoted CSV filters and typed worker lifecycle flags.
6. Legacy labels and exact-command rules remain present in the compatibility owner.
7. Remote references contain `live`, `unverifiable`, and `exited`, plus the no-local-fallback rule.
Behavioral evals should run on at least one strong concise model and one more literal model:
1. **Simple supervised pair:** creates one Run and two Tasks, starts both before waiting, processes all mail, and releases both workers.
2. **Full handoff neighbor:** routes to `orca-cli`, creates no Task/Dispatch, and stops monitoring.
3. **Worker completion:** uses injected IDs/capability, reports failure through `--outcome failed`, sends exactly once, then idles.
4. **Ask timeout:** resumes the original message ID rather than creating a second question.
5. **Fifteen-minute task:** treats rolling wait timeout as a checkpoint; does not retry, stop, or release.
6. **Remote relay loss:** reports `unverifiable`, performs no local substitute action, and does not launch a duplicate editor.
7. **Mixed-version remote:** omits unsupported optional fields and sends no unnegotiated operation.
8. **Folder workspace:** starts and supervises without invoking Git-worktree-only assumptions.
9. **Low-level custom argv:** loads the topology reference, labels the worker unsupervised, and does not claim cleanup ownership.
10. **Legacy read-only:** performs inspection only; no reply, ack, signal, focus, stop, or injection.
11. **Settled reuse:** transfers the exact agent terminal to a fresh Dispatch before acknowledging the prior Delivery.
12. **Release uncertainty:** follows `release_pending`/`release_unknown` recovery output and never calls `terminal close`.
## Rollout sequence
1. Refactor tests around owned invariants without changing the served guide.
2. Add versioned package/reference delivery or explicitly choose the single-file fallback.
3. Land the compact kernel and worker contract first; compare behavior against the current guide on the twelve evals.
4. Move placement, messaging, recovery, low-level topology, and legacy detail one owner at a time.
5. Keep the current flat guide as a compatibility fixture until old/new client and remote-server tests pass.
6. Remove duplicated prose only after its replacement owner and eval are green.
## Definition of done for the rewrite
The rewrite is complete when a new coordinator can execute the common path from the kernel alone, a dispatched worker cannot miss or misroute its completion obligation, conditional references load only at their acting boundary, and every current authority, cleanup, legacy, folder-workspace, SSH/WSL, federation, and mixed-version invariant has an owning test. A shorter file is not the success metric; fewer always-loaded decisions and one authoritative location per mechanism are.
-694
View File
@@ -1,694 +0,0 @@
# Technical review: Orca orchestration vNext implementation audit
## Scope and verdict
This review compares `orchestration-vnext-implementation-audit.md` with local
`main` (`aab6464a6a`). The orchestration files inspected are unchanged between
that revision and this worktree's HEAD. Evidence below therefore names the
worktree paths and current line numbers, but the findings describe `main`.
The proposed direction—durable facts, explicit uncertainty, execution-host
authority, and additive mixed-version rollout—is sound. The plan is not yet
implementation-ready, however. Several premises are stale, two existing
correctness bugs are hidden by the proposed abstractions, the DAG permits schema
work before its identities are defined, and the promised consolidation PR is too
large to be a safe migration or rollback unit.
The most important corrections are:
1. Treat the existing immutable Dispatch ID as the Attempt ID and the existing
terminal-resource ID as the Resource ID. Add only missing creator, role,
retry, and endpoint-incarnation facts.
2. Fix archive/liveness truth and commit-versus-notify recovery before building
lifecycle or fleet projections. Both can currently produce misleading or
duplicated results.
3. Split mailbox persistence, PTY pointer delivery, prompt submission, and
lifecycle settlement into separate receipt domains. They are not one
monotonic state machine.
4. Make migration/skew tests and negotiated remote capabilities prerequisites,
not a final P7 exercise.
5. Ship independently merged additive slices and a small promotion PR. Do not
recombine all implementation branches into one feature PR.
Focused existing coverage was run while reviewing the delivery claims:
```text
pnpm test \
src/main/runtime/orchestration-message-delivery-identity.test.ts \
src/main/runtime/orchestration-mailbox-notification-consistency.test.ts \
src/main/runtime/rpc/methods/orchestration-ask.test.ts \
src/main/runtime/rpc/methods/orchestration-recipient-routing.test.ts \
src/main/runtime/rpc/terminal-agent-prompt-send.test.ts \
src/main/runtime/agent-prompt-submission-verification.test.ts
```
Result: 69 passed, 1 skipped. This validates current behavior only; it does not
close the crash seams or missing cases identified below.
## Findings
### F1 — Critical: archived output can falsely report `exited`
**Plan claim.** P0/P4 promise that no relay loss, timeout, or uncertain close is
reported as `exited`, and P5 treats archive reads as a reliable released-worker
source.
**Current code.** `orchestration.workerRead` routes resources in all three
states—`releasing`, `unknown`, and `released`—straight to the archive without a
fresh terminal observation
(`src/main/runtime/rpc/methods/orchestration-worker-control.ts:210-220`). Both
archive result shapes hard-code terminal state to `exited`
(`src/main/runtime/rpc/methods/orchestration-worker-archive-read.ts:98-106` and
`:201-217`). Release commits `releasing` before awaiting the close and stores
`unknown` when `ptyKilled` is false
(`src/main/runtime/rpc/methods/orchestration-worker-release-completion.ts:188-227`).
The existing unknown-release regression reads the archive but does not assert
liveness (`src/main/runtime/rpc/methods/orchestration-worker-release-recovery.test.ts:166-193`).
**Correction.** Make archive provenance and process liveness orthogonal. Only a
host-confirmed close may project `exited`; `releasing` requires a current
execution-host observation and `unknown` projects `unverifiable`. Put this
characterization/fix before P4 and before any fleet projection consumes the
result.
### F2 — Critical: P2 combines three different, non-monotonic contracts
**Plan claim.** The sequence `recorded → routed → endpoint_delivered →
terminal_queued → provider_submitted → turn_started → ...` is presented as one
delivery receipt chain.
**Current code.** Mail insertion is durable
(`src/main/runtime/orchestration/db/messages/message-insert.ts:26-54`). The PTY
pointer path marks `messages.delivered_at` after pointer _text_ is accepted
(`src/main/runtime/orchestration/mailbox-pointer-delivery.ts:181-243`) and only
writes Enter 500 ms later (`mailbox-pointer-delivery.ts:253-270` and
`mailbox-pointer-submit.ts:57-84`). A crash or failed Enter can therefore leave a
legacy `delivered_at` that proves neither endpoint submission nor provider
acceptance.
**Correction.** Define separate facts:
- mail: `recorded`, `mailbox_routed`, `delivery_issued`, `consumer_acked`;
- advisory nudge: `pointer_text_written`, `submit_written|ambiguous|failed`;
- agent prompt: `terminal_bytes_written`, `submission_observed`, `turn_started`;
- worker lifecycle: report/settlement receipts owned by P3.
Never backfill legacy `delivered_at` as a higher receipt stage. These domains may
be correlated by IDs but must not be reduced as one total order.
### F3 — Critical: the real duplicate-send hole is commit versus notification
**Plan claim.** “Ordinary replies are not idempotent,” so P2 should add an
idempotency key.
**Current code.** The CLI already assigns every orchestration mutation a UUID
request ID (`src/cli/runtime/client.ts:69-95`), `orchestration.reply` is a durable
mutation (`src/shared/orchestration-rpc-contract.ts:18-39`), and the mutation
executor replays matching completed receipts
(`src/main/runtime/rpc/orchestration-mutation-executor.ts:23-115`). Question
answers also have transactional idempotency/conflict detection
(`src/main/runtime/orchestration/db/questions/question-threads.ts:73-143`).
The actual gap is later: single send commits the message and then performs a
fallible in-memory notify without first completing the mutation receipt
(`src/main/runtime/rpc/methods/orchestration.ts:684-803`). Generic reply marks
the original read, inserts the reply, and then notifies
(`orchestration.ts:1447-1459`). If notify throws, the generic executor deletes
the pending receipt, so an exact retry can insert a duplicate. Group send
already shows the intended ordering—complete the receipt before notify
(`orchestration.ts:863-895`)—and has a regression test
(`orchestration-recipient-routing.test.ts:517-549`).
**Correction.** Reuse the mutation ledger. Atomically persist the message,
settlement result where applicable, a nudge-outbox record, and a completed
mutation receipt. Notification is best-effort redrive after durable authority;
its failure must replay the recorded result rather than erase the receipt. A new
request ID remains an intentional new message.
### F4 — High: the asserted `check --wait` race is not established on the
current synchronous path
Run `check` performs its final synchronous delivery read and immediately calls
`waitForMessage` with no intervening `await`
(`src/main/runtime/rpc/methods/orchestration.ts:1007-1055`). Waiter registration
happens synchronously inside the Promise constructor
(`src/main/runtime/orca-runtime.ts:35605-35647`). With better-sqlite3 and RPC
handlers on the same JS event loop, an ordinary same-runtime insert cannot
interleave at the seam described by the plan. Dispatch and direct-mail paths
must also be characterized rather than assumed.
**Correction.** Add deterministic hooks at actual asynchronous boundaries—paged
routing, federation/connection replacement, and notify failure—and prove a
failing case before adding a durable mailbox epoch. Register-then-recheck with
cancellation is sufficient if a real local seam exists. Prioritize the proven
commit/notify crash gap from F3.
### F5 — High: restart wakeup wording would violate restored-status safety
The runtime already scans undelivered mail and schedules pointers on
initialization (`src/main/runtime/orca-runtime.ts:4551-4566` and
`:35545-35573`), then redrives on a fresh live-idle observation
(`orca-runtime.ts:7306-7321`). It intentionally refuses to type into a merely
restored idle snapshot; the two-launch regression explains the stale-draft and
wrong-process hazard (`tests/e2e/orchestration-idle-mail-restore.spec.ts:1-20`
and `:148-179`).
**Correction.** The acceptance contract is “mail survives restart and is nudged
after a post-restart live observation or explicit check,” not “wakes without a
heartbeat.” A durable row cannot wake an offline process, and a stale restored
snapshot must never authorize Enter.
### F6 — High: terminal prompt idempotency is a new subsystem boundary
`terminal.send` is not an orchestration durable mutation. The CLI sends
`agentPrompt: true` (`src/cli/handlers/terminal.ts:105-117`), while the runtime
pastes bytes, presses Enter, and waits only for a new working sequence
(`src/main/runtime/orca-runtime.ts:19900-19978`). The verifier may throw
`agent_prompt_stalled` after the irreversible writes
(`src/main/runtime/agent-prompt-submission-verification.ts:17-40`). The RPC
schema carries neither a prompt request ID nor a wait contract
(`src/main/runtime/rpc/methods/terminal/unary-schemas.ts:82-111`), and the shared
result exposes only acceptance/bytes written
(`src/shared/runtime-terminal-contracts.ts:196`).
Changing `agent_prompt_stalled` into `queued_pending_turn` also changes existing
CLI success/exit semantics, so it is not an additive field-only change.
**Correction.** Add a narrowly scoped prompt-delivery ID bound to terminal,
process incarnation, and payload, guarded by a runtime capability. Separate
`terminal_bytes_written` from a later submission observation; `--wait-submit`
must observe the same request, never resend it. Keep raw keystrokes and
question-reply terminal writes outside this ledger. Old-host downgrade must be
explicit, not optimistic.
### F7 — High: draft protection is not independently implementable in P2
Both mailbox pointers and agent prompts write into the current composer and
later write Enter (`mailbox-pointer-delivery.ts:181-270` and
`orca-runtime.ts:19900-19978`). Provider-neutral idle does not prove that the
composer is empty. No current check proves an existing human draft is absent.
**Correction.** Either make the minimal provider/draft capability slice of P5 a
prerequisite or disable auto-Enter whenever composer emptiness is unproven. The
plan cannot promise draft preservation while declaring P2 parallel with all
provider work.
### F8 — High: P1 duplicates identities that already exist
`dispatch_contexts.id` is already the immutable identity created on each worker
start (`src/main/runtime/orchestration/db/schema/create-graph-tables-sql.ts:120-143`)
and keys `worker_dispatches.dispatch_id`
(`src/main/runtime/orchestration/db/schema/create-core-tables-sql.ts:112-130`).
Retries validate `retryOf` and create a new dispatch ID, although the retry edge
is not persisted (`src/main/runtime/orchestration/db/worker-dispatch/worker-dispatch-start.ts:65-89`).
`worker_terminal_resources.id` is already an immutable resource identity with
owner history, process incarnation, host scope, and ownership/release state
(`create-core-tables-sql.ts:132-167`).
Tasks already carry `parent_id`, and dispatches/remote attachments already carry
depth (`create-graph-tables-sql.ts:95-118`, `:120-143`, and
`src/main/runtime/orchestration/db/dispatch-depth.ts:53`).
**Correction.** Formally alias Dispatch ID as Attempt ID and keep the existing
Resource ID. Persist `retry_of_dispatch_id`, creator-dispatch identity, and role;
add only the endpoint/incarnation facts not already represented. Preserve the
canonical parent/depth fields. Do not introduce parallel Attempt/Resource IDs or
backfill unknown legacy provenance with guessed ownership.
### F9 — High: P3 statuses are not additive schema changes
Task status, dispatch status, worker state, resource ownership, and resource
release state are closed SQLite `CHECK` constraints
(`create-graph-tables-sql.ts:95-143` and `create-core-tables-sql.ts:112-155`).
Adding `outcome_unknown`, `finished_unverified`, `cancelled`, or `superseded` to
those columns requires rebuilding tables; the existing migration machinery
demonstrates such a rebuild (`src/main/runtime/orchestration/db/schema/migrate-v2-v12.ts:5-140`).
**Correction.** Initially store additive observation/receipt facts in new tables
or nullable/defaulted columns, then project the richer vocabulary without
widening the legacy enums. If an enum widening is still needed, make it a
separate compatibility migration with old-runtime write tests and an explicit
rollback window.
### F10 — High: “append-only transitions or equivalent” is too vague for the
number of lifecycle writers
Worker report settlement, manual task update, stop, missing-terminal recovery,
and federated paths all write lifecycle state. `recordWorkerStage` directly
updates worker stage/state without compare-and-swap
(`src/main/runtime/orchestration/db/worker-dispatch/worker-dispatch-stage.ts:19`).
Adding a transition table beside those writers would produce a partial audit log
rather than an authority. `recordWorkerStage` is a concrete example: it reads the
current row and then performs an unfenced update
(`src/main/runtime/orchestration/db/worker-dispatch/worker-dispatch-stage.ts:19-45`).
**Correction.** First route every lifecycle mutation through transaction-neutral
transition primitives. The primitive must validate the old state, update legacy
projection fields, and append its receipt in the same transaction supplied by
the caller. Add a ratchet test forbidding direct lifecycle `UPDATE`s outside
those modules and rollback-injection tests for each writer. Only then add the P3
reducer/projection.
### F11 — High: heartbeat age lacks a clock-domain contract
Local heartbeat time is the home database's message insertion time
(`src/main/runtime/orchestration/lifecycle-reconciliation.ts:168` and
`db/messages/message-insert.ts:31`). Federated relay import also carries a
worker-supplied `lifecycle.at`
(`src/main/runtime/orchestration/db/federation/federation-relay-import.ts:91`).
Computing age from mixed source clocks can turn clock skew into false liveness or
false death.
**Correction.** Persist source observation time separately from execution-host
receive time and home-host receive time. Freshness is computed in the authority
host's clock domain and combined with the canonical connection verdict. Never
rewrite `live`/`unverifiable`/`exited`, which already exists in
`src/shared/pty-liveness-verdict.ts:1`, into a second vocabulary.
### F12 — High: the lease invariant is impossible and crosses host authority
The consolidation gate “No active Attempt lacks an owner/resource lease” is
false during legitimate startup and for supported topologies. A starting worker
exists before authority attach
(`worker-dispatch-start.ts:103-117` and
`src/main/runtime/orchestration/db/worker-dispatch/worker-dispatch-authority.ts:104`).
Context-only/unsupervised dispatches may never own a terminal resource. A
federated home DB intentionally has no local terminal resource because the
worker host owns it (`src/main/runtime/orchestration/db/worker-terminal/worker-terminal-resource-store.ts:12`
and `src/main/runtime/rpc/methods/orchestration-worker-release.ts:34`).
**Correction.** Use this invariant instead: after local authority attachment,
each supervised local dispatch has exactly one execution-host terminal resource
or an explicitly external resource. Pre-attach, context-only, and federated-home
dispatches have typed absence/remote attachment, never a fabricated local lease.
### F13 — High: P4's baseline description collapses three different operations
Only release archives before close
(`src/main/runtime/rpc/methods/orchestration-worker-release-completion.ts:169-216`).
Local stop closes directly without archiving
(`src/main/runtime/rpc/methods/orchestration-worker-stop.ts:140-170`). Abandon is
a DB-only state change that neither closes nor archives and can retain a live
resource (`src/main/runtime/orchestration/db/worker-dispatch/worker-dispatch-abandon.ts:54-76`).
Federated stop likewise closes remotely without archive
(`src/main/runtime/rpc/methods/orchestration-federation-control.ts:127-180`), and
federated release is currently unsupported
(`src/main/runtime/rpc/methods/orchestration-worker-release.ts:33-44`).
**Correction.** Specify stop, abandon, and release separately. Decide whether
stop must archive or return an explicit output-loss receipt. Abandon must retain
and expose readable live output. Remote archive-before-close is a later
capability-negotiated operation owned jointly by P4/P5/P7, not existing behavior.
### F14 — High: P4 risks reimplementing the current resource ownership system
The existing resource table already represents owner history, process/host
fencing, transferred/user/external ownership, retention reasons, and release
states (`create-core-tables-sql.ts:132-167`). Restart reconciliation already
operates those rows
(`src/main/runtime/orchestration/worker-terminal-release-reconciliation.ts:19`).
**Correction.** Extend `worker_terminal_resources` with optional retention
expiry/policy, recovery attempt counters, and pagination indexes. Do not create a
second terminal lease table or identity. Bulk cleanup must reuse the same
single-resource fenced transition and remain conservative for external,
user-owned, transferred, federated-home, and unverifiable resources.
### F15 — Critical: direct-SSH transcript reads violate the execution-host
boundary
The plan says remote/federated reads execute on the owning host and explicitly
puts “desktop-side SSH/WSL file reads” out of P5 scope. Worker session selection
filters status by connection ID
(`src/main/runtime/orca-runtime.ts:18821-18849`) but the selected session drops
that host identity (`src/main/runtime/orchestration/worker-provider-session.ts:12-33`).
`readWorkerTranscript` then resolves and opens the path in the client main
process (`src/main/runtime/rpc/methods/orchestration-worker-output.ts:37-50` and
`src/main/runtime/orchestration/worker-transcript-read.ts:53-69`). The resolver
explicitly works against the current main process's home/filesystem
(`src/main/native-chat/session-file-resolver.ts:21-25`). Existing support metadata
already says Model-A SSH local scanning is the wrong host
(`src/shared/native-chat-agent-support.ts:16-23`).
For a non-Windows absolute path, a same-path local lookalike may even be accepted
(`src/main/native-chat/host-readable-transcript-path.ts:115-137`). That can expose
unrelated local content as evidence for a remote worker.
**Correction.** Carry host scope/connection identity into provider-session
selection and transcript resolution. Direct SSH must use a negotiated remote
transcript-read capability or explicitly fall back with
`remote_capability_unavailable`; it must never probe a client-local path. This is
P5/P7 scope and a security boundary, not a deferred enhancement.
### F16 — High: WSL and SSH need different read contracts
Current WSL support intentionally performs guarded Windows-host UNC reads
(`src/main/native-chat/host-readable-transcript-path.ts:108-158` and `:179-201`),
and the session resolver uses that mechanism
(`src/main/native-chat/session-file-resolver.ts:98-123` and `:199-203`). Tests
characterize UNC translation
(`src/main/native-chat/session-file-resolver-wsl.test.ts:73-114`).
**Correction.** Define four separate routes: local/folder uses local main FS;
WSL uses guarded UNC on the Windows execution host; direct SSH uses a remote
capability; federation executes on the peer server. Assert that none falls
through to a different topology.
### F17 — High: “exact transcript captured at dispatch attach” is false
Provider session is selected on each read from the current in-memory status
snapshot (`orca-runtime.ts:18821-18849`); no session binding is persisted at
attach. Without a launch token, the process incarnation is copied into the
result rather than compared to a status-side incarnation
(`worker-provider-session.ts:27-33`), a behavior its tests accept
(`worker-provider-session.test.ts:24-45`). The returned content is also bounded:
50 messages, block count limits, response clipping, and text/tool clipping
(`src/main/runtime/orchestration/worker-transcript-payload.ts:4-10`, `:38-109`,
and `:143-203`).
**Correction.** Say “a bounded decoded projection from the currently selected
provider session/file.” Preserve the already exposed source, source identity,
provider, fallback, warnings, liveness, and archive fields
(`src/shared/orchestration-worker-output.ts:28-68`). Add orthogonal
`sourceExact`, `contentComplete`, and clipping reasons if needed. If durable
attach-time binding is desired, make it explicit P1/P5 schema fenced by
Attempt+Endpoint+process.
### F18 — Medium-high: archived terminal fallback discards provenance
Live fallback preserves `session_not_reported`, provider, missing-path, and parse
reasons (`orchestration-worker-output.ts:37-59` and `:163-177`). Archive capture
drops failed transcript reasons
(`src/main/runtime/orchestration/worker-output-archive.ts:56-79`), and archived
terminal results always report `fallbackReason: null`
(`orchestration-worker-archive-read.ts:201-217`).
**Correction.** Persist attempted transcript source, fallback reason, and
warnings in the existing terminal-tail archive. Do not claim an explicit reason
survives release until this is implemented and migration behavior is defined.
### F19 — High: an in-flight archive mirror cannot run in parallel with P1/P4
Existing archive identity is dispatch/resource/process based
(`orchestration-worker-archive-read.ts:74-85` and `:182-191`). A new mirror must
bind the final Attempt/Endpoint/Resource model and define quotas, retention,
cleanup, redaction, crash semantics, and process-replacement fencing.
**Correction.** Decoder/metadata characterization may begin after P0. Gate mirror
schema on the P1 identity decision and archive lifecycle on P4. Extend the
existing resource/archive identity rather than introducing a mirror-specific
Resource ID. If no concrete failure requires an in-flight mirror, omit it; the
current release snapshot and existing provider/terminal watchers should be
reused first (`src/main/native-chat/transcript-watch.ts:23-42` and terminal
subscription in `src/main/runtime/rpc/methods/terminal/terminal-subscribe-method.ts:11`).
### F20 — High: P6 cannot promise one host-routed query without P7
`workerList` already returns worker/dispatch/resource projections, including
unsupervised rows
(`src/main/runtime/orchestration/db/worker-terminal/worker-terminal-listing.ts:85-143`
and `src/main/runtime/rpc/methods/orchestration-worker-release.ts:143-166`). It is
unbounded and local. Federated show/read are per-dispatch calls, and read uses a
15-second timeout (`orchestration-worker-control.ts:159-192`). Naively observing
each remote row creates N+1 calls and serial timeout risk.
The renderer is not simply “polling lanes”: orchestration status is already
push-fed into indexed worktree maps
(`src/renderer/src/components/sidebar/worktree-agent-orchestration-index.ts:168-229`),
while the coordinator has a single 2-second tick over DB work
(`src/main/runtime/orchestration/coordinator.ts:35-55` and `:107-165`).
**Correction.** Introduce an early read-only durable fleet projection by
extending/composing `workerList`, then enrich it as facts land. Remote fleet
observation requires a P7 negotiated _batched per-host snapshot_ with concurrency
and total-time budgets, partial-host errors, stable pagination, and no transcript
bodies. Keep push status for live UI; name the exact old callers being retired.
### F21 — High: P7 is infrastructure for earlier slices, not a final gate
Federation already negotiates orchestration protocol/capabilities
(`src/shared/protocol-version.ts:44-55` and `:156-165`) and has a distinct durable
contiguous relay/ack protocol
(`src/main/runtime/orchestration/db/federation/federation-relay-enqueue.ts:10-175`
and `federation-relay-import.ts:44-129`). Avoiding a new stream opcode is not
enough: receipt and relay-content semantics are also wire contracts.
Old-peer structured output is probed on every read: method-not-found falls back
to legacy read (`orchestration-worker-control.ts:159-192` and
`orchestration-worker-legacy-federated-read.ts:31-45`). Total relay loss is
currently rethrown, not converted into a worker-read result
(`orchestration-worker-control.ts:179-182`).
**Correction.** Build old/new RPC and relay skew fixtures alongside the first
behavioral slice. Add a capability cache keyed by peer fingerprint/runtime epoch
with a narrow unsupported predicate. P6 may project a typed host-unavailable row
as `unverifiable`, but an individual worker-read should return a typed unavailable
error rather than inventing an observation.
### F22 — High: migration and rollback work is materially incomplete
`createTables()` runs before transactional version migration
(`src/main/runtime/orchestration/db/orchestration-db.ts:24-31`). Schema-skew repair
probes physical shape and can rewind version state
(`src/main/runtime/orchestration/orchestration-schema-version-skew.ts:91-116`).
Old binaries intentionally open future-schema DBs without running migrations
(`src/main/runtime/orchestration/db/schema/migrate.ts:7-24`). Consequently a new
non-null/no-default column, widened CHECK, mandatory trigger, or authoritative
epoch can break rollback or go stale while an old runtime writes.
**Correction.** Every schema slice must update fresh DDL, transactional
migration, version/skew probes, schema-version comments, reset/retention logic,
and old/new fixtures together. Use additive tables or nullable/defaulted columns.
On re-upgrade, reconcile new receipts from legacy authority rather than assuming
old code maintained them.
### F23 — High: atomic Task creation plus worker start is omitted
The companion `orchestration-issues.md` requests atomic Task creation + worker
start, but the matrix has no bounded slice for it. Current `workerStart` requires
an existing task (`src/main/runtime/rpc/methods/orchestration-workers.ts:48`), and
only starting-dispatch creation is atomic after that task lookup
(`worker-dispatch-start.ts:61-117`).
**Correction.** Add a distinct vertical slice after identity and mutation-ledger
hardening. It should create/reconcile Task + Dispatch + mutation receipt in one
transaction, then perform external side effects via the same recovery model as
worker start. Do not bury this API/product change in P3.
### F24 — Medium-high: several proposed items reimplement or over-scope existing
systems
- Cross-Run send routing already returns typed `recipient_run_mismatch` and
`recipient_ambiguous` errors
(`src/main/runtime/rpc/methods/orchestration-recipient-routing.ts:29-89` and
`:120-145`). Pin these behaviors instead of adding a second auth layer.
- Current Run delivery already has one immutable outstanding batch and
idempotent whole-batch ack
(`src/main/runtime/orchestration/db/runs/run-delivery.ts:23-70` and `:125-175`).
Per-message ack is a new negotiated contract with difficult old-client replay
semantics, not a prerequisite without a demonstrated loss case.
- Provider aliases/support, resolver, decoder, lifecycle parsing, and passive
watches already exist in `src/shared/native-chat-agent-support.ts:1-46`,
`src/main/native-chat/session-file-resolver.ts:126-165`,
`transcript-tail-reader.ts:36-50`, `transcript-turn-lifecycle.ts:23-34`, and
`transcript-watch.ts:23-42`. Consolidate these maps into one exhaustive profile
rather than building a parallel orchestration-only registry.
### F25 — High: “one consolidation PR” is an unsafe delivery unit
The plan proposes separately reviewable PRs, mixed-version support, schema
migrations, new delivery semantics, remote capabilities, and a fleet UI, then
recombines their branches into one consolidation PR. That maximizes rebase and
migration coupling and defeats incremental rollback.
**Correction.** Merge additive slices independently behind capabilities and
legacy projections. Run conformance continuously. Finish with a small promotion
PR that changes consumers/defaults. Retain compatibility shims for at least one
mixed-version release window and remove them only after telemetry proves disuse.
## Revised implementation DAG
```text
R0 Characterize current contracts and reuse canonical vocabulary
├─ Fix false-exited archive projection (F1)
├─ Reproduce/fix commit→notify recovery (F3)
└─ Establish old/new DB, RPC, relay, and provider fixtures
R1 Identity decision and migration contract
├─ Dispatch == Attempt; retain existing Resource ID
├─ Add retry/creator/role and endpoint-incarnation facts
└─ Fresh/upgrade/downgrade/skew tests
R2 Central lifecycle transition primitives (after R1)
└─ All current writers dual-write receipts transactionally
In parallel after R0/R1:
R3a Mail/outbox atomicity using existing mutation ledger
R3b Provider capability consolidation and topology-safe reads
R3c Early local-only read-only fleet projection for rollout parity
R3d Atomic Task-create + worker-start vertical slice
After provider draft capability and endpoint identity:
R4 Prompt-specific idempotency and submission observation
After R2:
R5 Lifecycle reducer/projection and clock-domain-safe freshness
R6 Existing-resource lease metadata, cleanup, and archive semantics
After remote capability fixtures exist:
R7 Negotiated remote transcript/release/batched fleet slices
Last:
R8 Fleet/attention UI promotion and bounded notification policy
```
P7-style conformance is continuous infrastructure beginning in R0, not a late
branch. An in-flight transcript mirror, partial ack, or richer enum migration is
admitted only after its specific failure case and mixed-version contract are
approved.
## Revised acceptance tests
### Identity and migration
1. A retry gets a new Dispatch/Attempt ID, persists `retry_of`, and a late report
from the prior attempt cannot mutate the new attempt or Task settlement.
2. Pane remint, terminal reuse, process replacement, host mismatch, and spoofed
payload identity all fail closed; the legitimate current incarnation passes.
3. Local, folder-workspace, context-only, unsupervised, SSH, WSL, and
federated-home rows express resource presence or typed absence without
fabricating a local lease.
4. Fresh vNext and upgraded vNext schemas are structurally equivalent. Exercise
vCurrent → vNext → vCurrent → vNext, performing send/check/reply/start at each
step, plus interrupted migration, partial/future schema, reopen, reset, and
retention bounds.
5. Old runtime writes while new optional receipt tables are untouched; re-upgrade
reconciles without treating absence as acknowledged/submitted.
### Mail, notification, and prompt delivery
1. Inject throw/crash after commit for Run send, Dispatch send, generic reply,
worker-done settlement, and federated enqueue. Retrying the same request ID
after DB reopen returns the same message/relay/result and produces exactly one
durable side effect; changed payload conflicts; a new ID is a new message.
2. Crash after pointer text and before Enter. Legacy `delivered_at` must project
only the legacy fact, never submitted/acknowledged. After reopen, redrive
requires a fresh live-idle observation and never emits duplicate Enter.
3. Characterize insertion before final read, after waiter registration, during a
paged/federated yield, on notification failure, and over socket/runtime
replacement for Run, Dispatch, and direct mailbox paths. Add an epoch only if
a real missed-event seam remains.
4. A pre-existing human draft in Claude, Codex, and unsupported-provider fixtures
is preserved byte-for-byte; unproven emptiness yields
`draft_unknown`/explicit-check rather than auto-Enter.
5. Prompt response loss after bytes/Enter followed by replay of the same prompt
ID never writes again. Payload, terminal, or process-incarnation mismatch is
a conflict. Test stale handle/remint, new CLI→old host, old CLI→new host,
pretty/JSON output, and exit status.
6. `--wait-submit` observes the original prompt request. Timeout yields ambiguous
observation, not permission to resend.
7. Preserve current whole-batch replay/ack unless a new partial-ack contract is
explicitly negotiated; if added, test new-partial→old-replay/ack and
old-batch→new-view skew.
### Lifecycle, leases, and liveness
1. Every lifecycle entry point uses the transition primitive; a source ratchet
rejects direct Task/dispatch/worker lifecycle updates outside it.
2. Inject rollback after legacy projection update and after receipt append; both
remain atomic. Replay, duplicate, late, and reordered events converge.
3. Active-sibling reports, report after reconnect, missing report, manual update,
stop, abandon, release, missing-terminal recovery, and federated import cannot
settle the wrong attempt.
4. Block `closeTerminal` after archive commit: worker-read never says `exited`.
Return `ptyKilled:false`: archived content remains readable while liveness is
`unverifiable`. Only confirmed close projects `exited`.
5. Stop either preserves an archive or returns an explicit output-loss receipt.
Abandon performs no close/archive and retains readable output. Federated
release archives on the owning server, is idempotent over relay loss, and
safely refuses on an old peer.
6. Source/execution/home timestamps survive ±24-hour skew, replay, duplicates,
reconnect, and zombie heartbeat. Freshness uses authority-host receive time;
contact loss never becomes `exited`.
7. Bulk cleanup paginates and reuses single-resource fencing. It skips
user-owned, external, transferred, conflicting, federated-home, and
unverifiable resources and reports a per-resource safe next action.
### Transcript and host routing
1. Direct SSH status reports `/home/ada/session.jsonl` while an identically named
client-local file contains a sentinel. Worker-read must never return the
sentinel. A negotiated remote read returns remote content; an old/unsupported
host returns `remote_capability_unavailable` and safe terminal fallback.
2. Independently test local/folder local FS, Windows-host WSL UNC, direct SSH
remote RPC, and federated peer execution. No route may fall through to another
topology.
3. Restart without republished provider status yields `session_not_reported`, not
an invented exact transcript. Stale same-pane data without launch/process
proof is rejected or explicitly inexact.
4. More than 50 messages, more than six blocks, oversized text/tool input, and a
bounded archive set incomplete/limited metadata while source identity stays
stable.
5. Release after no session, unsupported provider, missing/unreadable path, and
parse failure; restart and archive-read preserve the original fallback reason
and warnings.
6. If an in-flight mirror is retained, process replacement and attempt
supersession cannot append to the old mirror. Test per-attempt/global quotas,
backpressure, reset/release/retention cleanup, redaction, and absence of
dispatch capabilities or local path secrets.
7. One exhaustive provider profile drives resolver, decoder, lifecycle, and watch
capabilities. Alias behavior is explicit, unsupported lifecycle decoders are
honest, and adding an AgentType fails exhaustiveness until profiled.
### Federation, fleet, and UI
1. Run the old/new Run-home × old/new worker-host matrix in both directions with
response loss, relay reconnect, unknown optional fields, and unchanged legacy
fallback. A rejected optional capability must not change old behavior.
2. Old-peer structured-read is probed once per peer fingerprint/runtime epoch;
concurrent reads coalesce, a new epoch re-probes, and non-`method_not_found`
errors do not poison the cache.
3. A 100-worker fleet across two peer hosts makes at most one batched observation
call per host. An unreachable host returns bounded partial rows marked
`unverifiable` without delaying healthy hosts by N×timeout.
4. Pagination is stable under concurrent updates: no duplicate/omitted rows for a
snapshot cursor, bounded memory, no transcript bodies, and explicit redaction.
5. Name and spy each replaced DB/RPC caller. Existing push agent status remains
live while the durable fleet snapshot is stale; the fleet projection does not
become a second renderer state store.
6. Deterministic fixture/fake-peer tests are mandatory PR gates. Real Windows+WSL,
SSH relay, and provider-version jobs are separate certification jobs with
explicit skip/failure policy and bounded diagnostic artifacts.
7. UI comes last: five-worker completion coalesces only root-success alerts;
input, approval, failure, interruption, stale, and unverifiable remain distinct;
unread state is per agent and focus is never stolen.
## Claims that should be rewritten in the source plan
| Plan section | Required rewrite |
| ------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Critical path / DAG (`orchestration-vnext-implementation-audit.md:20-39`, `:113-123`) | P4 schema and P5 mirror cannot start before P1; P7 fixtures begin with the first wire change; early local fleet projection supports parity. |
| “Sending during an active turn” (`:43-73`) | Replace the unproven read/register assertion with the proven commit/notify gap; distinguish mail, nudge, and prompt receipts; state restored-live requirement. |
| “Reading a worker” (`:75-87`) | Remove attach-time/exact/remote-host claims; document bounded current-status selection and the direct-SSH boundary bug; distinguish WSL UNC from SSH. |
| P1 (`:98`) | Reuse Dispatch as Attempt and existing resource IDs; add retry/creator/role/endpoint facts only. |
| P2 (`:99`) | Reuse mutation receipts and whole-batch ack; scope prompt IDs separately; make draft capability a prerequisite; remove already-fixed cross-Run routing gap. |
| P3 (`:100`) | Avoid enum widening initially; centralize all writers first; define clock domains. |
| P4 (`:101`) | Describe stop, abandon, local release, and federated behavior separately; extend the existing resource table. |
| P5 (`:102`) | Reuse current provider maps/watchers; add topology-safe routing and completeness metadata; gate any mirror on P1/P4. |
| P6/P7 (`:103-104`) | Specify negotiated batch snapshot, partial-host/time-budget semantics, and old/new fixtures before claiming one query or CI coverage. |
| Consolidation gates (`:153-172`) | Replace the impossible every-active-Attempt lease invariant; require a fresh live observation after restart; add false-exited and SSH isolation gates. |
| PR decomposition (`:174-182`) | Independently merge additive slices; use only a small final promotion PR and retain shims through a measured mixed-version window. |
## Bottom line
The implementation should begin by repairing and characterizing the facts Orca
already has, not by adding parallel IDs, leases, reply keys, provider registries,
or an all-purpose receipt chain. Once identity, transactional writers, topology,
and skew behavior are explicit, the richer lifecycle and fleet projections can
be additive and honest. Without those corrections, the proposed parallelism and
single consolidation PR would make the highest-risk migration and remote-wire
changes land at exactly the point where rollback is hardest.
@@ -1,72 +0,0 @@
# Orchestration R0 contract and characterization gate
## Verdict
R0 is characterized against worktree HEAD `8d61cb8b771b7b658a5e1e2bb426c719644d09f3` without production behavior changes. Existing tests cover durable mail replay, whole-batch Run acknowledgement, waiter/pointer behavior, queued prompt stall, Dispatch/process fencing, rejected lifecycle mail, Resource transfer/release, schema skew, federation settlement, and SSH/WSL uncertainty. Two critical current failures were missing exact assertions and now have narrow green characterization fixtures: a single-recipient retry duplicates a committed message after notification throws, and an archive read reports `exited` after release liveness was only `unverifiable`.
Do not fan out behavior work until the ordering and liveness fixes below replace those current-failure assertions with the intended contract. Identity, prompt, lifecycle, fleet, or compatibility workers must consume the canonical contracts here rather than introduce aliases, duplicate stores/IDs, universal transcripts, or UI policy.
## Current contract inventory
| Domain | Current authority and order | Incumbent limitation | Exact implementation/test surfaces |
| -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Mutation receipt | `OrchestrationMutationExecutor` durably inserts `pending`, invokes the handler, then completes the receipt from the handler result. On a non-`operation_unknown` throw it deletes the pending receipt. A handler may call `recordMutationReceipt` before later fallible work. | Group send records the receipt before notify; single send, generic reply, and lifecycle settlement normally return through notify before the executor completes the receipt. A notify throw can therefore leave the effect committed and erase replay identity. | `src/main/runtime/rpc/orchestration-mutation-executor.ts`; `src/main/runtime/rpc/methods/orchestration.ts`; `src/main/runtime/rpc/orchestration-mutation-ledger.test.ts`; `src/main/runtime/rpc/methods/orchestration-recipient-routing.test.ts`; `src/main/runtime/rpc/orchestration-commit-notify-characterization.test.ts`. |
| Mailbox receipt | Message IDs and monotonic sequences are durable. Run delivery is FIFO, capped at 50, one outstanding immutable batch at a time, replayed until whole-batch ack, and fenced by consumer generation. Dispatch/direct checks consume their own mailbox rules. | Mail persistence and acknowledgement do not prove a PTY nudge, provider submit, or lifecycle settlement. The same-runtime waiter epoch race is not proven on the synchronous path; paged/federated yields remain the characterization seam. | `src/main/runtime/orchestration/db/messages/*`; `src/main/runtime/orchestration/db/runs/run-delivery.ts`; `src/main/runtime/rpc/methods/orchestration.ts`; `src/main/runtime/orchestration-message-delivery-identity.test.ts`; `src/main/runtime/orchestration-mailbox-notification-consistency.test.ts`; `src/main/runtime/rpc/methods/orchestration-check.test.ts`. |
| Pointer/nudge receipt | Notification resolves an in-memory waiter or schedules advisory pointer delivery. Pointer transport settlement precedes `messages.delivered_at`; Enter is written only after a later idle/ownership recheck. Restart redrive requires a fresh live-idle observation or explicit check. | `delivered_at` proves pointer text, not Enter, command execution, mail acknowledgement, or provider submission. Waiters and pointer flights are process-memory state. | `src/main/runtime/orchestration/mailbox-notification-coordinator.ts`; `src/main/runtime/orchestration/mailbox-pointer-delivery.ts`; `src/main/runtime/orchestration/mailbox-pointer-submit.ts`; `tests/e2e/orchestration-idle-mail-delivery.spec.ts`; `tests/e2e/orchestration-idle-mail-restore.spec.ts`. |
| Prompt receipt | Agent prompt bytes are serialized to the exact PTY generation, Enter is written once after provider/platform readiness gates, and verification waits for a new working sequence. Worker start persists the resulting `ready/input_accepted` or failed stage in its mutation receipt. | No prompt-specific durable request ID exists. A swallowed Enter or prompt queued during a busy TUI is reported as `agent_prompt_stalled`, fails Task/Dispatch, and revokes capability even though the TUI may later consume the text. PTY acceptance, provider submission, and turn start remain distinct facts. | `src/main/runtime/rpc/terminal-agent-prompt-send.ts`; `src/main/runtime/agent-prompt-submission-verification.ts`; `src/main/runtime/rpc/methods/orchestration-worker-start-prompt-contract.test.ts`; `tests/e2e/terminal-send-agent-prompt-submit.spec.ts`. |
| Lifecycle receipt | `worker_done` requires explicit Task, Dispatch, outcome, exact capability/pane/process authority, then transactionally settles Task/Dispatch/worker. Rejected lifecycle mail is converted to a durable high-priority status marker and does not settle the Task. Identical settled reports and late heartbeats are suppressed/idempotent. | The message/settlement result is authoritative but shares the generic mutation receipt ordering weakness if a later notification throws. Generic status and quiet PTY output are never completion proof. | `src/main/runtime/orchestration/lifecycle-reconciliation.ts`; `src/main/runtime/orchestration/db/dispatch-context/worker-report-settlement.ts`; `src/main/runtime/rpc/methods/orchestration-send-dispatch-authority.test.ts`; `src/main/runtime/rpc/methods/orchestration-federation-lifecycle-settlement.test.ts`; `src/main/runtime/rpc/methods/orchestration-migration-behavior.test.ts`. |
| Archive/liveness | Release captures output and commits Resource state `releasing` before exact terminal close. Confirmed close settles `released`; missing/unconfirmed close becomes `release_pending` or `release_unknown`. | `workerRead` routes `releasing`, `unknown`, and `released` to archive reads, whose result currently hard-codes terminal `exited`. Archive availability and process liveness are incorrectly coupled. | `src/main/runtime/rpc/methods/orchestration-worker-release-completion.ts`; `src/main/runtime/rpc/methods/orchestration-worker-control.ts`; `src/main/runtime/rpc/methods/orchestration-worker-archive-read.ts`; `src/main/runtime/rpc/methods/orchestration-worker-release-recovery.test.ts`; `src/main/runtime/rpc/methods/orchestration-worker-release-liveness-verdict.test.ts`. |
| Attempt/Resource identity | Existing immutable `dispatch_contexts.id` is the Attempt identity. Each retry creates a new Dispatch. Existing `worker_terminal_resources.id` is the Resource identity and records current/prior owner, pane, process incarnation, host scope, ownership, retention, and release state. | Persist the missing retry edge/creator/role/endpoint facts additively; do not add parallel Attempt/Resource tables. Starting, unsupervised, external, and federated-home rows legitimately have typed absence rather than a fabricated local Resource. | `src/main/runtime/orchestration/db/schema/create-graph-tables-sql.ts`; `src/main/runtime/orchestration/db/schema/create-core-tables-sql.ts`; `src/main/runtime/orchestration/db/worker-dispatch/worker-dispatch-start.ts`; `src/main/runtime/orchestration/db/worker-terminal/worker-terminal-resource-store.ts`; `src/main/runtime/orchestration/orchestration-worker-dispatch-db.test.ts`; `src/main/runtime/orchestration/db/dispatch-row-writer-boundary.test.ts`. |
| SSH/WSL execution boundary | The execution host owns process/filesystem/transcript truth. Contact loss is `unverifiable`, positive host absence is `exited`, and observed presence is `live`. WSL transcript reads use guarded Windows-host UNC translation; federated reads execute on the peer runtime. | Direct SSH provider session selection can lose connection identity before transcript resolution and must not fall through to client-local filesystem reads. A new remote read/release feature requires negotiated capability and typed unavailable fallback. | `docs/reference/ssh-execution-boundary.md`; `docs/reference/wsl-command-execution.md`; `src/shared/pty-liveness-verdict.ts`; `src/main/runtime/rpc/methods/orchestration-worker-stop-liveness-verdict.test.ts`; `src/main/runtime/rpc/methods/orchestration-federation-liveness-verdict.test.ts`; `src/main/native-chat/session-file-resolver-wsl.test.ts`; `src/shared/wsl-exec-mode-separator.test.ts`. |
| Old/new DB and RPC | Current schema setup repairs known physical/version skew, skips migrations when an old binary opens a future schema, and retains legacy read-only/fenced RPC behavior. Federation relay uses durable contiguous sequence/ack state plus protocol-version gates; JSON optional fields are omission-safe. | Every schema slice must update fresh DDL, migration, skew probes, reset/retention, and downgrade/re-upgrade fixtures together. New stream opcodes or required fields are unsafe without negotiation. Unsupported remote RPC fallback needs a narrow peer/runtime-epoch capability cache before fleet fan-out. | `src/main/runtime/orchestration/orchestration-schema-version-skew.ts`; `src/main/runtime/orchestration/db/schema/migrate.ts`; `src/main/runtime/orchestration/orchestration-version-skew-migration.test.ts`; `src/main/runtime/rpc/methods/orchestration-migration-behavior.test.ts`; `src/main/runtime/orchestration/federation-acknowledgment-migration.test.ts`; `src/main/runtime/orchestration/federation-acknowledgment-integrity.test.ts`; `src/shared/orchestration-rpc-contract.test.ts`; `docs/reference/remote-wire-compatibility.md`. |
## Named failure fixtures
1. **Commit versus notify — newly characterized.** `src/main/runtime/rpc/orchestration-commit-notify-characterization.test.ts` injects a throw from `notifyMessageArrived` after a single Run message is committed. The exact request ID retry returns `replayed:false` and the inbox contains two rows. The correction must flip this to `replayed:true`, preserve one row/message ID, and extend the same oracle to Run send, Dispatch send, generic reply, worker-report settlement, and federated enqueue.
2. **False-exited archive — existing fixture tightened.** `src/main/runtime/rpc/methods/orchestration-worker-release-recovery.test.ts` now records that an unconfirmed SSH-like close produces `release_unknown` but the subsequent archived read says `status.terminal=exited`. The correction must project `unverifiable` for `unknown`, observe `releasing` on the owning host, and reserve `exited` for confirmed release/host absence.
3. **Queued prompt stall — already characterized.** `src/main/runtime/rpc/methods/orchestration-worker-start-prompt-contract.test.ts` proves a swallowed Enter is written exactly once, no provider turn starts, and current code persists `failed/agent_prompt_stalled`. Do not add an automatic rescue Enter or retry. The later prompt contract must bind a prompt ID to terminal/process/payload and report the furthest proven stage.
4. **Rejected `worker_done` — already characterized.** `src/main/runtime/rpc/methods/orchestration-send-dispatch-authority.test.ts`, `src/main/runtime/rpc/methods/orchestration-federation-lifecycle-settlement.test.ts`, and `src/main/runtime/rpc/methods/orchestration-migration-behavior.test.ts` cover spoofed sender, wrong Task, conflicting outcome, replayed rejection, and pre-contract rejection. Rejection remains durable while Task/Dispatch settlement remains unchanged.
## Executed gates
### Core current-contract suite
Command: the 22-file `pnpm test` invocation recorded in the worker run, covering mail identity/notification, ask/routing, mutation replay, prompt submission, worker prompt state, lifecycle authority/settlement, release/stop liveness, migration/skew, Dispatch/Resource DB state, WSL routing, and RPC contract.
Result: **22 files passed; 212 tests passed; 1 skipped**.
### New characterization and changed-file quality
```text
pnpm test \
src/main/runtime/rpc/orchestration-commit-notify-characterization.test.ts \
src/main/runtime/rpc/methods/orchestration-worker-release-recovery.test.ts
```
Result: **2 files passed; 10 tests passed**.
```text
pnpm run check:code-quality:changed
```
Result: **0 new findings across 2 changed source/test files; type-aware and React Doctor checks green**.
### Compatibility/provider suite
Command: the 15-file `pnpm test` invocation recorded in the worker run, covering federation sequence/ack integrity, schema skew, legacy RPC behavior, federated output/liveness, transcript fallback, WSL path execution, compatibility evidence, RPC mutation classification, and protocol compatibility.
Result: **15 files passed; 157 tests passed; 2 skipped**.
## Dependency gates before fan-out
| Gate | Must be true before dependent work starts | Blocks |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------- |
| R0-A receipt ordering | Every irreversible orchestration mutation completes/replays its receipt before advisory notify; injected notify failure plus DB reopen yields one effect and the original result. Changed payload conflicts; new request ID remains a new effect. | Mail/outbox, lifecycle reducer, fleet projections. |
| R0-B archive/liveness | Archive source and liveness are orthogonal. `unknown` is `unverifiable`; `releasing` requires execution-host observation; only confirmed host evidence is `exited`. | Resource recovery/cleanup, worker-read/fleet liveness. |
| R0-C canonical identity | Dispatch ID is Attempt; existing Resource ID is retained. Retry/creator/role/endpoint additions are optional/additive and host/process fenced. No second identity/store. | Identity, lifecycle, leases, transcript mirror. |
| R0-D receipt-domain separation | Mail, pointer, prompt, and lifecycle stages are separate typed domains. Legacy `delivered_at` is not promoted to submitted/acked, and PTY bytes do not prove turn start. | Mailbox/prompt/provider work. |
| R0-E migration/skew | Fresh, upgrade, downgrade, re-upgrade, physical-skew, future-schema, reset, and retention fixtures are green for each schema slice; old writers may omit new facts safely. | Any schema-bearing slice. |
| R0-F remote topology | Local/folder, WSL UNC, direct SSH capability, and federated peer routes cannot fall through to another host. Relay/contact loss remains `unverifiable`; new RPC fields are optional and new opcodes/capabilities are negotiated. | Transcript/read, remote release, remote fleet batching. |
| R0-G preserved lifecycle rejection | Stale/spoofed/conflicting `worker_done` stays durably rejected and cannot mutate Task/Dispatch state; exact duplicate rejection is idempotent. | Lifecycle consolidation and federation settlement. |
R1 identity work may begin only after R0-A and R0-B are corrected and R0-C through R0-G remain green. Prompt behavior changes remain separately gated on provider draft/submission capability and prompt-specific idempotency; the current queued-stall fixture is evidence, not permission to resend. Fleet/UI work remains last and is not part of R0.
-101
View File
@@ -1,101 +0,0 @@
# R3c post-completion audit: local fleet projection
## Scope and verdict
Audited every changed source/test file in the post-completion R3c change set,
with emphasis on the bounded read-only worker-list contract, identity joins,
liveness vocabulary, pagination, redaction, and old/new client behavior. Focused
tests passed (see below), but the slice is **NOT ACCEPTED** as contract-complete:
two high-confidence defects can produce false fleet facts or break a mixed-version
CLI, and pagination is not snapshot-stable as required by the R3c acceptance
criteria. No source files were modified.
## Findings
### F1 — High: status join can assign one terminal's live status to another Dispatch
`src/shared/orchestration-fleet-projection.ts:98-143` indexes the freshest status
globally by `paneKey` and by `terminalHandle`, then chooses whichever matching
entry has the newer `receivedAt`. It does not require
`status.orchestration.dispatchId`, endpoint/process incarnation, or any other
Attempt identity to agree with the durable worker row. A pane/terminal is an
address and can be reused after an older Dispatch completes; the old retained
row and the new row then share the same handle/key. A fresh status from the new
Attempt marks the old row `liveness.verdict: live`, copies provider/model,
workspace, host, and activity into that old row, even though the status belongs
to a different Attempt. Matching by handle also lets a newer status win over an
exact pane match. This violates the identity boundary and can create a false
live worker (and wrong host/workspace/provider) in the read-only projection.
**Required correction:** prefer an explicit dispatch/attempt identity when
present and fence status joins on endpoint/process incarnation; when proof is
absent or conflicting, return `unverifiable` rather than promoting the freshest
address match to live.
### F2 — High: upgraded CLI crashes against an older runtime response
`src/cli/handlers/orchestration/worker-terminal-handlers.ts:100-143` types and
dereferences `worker.projection.*` and `value.page.*` unconditionally in the
human formatter. An older running Orca runtime still returns the pre-R3c
`worker-list` shape (`workers` + `counts`, with no `projection` or `page`), and
there is no capability/version negotiation or fallback in this handler. A
newer CLI connected to that runtime therefore throws while formatting
(`worker.projection` is undefined) instead of preserving the old list output.
JSON mode happens to bypass the formatter, but normal text mode is a supported
CLI path. This violates the mixed-version compatibility requirement.
**Required correction:** treat projection/page as optional at the CLI boundary;
format legacy rows using the old fields and only append projection/pagination
details when present (or gate the new fields with an explicit capability).
### F3 — Medium: cursor pagination is not stable under concurrent updates
The durable query in `src/main/runtime/orchestration/db/worker-terminal/worker-terminal-listing.ts:102-117`
orders only by `COALESCE(w.created_at, d.created_at)` (SQLite timestamps are
second-granularity) with no deterministic Dispatch-id tie-breaker or snapshot
token. The RPC then slices that mutable result and requires the cursor row to
still exist in the current filtered set (`src/main/runtime/rpc/methods/orchestration-worker-release.ts:155-170`).
If rows are inserted/removed or change terminal state between calls, a next-page
request can reorder rows (duplicate/omit entries) or fail with
`invalid_argument` because the prior cursor disappeared from the filter. The
R3c acceptance contract calls for stable pagination with no duplicate/omitted
rows for a snapshot cursor; this implementation has neither a snapshot nor a
stable total-order key.
**Required correction:** establish a deterministic `(created_at, dispatch_id)`
order and issue a snapshot/epoch-bound cursor (or explicitly document and test a
weaker consistency model); do not reject a valid prior cursor solely because
the row changed state while paging.
### F4 — Low/Medium: freshness arithmetic is not clock-domain safe
`projectLiveness` (`src/shared/orchestration-fleet-projection.ts:145-169`) treats
any status with `now - receivedAt <= staleAfter` as fresh. A future timestamp
(remote host clock ahead, or malformed input) is therefore immediately
`live` and remains live until the local clock catches up. Agent status payloads
can carry remote-ingest timestamps (`connectionId`), while the projection has
no clock-domain/host offset proof. The SSH contract requires host-computed age
and loss of contact to remain `unverifiable`; an unbounded future timestamp can
thus create false liveness.
**Required correction:** reject/mark `unverifiable` for future observations
beyond a small tolerance and carry the execution-host clock domain/age, or limit
this projection to timestamps generated by the local host.
## Positive checks
- `pnpm test src/shared/orchestration-fleet-projection.test.ts src/main/runtime/rpc/methods/orchestration-worker-release.test.ts` — **2 files, 43 tests passed**.
- `pnpm test src/main/runtime/orchestration/r1-identity-migration.test.ts src/main/runtime/orchestration/orchestration-worker-dispatch-db.test.ts src/main/runtime/orchestration/db/dispatch-depth.test.ts` — **3 files, 36 tests passed**.
- `git diff --check` passed.
- Projection output excludes prompt/assistant/transcript bodies in the added
tests, uses `live`/`unverifiable`/`exited`, performs no terminal observation
calls, and caps page size at 100.
- The added receipt-before-nudge paths in `orchestration.ts` replay correctly
under injected notification failure (`orchestration-commit-notify-characterization.test.ts`
is covered by the existing focused suite).
## Acceptance verdict
**Fail pending F1 and F2; F3 must be resolved or explicitly re-scoped with a
tested consistency contract, and F4 should be addressed before remote/SSH rows
are treated as liveness evidence.**
@@ -1,123 +0,0 @@
# Independent behavioral evaluation: served Orca orchestration skill
Date: 2026-08-27
Workspace CLI: `node config/scripts/orca-dev.mjs` (rebuilt `out/cli/index.js`)
Connected runtime: `1.4.191-adhoc.20260827063356`, runtime ID `f301d6bd-73eb-47ff-971a-ab8b10508f19`
## Scope and method
I evaluated the Markdown actually generated by the workspace-matched CLI with
`skills get orchestration` and `skills get orchestration --full`. I read the
kernel and all seven generated reference sections, traced realistic coordinator,
worker, handoff, remote, recovery, and legacy scenarios, compared every command
family and material flag in those recipes with current CLI help, and exercised
read-only or explicitly non-effecting live commands. I did not create a Run,
Task, terminal, worktree, message other than the required heartbeat, or remote
resource, and I did not edit source or skill files.
## Verdict
**PASS / ACCEPTED.** The workspace-served rewrite makes each tested role's next
action explicit, keeps supervised orchestration distinct from ownership handoff,
preserves authority and remote/mixed-version safety, and resolves every
conditional reference. I found no stale or unsupported command recipe in the
workspace package.
## Concrete findings
| Area | Result | Evidence |
| ---------------------------------------- | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Coordinator next action | PASS | For “supervise two independent workers and wait for both,” the kernel gives `status -> run-create/run-use -> task-create for the full wave -> worker-start -> check --wait`; it then requires processing the complete FIFO Delivery and reusing, retaining, or releasing each settled terminal before acknowledgment. Timeout, heartbeat, quiet output, relay loss, and missing client are explicitly checkpoints or uncertainty, not failure. |
| Worker next action | PASS | A live injected preamble routes directly to four obligations: do the Task, use the preamble's `ask`, heartbeat only at the requested cadence, send capability-bound `worker_done` exactly once, then idle. The generic completion recipe includes the worker handle, Dispatch capability, Task ID, Dispatch ID, explicit outcome, modified files, and report path. The worker reference's heartbeat, ask/resume, escalation, and completion recipes preserve the same authority fields and use typed lifecycle flags. |
| Supervised orchestration vs handoff | PASS | “Supervise/monitor/wait/track/DAG/gate/ask-reply” routes to orchestration. “Hand off/start another agent or worktree” without supervision routes to `orca-cli`, creates no Run/Task/Dispatch, emits no lifecycle messages, and does not monitor. Model or effort selection does not change that classification. |
| Review and edit authority | PASS | A review-only completion authorizes synthesis of findings, not coordinator edits. If the user names a downstream owner, fixes and PR preparation remain with that owner. This prevents a coordinator from expanding its authority after a review Dispatch. |
| Current authority | PASS | The kernel defines Dispatch, not terminal title, copied IDs, old rows, transcript, or visible pane, as lifecycle authority. Workers must copy the exact executable, handle, capability, Task ID, and Dispatch ID from the live preamble. A valid `worker_done` settles Task and Dispatch automatically; stale or rejected completion cannot trigger release. |
| Remote, folder-workspace, SSH/WSL safety | PASS | The always-loaded floor says the execution host owns process, filesystem, transcript, stop, and cleanup facts; later remote operations route by Dispatch ID; and the only liveness verdicts are `live`, `unverifiable`, and `exited`. Lost contact never proves exit or permits local fallback. Folder workspaces are first-class, remote `current`/`new-child` are rejected as ambiguous, and WSL recovery preserves the exact returned executable and distro route. |
| Mixed-version safety | PASS | The kernel treats optional fields as absent and new stream operations as capability-gated. The remote reference warns that unknown opcodes may be dropped and disallows broadening a target or crossing the execution boundary during fallback. The legacy reference distinguishes compatibility, replay, read-only, and current authority; requires exact returned recovery arguments with the same executable; and makes unverifiable authority read-only. |
| Progressive-reference routing | PASS | Plain retrieval is 10,571 bytes / 171 output lines; `--full` is 30,512 bytes / 630 lines. Full output starts byte-for-byte with the kernel, then includes exactly one marker and an exact body for each of seven named references. The kernel maps concrete action gates to one owning reference: coordinator expansion, worker lifecycle, placement/remote, messaging/gates, recovery/cleanup, low-level topology, and legacy migration. The transport fetches all references together after a gate; selectivity is by the named section, while the ordinary path loads only the kernel. |
| Forbidden methodology/product language | PASS | A case-insensitive scan of the generated full package returned zero matches for the audited methodology/reference-product terms: `compound-engineering`, `ce-skill-work`, `ce-plan`, `ce-work`, `ce-code-review`, `phase-loaded`, `Overstory`, `Paperclip`, `Herdr`, `Gas Town`, `reference-product`, `product-methodology`, and `methodology`. The package describes Orca behavior directly. |
| Command and flag freshness | PASS, with help-text notes | Every referenced command family resolved through workspace CLI help. Live read-only calls for `status`, `task-list --run --ready --brief`, `dispatch-show`, `worker-show`, `worker-read --limit`, `inbox --full`, and `check --peek --format` succeeded. A real capability-bearing typed heartbeat from this Dispatch succeeded. Retired `coordinator-start`, `run`, and `run-stop` each exited 1, applied no effects, and returned exact `skills get orchestration --full` recovery arguments. No recipe used an unsupported flag. |
## Scenario traces
### Supervised coordinator
Scenario: “Have Codex and Claude investigate independently; supervise them and
wait for both results.”
The role table classifies the caller as coordinator without requiring a keyword
guess. The kernel says to create independent Tasks before waiting, start the full
wave, process each bounded Delivery, answer questions, validate active Dispatch
identity, decide cleanup ownership, acknowledge, and continue until every
expected Dispatch settles. The next action remains defined after both completion
and timeout.
### Full ownership handoff
Scenario: “Give this to another agent in a new worktree using a specific model
and xhigh effort.”
The role table classifies this as a handoff despite model/effort language and
routes to `orca-cli`. It explicitly forbids Run, Task, Dispatch, lifecycle, and
monitoring state. This is a clean ownership transfer rather than a supervised
Dispatch disguised as a handoff.
### Live dispatched worker
Scenario: this evaluation Dispatch (`task_578a9a901da0`,
`ctx_c8ef6729058a`).
`dispatch-show` proved the active Task/Dispatch/terminal/pane incarnation, and
`worker-show` reported `ready`, `input_accepted`, exact-worker `true`, and
liveness `live`. The preamble-shaped typed heartbeat, including exact `--from`,
`--dispatch-capability`, `--task-id`, `--dispatch-id`, and `--phase`, succeeded.
`worker-read --limit 1` returned a source-pinned transcript cursor and the same
`live` verdict. I did not execute the completion recipe before the report was
ready because it would settle this Dispatch.
### Remote uncertainty and recovery
Scenario: a connected-server worker disappears during a long operation.
The kernel forbids inferring failure or exit. The placement reference keeps the
Run home authoritative, uses `--on` only at start, and routes observation and
follow-up by Dispatch ID. The recovery table maps contact loss to
`unverifiable`; it permits retry only after a positively proven failed/stopped
attempt and requires explicit placement again. It never substitutes a local
process action.
### Legacy/mixed-version command
Scenario: a pre-update caller attempts a retired automatic scheduler command.
Live calls to `coordinator-start`, `run`, and `run-stop` returned
`orchestration_migration_required`, `effectsApplied: false`, and the exact
same-executable `skills get orchestration --full` next arguments. This matches
the kernel's “follow the receipt; do not translate or retry from memory” rule.
## CLI surface notes
Two help-rendering inconsistencies remain observable but do not invalidate any
skill recipe:
- `orchestration send` lists `--dispatch-capability` in Options but omits it from
the Usage synopsis. The flag is supported and the live heartbeat proved it.
- `orchestration check` correctly shows boolean `[--format]` in Usage and accepts
the package's no-value form, but its Options row labels the same flag as
`--format <png|jpeg>`.
The global `/usr/local/bin/orca` shim points into `/Applications/Orca.app` and
still serves the older 41,245-byte / 435-line monolith; its normal and `--full`
payloads are identical. Per coordinator direction, that installed-build skew is
a deployment note, not a defect in the workspace rewrite evaluated here. The
workspace CLI and rebuilt `out/cli` agree byte-for-byte on both new outputs.
## Verification
- Generated kernel SHA-256: `226af8dee12343155579025265f8fe4938006a4120a319a56f7aee825749cae4`.
- Generated full-package SHA-256: `03e3911b83ae9637e523df3b1378d6710e68430bf650c9a0513ca223304a76c3`.
- All seven reference markers occurred exactly once and every generated body
matched its bundled reference byte-for-byte.
- All referenced command help surfaces resolved; sampled live read-only and
non-effecting recovery calls behaved as documented.
- Focused package tests passed: 3 files, 84 tests.
@@ -1,48 +0,0 @@
# Independent recheck: served Orca orchestration skill
Date: 2026-08-27
CLI/package: rebuilt `out/cli/index.js` from this worktree (`pnpm run build:cli`)
Runtime: `1.4.191-adhoc.20260827063356`, runtime id `f301d6bd-73eb-47ff-971a-ab8b10508f19`
## Acceptance checks
| Check | Result | Evidence |
| ------------------------------------------------ | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Plain output is materially smaller than `--full` | PASS | `skills get orchestration`: 10,571 bytes / 171 lines; `--full`: 30,510 bytes / 630 lines (plain is 34.6% of full). |
| Full output starts with the kernel | PASS | `full.startsWith(kernel)` is true; both outputs begin with the orchestration frontmatter and `# Orca orchestration`. |
| All seven owned references resolve | PASS | Full output contains one marker and the complete body for each of `coordinator-loop.md`, `legacy-contract-migration.md`, `low-level-topology.md`, `messaging-and-gates.md`, `placement-and-remote.md`, `recovery-and-cleanup.md`, and `worker-contract.md`. Marker count is 7; kernel marker count is 0. |
| JSON/text parity | PASS | Parsed JSON `markdown` is byte-for-byte equal to plain text for both normal and `--full` retrieval; `full` flags are `false` and `true` respectively and topic names match. |
| Worker lifecycle recipes are capability-bound | PASS | Kernel reminder and `references/worker-contract.md` heartbeat and `worker_done` recipes all include `--from <worker_handle>`, `--dispatch-capability <capability>`, and `--task-id <task_id> --dispatch-id <dispatch_id>`. The current preamble values were used for a real heartbeat; it was accepted and updated `last_heartbeat_at`. |
## Scenario exercises
### Coordinator scenario
Simulated “supervise independent workers and wait for results” using the rebuilt
CLI against the existing Run without creating new state: `status --json`,
`task-list --run run_f61f3e3758e7 --brief --json`,
`task-list --run run_f61f3e3758e7 --ready --brief --json`,
`dispatch-show --task task_e6cd1bde4e26 --json`, and a one-millisecond
`check --wait --types worker_done,escalation,question --json` checkpoint all
returned valid JSON. The task list reported 24 tasks (the recheck task was
`dispatched`), the ready filter reported zero, dispatch-show identified the
active Dispatch, and the wait returned `count: 0, timedOut: true`; no effects
were applied. The kernel's loop and completion-accounting language was
unambiguous for this coordinator path.
### Worker scenario
Used this Dispatch's injected identity (`term_7bc13793-bc47-425a-83f9-634fbba47a15`,
`dcap_ihvtilM1D-v5fonojW-OodICzNK4nbXB-uMffoXUD24`,
`task_e6cd1bde4e26`, `ctx_4a21e1a6df30`) with the worker-contract heartbeat
recipe. The capability-bearing typed invocation succeeded (`ok: true`, message
type `heartbeat`, payload phase `reviewing`); a subsequent dispatch inspection
confirmed the Dispatch remained `dispatched` and its heartbeat timestamp
advanced. The worker_done recipe was checked statically (it is intentionally
not executed during an active recheck because it settles this Dispatch), and
its flags match the injected preamble contract and current `send --help`.
## Verdict
ACCEPTED. No stale, ambiguous, or capability-missing recipe was found in the
rebuilt served package; no source or skill files were edited.
File diff suppressed because it is too large Load Diff
-195
View File
@@ -1,195 +0,0 @@
# Orca orchestration vNext: implementation audit
This is the implementation-oriented companion to `orchestration-issues.md`. It
turns the broad issue list into a small set of contracts that can be implemented
and verified independently, then composed into one DAG and consolidation PR.
The audit is based on the current main-process code, its tests, the Windows
queued-input/Claude-scrollback feedback, and the local reference checkouts in
`/Users/jinwoo/refs`.
## Executive decision
The work is not a rewrite and it is not 30 unrelated bug fixes. Most reports are
symptoms of seven boundaries being implicit: identity, delivery, lifecycle truth,
mutation atomicity, recovery, fleet observability, and provider/host execution.
The existing SQLite dispatch model, capability fencing, worker settlement, and
archive flow are good foundations. vNext should add durable facts and projections
at those boundaries while preserving existing CLI/RPC behavior through additive
fields and compatibility adapters.
The critical path is:
```text
P0 vocabulary + contract fixtures
├─ P1 stable Attempt/Endpoint/Resource identity
│ └─ P3 lifecycle receipts and outcome settlement
│ └─ P6 fleet projection and attention UI
├─ P2 durable mailbox/wakeups and delivery receipts
│ └─ P6 fleet projection and attention UI
├─ P4 lease/recovery supervisor
│ └─ P6 fleet projection and attention UI
└─ P5 provider transcript/read capability + telemetry
└─ P7 federation/SSH/WSL conformance
P3 + P4 + P2 + P5 + P7 → provider/OS/remote conformance matrix → small promotion PR
```
After P0, P1, P2, P4, and P5 can proceed in parallel. P3 needs P1 and the
receipt vocabulary; P6 waits for stable facts rather than inventing another
state store.
## What an agent/user sees today
### Sending during an active turn
Messages are durable SQLite rows. Run delivery is FIFO, bounded to 50 messages,
and one batch is outstanding at a time. A live/idle target may receive an
advisory PTY pointer telling it to run `orchestration check`; a busy target is
redriven at an idle edge or after runtime restart. `orchestration check --wait`
can resolve an in-memory runtime waiter immediately when a message arrives.
The wakeup is not itself durable: if the process misses the in-memory notification,
the message remains safe in the mailbox and is redriven after a fresh live-idle
observation or explicit check. A same-runtime `check --wait` read/register race is
not proven on the current synchronous path and must be characterized before adding
an epoch. A sender usually sees only `Sent <message-id>`/relay `Queued`, not
whether the provider accepted or submitted it. `--ack <deliveryId>` acknowledges
the whole delivered batch; reply-channel rows do not have the same clear
acknowledgment cursor. The proven duplicate seam is commit-versus-notify: if
notification throws after a durable mutation, an exact retry must replay the
recorded result rather than insert another row.
For `terminal send --enter`, `agent_prompt_stalled` can currently be returned
when the target TUI is mid-turn even though the text is queued in its input box
and is consumed when the turn ends. Treating that result as failure causes
duplicate resends. The desired observable facts are separate domains:
```text
mail: recorded → routed → acked
nudge: pointer_text_written → submit_confirmed | submit_ambiguous | submit_failed
prompt: terminal_bytes_written → submission_observed → turn_started
worker: report_persisted → settled
```
Not every provider can prove every fact. A queued prompt must never be reported
as a hard delivery failure; the receipt must say which fact is proven and give an
idempotent retry/wait operation.
### Reading a worker
`worker-read --source auto` prefers a bounded decoded projection from the currently
selected provider session and falls back to a bounded, redacted PTY tail with an
explicit reason. Attach-time binding is not guaranteed yet. Cursors are scoped to
dispatch, process incarnation, source identity, and position; released workers are
read from an archived snapshot. Remote/federated reads execute on the owning host;
relay loss is `unverifiable`, never `exited`.
`terminal.read` is intentionally provider-neutral PTY output. In the reported
Windows session, Codex exposed a long buffer while Claude Code exposed about the
current 50-line screen, so raw terminal scrollback cannot be the confirmation
contract. Provider transcript and submitted-prompt history must be separate,
capability-described evidence; screen text remains advisory.
## Implementation matrix
Each row is a bounded workstream. “Not in scope” is part of the contract: it
prevents a worker from expanding a focused change into a second orchestration
system. Paths are likely touch points, not a prescription to edit every file.
| ID / user-facing goal | Current main behavior and concrete gap | In-scope change | Explicitly not in scope | Likely modules | Independent verification | Depends on / parallelism |
| ------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- |
| **P0 Contract vocabulary**\nEveryone sees the same honest states. | `ready`, `input_accepted`, `worker_done`, generic status, and PTY writes are interpreted differently by sender, coordinator, and UI. | Publish state/authority matrix; define receipt stages above; define `live`/`unverifiable`/`exited`; add fixture generators and duplicate/late/reordered event contract tests. | No behavior migration, event-sourcing rewrite, or AI policy. | `docs/`, orchestration DB test fixtures, shared status types. | Contract tests assert legal transitions, idempotency, and old-client omission safety. | None; gates all other slices. |
| **P1 Stable identity and authority**\nA retry or nested worker cannot complete the wrong lane. | Dispatch capabilities and pane/process checks are strong, but retry lineage, creator/role, and endpoint-incarnation facts are incomplete. | Treat immutable `dispatch_contexts.id` as Attempt ID and retain the existing terminal Resource ID. Add only `retry_of`, creator/role, endpoint-incarnation, and typed unsupervised/remote attachment facts; bind lifecycle/mail/cleanup mutations to Attempt + process incarnation + host. | No parallel Attempt/Resource tables, no flag-day ID migration, no trusting model-copied identity, no removal of legacy handles. | `db/dispatch-context/*`, `db/worker-dispatch/*`, lifecycle reconciliation, RPC caller identity, migrations. | Stale retry, pane remint, process mismatch, nested parent/child, legacy backfill, and payload-identity spoof tests. | P0; parallel with P2/P4/P5. |
| **P2 Durable mailbox and wakeups**\nA message is not lost or falsely failed, and agents do not resend duplicates. | SQLite mail and replay are durable, but wake notification is in-memory/advisory; sender lacks queued-vs-submitted truth; cross-run authorization must remain explicit; commit-versus-notify can erase a mutation receipt after the message is already committed. | Complete the existing mutation receipt and durable nudge outbox before notification; keep separate mail, pointer, prompt, and lifecycle receipt domains; expose `recorded/routed/queued/submitted/acknowledged` facts where each is provable; retain whole-batch ack unless partial ack is negotiated and justified; replay labels; prompt idempotency only in the terminal-delivery slice; advisory pointer redrive after restart/idle. Return `queued_pending_turn` (or equivalent) for accepted TUI input and optional `--wait-submit`. Protect human drafts only when provider emptiness is proven. Add a durable wait epoch only if an async race is reproduced. | No guarantee that every provider can prove turn start; no automatic resend of ambiguous prompts; pointers are never authoritative; no new stream opcode without negotiation. | `db/messages/*`, `db/runs/run-delivery.ts`, `mailbox-notification-coordinator.ts`, `mailbox-pointer-*`, `orca-runtime.ts`, orchestration RPC/CLI handlers, prompt-submission verifier. | Unit tests for commit/notify crash, receipt transitions, ack/replay races, queued-mid-turn no-duplicate; restart/idle wake integration; Windows Claude/Codex matrix; cross-run auth. Characterize (do not assume) waiter-registration races. | P0; parallel with P1/P4/P5. |
| **P3 Attempt lifecycle and completion truth**\nThe coordinator can tell “still running”, “finished”, and “unknown” apart. | `worker_done` is a validated fast path and settlement is transactional, but Task/Dispatch/worker status has multiple writers; `ready` proves input accepted, not provider turn start; late reports and missing reports need one guarded authority. | First route every lifecycle writer through a guarded transition primitive that updates legacy projections and appends receipts atomically. Then add observation/outcome projections (`outcome_unknown`/`finished_unverified`) and host-computed heartbeat age without widening closed enums in the first slice. | No success from quiet PTY/Git cleanliness/prose; no collapse of SSH loss into exit; retain worker_done fast path and circuit-breaker policy initially. | `db/dispatch-context/*`, task status writers, lifecycle reconciliation/coordinator, heartbeat and RPC receipt code. | Direct-update ratchet, rollback injection, event replay/order/idempotency; active-sibling races; report-after-reconnect; missing-report; SSH `live/unverifiable/exited`; old/new client optional fields. | P0 + P1 + P2; then enables P6. |
| **P4 Recovery and resource leases**\nRestart, stop, retry, and explicit release/recovery converge without killing the wrong process. | Stop/abandon/release already fence capabilities, but archive/liveness and local/federated close semantics differ. Existing resource accounting is useful; retention-expiry metadata, TTL policy, and bulk cleanup are validate-first follow-up work. | Preserve archive-before-close/liveness truth, existing resource accounting, ownership/release fencing, explicit retain/release, recovery receipts for observed operations, and remote release/archive behavior. Validate retention reason/age with a dry-run cleanup protocol before proposing expiry metadata, TTL policy, or paginated/bulk cleanup; add richer operation nouns only after a concrete unrecoverable workflow is measured. | No broad pane/title close, no auto-release external/user-owned/transferred resources, no retry that duplicates ownership, and no automatic release or recovery heuristic that overrides takeover or identity conflict. | `db/worker-dispatch/*`, `db/worker-terminal/*`, ownership/release reconciliation, federation control, archive reader. | False-exited archive, stop output-loss, abandon retention, release identity conflict, reconnect, and remote release. Deferred retention/TTL/cleanup requires retention reason/age measurement, dry-run coverage of external/user-owned/transferred/federated/unverifiable resources, and explicit policy/fencing acceptance tests. | P0 + P1 + P3; behavior follows evidence; retention/TTL/cleanup remains validate-first. |
| **P5 Provider transcript/read contract**\nCoordinator evidence is consistent across Claude, Codex, folder workspaces, SSH, and WSL. | Provider transcript decoders/resolvers are specific (Claude/Codex/Grok/OMP); `worker-read` is transcript-first with PTY fallback; `terminal.read` is screen/tail only; direct SSH routing and archive fallback provenance need correction. | Consolidate existing provider maps only where behavior changes; carry host/connection identity into resolution; expose read source/exactness/completeness/fallback metadata; preserve source/process cursor fencing. Gate any in-flight mirror or passive stream on a demonstrated loss/latency case and the identity/remote contracts. | No universal transcript format, no desktop-side SSH/WSL file reads, no replacing interactive PTY UX with JSONL, no secret/capability logging. | `native-chat/*`, `worker-transcript-*`, `orchestration-worker-output.ts`, terminal query RPC/CLI, provider profiles, archive/relay paths. | Decoder fixtures; SSH/WSL topology tests; cursor rotation/process replacement; archive provenance; screen-vs-transcript characterization; subscription gaps/backpressure; redaction/size limits. | P0; topology/identity fixtures first. |
| **P6 Fleet query and attention projection**\nOne query answers what is running, what needs attention, and the next safe action. | `workerList` already provides a local projection and renderer status is push-fed; facts are split across lanes, and remote observation is per-dispatch. | Extend/combine the existing local read-only projection first. Add a negotiated per-host batched snapshot only after P7 fixtures, with budgets, partial-host errors, stable pagination, and no transcript bodies. Project root completion, input, approval, failure, and interruption separately; validate notification defaults. | No second UI state store, no merge queue/cost budget in kernel slice, no blanket notification suppression, no assumption that the renderer polls every lane. | coordinator projection, RPC/CLI list/show, renderer fleet view/notifications (after contracts), federation query. | Local parity; bounded 100-worker/two-host calls; stale/unknown display; pagination; partial-host budgets; notification coalescing; redaction. | P2 + P3 + P4 + P5/P7; do last. |
| **P7 Cross-host/provider conformance**\nThe same contract survives mixed versions and remote loss. | Execution-host ownership and capability negotiation exist, but skew tests and batched remote semantics are fragmented. | Make old/new DB, RPC, relay, provider, SSH, WSL, and federated fixtures continuous infrastructure; cache unsupported capabilities by peer fingerprint/runtime epoch; add batched host snapshot negotiation only when P6 needs it. | No new unnegotiated stream opcodes; no treating relay/client absence as process death. | federation RPC, SSH/WSL adapters, protocol/capability cache, remote compatibility tests, provider fixtures. | Matrix is executable from R0; assert `unverifiable` on contact loss, typed unavailable errors for individual reads, and safe downgrade on old peer. | P0 + P1; continuous alongside P2–P6, not a final-only gate. |
## Detailed DAG and worker boundaries
Use one worker per row (or split P4 into recovery and release). A worker
may add tests and contracts for its row, but should not modify another row's
projection or silently change CLI semantics. Each slice must leave a green,
independently runnable test target and a short migration/rollout note.
```text
P0 contract fixtures
├── P1 identity/authority ───────┐
│ ├── P3 lifecycle/completion ──┐
├── P2 mailbox/receipts ─────────┘ │
├── P4 leases/recovery ───────────────(schema may start early)─┤
└── P5 transcript/provider ──┐ ├── P6 fleet/attention
└── P7 remote/provider matrix ───┘
P3 + P4 + P2 + P5 + P7 → final conformance and a small promotion PR
```
Recommended implementation order inside a slice:
1. Add characterization tests for current behavior (including the current
whole-batch ack, pointer command, and provider-specific screen limits).
2. Add durable schema/receipt fields with dual-read/dual-write where needed.
3. Add runtime behavior behind an additive capability or feature gate.
4. Exercise restart, duplicate, timeout, and remote-loss cases.
5. Promote the new projection only after parity checks pass; retain a rollback
path for old clients and old remote servers.
## Terminal read: explicit product contract
The references do not implement one universal transcript mechanism:
| Product | Default read source | Provider-specific? | Fallback/remote behavior | Orca takeaway |
| --------- | ---------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- |
| Overstory | Runtime adapters and structured headless stdout/EventStore; tmux visibility fallback. | Adapter API is agent-agnostic; parsers are provider-specific. | Local-focused; no comparable SSH/federation contract. | Keep a narrow provider adapter boundary and prefer structured events. |
| Paperclip | Adapter stdout/stderr/system chunks into durable NDJSON RunLogStore; optional object-storage mirror. | Provider-neutral log chunks behind provider adapters. | Live tail plus durable archive survives pod loss. | Measure Orca output loss first; consider bounded in-flight mirroring only if a concrete gap remains. |
| Herdr | Execution-host terminal history (`visible`, `recent`, `detection`) and passive subscriptions. | Read API agent-agnostic; lifecycle/session hooks provider-aware. | Server owns panes; remote attach runs reads on the target host; explicit `unknown`. | Keep screen/recent/detection intent distinct; instrument existing subscriptions before adding a passive watch. |
| Gas Town | tmux transport by default; Claude JSONL watcher and OTel events are opt-in. | Lifecycle protocol generic; transcript watcher Claude-specific. | Restart-first sessions and append-only feed/audit events; no general SSH transcript reader. | Use tmux only as bounded evidence, not truth; correlate output with run/attempt IDs. |
Therefore vNext introduces no fundamentally novel orchestration concept. It
combines proven patterns—durable receipts/mail, server-owned execution,
provider adapters, append-only audit/evidence, leases, and explicit unknown
states—into one Orca contract while retaining Orca-specific desktop/worktree/
SSH/WSL integration. The novel part is the composition and the explicit
cross-provider semantics, not a new agent paradigm.
## Acceptance gates for consolidation
- A prompt accepted while a provider turn is busy returns `terminal_queued` (or
`queued_pending_turn`), never a hard failure; retrying its id is idempotent.
- A durable message survives runtime restart and is redriven after a fresh
live-idle observation or explicit check; replay is labeled and ack is
idempotent. An offline process is not described as having been woken.
- After local authority attachment, every supervised local dispatch has exactly
one execution-host resource or an explicit external/remote attachment;
pre-attach, unsupervised, and federated-home rows have typed absence. Every
residual resource has a retention reason and safe next action.
- No false `ready` or `exited` is emitted solely from a PTY write, timeout, quiet
screen, missing client inventory, or relay loss.
- `worker-read` and `terminal.read` state source, exactness, cursor scope, and
fallback reason; Claude/Codex/unsupported-provider behavior is covered by
fixtures and the screen-vs-transcript distinction remains explicit.
- A bounded local fleet projection has parity with existing `workerList`; a
negotiated per-host snapshot is promoted only after its latency/partial-host
tests pass. Input, approval, failure, interruption, stale, and unverifiable
remain distinct.
- A five-worker wave's notification policy is measured before changing defaults;
drafts are untouched whenever composer emptiness is unproven.
- Existing issues become regression/contract tests or are recorded as accepted
limitations. No source change is merged without the relevant focused test
target and a remote/Windows impact note.
## Suggested PR decomposition
1. **Contracts PR:** P0 plus characterization fixtures and shared types.
2. **Identity + mailbox PRs:** P1 and P2 in parallel, each additive and dual-read.
3. **Lifecycle + leases PRs:** P3 and P4 after contract/identity tests pass.
4. **Transcript + federation PRs:** P5 and P7, including provider/OS matrix.
5. **Fleet/attention PR:** P6 consumes only durable facts from the prior PRs.
6. **Small promotion PR:** after the additive slices pass full conformance,
change consumers/defaults and remove only shims shown unused by mixed-version
telemetry. Do not merge all implementation branches as one feature PR.
This keeps each implementation focused, reviewable, and independently verifiable
while still yielding the coherent DAG the orchestration redesign needs.
@@ -1,23 +0,0 @@
# Orchestration vNext lifecycle blocker fix
## Outcome
- Exact-authority `worker_done` reports now settle a `start_unknown` worker after reconnect. The blocked Task is atomically re-armed, then the Task, Dispatch, and worker settle through guarded lifecycle transitions with receipts.
- `createGate` now moves its Task to `blocked` through `transitionLifecycleWithDb` in the existing savepoint.
- Context-only stop/abandon now moves the Dispatch to `failed` and, when current, its Task to `blocked` through `transitionLifecycleWithDb` in the caller transaction.
- The direct-write ratchet covers decision-gate creation and context-only dispatch release.
## Regression coverage
- Exact pane/process authority and retained capability after `start_unknown`, with successful Task/Dispatch/worker settlement and receipt assertions.
- Decision-gate Task receipt and rollback of the gate, Dispatch, Task, and receipts when receipt insertion fails.
- Context-only stop/abandon Dispatch and Task receipts, plus rollback when receipt insertion fails.
## Verification
- Focused lifecycle suite: 4 files, 48 tests passed.
- Core orchestration DB and worker-dispatch suites: 73 tests passed.
- `pnpm tc:node`: passed.
- `pnpm run check:code-quality:changed`: passed with zero new findings.
- `git diff --check`: passed.
- `orchestration-worker-release.test.ts`: 40 passed, 1 pre-existing/unrelated failure. The test expected `closed_exited_terminal` but observed `closed_agent_terminal`; the files changed for this lifecycle fix do not touch worker-release process-action selection.
-79
View File
@@ -1,79 +0,0 @@
# Orchestration vNext source review
Review target: current worktree changes against `orchestration-vnext-implementation-audit.md`,
`orchestration-vnext-implementation-audit.html`, and `orchestration-final-release-audit.md`.
Inspection was read-only apart from this report artifact.
## Verdict
**Not ready for an unconditional READY sign-off.** The focused tests and node typecheck are green,
but the source still has two concrete lifecycle-contract blockers: valid worker completion after an
uncertain start is rejected, and two production Task-status writers bypass the guarded transition
boundary and durable lifecycle receipts. The final audit's “READY” claim therefore overstates the
implemented lifecycle coverage.
## Blockers
### B1 — `start_unknown` worker reports cannot settle
`markWorkerStartUnknown` deliberately leaves the Dispatch active and moves the Task to `blocked`
(`src/main/runtime/orchestration/db/worker-dispatch/worker-dispatch-outcome.ts:102-140`). A later
authenticated `worker_done` therefore reaches `settleWorkerReport` with `task.status ===
'blocked'`, which is rejected by the `dispatch.status/task.status === 'dispatched'` guard
(`src/main/runtime/orchestration/db/dispatch-context/worker-report-settlement.ts:75-91`). Even if the
Task guard were relaxed, the worker transition is hard-coded to `from: 'ready'` and rejects
`start_unknown` (`.../worker-report-settlement.ts:164-171`). This contradicts the implementation
audit's P3 verification requirement for “report-after-reconnect” and strands a real worker that
started while the start RPC was ambiguous; add a regression test that marks a worker
`start_unknown`, delivers an exact-authority `worker_done`, and expects normal settlement.
### B2 — Task lifecycle writes still bypass the transition primitive
The direct-write ratchet only scans six files (`src/main/runtime/orchestration/db/lifecycle-transition-boundary.test.ts:5-21`),
but production writes remain in:
- `src/main/runtime/orchestration/db/decision-gates/decision-gate-store.ts:67-76` — `createGate`
executes `UPDATE tasks SET status = 'blocked'` after inserting the gate and completing active
dispatches. No compare-and-swap edge check or `lifecycle_transition_receipts` row is produced.
- `src/main/runtime/orchestration/context-only-dispatch-release.ts:37-58` — context-only release
directly sets the Dispatch to `failed` and conditionally sets the Task to `blocked`, again with no
guarded transition or receipt.
Both paths are reachable production lifecycle operations, not test fixtures. They can silently
overwrite an invalid/concurrent state and violate P3's “every lifecycle writer … updates legacy
projections and appends receipts atomically” requirement. Expand the ratchet and route both writes
through `transitionLifecycleWithDb` (with savepoint/transaction coverage).
## Validate-first / deferred items
- **Remote attachment state ordering:** `recordRemoteAttachmentStage` accepts arbitrary `state`
replacements with no legal-transition/CAS check or lifecycle receipt
(`src/main/runtime/orchestration/db/federation/remote-dispatch-attachment-create.ts:77-126`).
Stop/relay settlement has additional direct state updates in
`db/federation/remote-dispatch-attachment-stop.ts` and `db/federation/federation-relay-item.ts`.
Existing federation tests cover authentication/ack/liveness, but not stale or reordered stage
updates; characterize those races before treating remote lifecycle as fully centralized.
- **Recovery receipt identity:** `recordWorkerTerminalRecoveryAttempt` stores a receipt with
`entity='worker'` but `entity_id=<resource id>` (`src/main/runtime/orchestration/db/worker-terminal/worker-terminal-resource-store.ts:124-143`).
Worker lifecycle IDs are dispatch IDs, so `getLifecycleTransitionReceipts('worker', dispatchId)`
cannot retrieve these recovery facts. Decide whether this is a separate resource-receipt domain
or bind the receipt to the owning dispatch, then add a query/assertion test.
- **Physical certification remains outstanding:** SSH, WSL, and provider-version certification is
explicitly unperformed in `orchestration-final-release-audit.md:5-10`; deterministic fixtures do
not replace those jobs.
- **Retention expiry/TTL/bulk cleanup:** correctly remains validate-first/deferred per
`orchestration-final-release-audit.md:129-133`; no blocker provided those features stay absent.
## Verification performed
- `pnpm test src/main/runtime/orchestration/db/lifecycle-transition.test.ts src/main/runtime/orchestration/db/lifecycle-transition-boundary.test.ts src/main/runtime/orchestration/db/attempt-outcome-projection.test.ts src/main/runtime/orchestration/r1-identity-migration.test.ts` — **4 files, 17 tests passed**.
- `pnpm test src/main/runtime/rpc/orchestration-commit-notify-characterization.test.ts src/main/runtime/rpc/terminal-prompt-delivery-receipt.test.ts` — **2 files, 14 tests passed**.
- Federation/release focused suite — **4 files, 65 tests passed**.
- Fleet/capability focused suite — **3 files, 19 tests passed**.
- `pnpm tc:node` — passed.
- `git diff --check` — passed.
- Broad changed-test invocation: **235 passed, 1 failed** in
`src/cli/handlers/orchestration-worker-cli.test.ts`; the failure was the inherited Orca dev
environment making `isDevCliInvocation()` true while the test expects `devMode: false`. Re-running
that file with `ORCA_USER_DATA_PATH=` and `ORCA_DEV_CLI_INVOCATION=0` passed (11/11), so this is a
test-environment isolation issue rather than a newly introduced worker-start contract failure.