docs(orchestration): scope worker-list recipes and rank liveness evidence

Every worker-list recipe now passes --run <run_id>. Unscoped, a real profile
returned 2103 rows and none of the live workers, so the stall recipe the kernel
prescribes could not find them.

Reword the disagreement rule: the fleet verdict still decides, except against an
unverifiable reason that names a client-side gap (missing_status,
host_unavailable), where a worker-show verdict from the execution host is the
better evidence. Note that a peer lacking the fleet-snapshot capability also
reports host_unavailable on the row while the host warning names
capability_unsupported. unverifiable still authorizes nothing.
This commit is contained in:
Jinwoo-H
2026-09-05 03:44:01 -04:00
parent 2bd094c103
commit 71d2bc6d4b
5 changed files with 53 additions and 29 deletions
@@ -163,7 +163,9 @@ describe('orchestration kernel', () => {
expect(kernel).toContain("`worker-list`'s `projection.liveness` is the fleet verdict")
expect(kernel).toContain("`worker-show`'s `observation.status` is PTY liveness only")
expect(kernel).toContain('After three consecutive empty waits')
expect(kernel).toContain('`ORCA orchestration worker-list --include-remote --json`')
expect(kernel).toContain(
'`ORCA orchestration worker-list --run <run_id> --include-remote --json`'
)
expect(kernel).toContain(
'`projection.attention` categories, `projection.attention.requiresAction`, and literal `projection.nextAction` argv'
)
@@ -214,7 +216,9 @@ describe('orchestration kernel', () => {
)
expect(kernel).toContain('worker-release --dispatch <dispatch_id>')
expect(kernel).toContain('check --ack <delivery_id> --wait')
expect(squash(kernel)).toContain('`worker-list --terminal-state reclaimable --json`')
expect(squash(kernel)).toContain(
'`worker-list --run <run_id> --terminal-state reclaimable --json`'
)
expect(squash(kernel)).toContain('do not follow it with `task-update --status completed`')
})
@@ -356,7 +360,9 @@ describe('owned orchestration references', () => {
'ORCA project setup-existing-folder --project <project_id> --host <host_id> --path <abs_path> --kind folder --json'
)
expect(squash(reference)).toContain('and rejects a plain directory')
expect(reference).toContain('ORCA orchestration worker-list --include-remote --json')
expect(reference).toContain(
'ORCA orchestration worker-list --run <run_id> --include-remote --json'
)
expect(squash(reference)).toContain(
'enumerate remote workers with `--include-remote` or every one of them reads `unverifiable`'
)
@@ -413,14 +419,16 @@ describe('owned orchestration references', () => {
it('names worker-list as the enumerating command and the agent-liveness authority', () => {
const reference = squash(readReference('recovery-and-cleanup.md'))
expect(reference).toContain('ORCA orchestration worker-list --json')
expect(reference).toContain('ORCA orchestration worker-list --run <run_id> --json')
expect(reference).toContain("`worker-show`'s `observation.status` is PTY liveness only")
expect(reference).toContain(
'`projection.attention.categories`, `projection.attention.requiresAction`'
)
expect(reference).toContain('`projection.nextAction` argv')
expect(reference).toContain('the fleet verdict decides')
expect(reference).toContain('ORCA orchestration worker-list --include-remote --json')
expect(reference).toContain(
'ORCA orchestration worker-list --run <run_id> --include-remote --json'
)
expect(reference).toContain('reads `unverifiable` until you enumerate with `--include-remote`')
expect(reference).toContain('follow `page.nextCursor` with `--cursor <value>`')
})
+13 -13
View File
@@ -134,17 +134,17 @@ a checkpoint, not a failure. Do not stop, retry, release, or launch a duplicate
editor without the positive proof `## Outcome` requires.
After three consecutive empty waits, stop waiting blindly and enumerate with
`ORCA orchestration worker-list --include-remote --json`, acting on each row's
`projection.attention` categories, `projection.attention.requiresAction`, and
literal `projection.nextAction` argv. An `inspect` `nextAction` on a `live` row
with `attention.requiresAction` false is informational, not a command to re-run:
keep waiting with `check --wait`. Leave the wait only on positive proof the
agent stopped: `exited` liveness, the worker's own observation of process exit,
or a transcript whose final agent turn sent no `worker_done`. Then load
`references/recovery-and-cleanup.md` and choose `worker-stop` or
`worker-abandon` explicitly. `unverifiable` is absence, including when
`worker-show` reports `agentWait` null. Absence never authorizes stop, abandon,
retry, or release; keep waiting or inspect.
`ORCA orchestration worker-list --run <run_id> --include-remote --json`, acting
on each row's `projection.attention` categories,
`projection.attention.requiresAction`, and literal `projection.nextAction` argv.
An `inspect` `nextAction` on a `live` row with `attention.requiresAction` false
is informational, not a command to re-run: keep waiting with `check --wait`.
Leave the wait only on positive proof the agent stopped: `exited` liveness, the
worker's own observation of process exit, or a transcript whose final agent turn
sent no `worker_done`. Then load `references/recovery-and-cleanup.md` and choose
`worker-stop` or `worker-abandon` explicitly. `unverifiable` is absence,
including when `worker-show` reports `agentWait` null. Absence never authorizes
stop, abandon, retry, or release; keep waiting or inspect.
`worker-start` is the normal path, composing placement, terminal readiness,
prompt injection, and supervised resource ownership. `dispatch --inject` leaves
@@ -174,8 +174,8 @@ follow its exact recovery receipt and never substitute `terminal close`.
A valid `worker_done` settles the Task and Dispatch automatically; do not follow
it with `task-update --status completed`. Enumerate the terminals still owing a
decision with `worker-list --terminal-state reclaimable --json`, and do not end
the coordinator turn until it returns none.
decision with `worker-list --run <run_id> --terminal-state reclaimable --json`,
and do not end the coordinator turn until it returns none.
## Conditional references
@@ -62,11 +62,13 @@ terminal handle.
ORCA orchestration worker-show --dispatch <dispatch_id> --json
ORCA orchestration worker-read --dispatch <dispatch_id> --limit 50 --json
ORCA orchestration send --to dispatch:<dispatch_id> --subject "Follow-up" --body "<guidance>" --json
ORCA orchestration worker-list --include-remote --json
ORCA orchestration worker-list --run <run_id> --include-remote --json
```
`worker-list` reads local fleet state only; enumerate remote workers with
`--include-remote` or every one of them reads `unverifiable`.
`--include-remote` or every one of them reads `unverifiable`. Scope every list
with `--run <run_id>`: unscoped, it reports every Dispatch this runtime has
recorded, and the workers you are waiting on are lost in that history.
## Execution-host and mixed-version floor
@@ -16,8 +16,8 @@ decision, stop/abandon request, retention request, or uncertain release.
## Inspect before acting
```text
ORCA orchestration worker-list --json
ORCA orchestration worker-list --include-remote --json
ORCA orchestration worker-list --run <run_id> --json
ORCA orchestration worker-list --run <run_id> --include-remote --json
ORCA orchestration worker-show --dispatch <dispatch_id> --json
ORCA orchestration worker-read --dispatch <dispatch_id> --limit 50 --json
```
@@ -25,9 +25,23 @@ ORCA orchestration worker-read --dispatch <dispatch_id> --limit 50 --json
`worker-list` is the enumerating command and the authority on agent liveness:
each row carries `projection.liveness`, `projection.attention.categories`,
`projection.attention.requiresAction`, and a literal `projection.nextAction`
argv to run. `worker-show`'s `observation.status` is PTY liveness only, so a
`live` terminal whose agent died at a trust prompt still reads `live` there.
When the two disagree, the fleet verdict decides.
argv to run. Always scope it with `--run <run_id>`; an unscoped list reports
every Dispatch this runtime has ever recorded and buries the live ones.
`worker-show`'s `observation.status` is PTY liveness only, so a `live` terminal
whose agent died at a trust prompt still reads `live` there.
When the two disagree, the fleet verdict decides — unless the fleet row is
`unverifiable` for a reason that names a gap on this client rather than a fact
about the worker. `missing_status` and `host_unavailable` are such gaps: the
first means this runtime holds no status row, the second that it could not ask
the execution host at all. Against either, a `worker-show` verdict sourced from
the execution host is the better evidence and outranks the row. A stale peer
that simply lacks the fleet-snapshot capability also reports `host_unavailable`
on the row; the accompanying host warning names `capability_unsupported`, so
read the warning before treating the row as contact loss.
This never promotes absence. `unverifiable` from either command still authorizes
nothing — only a positive `live` or `exited` verdict does.
A worker started with `--on <environment>` reads `unverifiable` until you
enumerate with `--include-remote`, which asks its execution host for the
File diff suppressed because one or more lines are too long