fix(orchestration): keep published refusal messages and type the DB claim guards

Codex review of #18902:

- Every call site keeps the exact message it published on main
  ("Task not found: <id>", "only a ready Task can start.", "cannot retry
  from Dispatch"); the shared contract now takes the message per site and
  only owns the code and data. Baseline strings are pinned as literals.
- createDispatchContext's own missing/non-ready guards, including the
  atomic-claim loser, now emit the same typed receipt instead of a bare
  Error, so a dispatch that races a status change no longer flattens to
  runtime_error. Covered by a dispatcher-level race test.
- Invalid --retry-of keeps task_not_startable but now carries status,
  unmetDependencies, retryOf, and a retry-specific next step.
- Dependency recovery text distinguishes waiting on running deps from
  retrying/unblocking failed ones.
- CLI test adds an unknown-code case so the old-client claim rests on an
  assertion, not a comment; SSH test asserts exact stdout.
- Guide table narrowed to the covered preflight cases; occupancy stays
  runtime_error and is named as such.
This commit is contained in:
Jinwoo-H
2026-09-05 17:53:22 -04:00
parent 9a05f4828c
commit a6438ba0ff
12 changed files with 237 additions and 73 deletions
+7 -7
View File
@@ -180,14 +180,14 @@ Dispatch rules:
- After 3 consecutive failures on one task, the dispatch context circuit-breaks and the task is marked failed.
- Use `task-list --brief --json` for coordinator sweeps; it collapses whitespace and caps each echoed spec at 160 characters (`spec_truncated` marks shortened rows). Omit `--brief` when the full spec is required, or when an older CLI rejects it as an unknown flag.
Dispatch and `worker-start` refuse with a stable `error.code`; read it before choosing a recovery, and treat `error.data.nextSteps` as the exact recovery text:
`dispatch` and `worker-start` refuse the following preflight cases with a stable `error.code`; read it before choosing a recovery, and treat `error.data.nextSteps` as the exact recovery text. Older hosts may omit `data`, so treat every field as optional.
| Code | Meaning | Recovery |
| -------------------- | ------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------- |
| `task_not_found` | No Task with that id in the bound Run (`data.taskId`) | Check `task-list --json`; create the Task with `task-create` if it does not exist |
| `task_not_startable` | Task is not `ready` (`data.status`, `data.unmetDependencies`) | Wait for the listed dependencies with `check --wait`, or inspect `dispatch-show` if already dispatched |
| `inject_rejected` | Target terminal refused injection (`data.terminal`, `data.reason`) | Start a recognized agent there or pick another terminal; or dispatch without `--inject` and `terminal send` |
| `runtime_error` | Unexpected failure; nothing about the Task or terminal is implied | Read the message, inspect state, and do not retry unchanged |
| Code | Meaning | Recovery |
| -------------------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| `task_not_found` | No Task with that id, or not in the bound Run (`data.taskId`, `data.runId`) | Check `task-list --json`; create the Task with `task-create` if it does not exist |
| `task_not_startable` | Task cannot start now: not `ready`, or invalid `--retry-of` (`data.status`, `data.unmetDependencies`, `data.retryOf`) | Wait for running dependencies with `check --wait`; retry or unblock failed ones; inspect `dispatch-show` if already dispatched |
| `inject_rejected` | `--inject` refused because no recognized agent runs in the target (`data.terminal`, `data.reason`) | Start a recognized agent there or pick another terminal; or dispatch without `--inject` and use `terminal send` |
| `runtime_error` | Any other failure, including a target terminal that already owns an active Dispatch | Read the message, inspect state, and do not retry unchanged |
## How deep workers can nest
File diff suppressed because one or more lines are too long
@@ -12,20 +12,20 @@ afterEach(() => {
vi.restoreAllMocks()
})
// Why: these are the exact envelopes the RPC dispatcher test proved the runtime emits. The CLI
// never enumerates codes, so an older CLI prints the same recovery for a code it has never seen.
// Why: these are the exact envelopes the RPC dispatcher test proved the runtime emits. This
// checkout's formatter never enumerates codes (verified below with a code no build has defined),
// which is what lets a client that predates a new code still print its message and nextSteps.
describe('orchestration dispatch refusals through the CLI error boundary', () => {
it.each([
{
receipt: taskNotFoundRefusal('task_missing'),
receipt: taskNotFoundRefusal('Task not found: task_missing', { taskId: 'task_missing' }),
recovery: /task-create|task-list/
},
{
receipt: taskNotStartableRefusal({
taskId: 'task_child',
status: 'pending',
unmetDependencies: ['task_parent']
}),
receipt: taskNotStartableRefusal(
'Task task_child is pending; only ready tasks can be dispatched',
{ taskId: 'task_child', status: 'pending', unmetDependencies: ['task_parent'] }
),
recovery: /task_parent/
},
{
@@ -49,6 +49,29 @@ describe('orchestration dispatch refusals through the CLI error boundary', () =>
})
})
// Why: a code this build has never defined stands in for a future host's new code; if the
// formatter ever starts gating on known codes, this is the assertion that catches it.
it('prints an unknown code with its message and nextSteps unchanged', () => {
const failure: RuntimeRpcFailure = {
id: 'rpc_1',
ok: false,
error: {
code: 'code_from_a_newer_host',
message: 'Refused for a reason this CLI has never heard of.',
data: { nextSteps: ['Do the thing the newer host suggested.'] }
},
_meta: { runtimeId: 'runtime_1' }
}
const error = new RuntimeRpcFailureError(failure)
expect(formatCliError(error, { commandPath: ['orchestration', 'dispatch'] })).toBe(
'Refused for a reason this CLI has never heard of.\nNext step: Do the thing the newer host suggested.'
)
const log = vi.spyOn(console, 'log').mockImplementation(() => {})
reportCliError(error, true, { commandPath: ['orchestration', 'dispatch'] })
expect(JSON.parse(log.mock.calls[0]?.[0] as string)).toEqual(failure)
})
function envelope(receipt: DispatchRefusalReceipt): RuntimeRpcFailure {
return { id: 'rpc_1', ok: false, error: receipt, _meta: { runtimeId: 'runtime_1' } }
}
@@ -7,6 +7,7 @@ import { paneKeyMatchSuffix } from '../pane-key-match'
import { claimDispatchContextRow } from '../dispatch-row-writer'
import type { DispatchCreator } from '../dispatch-depth'
import type { OrchestrationDb } from '../orchestration-db'
import { taskNotFoundError, taskNotStartableError } from '../../task-dispatch-refusal'
export function createDispatchContext(
this: OrchestrationDb,
@@ -26,10 +27,14 @@ export function createDispatchContext(
const depth = this.resolveChildDispatchDepth(params.creator, params.maxDepth)
const task = this.getTask(taskId)
if (!task) {
throw new Error(`Task not found: ${taskId}`)
throw taskNotFoundError(`Task not found: ${taskId}`, { taskId })
}
if (task.status !== 'ready') {
throw new Error(`Task ${taskId} is ${task.status}; only ready tasks can be dispatched`)
throw taskNotStartableError(
this,
`Task ${taskId} is ${task.status}; only ready tasks can be dispatched`,
task
)
}
// Why: lock on pane identity too, so a reminted handle can't open a second concurrent dispatch on the same pane.
@@ -72,9 +77,12 @@ export function createDispatchContext(
`Terminal ${assigneeHandle} already has an active dispatch (${occupied.id} for task ${occupied.task_id})`
)
}
throw new Error(
`Task ${taskId} is ${current?.status ?? 'missing'}; only ready tasks can be dispatched`
)
// Why: the atomic claim lost to a concurrent status change; report it with the same
// typed receipt as the precheck so the loser can recover instead of reading runtime_error.
const message = `Task ${taskId} is ${current?.status ?? 'missing'}; only ready tasks can be dispatched`
throw current
? taskNotStartableError(this, message, current)
: taskNotFoundError(message, { taskId })
}
this.db.prepare("UPDATE tasks SET status = 'dispatched' WHERE id = ?").run(taskId)
const dispatch = this.db
@@ -61,7 +61,7 @@ export function createStartingWorkerDispatch(
}
const task = this.getTask(params.taskId)
if (!task) {
throw taskNotFoundError(params.taskId)
throw taskNotFoundError(`Task ${params.taskId} was not found.`, { taskId: params.taskId })
}
if (params.retryOf) {
const prior = this.getDispatchContextById(params.retryOf)
@@ -75,13 +75,19 @@ export function createStartingWorkerDispatch(
!['failed', 'stopped', 'abandoned'].includes(priorWorker.state) ||
!['failed', 'blocked'].includes(task.status)
) {
throw new OrchestrationError(
'task_not_startable',
`Task ${task.id} cannot retry from Dispatch ${params.retryOf}.`
throw taskNotStartableError(
this,
`Task ${task.id} cannot retry from Dispatch ${params.retryOf}.`,
task,
params.retryOf
)
}
} else if (task.status !== 'ready') {
throw taskNotStartableError(this, task)
throw taskNotStartableError(
this,
`Task ${task.id} is ${task.status}; only a ready Task can start.`,
task
)
}
const id = generateId('ctx')
@@ -9,16 +9,27 @@ import {
type InjectRejectionReason
} from '../../../shared/orchestration-dispatch-refusal-contract'
export function taskNotFoundError(taskId: string, runId?: string): OrchestrationError {
return toError(taskNotFoundRefusal(taskId, runId))
// Why: each site keeps the exact message it published before; only the code and data are shared.
export function taskNotFoundError(
message: string,
detail: { taskId: string; runId?: string }
): OrchestrationError {
return toError(taskNotFoundRefusal(message, detail))
}
export function taskNotStartableError(db: OrchestrationDb, task: TaskRow): OrchestrationError {
export function taskNotStartableError(
db: OrchestrationDb,
message: string,
task: TaskRow,
retryOf?: string
): OrchestrationError {
return toError(
taskNotStartableRefusal({
taskNotStartableRefusal(message, {
taskId: task.id,
status: task.status,
unmetDependencies: unmetTaskDependencies(db, task)
unmetDependencies: unmetTaskDependencies(db, task),
...(retryOf ? { retryOf } : {})
})
)
}
@@ -1,6 +1,7 @@
import { afterEach, describe, expect, it, vi } from 'vitest'
import { ORCHESTRATION_CONTRACT_VERSION } from '../../../../shared/protocol-version'
import {
buildInjectRejectionMessage,
injectRejectedRefusal,
taskNotFoundRefusal,
taskNotStartableRefusal
@@ -36,7 +37,9 @@ describe('orchestration dispatch failure codes through RpcDispatcher', () => {
const response = await dispatch(harness, { task: 'task_missing', to: WORKER_HANDLE })
expect(expectFailure(response).error).toEqual(taskNotFoundRefusal('task_missing'))
expect(expectFailure(response).error).toEqual(
taskNotFoundRefusal('Task not found: task_missing', { taskId: 'task_missing' })
)
})
it('reports task_not_startable with the unmet dependencies for a pending task', async () => {
@@ -47,7 +50,7 @@ describe('orchestration dispatch failure codes through RpcDispatcher', () => {
const response = await dispatch(harness, { task: child.id, to: WORKER_HANDLE })
expect(expectFailure(response).error).toEqual(
taskNotStartableRefusal({
taskNotStartableRefusal(`Task ${child.id} is pending; only ready tasks can be dispatched`, {
taskId: child.id,
status: 'pending',
unmetDependencies: [parent.id]
@@ -64,7 +67,11 @@ describe('orchestration dispatch failure codes through RpcDispatcher', () => {
const response = await dispatch(harness, { task: task.id, to: WORKER_HANDLE })
expect(expectFailure(response).error).toEqual(
taskNotStartableRefusal({ taskId: task.id, status: 'completed', unmetDependencies: [] })
taskNotStartableRefusal(`Task ${task.id} is completed; only ready tasks can be dispatched`, {
taskId: task.id,
status: 'completed',
unmetDependencies: []
})
)
})
@@ -78,6 +85,7 @@ describe('orchestration dispatch failure codes through RpcDispatcher', () => {
expect(expectFailure(response).error).toEqual(
injectRejectedRefusal(WORKER_HANDLE, 'no_agent_detected')
)
expect(expectFailure(response).error.message).toBe(buildInjectRejectionMessage(WORKER_HANDLE))
expect(harness.db.getTask(task.id)?.status).toBe('ready')
expect(harness.db.getDispatchContext(task.id)).toBeUndefined()
})
@@ -97,7 +105,7 @@ describe('orchestration dispatch failure codes through RpcDispatcher', () => {
)
expect(expectFailure(response).error).toEqual(
taskNotStartableRefusal({
taskNotStartableRefusal(`Task ${child.id} is pending; only a ready Task can start.`, {
taskId: child.id,
status: 'pending',
unmetDependencies: [parent.id]
@@ -106,6 +114,56 @@ describe('orchestration dispatch failure codes through RpcDispatcher', () => {
expect(harness.db.getTask(child.id)?.status).toBe('pending')
})
it('reports task_not_startable with retry detail for an invalid --retry-of', async () => {
const harness = createHarness()
const task = harness.db.createTask({ spec: 'work' })
mockWorkerStartTopology(harness.runtime)
const response = await harness.dispatcher.dispatch(
request('orchestration.workerStart', {
task: task.id,
from: COORDINATOR_HANDLE,
agent: 'claude',
retryOf: 'ctx_missing'
})
)
expect(expectFailure(response).error).toEqual(
taskNotStartableRefusal(`Task ${task.id} cannot retry from Dispatch ctx_missing.`, {
taskId: task.id,
status: 'ready',
unmetDependencies: [],
retryOf: 'ctx_missing'
})
)
})
it('types the atomic claim loser when the task changes after the ready precheck', async () => {
const harness = createHarness()
const task = harness.db.createTask({ spec: 'raced' })
// Why: the pane lookup runs after the ready precheck and before the DB claim, so failing the
// task there is the interleaving a concurrent status change produces. The loser's DB refusal
// must carry the same typed receipt instead of the bare Error it used to throw.
vi.mocked(harness.runtime.getTerminalPaneKey).mockImplementation((handle) => {
if (handle === WORKER_HANDLE) {
harness.db.updateTaskStatus(task.id, 'failed', 'raced out')
return WORKER_PANE
}
return handle === COORDINATOR_HANDLE ? COORDINATOR_PANE : null
})
const response = await dispatch(harness, { task: task.id, to: WORKER_HANDLE })
expect(expectFailure(response).error).toEqual(
taskNotStartableRefusal(`Task ${task.id} is failed; only ready tasks can be dispatched`, {
taskId: task.id,
status: 'failed',
unmetDependencies: []
})
)
expect(harness.db.getDispatchContext(task.id)).toBeUndefined()
})
it('keeps runtime_error for a genuinely unexpected dispatch failure', async () => {
const harness = createHarness()
const task = harness.db.createTask({ spec: 'work' })
@@ -26,7 +26,7 @@ export const ORCHESTRATION_DISPATCH_METHODS: RpcMethod[] = [
const db = runtime.getOrchestrationDb()
const task = db.getTask(params.task)
if (!task) {
throw taskNotFoundError(params.task)
throw taskNotFoundError(`Task not found: ${params.task}`, { taskId: params.task })
}
const run = resolveRunScope(runtime, {
runId: params.run,
@@ -36,7 +36,10 @@ export const ORCHESTRATION_DISPATCH_METHODS: RpcMethod[] = [
callerEvidence: orchestrationCompatibilityEvidence
})
if (task.run_id !== run.id) {
throw taskNotFoundError(task.id, run.id)
throw taskNotFoundError(`Task ${task.id} was not found in Run ${run.id}.`, {
taskId: task.id,
runId: run.id
})
}
// Why: dry-run previews the preamble without mutating state, so it skips the ready-status check and uses a placeholder dispatchId.
@@ -67,7 +70,11 @@ export const ORCHESTRATION_DISPATCH_METHODS: RpcMethod[] = [
const to = params.to
if (task.status !== 'ready') {
throw taskNotStartableError(db, task)
throw taskNotStartableError(
db,
`Task ${params.task} is ${task.status}; only ready tasks can be dispatched`,
task
)
}
// Why: injecting the preamble into a bare shell dumps it as shell commands (gibberish), so require a detected agent first.
@@ -59,7 +59,10 @@ export const ORCHESTRATION_WORKER_START_METHODS: RpcMethod[] = [
}
const task = db.getTask(params.task)
if (!task || task.run_id !== run.id) {
throw taskNotFoundError(params.task, run.id)
throw taskNotFoundError(`Task ${params.task} was not found in Run ${run.id}.`, {
taskId: params.task,
runId: run.id
})
}
if (params.on) {
@@ -14,8 +14,8 @@ vi.mock('../persistence', () => ({
import { OrcaRuntimeService } from '../runtime/orca-runtime'
import { runRemoteOrcaCli } from './ssh-remote-orca-cli'
// Why: the SSH bridge relays the host CLI's stdout byte-for-byte; a typed dispatch refusal must
// arrive at the remote agent with the same code and data it would see locally.
// Why: the SSH bridge captures the host CLI child's stdout and exit code without reparsing; this
// pins that a typed refusal envelope and its nonzero exit reach the remote agent unchanged.
it('relays typed dispatch refusal codes from the host CLI unchanged', async () => {
const child = new EventEmitter() as EventEmitter & {
stdout: EventEmitter
@@ -55,11 +55,10 @@ it('relays typed dispatch refusal codes from the host CLI unchanged', async () =
}
)
const stdout = `${JSON.stringify(refusal, null, 2)}\n`
await Promise.resolve()
child.stdout.emit('data', Buffer.from(`${JSON.stringify(refusal, null, 2)}\n`))
child.stdout.emit('data', Buffer.from(stdout))
child.emit('close', 1)
const result = await resultPromise
expect(result.exitCode).toBe(1)
expect(JSON.parse(result.stdout)).toEqual(refusal)
expect(await resultPromise).toEqual({ stdout, stderr: '', exitCode: 1 })
})
@@ -1,5 +1,9 @@
import { describe, expect, it } from 'vitest'
import { buildInjectRejectionMessage } from './orchestration-dispatch-refusal-contract'
import {
buildInjectRejectionMessage,
taskNotFoundRefusal,
taskNotStartableRefusal
} from './orchestration-dispatch-refusal-contract'
import { TUI_AGENT_CONFIG } from './tui-agent-config'
import { recognizeAgentProcess } from './agent-process-recognition'
@@ -29,3 +33,35 @@ describe('buildInjectRejectionMessage', () => {
}
})
})
// Why: these strings are published receipts; they are pinned as literals, independent of the
// builders, so a refactor cannot silently rewrite them together with the expectation.
describe('dispatch refusal receipts keep their published messages', () => {
it('leaves the message exactly as each call site supplies it', () => {
expect(taskNotFoundRefusal('Task not found: task_1', { taskId: 'task_1' }).message).toBe(
'Task not found: task_1'
)
expect(
taskNotStartableRefusal('Task task_1 is pending; only a ready Task can start.', {
taskId: 'task_1',
status: 'pending',
unmetDependencies: []
}).message
).toBe('Task task_1 is pending; only a ready Task can start.')
})
it('tailors nextSteps to retry, dependency, occupancy, and terminal-status refusals', () => {
const base = { taskId: 'task_1', status: 'failed', unmetDependencies: [] }
expect(taskNotStartableRefusal('m', { ...base, retryOf: 'ctx_1' }).data.nextSteps[0]).toMatch(
/--retry-of.*ctx_1/
)
expect(
taskNotStartableRefusal('m', { ...base, status: 'pending', unmetDependencies: ['task_0'] })
.data.nextSteps[0]
).toMatch(/task_0.*unblock failed/)
expect(
taskNotStartableRefusal('m', { ...base, status: 'dispatched' }).data.nextSteps[0]
).toMatch(/dispatch-show --task task_1/)
expect(taskNotStartableRefusal('m', base).data.nextSteps[0]).toMatch(/failed Task cannot/)
})
})
@@ -1,10 +1,8 @@
import { TUI_AGENT_CONFIG } from './tui-agent-config'
// Why: dispatch and worker-start used to flatten these refusals to runtime_error, so an agent
// reading the receipt could not tell "create the task" from "wait on deps" from "pick another
// terminal". This module is the single source of each receipt's code, message, and data so the
// runtime emits it and the CLI test formats exactly the same envelope. Recovery rides in
// data.nextSteps, which every shipped CLI already prints for any code.
// Why: one source for each dispatch refusal's code, message, and data, so the runtime emits and
// the CLI test formats the identical envelope. Messages are supplied per call site because each
// existing string is a published receipt an old consumer may match on.
export type DispatchRefusalReceipt = {
code: 'task_not_found' | 'task_not_startable' | 'inject_rejected'
@@ -12,15 +10,15 @@ export type DispatchRefusalReceipt = {
data: Record<string, unknown> & { nextSteps: string[] }
}
export function taskNotFoundRefusal(taskId: string, runId?: string): DispatchRefusalReceipt {
export function taskNotFoundRefusal(
message: string,
detail: { taskId: string; runId?: string }
): DispatchRefusalReceipt {
return {
code: 'task_not_found',
message: runId
? `Task ${taskId} was not found in Run ${runId}.`
: `Task ${taskId} was not found.`,
message,
data: {
taskId,
...(runId ? { runId } : {}),
...detail,
nextSteps: [
'Run orca orchestration task-list --json in the bound Run to find the intended Task id.',
'If the Task does not exist yet, create it with orca orchestration task-create --spec <text> --json.'
@@ -29,30 +27,45 @@ export function taskNotFoundRefusal(taskId: string, runId?: string): DispatchRef
}
}
export function taskNotStartableRefusal(task: {
export type TaskNotStartableDetail = {
taskId: string
status: string
unmetDependencies: string[]
}): DispatchRefusalReceipt {
const nextSteps =
task.unmetDependencies.length > 0
? [
`Wait for ${task.unmetDependencies.join(', ')} to complete (orca orchestration check --wait --json), then dispatch again.`
]
: task.status === 'dispatched'
? [
'The Task already has an active Dispatch; inspect it with orca orchestration dispatch-show --task <task_id> --json.'
]
: [
`A ${task.status} Task cannot be dispatched; create a new Task or use worker-start --retry-of for a failed attempt.`
]
retryOf?: string
}
export function taskNotStartableRefusal(
message: string,
detail: TaskNotStartableDetail
): DispatchRefusalReceipt {
return {
code: 'task_not_startable',
message: `Task ${task.taskId} is ${task.status}; only ready tasks can be dispatched`,
data: { ...task, nextSteps }
message,
data: { ...detail, nextSteps: taskNotStartableNextSteps(detail) }
}
}
function taskNotStartableNextSteps(detail: TaskNotStartableDetail): string[] {
if (detail.retryOf) {
return [
`--retry-of must name the latest settled Dispatch of a failed or blocked Task; check orca orchestration dispatch-show --task ${detail.taskId} --json and orca orchestration worker-show --dispatch ${detail.retryOf} --json.`
]
}
if (detail.unmetDependencies.length > 0) {
return [
`Dependencies ${detail.unmetDependencies.join(', ')} are not completed. Wait for running ones with orca orchestration check --wait --json; retry or unblock failed ones before dispatching again.`
]
}
if (detail.status === 'dispatched') {
return [
`The Task already has an active Dispatch; inspect it with orca orchestration dispatch-show --task ${detail.taskId} --json.`
]
}
return [
`A ${detail.status} Task cannot be dispatched; create a new Task or use worker-start --retry-of for a failed attempt.`
]
}
// Why: the old five-name example read as an allowlist (#15125); derive from the field detection keys on so it cannot drift.
// Not filtered by `disabledTuiAgents` — that gates Orca's launchers, not detection, so a hand-started disabled agent still injects.
const RECOGNIZED_AGENT_PROCESS_NAMES = [