mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 00:02:29 +00:00
* feat(orchestration): deliver worker results to a structured chat coordinator
Resolve a Run's handle-less session coordinator and a session:<id> mailbox to
the live session, wake an evicted session for the delivery, redrive a session's
own mail on its idle edge, route a terminal view's pointer through its PTY, and
accept session:<id> (or a bare Orca session id) as a recipient.
* test(orchestration): pin coordinator delivery through the real session host, wake, idempotence and session addresses
* test(orchestration): type the coordinator mail fixture's attach params
* fix(orchestration): refuse session recipients with the caller codes, and treat a worker without its identity as undeliverable
* test(orchestration): pin a chat's terminal view reading the chat's coordinator mail
* test(orchestration): read coordinator mail fixtures through checked guards instead of assertions
* fix(orchestration): name a structured session's CLI by $ORCA_CLI_COMMAND in its pointer turn
* fix(orchestration): render the pointer's CLI invocation for the shell the session runs in
* fix(orchestration): address a session recipient where its check reads, so a structured worker gets its mail
* test(orchestration): pin the runtime's own idle-edge redrive wiring; say a released session is not running, not ended
* fix(orchestration): point mail that has not been pointed, not mail nobody has acked, naming the ack a held batch needs
* feat(orchestration): hand a /clear-replaced chat's Runs and unread mail to the session that replaced it
* feat(orchestration): adopt a /clear predecessor's Runs at the clear's commit, the edge its replacement's own status misses
* fix(orchestration): log a wake that could not resume a session, instead of retaining silently
* test(orchestration): give coordinator mail waits a budget that holds under a loaded parallel run
* refactor(orchestration): narrow a retained pointer's dispatch state by type guard instead of a cast
* fix(orchestration): give back a pointer whose admitted turn never ran
A pending send stamped its rows delivered and dropped its operation row, so a
provider that died before echoing left the last result pointed at nobody. The
lane now awaits the admitted turn's settlement: accepted consumes the claim,
anything else returns the rows and drops this send's operation row, so the
re-point on the next edge is a new send rather than a replay of unknown.
* fix(native-chat): keep a committed /clear from failing on its replacement observer
The observer runs after the clear's durable commit; a throwing adoption turned a
committed clear into a failed RPC. It is now best-effort and logged, like the
status feed's observer, and the successor's idle edges re-derive the adoption.
* test(orchestration): pin /clear adoption across a chain of clears and a predecessor's check
A session cleared twice before any edge hands both predecessors' Runs and mail
to the end of the chain. A predecessor still live in the clear's tail reads
none of the re-addressed mail: a session's direct mailbox is consume-on-read
and holds no replayable batch.
* test(orchestration): type the pending-settlement host test's send input without an assertion
* fix(orchestration): keep a chat's orchestration address across /clear by deriving its lineage
A chat's orchestration address is now the first session of its /clear lineage. Every session of the
lineage resolves to that one actor when it acts and when it is reached, and delivery goes to the
lineage's live session. Nothing is rewritten at a clear, so the predecessor-adoption path is gone:
the commit-edge observer, the idle-edge rebind, and the unread-mail re-address. That path could
unbind the successor's own Run and orphan all but one Run of a chain.
The idle edge now opens the orchestration database through its lazy getter and logs when it cannot,
instead of reading a field that stays null until the first orchestration call after a restart.
* fix(orchestration): give back a structured pointer claim an earlier process left open
A pointer the host admits as pending stamps its batch delivered, and only an in-memory settlement
waiter gives it back if no turn ran. A process that died in that window left the batch stamped with
nothing to release it, so the last result on that mailbox was never pointed again. When the
database opens, every surviving pointer operation row is from an earlier process: its stamped batch
is found by the row's fingerprint, released, and the row dropped, before the restored-mailbox scan
points it again. A row whose batch was never stamped keeps its id for the retry.
* test(orchestration): pin that a cleared chat's sends carry its conversation's address
* fix(orchestration): open the orchestration database at an idle edge only when it already exists
A profile with no orchestration database has no mail to redrive, so a structured chat's idle edge no
longer creates one, and says nothing. An existing database is still opened lazily there.
* fix(orchestration): read a chat-coordinated Run's session through the actor's generation
A handle-less Run names its coordinator session only by an actor that still counts at the Run's
current generation, the same rule every other binding read uses; an actor an older binary's rebind
or unbind left behind no longer routes the Run's mail. A test pins that a cleared chat's run-create
and run-use write its conversation's root actor at the Run's current generation.
* refactor(orchestration): address a session's conversation by its bare root Orca session id
Carries the Orca session id rename into coordinator delivery. A session's orchestration
identity holds its conversation's bare Orca session id, the /clear lineage root's, and
its mail address is derived from it by formatOrcaSessionAddress; a Run a cleared chat
creates or uses stores that bare root id. A session recipient carries the parsed id and
its address as distinct types, a handle-less coordinator's session is read through
currentRunCoordinatorOrcaSessionId, and session ids read from records or PTY bindings
are checked with isOrcaSessionId before they become an identity. The session address
prefix comes from the one exported constant.
* refactor(orchestration): canonicalize a cleared session through the one id hook and the party resolver
- canonicalOrcaSessionId now walks a session's /clear lineage to its root; the
parallel session identity and lost-worker rule are deleted, so the caller
resolver, recipient routing, reach and idle-edge mailboxes all resolve a
session through resolveOrcaSessionParty.
- A Dispatch row's assignee_orca_session_id goes through the same hook.
- The terminal-view delivery lane is gone with the terminal handoff: no PTY is
bound to a session, so a chat's mail is always a session turn.
- Pins a send to a Run-less chat's session address after a restart, when the
send must start the agent-session host before routing reads its record.
* fix(orchestration): point a structured session with the PTY lane's exact text
A chat or structured worker is now told what a terminal agent is told: the pointer is
formatMessagePointer with the CLI name the PTY lane resolves for a local terminal (orca, or
orca-dev in a dev build), with no shell-specific invocation and no ack lesson. The lane still
excludes the batch a reader holds unacknowledged and points newer mail; that stays host-side,
and the reader's own check replays the held batch and names its ack as it does for a terminal.
* fix(orchestration): deliver a cleared structured worker's mail to its live successor
A terminal keeps its handle across /clear; a structured worker's successor now does the same.
Mail at the worker's handle, its dispatch mailbox, and a Run it coordinates resolved to the
session minted for the worker, which /clear replaced. Each now walks the /clear lineage forward
to the live session, which the caller resolver already treats as the worker.
* test(orchestration): compare a chat's pointer turn to the PTY lane's text for this build's CLI name
* refactor(orchestration): resolve the local CLI name once, for the PTY lane and the structured lane alike
* fix(orchestration): let a /clear-ed chat restate its address, placed by the host's lineage
The CLI entry compared a restated --from/--terminal with the injected session id as a
plain string, so a cleared chat restating the address it had before the clear (its lineage
root, the address it keeps) was refused. Only the host's session records know the lineage,
so a session address that is not this session's own spelling is now sent as the caller
param, where the host's canonical-id check accepts it or refuses it before any effect.
Plain restatements are still dropped and any other name is still refused at the entry.
* test(orchestration): read the coordinator journal through the async snapshot, and wait for the held turn's handover
Main made journalSnapshot async and delivers an accepted send once the host hands it over, so the fixture awaits the snapshot and waits for the provider's turn before echoing it.
* refactor(orchestration): point a chat whose agent is not running through the plain send
Main's host no longer has hold/release: an accepted send starts the agent itself. The
pointer lane's wake step called host.hold, which no longer exists, so it is deleted
from the pointer host, the delivery lane and their tests. An idle or evicted chat gets
its pointer through the same send a user message takes; the claim is still consumed
only on accepted and given back otherwise.
* fix(orchestration): leave a chat whose provider died stopped instead of respawning it for mail
A pointer whose provider died before echoing it is given back. The death's own status
edge then redrove the mail, and since a send starts the agent, a provider that died on
every turn was restarted about once a second for as long as the mail was unread. The
pointer lane now reads, on every attempt, whether the session's latest send never ran
because its provider exited or could not start, and holds the mail until a later send
runs. Every trigger passes through that gate: a parked retry, the idle edge's re-derive
and new mail. A rejection for any other reason still points at the next idle edge.
* fix(orchestration): hold mail after a start the person must fix, placing every failure kind
A start refused for a reason only the person can fix (not signed in, history too large,
a managed-account problem) or one the host stopped because it never came is held like a
failed start: the mail waits for the person's next message, which also retries the start.
An account switch still in progress is transient, so it stays ungated and the first edge
after the switch settles points the mail. One exhaustive record places every rejection
kind, so a new kind does not compile until it is placed.
* fix(orchestration): retry a structured pointer under its own id instead of gating on the failure reason
A pointer send that failed was given back and re-sent under a fresh operation id,
so every status edge after a provider death was a new send that started the provider
again. A reason-string gate held some of those deaths, but missed a Codex crash with
turn/start in flight (it settles unknown with the connection's error), latched all later
mail after one transient death, and did not keep a user's Stop.
A retry now reuses the mailbox's operation id, which the host answers by replaying the
recorded verdict without reaching the provider. The id is re-minted only for new mail,
after a later send ran, for a row an earlier process left, or once an account switch
settles. Rows are stamped only on accepted, so the admitted-stage claim, its give-back
and the restart claim-release scan are gone.
* test(orchestration): pin that a Stop keeps a pointer unsent before the next status edge, too
* fix(orchestration): point a structured chat's mail by the same rule as a terminal's
The chat lane pointed newer mail past a batch its reader had checked and not
acknowledged, while the terminal lane skips a mailbox until that batch is acked. A chat
now waits for the ack the same way a terminal does, and the lookup that let the chat lane
filter the held batch out is removed.
* fix(orchestration): refuse a session address that gate-list or task-list --run would drop
The CLI entry lets a `session:` address that is not the session's own spelling through
for the host to place, but gate-list and task-list send no caller when --run names the
Run, so `--from session:<other>` was silently ignored. With --run they now refuse it
(consumer_fenced) before any request, the same as a conflicting terminal handle.
* fix(orchestration): keep a failed mail redrive from skipping a chat's first-turn workspace rename
A structured session's status callback redrives its mail before the first-turn workspace
auto-rename, outside any try, so a database error there threw past the rename. The redrive
now logs its failure and returns.
* chore: take main's pnpm-lock.yaml the merge of origin/main left stale
* fix(orchestration): give a stamp or park decision its own variant so the send branch narrows
* fix(orchestration): re-mint a held structured pointer from facts that cannot strand it
A pointer row whose id had no submission in the journal was always resent under that id,
before any re-mint test ran. After a rewind rebuilt the journal, or a send refused before it
was recorded aged past the host's 24h admission window, the mail was held forever: neither
the person's next turn nor a restart pointed it again.
The re-mint tests now run first. "The agent has run since" is any accepted send submitted
after the row was minted; "an earlier process minted it" is a row whose id this lane did not
send, not a wall-clock comparison a clock step could fool; and a row the host never recorded
is re-minted once it is too old for the host to admit. The account-switch exception is gone:
nothing tells the lane when a switch ends, and each outside edge during one added another
pointer and failure to the chat. A parked pointer now retains as turn-unsettled.
* fix(orchestration): let the host check a session caller that gate-list or task-list --run names
5f753af0a7 refused any `--from session:<x>` beside --run that was not the session's own
spelling, which also refused a /clear-ed chat restating its lineage root, an address the host
accepts everywhere else. The CLI now sends that address with --run, and the host's declared
caller check accepts the root and refuses anyone else (consumer_fenced) before any effect.
* fix(orchestration): date a pointer on the journal's clock so a backward clock step cannot re-mint it every edge
669 lines
23 KiB
TypeScript
669 lines
23 KiB
TypeScript
/**
|
|
* A command that runs inside a structured agent session is that session: the injected
|
|
* `ORCA_AGENT_SESSION_ID` names the caller, and nothing resolves or guesses a terminal for it.
|
|
*
|
|
* One rule for every verb that names a caller: a caller flag may restate the session, but a flag
|
|
* naming anyone else is refused — never silently dropped, never allowed to win. A session address
|
|
* the CLI cannot place is left to the host, whose `/clear` lineage decides.
|
|
* The #21097 accident was a chat that named a sibling's terminal and consumed that sibling's mail.
|
|
*
|
|
* The session env here is the hardest case, a chat that inherited a pane's `ORCA_TERMINAL_HANDLE`
|
|
* and `ORCA_PANE_KEY` (an Orca launched from an Orca terminal), and the implicit-terminal guess has
|
|
* a sibling to find.
|
|
*/
|
|
|
|
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
|
|
|
|
const callMock = vi.hoisted(() => vi.fn())
|
|
const getTerminalHandleMock = vi.hoisted(() => vi.fn())
|
|
|
|
vi.mock('./format', () => ({ printResult: vi.fn() }))
|
|
vi.mock('./selectors', () => ({ getTerminalHandle: getTerminalHandleMock }))
|
|
|
|
import { ORCHESTRATION_HANDLERS } from './handlers/orchestration'
|
|
import { findCommandSpec } from './args'
|
|
import { COMMAND_SPECS } from './specs'
|
|
import { refuseConflictingSessionCallerFlags } from './session-caller-flags'
|
|
import { createOrchestrationCompatibilityEnvelope } from './runtime/orchestration-compatibility-envelope'
|
|
import { getDefaultUserDataPath } from './runtime/metadata'
|
|
import { formatCliError, reportCliError } from './cli-error'
|
|
import { RuntimeRpcFailureError } from './runtime/types'
|
|
|
|
const SESSION = 'f7a1c0de-1111-4222-8333-444455556666'
|
|
/** The session a `/clear` continued as SESSION: the chat's orchestration address stays this one's. */
|
|
const ROOT = '0b5e2d7c-9a41-4c3e-8f62-7d1a3e5b9c08'
|
|
const IDENTITY_ENV = [
|
|
'ORCA_AGENT_SESSION_ID',
|
|
'ORCA_TERMINAL_HANDLE',
|
|
'ORCA_PANE_KEY',
|
|
'ORCA_STRUCTURED_SESSION'
|
|
] as const
|
|
const originalEnv = Object.fromEntries(IDENTITY_ENV.map((name) => [name, process.env[name]]))
|
|
|
|
/** Enough of every receipt shape that each handler finishes after its RPC. */
|
|
const RESULT = {
|
|
result: {
|
|
run: { id: 'run_1', objective: 'o', consumer_generation: 1 },
|
|
runs: [],
|
|
nextCursor: null,
|
|
messages: [],
|
|
count: 0,
|
|
message: { id: 'msg_1' },
|
|
lifecycle: { action: 'completed' },
|
|
dispatch: { id: 'dispatch_1', task_id: 'task_1', status: 'dispatched' },
|
|
gate: { id: 'gate_1', task_id: 'task_1', status: 'pending', resolution: 'r' },
|
|
gates: [],
|
|
task: { id: 'task_1', status: 'pending' },
|
|
tasks: [],
|
|
answer: 'yes',
|
|
messageId: 'msg_1',
|
|
threadId: 'thread_1',
|
|
timedOut: false,
|
|
state: 'ready',
|
|
runId: 'run_1',
|
|
taskId: 'task_1',
|
|
dispatchId: 'dispatch_1',
|
|
effects: [],
|
|
residualResources: [],
|
|
workers: [],
|
|
counts: {}
|
|
}
|
|
}
|
|
|
|
type Verb = {
|
|
command: string
|
|
flags: Record<string, string | true>
|
|
/** The flag that names the caller, when the verb has one. */
|
|
callerFlag?: 'from' | 'terminal'
|
|
method: string
|
|
callerParam: 'from' | 'terminal' | 'callerTerminalHandle'
|
|
}
|
|
|
|
/** Every verb whose request names its caller: the host's caller-param map, from the CLI side.
|
|
* `dispatch-show` is not one: its --from only fills preview text, so it is pinned on its own. */
|
|
const CALLER_VERBS: Verb[] = [
|
|
{
|
|
command: 'run-create',
|
|
flags: { objective: 'o' },
|
|
callerFlag: 'from',
|
|
method: 'runCreate',
|
|
callerParam: 'from'
|
|
},
|
|
{
|
|
command: 'run-use',
|
|
flags: { id: 'run_1' },
|
|
callerFlag: 'from',
|
|
method: 'runUse',
|
|
callerParam: 'from'
|
|
},
|
|
{
|
|
command: 'run-current',
|
|
flags: {},
|
|
callerFlag: 'from',
|
|
method: 'runCurrent',
|
|
callerParam: 'from'
|
|
},
|
|
{ command: 'check', flags: {}, callerFlag: 'terminal', method: 'check', callerParam: 'terminal' },
|
|
{
|
|
command: 'send',
|
|
flags: { to: 'term_worker', subject: 's', body: 'b' },
|
|
callerFlag: 'from',
|
|
method: 'send',
|
|
callerParam: 'from'
|
|
},
|
|
{
|
|
command: 'reply',
|
|
flags: { id: 'msg_1', body: 'b' },
|
|
callerFlag: 'from',
|
|
method: 'reply',
|
|
callerParam: 'from'
|
|
},
|
|
{
|
|
command: 'ask',
|
|
flags: { to: 'term_worker', question: 'q' },
|
|
callerFlag: 'from',
|
|
method: 'ask',
|
|
callerParam: 'from'
|
|
},
|
|
{
|
|
command: 'dispatch',
|
|
flags: { task: 'task_1', to: 'term_worker' },
|
|
callerFlag: 'from',
|
|
method: 'dispatch',
|
|
callerParam: 'from'
|
|
},
|
|
{
|
|
command: 'gate-create',
|
|
flags: { task: 'task_1', question: 'q' },
|
|
callerFlag: 'from',
|
|
method: 'gateCreate',
|
|
callerParam: 'from'
|
|
},
|
|
{
|
|
command: 'gate-resolve',
|
|
flags: { id: 'gate_1', resolution: 'r' },
|
|
callerFlag: 'from',
|
|
method: 'gateResolve',
|
|
callerParam: 'from'
|
|
},
|
|
{ command: 'gate-list', flags: {}, callerFlag: 'from', method: 'gateList', callerParam: 'from' },
|
|
{
|
|
command: 'task-create',
|
|
flags: { spec: 's' },
|
|
callerFlag: 'from',
|
|
method: 'taskCreate',
|
|
callerParam: 'callerTerminalHandle'
|
|
},
|
|
{
|
|
command: 'task-list',
|
|
flags: {},
|
|
callerFlag: 'from',
|
|
method: 'taskList',
|
|
callerParam: 'callerTerminalHandle'
|
|
},
|
|
{
|
|
command: 'task-update',
|
|
flags: { id: 'task_1', status: 'completed' },
|
|
callerFlag: 'from',
|
|
method: 'taskUpdate',
|
|
callerParam: 'callerTerminalHandle'
|
|
},
|
|
{
|
|
command: 'worker-start',
|
|
flags: { spec: 's' },
|
|
callerFlag: 'from',
|
|
method: 'workerStart',
|
|
callerParam: 'from'
|
|
},
|
|
// No caller flag: its spec takes none. It asks runCurrent for the caller's Run.
|
|
{ command: 'worker-list', flags: {}, method: 'runCurrent', callerParam: 'from' }
|
|
]
|
|
|
|
/** Enough flags for any verb to get past its own validation to identity resolution. */
|
|
const EVERY_REQUIRED_FLAG = {
|
|
objective: 'o',
|
|
id: 'id_1',
|
|
task: 'task_1',
|
|
spec: 's',
|
|
question: 'q',
|
|
resolution: 'r',
|
|
subject: 's',
|
|
body: 'b',
|
|
to: 'term_worker',
|
|
status: 'completed',
|
|
preamble: true,
|
|
request: 'req_1'
|
|
} as const
|
|
|
|
function flagMap(flags: Record<string, string | true>): Map<string, string | boolean> {
|
|
return new Map(Object.entries(flags))
|
|
}
|
|
|
|
/** What `main()` does between parsing and dispatch: the spec-driven caller check, then the handler. */
|
|
async function invoke(
|
|
command: string,
|
|
flags: Map<string, string | boolean>,
|
|
json = true
|
|
): Promise<void> {
|
|
const handler = ORCHESTRATION_HANDLERS[`orchestration ${command}`]
|
|
if (!handler) {
|
|
throw new Error(`no handler for ${command}`)
|
|
}
|
|
refuseConflictingSessionCallerFlags(
|
|
findCommandSpec(COMMAND_SPECS, ['orchestration', command]),
|
|
flags
|
|
)
|
|
await handler({
|
|
flags,
|
|
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: these handlers read only `call`; RuntimeClient is a class, so a structural double cannot satisfy it without the cast.
|
|
client: { call: callMock } as never,
|
|
cwd: '/tmp/repo',
|
|
json
|
|
})
|
|
}
|
|
|
|
function isParams(value: unknown): value is Record<string, unknown> {
|
|
return typeof value === 'object' && value !== null && !Array.isArray(value)
|
|
}
|
|
|
|
function callsTo(method: string): Record<string, unknown>[] {
|
|
return callMock.mock.calls
|
|
.filter(([name]) => name === `orchestration.${method}`)
|
|
.map(([, params]) => (isParams(params) ? params : {}))
|
|
}
|
|
|
|
function setEnv(env: Partial<Record<(typeof IDENTITY_ENV)[number], string>>): void {
|
|
for (const name of IDENTITY_ENV) {
|
|
const value = env[name]
|
|
if (value === undefined) {
|
|
delete process.env[name]
|
|
} else {
|
|
process.env[name] = value
|
|
}
|
|
}
|
|
}
|
|
|
|
/** A chat with its id, plus a pane identity it inherited from the Orca that launched it. */
|
|
function asSessionWithInheritedPane(): void {
|
|
setEnv({
|
|
ORCA_AGENT_SESSION_ID: SESSION,
|
|
ORCA_TERMINAL_HANDLE: 'term_inherited_pane',
|
|
ORCA_PANE_KEY: 'tab_inherited:11111111-1111-4111-8111-111111111111'
|
|
})
|
|
}
|
|
|
|
beforeEach(() => {
|
|
callMock.mockReset().mockResolvedValue(RESULT)
|
|
getTerminalHandleMock.mockReset().mockResolvedValue('term_sibling')
|
|
vi.spyOn(console, 'log').mockImplementation(() => {})
|
|
vi.spyOn(console, 'error').mockImplementation(() => {})
|
|
})
|
|
|
|
afterEach(() => {
|
|
vi.restoreAllMocks()
|
|
setEnv(originalEnv)
|
|
process.exitCode = undefined
|
|
})
|
|
|
|
describe.each(CALLER_VERBS)('orchestration $command run as an agent session', (verb) => {
|
|
beforeEach(asSessionWithInheritedPane)
|
|
|
|
it('acts as the session: no terminal is resolved, guessed or sent', async () => {
|
|
await invoke(verb.command, flagMap(verb.flags))
|
|
|
|
const [params] = callsTo(verb.method)
|
|
expect(params, 'the verb reached its method').toBeDefined()
|
|
expect(params?.[verb.callerParam]).toBeUndefined()
|
|
// An inherited pane is not the session's identity.
|
|
expect(params?.terminalPaneKey).toBeUndefined()
|
|
expect(params?.senderPaneKey).toBeUndefined()
|
|
expect(getTerminalHandleMock).not.toHaveBeenCalled()
|
|
expect(callMock.mock.calls.map(([name]) => name)).not.toEqual(
|
|
expect.arrayContaining([expect.stringMatching(/^terminal\./)])
|
|
)
|
|
})
|
|
|
|
it.runIf(verb.callerFlag !== undefined)(
|
|
'refuses a caller flag naming another caller, before any request',
|
|
async () => {
|
|
const flags = flagMap({ ...verb.flags, [verb.callerFlag ?? 'from']: 'term_sibling' })
|
|
|
|
await expect(invoke(verb.command, flags)).rejects.toMatchObject({
|
|
code: 'consumer_fenced',
|
|
message: expect.stringContaining(`agent session ${SESSION}`)
|
|
})
|
|
expect(callMock).not.toHaveBeenCalled()
|
|
expect(getTerminalHandleMock).not.toHaveBeenCalled()
|
|
}
|
|
)
|
|
|
|
it.runIf(verb.callerFlag !== undefined)(
|
|
'refuses an inherited pane handle too: the session, not the pane, is the caller',
|
|
async () => {
|
|
const flags = flagMap({ ...verb.flags, [verb.callerFlag ?? 'from']: 'term_inherited_pane' })
|
|
|
|
await expect(invoke(verb.command, flags)).rejects.toMatchObject({ code: 'consumer_fenced' })
|
|
expect(callMock).not.toHaveBeenCalled()
|
|
}
|
|
)
|
|
|
|
it.runIf(verb.callerFlag !== undefined).each([`session:${SESSION}`, SESSION])(
|
|
'accepts a caller flag that restates the session (%s)',
|
|
async (restated) => {
|
|
await invoke(verb.command, flagMap({ ...verb.flags, [verb.callerFlag ?? 'from']: restated }))
|
|
|
|
const [params] = callsTo(verb.method)
|
|
expect(params, 'the verb reached its method').toBeDefined()
|
|
expect(params?.[verb.callerParam]).toBeUndefined()
|
|
}
|
|
)
|
|
})
|
|
|
|
describe.each(CALLER_VERBS.filter((verb) => verb.callerFlag !== undefined))(
|
|
'orchestration $command run by a /clear-ed chat',
|
|
(verb) => {
|
|
beforeEach(asSessionWithInheritedPane)
|
|
|
|
it('sends the address it had before the clear for the host to place, instead of refusing it', async () => {
|
|
// The chat's address is its lineage root's, which only the host's session records know.
|
|
const flags = flagMap({ ...verb.flags, [verb.callerFlag ?? 'from']: `session:${ROOT}` })
|
|
await invoke(verb.command, flags)
|
|
|
|
const [params] = callsTo(verb.method)
|
|
expect(params?.[verb.callerParam]).toBe(`session:${ROOT}`)
|
|
expect(params?.terminalPaneKey).toBeUndefined()
|
|
expect(params?.senderPaneKey).toBeUndefined()
|
|
expect(getTerminalHandleMock).not.toHaveBeenCalled()
|
|
})
|
|
}
|
|
)
|
|
|
|
describe.each([
|
|
{ command: 'gate-list', method: 'gateList', callerParam: 'from' },
|
|
{ command: 'task-list', method: 'taskList', callerParam: 'callerTerminalHandle' }
|
|
])('orchestration $command --run run as an agent session', ({ command, method, callerParam }) => {
|
|
beforeEach(asSessionWithInheritedPane)
|
|
|
|
it('needs no caller, but refuses a --from naming another caller, before any request', async () => {
|
|
await invoke(command, flagMap({ run: 'run_1' }))
|
|
expect(callsTo(method)[0]).toMatchObject({ run: 'run_1' })
|
|
expect(callsTo(method)[0]?.[callerParam]).toBeUndefined()
|
|
|
|
callMock.mockClear()
|
|
await expect(
|
|
invoke(command, flagMap({ run: 'run_1', from: 'term_sibling' }))
|
|
).rejects.toMatchObject({ code: 'consumer_fenced' })
|
|
expect(callMock).not.toHaveBeenCalled()
|
|
})
|
|
|
|
it('accepts a --from that restates the session', async () => {
|
|
await invoke(command, flagMap({ run: 'run_1', from: `session:${SESSION}` }))
|
|
expect(callsTo(method)[0]?.[callerParam]).toBeUndefined()
|
|
})
|
|
|
|
it.each([`session:${ROOT}`, 'session:9d4c1b2a-3e5f-4a6b-8c7d-0e1f2a3b4c5d'])(
|
|
'sends a session address it cannot place (%s) for the host to check, never dropping it',
|
|
async (declared) => {
|
|
// A /clear-ed chat's root is accepted by the host; anyone else is refused there, before any
|
|
// effect. Dropped here, a conflicting caller would pass unchecked.
|
|
await invoke(command, flagMap({ run: 'run_1', from: declared }))
|
|
expect(callsTo(method)[0]).toMatchObject({ run: 'run_1', [callerParam]: declared })
|
|
}
|
|
)
|
|
})
|
|
|
|
describe('the identity a session presents', () => {
|
|
it("lets a structured worker restate its own minted handle, and nobody else's", async () => {
|
|
setEnv({ ORCA_AGENT_SESSION_ID: SESSION, ORCA_TERMINAL_HANDLE: 'structworker_self' })
|
|
|
|
await invoke('send', flagMap({ from: 'structworker_self', to: 'run:run_1', subject: 's' }))
|
|
expect(callsTo('send')[0]?.from).toBeUndefined()
|
|
|
|
callMock.mockClear()
|
|
await expect(
|
|
invoke('send', flagMap({ from: 'structworker_other', to: 'run:run_1', subject: 's' }))
|
|
).rejects.toMatchObject({ code: 'consumer_fenced' })
|
|
expect(callMock).not.toHaveBeenCalled()
|
|
})
|
|
|
|
it("sends a structured worker's lifecycle report as the session, not refused as identity-less", async () => {
|
|
setEnv({ ORCA_AGENT_SESSION_ID: SESSION })
|
|
|
|
await invoke(
|
|
'send',
|
|
flagMap({ to: 'run:run_1', subject: 'done', type: 'worker_done', outcome: 'succeeded' })
|
|
)
|
|
|
|
expect(callsTo('send')[0]).toMatchObject({ type: 'worker_done' })
|
|
expect(callsTo('send')[0]?.from).toBeUndefined()
|
|
})
|
|
|
|
it('never treats a session that has an id as identity-less, even beside the old marker', async () => {
|
|
setEnv({ ORCA_AGENT_SESSION_ID: SESSION, ORCA_STRUCTURED_SESSION: '1' })
|
|
|
|
await invoke('check', flagMap({}))
|
|
|
|
expect(callsTo('check')).toHaveLength(1)
|
|
expect(getTerminalHandleMock).not.toHaveBeenCalled()
|
|
})
|
|
|
|
it('keeps the identity-less refusal, without --from advice, for a child that has no id', async () => {
|
|
setEnv({ ORCA_STRUCTURED_SESSION: '1' })
|
|
|
|
await expect(invoke('reply', flagMap({ id: 'msg_1', body: 'b' }))).rejects.toMatchObject({
|
|
code: 'no_active_sender_terminal',
|
|
message: expect.not.stringContaining('Pass --from')
|
|
})
|
|
expect(getTerminalHandleMock).not.toHaveBeenCalled()
|
|
expect(callMock).not.toHaveBeenCalled()
|
|
})
|
|
|
|
it('leaves a terminal agent exactly as it was: its own handle is the caller', async () => {
|
|
setEnv({ ORCA_TERMINAL_HANDLE: 'term_pty', ORCA_PANE_KEY: 'tab_pty:1:2' })
|
|
callMock.mockImplementation(async (name: string) =>
|
|
name === 'terminal.resolveIdentity' ? { result: { identity: { live: true } } } : RESULT
|
|
)
|
|
|
|
await invoke('run-create', flagMap({ objective: 'o' }))
|
|
await invoke('check', flagMap({}))
|
|
|
|
expect(callsTo('runCreate')[0]?.from).toBe('term_pty')
|
|
expect(callsTo('check')[0]).toMatchObject({
|
|
terminal: 'term_pty',
|
|
terminalPaneKey: 'tab_pty:1:2'
|
|
})
|
|
})
|
|
|
|
it('previews a dispatch with the coordinator address the real dispatch would write', async () => {
|
|
const preview = async (flags: Record<string, string | true>) => {
|
|
callMock.mockClear()
|
|
await invoke('dispatch-show', flagMap({ task: 'task_1', preamble: true, ...flags }))
|
|
return callsTo('dispatchShow')[0]?.from
|
|
}
|
|
asSessionWithInheritedPane()
|
|
expect(await preview({})).toBe(`session:${SESSION}`)
|
|
// Not a caller flag: it names the text to preview, so it is never fenced.
|
|
expect(await preview({ from: 'term_sibling' })).toBe('term_sibling')
|
|
setEnv({ ORCA_AGENT_SESSION_ID: SESSION, ORCA_TERMINAL_HANDLE: 'structworker_self' })
|
|
expect(await preview({})).toBe('structworker_self')
|
|
expect(getTerminalHandleMock).not.toHaveBeenCalled()
|
|
})
|
|
|
|
it('resumes a timed-out ask as the session, without naming a terminal', async () => {
|
|
asSessionWithInheritedPane()
|
|
callMock.mockResolvedValue({ result: { ...RESULT.result, answer: null, timedOut: true } })
|
|
const errors = vi.mocked(console.error)
|
|
|
|
await invoke('ask', flagMap({ to: 'term_worker', question: 'q' }), false)
|
|
|
|
const advice = errors.mock.calls.map(([line]) => String(line)).join('\n')
|
|
expect(advice).toContain('--resume msg_1')
|
|
expect(advice).not.toContain('--from')
|
|
})
|
|
})
|
|
|
|
describe('a host refusal of the session', () => {
|
|
const WORKER_GONE = new RuntimeRpcFailureError({
|
|
id: 'rpc_1',
|
|
ok: false,
|
|
error: {
|
|
code: 'session_caller_not_live',
|
|
message: `Agent session ${SESSION} is a structured worker whose worker identity this host no longer has, so it cannot act in orchestration. No effects were applied.`,
|
|
data: { effectsApplied: false }
|
|
},
|
|
_meta: { runtimeId: 'runtime_1' }
|
|
})
|
|
|
|
it.each(['check', 'run-current', 'worker-list'])(
|
|
'surfaces from %s verbatim, never widened, retried or turned into a terminal guess',
|
|
async (command) => {
|
|
asSessionWithInheritedPane()
|
|
callMock.mockRejectedValue(WORKER_GONE)
|
|
|
|
await expect(invoke(command, flagMap({}))).rejects.toBe(WORKER_GONE)
|
|
expect(getTerminalHandleMock).not.toHaveBeenCalled()
|
|
expect(formatCliError(WORKER_GONE)).toBe(WORKER_GONE.message)
|
|
}
|
|
)
|
|
|
|
it('keeps the Orca id a provider-id refusal names, for a JSON reader to branch on', () => {
|
|
const providerId = new RuntimeRpcFailureError({
|
|
id: 'rpc_1',
|
|
ok: false,
|
|
error: {
|
|
code: 'session_caller_provider_id',
|
|
message: 'provider id',
|
|
data: { effectsApplied: false, orcaSessionId: SESSION }
|
|
},
|
|
_meta: { runtimeId: 'runtime_1' }
|
|
})
|
|
const printed: string[] = []
|
|
vi.mocked(console.log).mockImplementation((line: string) => {
|
|
printed.push(line)
|
|
})
|
|
|
|
reportCliError(providerId, true)
|
|
|
|
expect(JSON.parse(printed.join('\n'))).toMatchObject({
|
|
ok: false,
|
|
error: { code: 'session_caller_provider_id', data: { orcaSessionId: SESSION } }
|
|
})
|
|
})
|
|
})
|
|
|
|
describe('the orchestration envelope', () => {
|
|
it('carries the injected id beside whatever terminal evidence the process also has', () => {
|
|
const envelope = createOrchestrationCompatibilityEnvelope({
|
|
ORCA_AGENT_SESSION_ID: ` ${SESSION} `,
|
|
ORCA_TERMINAL_HANDLE: 'term_inherited_pane'
|
|
})
|
|
|
|
expect(envelope.orchestrationCompatibilityEvidence).toEqual({
|
|
terminalHandle: 'term_inherited_pane',
|
|
agentSessionId: SESSION
|
|
})
|
|
})
|
|
|
|
it('binds whichever current Orca CLI the agent reached, not only the one the session names', () => {
|
|
// A login shell can put another install's `orca` first, and the packaged Windows launcher
|
|
// rewrites ORCA_CLI_COMMAND in its own process. Neither matters: the id rides the envelope
|
|
// and the pinned instance is the one dialed.
|
|
vi.stubEnv('ORCA_USER_DATA_PATH', '/data/session-orca')
|
|
try {
|
|
const envelope = createOrchestrationCompatibilityEnvelope({
|
|
ORCA_AGENT_SESSION_ID: SESSION,
|
|
ORCA_CLI_COMMAND: 'orca',
|
|
ORCA_WINDOWS_PACKAGED_CLI_LAUNCHER: '1',
|
|
ORCA_USER_DATA_PATH: '/data/session-orca'
|
|
})
|
|
|
|
expect(envelope.orchestrationCompatibilityEvidence).toEqual({ agentSessionId: SESSION })
|
|
expect(getDefaultUserDataPath('linux', '/home/u')).toBe('/data/session-orca')
|
|
} finally {
|
|
vi.unstubAllEnvs()
|
|
}
|
|
})
|
|
|
|
it('claims no session without an injected id', () => {
|
|
expect(
|
|
createOrchestrationCompatibilityEnvelope({ ORCA_AGENT_SESSION_ID: ' ' })
|
|
.orchestrationCompatibilityEvidence
|
|
).toBeUndefined()
|
|
})
|
|
|
|
it('keeps a WSL stamp beside the id, so the host can refuse the cross-host claim', () => {
|
|
const envelope = createOrchestrationCompatibilityEnvelope({
|
|
ORCA_AGENT_SESSION_ID: SESSION,
|
|
ORCA_ORCHESTRATION_COMPATIBILITY_HOST_KIND: 'wsl',
|
|
ORCA_ORCHESTRATION_COMPATIBILITY_HOST_ID: 'local',
|
|
ORCA_ORCHESTRATION_COMPATIBILITY_HOST_INCARNATION: 'Ubuntu'
|
|
})
|
|
|
|
expect(envelope.orchestrationCompatibilityEvidence).toEqual({
|
|
agentSessionId: SESSION,
|
|
host: { kind: 'wsl', hostId: 'local', distro: 'Ubuntu' }
|
|
})
|
|
})
|
|
})
|
|
|
|
describe('which flag names the caller, declared on every spec', () => {
|
|
const ORCHESTRATION_SPECS = COMMAND_SPECS.filter((spec) => spec.path[0] === 'orchestration')
|
|
|
|
it('classifies every --from and --terminal an orchestration verb accepts', () => {
|
|
// A new verb cannot take either flag without saying whether it names the caller, so the entry
|
|
// check covers it by construction instead of each handler remembering to refuse.
|
|
const unclassified = ORCHESTRATION_SPECS.flatMap((spec) =>
|
|
(['from', 'terminal'] as const)
|
|
.filter((flag) => spec.allowedFlags.includes(flag) && !spec.identityFlagRoles?.[flag])
|
|
.map((flag) => `${spec.path.join(' ')} --${flag}`)
|
|
)
|
|
expect(unclassified).toEqual([])
|
|
})
|
|
|
|
const callerFlagVerbs = ORCHESTRATION_SPECS.flatMap((spec) =>
|
|
(['from', 'terminal'] as const)
|
|
.filter((flag) => spec.identityFlagRoles?.[flag] === 'caller')
|
|
.map((flag) => ({ command: spec.path[1] ?? '', flag }))
|
|
)
|
|
|
|
it('covers the verbs whose requests name a caller', () => {
|
|
expect(callerFlagVerbs.length).toBeGreaterThanOrEqual(CALLER_VERBS.length)
|
|
})
|
|
|
|
it.each(callerFlagVerbs)(
|
|
'$command refuses --$flag naming another caller, before any request',
|
|
async ({ command, flag }) => {
|
|
asSessionWithInheritedPane()
|
|
await expect(
|
|
invoke(command, flagMap({ ...EVERY_REQUIRED_FLAG, [flag]: 'term_sibling' }))
|
|
).rejects.toMatchObject({ code: 'consumer_fenced' })
|
|
expect(callMock).not.toHaveBeenCalled()
|
|
expect(getTerminalHandleMock).not.toHaveBeenCalled()
|
|
}
|
|
)
|
|
|
|
it.each(
|
|
ORCHESTRATION_SPECS.flatMap((spec) =>
|
|
(['from', 'terminal'] as const)
|
|
.filter((flag) => spec.identityFlagRoles?.[flag] === 'target')
|
|
.map((flag) => ({ command: spec.path[1] ?? '', flag }))
|
|
)
|
|
)('$command passes a --$flag target through unfenced', async ({ command, flag }) => {
|
|
asSessionWithInheritedPane()
|
|
await invoke(command, flagMap({ ...EVERY_REQUIRED_FLAG, [flag]: 'term_sibling' })).catch(
|
|
(error: unknown) => {
|
|
expect(error).not.toMatchObject({ code: 'consumer_fenced' })
|
|
}
|
|
)
|
|
expect(callMock.mock.calls.flatMap(([, params]) => Object.values(params ?? {}))).toContain(
|
|
'term_sibling'
|
|
)
|
|
})
|
|
})
|
|
|
|
describe('every orchestration verb, enumerated', () => {
|
|
/** Runs every verb once; returns the ones that guessed an implicit terminal. */
|
|
async function verbsThatGuess(): Promise<string[]> {
|
|
const guessed: string[] = []
|
|
for (const command of Object.keys(ORCHESTRATION_HANDLERS)) {
|
|
const verb = command.replace('orchestration ', '')
|
|
getTerminalHandleMock.mockClear()
|
|
callMock.mockClear()
|
|
await invoke(verb, flagMap(EVERY_REQUIRED_FLAG)).catch(() => undefined)
|
|
if (getTerminalHandleMock.mock.calls.length > 0) {
|
|
guessed.push(verb)
|
|
}
|
|
}
|
|
return guessed.sort()
|
|
}
|
|
|
|
it('guesses a terminal for no verb when the session id is present', async () => {
|
|
// Positive control first: with no identity at all the same harness sees the guess, so an empty
|
|
// result below is about the id, not a harness that cannot observe a guess.
|
|
setEnv({})
|
|
const population = await verbsThatGuess()
|
|
expect(population).toEqual([
|
|
'ask',
|
|
'check',
|
|
'dispatch',
|
|
'dispatch-show',
|
|
'gate-create',
|
|
'gate-list',
|
|
'gate-resolve',
|
|
'reply',
|
|
'run-create',
|
|
'run-current',
|
|
'run-use',
|
|
'send',
|
|
'task-create',
|
|
'task-list',
|
|
'task-update',
|
|
'worker-list',
|
|
'worker-start'
|
|
])
|
|
|
|
asSessionWithInheritedPane()
|
|
expect(await verbsThatGuess()).toEqual([])
|
|
})
|
|
})
|