Files
orca/tests/tools
Brennan Benson 6e7e964705 feat(orchestration): tell each agent its own orchestration address (#22636)
* feat(orchestration): report the caller's host-resolved orchestration address in orca status

orca status --json gains a caller block: the calling agent's address as the
host resolved it from the identity its environment carries. A structured
session is session:<id>; a terminal agent is its handle, with whether the host
still knows it. A session the host refuses reports that refusal instead.

The host answers through a new read-only orchestration.callerShow, so the
session claim runs through the same dispatch-entry resolver every verb uses.
An older host leaves caller unresolved. The help footer and the run/check
specs stop describing identity only in terminal terms.

* docs(orchestration): tell agents their address and give chat coordinators a non-waiting loop

The orchestration guide now states that a chat session's address is
session:<id> (never the provider's id), that orca status --json reports it,
and that no caller flag should name another agent. A consuming check no
longer tells every caller to name itself with --terminal. A chat coordinator
starts its wave, ends the turn, and on each turn Orca starts for new mail
runs a non-waiting check and ack; it never blocks in check --wait. The guide
also names ORCA_CLI_COMMAND as the executable in chat sessions.

* feat(native-chat): add Copy Orchestration Address to a structured chat's context menu

Copies session:<id>, the Orca-minted address other agents message the chat
by. The existing Copy Session ID still copies the provider's id and is left
as is; the new action is labelled so the two cannot be confused. Strings are
added to every locale catalog.

* feat(orchestration): tell every dispatched worker its own orchestration address

The worker preamble names the coordinator's address rather than a terminal
handle, and states the worker's own address. A structured worker is told it
is session:<id>, that its coordinator reaches it there or at its dispatch
mailbox, and that mail arriving while it is idle starts a new turn. Its
commands invoke the CLI through ORCA_CLI_COMMAND in its own shell's form, the
same rendering the pointer turn uses, because a bare orca in a login shell can
reach a different Orca.

* docs(orchestration): give the ORCA_CLI_COMMAND form for POSIX shells and PowerShell

A chat session's shell reads the variable as "$ORCA_CLI_COMMAND" in a POSIX
shell (Git Bash included) and as & $env:ORCA_CLI_COMMAND in PowerShell, the
same two forms the pointer turn and worker preamble render. The chat
coordinator loop now runs the check its pointer turn names.

* docs(orchestration): say that /clear gives a chat a new address and Orca moves its Runs

* fix(orchestration): keep CLI resolution in the shared skill stub and the orchestration kernel in budget

The guide-contract tests own two rules this PR broke: only the shared skill
stub may describe how to resolve the CLI, and the always-loaded orchestration
kernel stays within 202 lines. The ORCA_CLI_COMMAND text moves to the stub's
resolver block, which now covers chat sessions and login shells beside WSL and
gives the POSIX and PowerShell forms; every skill projection and the bundle
manifest are regenerated. The kernel keeps one line each for the caller's
address, the environment-resolved check caller and the chat coordinator's
non-waiting loop; the loop steps and the address details move to the
coordinator-loop and messaging references. The two kernel pins now assert the
new check contract and refuse the old --terminal <your_handle> shape.

* fix(orchestration): refuse a blocking check --wait from a native chat session

A chat runs turn by turn through a shell tool with its own timeout, so a
blocking wait is killed mid-wait and retried. The host now refuses it with
wait_requires_terminal and the turn-loop recovery, keyed on the session's
lease: a session a terminal view holds still runs in a PTY and may block.

* fix(orchestration): resolve orca status's caller with the verbs' ladder, host-side

callerShow now answers a terminal caller the way the coordinator verbs act:
the carried handle while it is live, else the handle its pane was reminted
as. The CLI always asks, so the host decides that a process has no identity
from the same envelope every verb sends; a pane key alone now resolves.

* fix(orchestration): show a structured worker as session:<id> wherever agents read mail

A structured worker was session:<id> in orca status and its preamble, but
structworker_<uuid> in check rows, banners, reply hints, its own check label
and a sub-worker's coordinator line. The minted handle is now only the
mailbox key: mailbox reads, the check label and preamble coordinator lines
spell the worker session:<id>, which the host binds back to that mailbox.
Send receipts still echo the stored row, whose sender key worker_done
settlement matches.

* fix(orchestration): teach a chat worker the turn loop and pin preamble parity at the contract

The worker preamble was byte-identical across modes except its address, so a
chat worker was taught a 600s blocking ask its shell tool kills before the
message ID for --resume prints, heartbeat exemptions for check --wait, and to
keep a shell open. Parity now pins the contract (sections, verbs, flags,
lifecycle ids); interaction discipline follows the mode: a chat asks with a
5s wait and ends its turn, owns sub-workers through the turn loop, and names
itself session:<id> in every command. The guide says a chat's address
survives /clear and that Orca refuses a chat's check --wait.

* test(orchestration): pass the db to preamble delivery and fence a terminal-view waiter

The coordinator line maps a structured coordinator's handle through the
orchestration db, so delivery takes it from its caller. The consumer-fencing
waiter test now waits as a terminal-view session, the only session kind that
may still block in check --wait.

* test(orchestration): read the Run id with the fixture's checked accessor

* feat(orchestration): copy a chat's conversation address, which /clear keeps

Copy Orchestration Address copied session:<live id>. A chat's address is its
conversation's, derived by the host from the session records, so the menu now
asks the host for it at copy time through orchestration.sessionAddress, the
same derivation a verb acting as that session binds to. A host that predates
the method has no /clear lineage, so there the live id is the address. The
guide's /clear text says the address survives and nothing moves.

* test(orchestration): pin that a cleared chat's successor copies its conversation's root address

* test(orchestration): give the mode-opacity fixture's record store the listing a lineage lookup reads

A structured worker's agent-visible address now resolves through its conversation's lineage,
which lists the session records; the fixture's partial store lacked that listing, so the
sub-worker start failed at dispatch input.

* refactor(orchestration): format a chat's copied and reported address from its root Orca session id

Carries the Orca session id rename into the self-address surfaces.
orchestration.sessionAddress, the copy action's fallback, callerShow and the address a
structured worker is shown now format `session:<id>` from the conversation's bare
root Orca session id with formatOrcaSessionAddress, and ids arriving as strings are
checked with isOrcaSessionId first. The CLI status line, the check caller label and
the dispatch preamble spell the prefix from the one exported constant.

* refactor(orchestration): resolve a session's reported address through the party resolver, and refuse every session's check --wait

- orchestration.sessionAddress, and the agent-visible spelling of a structured
  worker, resolve through the party resolver, so they format the lineage root
  the one id hook derives; sessionAddress.sessionId is classified as a target.
- With the terminal handoff gone every structured session runs turn by turn, so
  check --wait is refused for any session caller, a worker included, and the
  session caller no longer carries its lease's runtime kind.
- The coordinator loop no longer mentions a terminal view, and the messaging
  reference says a chat takes messages but is refused as a Dispatch assignee.

* fix(orchestration): cap a session caller's blocking wait below its shell tool instead of refusing it

A chat or structured worker runs each command under its provider's shell-tool timeout, so check
--wait was refused for every session caller and chats were taught a separate loop. The host now
caps check --wait and ask for a session caller below that timeout (Codex 10s one-shot exec
default, Claude Code Bash 120s) and answers the normal timed-out result, so the terminal
coordinator loop runs unchanged in a chat. A terminal caller's wait is untouched.

* refactor(orchestration): teach a chat worker the terminal worker's preamble, byte for byte but the address

One preamble for both modes: the chat variant (short ask, end your turn, this chat stays
available, ORCA_CLI_COMMAND invocation) is deleted. A structured worker's only difference is its
address, session:<id>; the byte-parity test between modes is restored with just that substituted.

* docs(orchestration): drop every chat-specific instruction; name the address once, generically

The guide, its references, the shared CLI-resolution stub and the help return to main's text,
with one kernel line saying `orca status --json` shows your address (the kernel stays at main's
length). The status caller block reports only the opaque address, the same shape for a chat and
a terminal agent. Guides regenerated.

* test(orchestration): pin that a chat and a terminal agent see the same preamble, pointer and guide

* test(orchestration): key the wait-cap fixture's records by plain session id strings

* chore(i18n): add the copy-address strings at the head of native-chat, clear of main's catalog edits

* test(orchestration): fail the capped-wait test on the settle, not on the test timeout

* test(orchestration): keep main's takeover assertions on a session coordinator's waiting check

With the session wait capped rather than refused, the test main extended runs as it is: the restack re-added the shorter pre-main version over it.

* fix(orchestration): show a /clear-ed chat its lineage root's address everywhere it reads its own

check labelled a session caller with session:<live id>, while orca status and the
preamble show the conversation's root. The CLI cannot read the lineage, so the label
now comes from the same host answer orca status prints (orchestration.callerShow),
asked only when there are messages to render, and falling back to the live id only
when the host cannot say. The host also spells a session address it shows an agent
with the lineage root: a dispatch preview filled in from the chat's own address, and
the provider-id refusal that names a session's address.

* test(orchestration): the parity test's gate facts resolve like the host's

* test(orchestration): the parity test's gate facts carry the submissions main's pointer lane reads

* fix(orchestration): wait a chat's check --wait and ask exactly as long as a terminal's

The host capped a session caller's blocking wait (Codex 6s, Claude 100s) so the
provider's shell tool would not kill it. Neither provider kills a long shell
call: default Codex's exec tool yields and keeps the command running, and Claude
Code moves a timed-out Bash call to the background. Terminal agents run the same
tools uncapped, so the cap only made a chat coordinator re-poll every few
seconds. A session caller's check --wait and ask now wait the budget asked for.

* fix(orchestration): show every agent one address, the mailbox address its mail is keyed by

A structured worker was told `session:<id>` in orca status and its preamble,
but its own send receipts, inbox, worker-list, dispatch previews and task rows
still showed the `structworker_` handle its mail is stored under; only some
reads were re-spelled. Instead of re-spelling reads, callerShow,
sessionAddress and the preamble now report the caller's stored mailbox
address (mailboxAddressOf): a terminal's handle, a structured worker's handle,
and a chat's `session:<lineage root>`. The read-side re-spelling layer
(withAgentVisibleAddresses and its check/banner/preamble call sites) is gone.

Dispatch previews spell the coordinator by its party's mailbox address, so a
`/clear`ed chat's dispatch-show still names its root.

* refactor(orchestration): label check output from what the CLI already knows

check asked the host for orchestration.callerShow after every non-empty check
by a session, only to fill a label used when a legacy row lacks to_handle,
which host rows never do. The label is again the caller's handle or its
injected mailbox address, with no second round trip after mail is consumed.

* refactor(native-chat): offer Copy Orchestration Address on chat tabs only

No mount passes both terminal-pane actions and an orchestration address: a
chat shown inside a terminal pane is that terminal's agent, copied by its
terminal ID. Drop the unreachable terminal-pane placement and its tests.

* fix(orchestration): have orca status report the handle the agent's own check reads

After a window reload a terminal agent keeps ORCA_TERMINAL_HANDLE=term_old while
its pane is reminted as term_new. callerShow reminted and advertised term_new,
but check, send and ask act as the carried handle and never remint, so mail
sent to the advertised address was never read by that agent. callerShow now
answers the carried handle with its liveness, and null for a pane key alone,
from which the mailbox verbs have no identity. Resolving terminal callers once
on the host for every verb is a separate follow-up.

* fix(orchestration): read the renamed coordinator line in the long-prompt repro, and trim round-one leftovers

The reliability repro's fake worker parsed "Your coordinator's terminal handle
is:", which the preamble now spells "Your coordinator's address is:", so it
silently skipped worker_done; it accepts both. Dispatch and its dry-run go back
to main's coordinator line (their `from` is already bound at the entry); only
dispatch-show, whose `from` is unbound, resolves it. Also drops a stale
status-caller comment, trims the wait test to its one uncapped-wait case, and
reverts comment-only churn in the worker opacity test.

* docs(orchestration): keep worker obligation 1 as main words it

The guide grows by the one caller.address line; the parity test bounds the
kernel at main's length plus that line instead of forcing a reword.

* test(native-chat): prove a structured chat tab offers Copy Orchestration Address

Renders the pane-commands hook as a structured chat tab and selects the item:
it asks orchestration.sessionAddress with the tab's target and session id.
Also corrects the menu item's comment to what it copies.

* test(orchestration): D5's tests expect the orca_session_id prefix and 'Orca session ID' wording

* fix(orchestration): name a session by its Orca session ID, and leave terminal agents as main has them

Terminal agents keep main's exact wording: a terminal worker's preamble is
byte-identical to main's, and orca status prints nothing new for them. A
session is named by its Orca session ID (orca_session_id:<id>, its /clear
root's): a structured worker's preamble says "Your Orca session ID is: …"
and its commands use that ID, and a session coordinator is "Your
coordinator's Orca session ID is: …". orca status shows a session caller's
`caller.orcaSessionId`; callerShow answers null for anyone else. The chat
tab menu item becomes "Copy Orca Session ID" with a tooltip saying what the
ID is, and its toasts match. No agent-read text calls this ID an address.
A structured worker's mail is still keyed by its minted handle.

* test(orchestration): check CLI help and status for "address" wording from a CLI test

The node project cannot compile src/cli, so the guard over CLI help, specs
and status text moves to src/cli; both halves share one pattern. Also brings
two comments and the long-prompt repro's coordinator-line regex to the Orca
session ID wording.

* fix(native-chat): keep the Orca session ID tooltip within the tooltip primitive's typography

Drops a restyle the design-system gate refuses on TooltipContent, keeps
"Agent" untranslated in the Japanese tooltip as that catalog does, and types
the test's tooltip mock without an assertion.

* fix(native-chat): the Orca session ID tooltip names the agent CLI's own session ID in the singular
2026-10-01 17:29:04 -07:00
..