mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 08:02:32 +00:00
0f5fd451def5278c5a95117c3a9202e2ddeb941d
66
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6e7e964705 |
feat(orchestration): tell each agent its own orchestration address (#22636)
* feat(orchestration): report the caller's host-resolved orchestration address in orca status orca status --json gains a caller block: the calling agent's address as the host resolved it from the identity its environment carries. A structured session is session:<id>; a terminal agent is its handle, with whether the host still knows it. A session the host refuses reports that refusal instead. The host answers through a new read-only orchestration.callerShow, so the session claim runs through the same dispatch-entry resolver every verb uses. An older host leaves caller unresolved. The help footer and the run/check specs stop describing identity only in terminal terms. * docs(orchestration): tell agents their address and give chat coordinators a non-waiting loop The orchestration guide now states that a chat session's address is session:<id> (never the provider's id), that orca status --json reports it, and that no caller flag should name another agent. A consuming check no longer tells every caller to name itself with --terminal. A chat coordinator starts its wave, ends the turn, and on each turn Orca starts for new mail runs a non-waiting check and ack; it never blocks in check --wait. The guide also names ORCA_CLI_COMMAND as the executable in chat sessions. * feat(native-chat): add Copy Orchestration Address to a structured chat's context menu Copies session:<id>, the Orca-minted address other agents message the chat by. The existing Copy Session ID still copies the provider's id and is left as is; the new action is labelled so the two cannot be confused. Strings are added to every locale catalog. * feat(orchestration): tell every dispatched worker its own orchestration address The worker preamble names the coordinator's address rather than a terminal handle, and states the worker's own address. A structured worker is told it is session:<id>, that its coordinator reaches it there or at its dispatch mailbox, and that mail arriving while it is idle starts a new turn. Its commands invoke the CLI through ORCA_CLI_COMMAND in its own shell's form, the same rendering the pointer turn uses, because a bare orca in a login shell can reach a different Orca. * docs(orchestration): give the ORCA_CLI_COMMAND form for POSIX shells and PowerShell A chat session's shell reads the variable as "$ORCA_CLI_COMMAND" in a POSIX shell (Git Bash included) and as & $env:ORCA_CLI_COMMAND in PowerShell, the same two forms the pointer turn and worker preamble render. The chat coordinator loop now runs the check its pointer turn names. * docs(orchestration): say that /clear gives a chat a new address and Orca moves its Runs * fix(orchestration): keep CLI resolution in the shared skill stub and the orchestration kernel in budget The guide-contract tests own two rules this PR broke: only the shared skill stub may describe how to resolve the CLI, and the always-loaded orchestration kernel stays within 202 lines. The ORCA_CLI_COMMAND text moves to the stub's resolver block, which now covers chat sessions and login shells beside WSL and gives the POSIX and PowerShell forms; every skill projection and the bundle manifest are regenerated. The kernel keeps one line each for the caller's address, the environment-resolved check caller and the chat coordinator's non-waiting loop; the loop steps and the address details move to the coordinator-loop and messaging references. The two kernel pins now assert the new check contract and refuse the old --terminal <your_handle> shape. * fix(orchestration): refuse a blocking check --wait from a native chat session A chat runs turn by turn through a shell tool with its own timeout, so a blocking wait is killed mid-wait and retried. The host now refuses it with wait_requires_terminal and the turn-loop recovery, keyed on the session's lease: a session a terminal view holds still runs in a PTY and may block. * fix(orchestration): resolve orca status's caller with the verbs' ladder, host-side callerShow now answers a terminal caller the way the coordinator verbs act: the carried handle while it is live, else the handle its pane was reminted as. The CLI always asks, so the host decides that a process has no identity from the same envelope every verb sends; a pane key alone now resolves. * fix(orchestration): show a structured worker as session:<id> wherever agents read mail A structured worker was session:<id> in orca status and its preamble, but structworker_<uuid> in check rows, banners, reply hints, its own check label and a sub-worker's coordinator line. The minted handle is now only the mailbox key: mailbox reads, the check label and preamble coordinator lines spell the worker session:<id>, which the host binds back to that mailbox. Send receipts still echo the stored row, whose sender key worker_done settlement matches. * fix(orchestration): teach a chat worker the turn loop and pin preamble parity at the contract The worker preamble was byte-identical across modes except its address, so a chat worker was taught a 600s blocking ask its shell tool kills before the message ID for --resume prints, heartbeat exemptions for check --wait, and to keep a shell open. Parity now pins the contract (sections, verbs, flags, lifecycle ids); interaction discipline follows the mode: a chat asks with a 5s wait and ends its turn, owns sub-workers through the turn loop, and names itself session:<id> in every command. The guide says a chat's address survives /clear and that Orca refuses a chat's check --wait. * test(orchestration): pass the db to preamble delivery and fence a terminal-view waiter The coordinator line maps a structured coordinator's handle through the orchestration db, so delivery takes it from its caller. The consumer-fencing waiter test now waits as a terminal-view session, the only session kind that may still block in check --wait. * test(orchestration): read the Run id with the fixture's checked accessor * feat(orchestration): copy a chat's conversation address, which /clear keeps Copy Orchestration Address copied session:<live id>. A chat's address is its conversation's, derived by the host from the session records, so the menu now asks the host for it at copy time through orchestration.sessionAddress, the same derivation a verb acting as that session binds to. A host that predates the method has no /clear lineage, so there the live id is the address. The guide's /clear text says the address survives and nothing moves. * test(orchestration): pin that a cleared chat's successor copies its conversation's root address * test(orchestration): give the mode-opacity fixture's record store the listing a lineage lookup reads A structured worker's agent-visible address now resolves through its conversation's lineage, which lists the session records; the fixture's partial store lacked that listing, so the sub-worker start failed at dispatch input. * refactor(orchestration): format a chat's copied and reported address from its root Orca session id Carries the Orca session id rename into the self-address surfaces. orchestration.sessionAddress, the copy action's fallback, callerShow and the address a structured worker is shown now format `session:<id>` from the conversation's bare root Orca session id with formatOrcaSessionAddress, and ids arriving as strings are checked with isOrcaSessionId first. The CLI status line, the check caller label and the dispatch preamble spell the prefix from the one exported constant. * refactor(orchestration): resolve a session's reported address through the party resolver, and refuse every session's check --wait - orchestration.sessionAddress, and the agent-visible spelling of a structured worker, resolve through the party resolver, so they format the lineage root the one id hook derives; sessionAddress.sessionId is classified as a target. - With the terminal handoff gone every structured session runs turn by turn, so check --wait is refused for any session caller, a worker included, and the session caller no longer carries its lease's runtime kind. - The coordinator loop no longer mentions a terminal view, and the messaging reference says a chat takes messages but is refused as a Dispatch assignee. * fix(orchestration): cap a session caller's blocking wait below its shell tool instead of refusing it A chat or structured worker runs each command under its provider's shell-tool timeout, so check --wait was refused for every session caller and chats were taught a separate loop. The host now caps check --wait and ask for a session caller below that timeout (Codex 10s one-shot exec default, Claude Code Bash 120s) and answers the normal timed-out result, so the terminal coordinator loop runs unchanged in a chat. A terminal caller's wait is untouched. * refactor(orchestration): teach a chat worker the terminal worker's preamble, byte for byte but the address One preamble for both modes: the chat variant (short ask, end your turn, this chat stays available, ORCA_CLI_COMMAND invocation) is deleted. A structured worker's only difference is its address, session:<id>; the byte-parity test between modes is restored with just that substituted. * docs(orchestration): drop every chat-specific instruction; name the address once, generically The guide, its references, the shared CLI-resolution stub and the help return to main's text, with one kernel line saying `orca status --json` shows your address (the kernel stays at main's length). The status caller block reports only the opaque address, the same shape for a chat and a terminal agent. Guides regenerated. * test(orchestration): pin that a chat and a terminal agent see the same preamble, pointer and guide * test(orchestration): key the wait-cap fixture's records by plain session id strings * chore(i18n): add the copy-address strings at the head of native-chat, clear of main's catalog edits * test(orchestration): fail the capped-wait test on the settle, not on the test timeout * test(orchestration): keep main's takeover assertions on a session coordinator's waiting check With the session wait capped rather than refused, the test main extended runs as it is: the restack re-added the shorter pre-main version over it. * fix(orchestration): show a /clear-ed chat its lineage root's address everywhere it reads its own check labelled a session caller with session:<live id>, while orca status and the preamble show the conversation's root. The CLI cannot read the lineage, so the label now comes from the same host answer orca status prints (orchestration.callerShow), asked only when there are messages to render, and falling back to the live id only when the host cannot say. The host also spells a session address it shows an agent with the lineage root: a dispatch preview filled in from the chat's own address, and the provider-id refusal that names a session's address. * test(orchestration): the parity test's gate facts resolve like the host's * test(orchestration): the parity test's gate facts carry the submissions main's pointer lane reads * fix(orchestration): wait a chat's check --wait and ask exactly as long as a terminal's The host capped a session caller's blocking wait (Codex 6s, Claude 100s) so the provider's shell tool would not kill it. Neither provider kills a long shell call: default Codex's exec tool yields and keeps the command running, and Claude Code moves a timed-out Bash call to the background. Terminal agents run the same tools uncapped, so the cap only made a chat coordinator re-poll every few seconds. A session caller's check --wait and ask now wait the budget asked for. * fix(orchestration): show every agent one address, the mailbox address its mail is keyed by A structured worker was told `session:<id>` in orca status and its preamble, but its own send receipts, inbox, worker-list, dispatch previews and task rows still showed the `structworker_` handle its mail is stored under; only some reads were re-spelled. Instead of re-spelling reads, callerShow, sessionAddress and the preamble now report the caller's stored mailbox address (mailboxAddressOf): a terminal's handle, a structured worker's handle, and a chat's `session:<lineage root>`. The read-side re-spelling layer (withAgentVisibleAddresses and its check/banner/preamble call sites) is gone. Dispatch previews spell the coordinator by its party's mailbox address, so a `/clear`ed chat's dispatch-show still names its root. * refactor(orchestration): label check output from what the CLI already knows check asked the host for orchestration.callerShow after every non-empty check by a session, only to fill a label used when a legacy row lacks to_handle, which host rows never do. The label is again the caller's handle or its injected mailbox address, with no second round trip after mail is consumed. * refactor(native-chat): offer Copy Orchestration Address on chat tabs only No mount passes both terminal-pane actions and an orchestration address: a chat shown inside a terminal pane is that terminal's agent, copied by its terminal ID. Drop the unreachable terminal-pane placement and its tests. * fix(orchestration): have orca status report the handle the agent's own check reads After a window reload a terminal agent keeps ORCA_TERMINAL_HANDLE=term_old while its pane is reminted as term_new. callerShow reminted and advertised term_new, but check, send and ask act as the carried handle and never remint, so mail sent to the advertised address was never read by that agent. callerShow now answers the carried handle with its liveness, and null for a pane key alone, from which the mailbox verbs have no identity. Resolving terminal callers once on the host for every verb is a separate follow-up. * fix(orchestration): read the renamed coordinator line in the long-prompt repro, and trim round-one leftovers The reliability repro's fake worker parsed "Your coordinator's terminal handle is:", which the preamble now spells "Your coordinator's address is:", so it silently skipped worker_done; it accepts both. Dispatch and its dry-run go back to main's coordinator line (their `from` is already bound at the entry); only dispatch-show, whose `from` is unbound, resolves it. Also drops a stale status-caller comment, trims the wait test to its one uncapped-wait case, and reverts comment-only churn in the worker opacity test. * docs(orchestration): keep worker obligation 1 as main words it The guide grows by the one caller.address line; the parity test bounds the kernel at main's length plus that line instead of forcing a reword. * test(native-chat): prove a structured chat tab offers Copy Orchestration Address Renders the pane-commands hook as a structured chat tab and selects the item: it asks orchestration.sessionAddress with the tab's target and session id. Also corrects the menu item's comment to what it copies. * test(orchestration): D5's tests expect the orca_session_id prefix and 'Orca session ID' wording * fix(orchestration): name a session by its Orca session ID, and leave terminal agents as main has them Terminal agents keep main's exact wording: a terminal worker's preamble is byte-identical to main's, and orca status prints nothing new for them. A session is named by its Orca session ID (orca_session_id:<id>, its /clear root's): a structured worker's preamble says "Your Orca session ID is: …" and its commands use that ID, and a session coordinator is "Your coordinator's Orca session ID is: …". orca status shows a session caller's `caller.orcaSessionId`; callerShow answers null for anyone else. The chat tab menu item becomes "Copy Orca Session ID" with a tooltip saying what the ID is, and its toasts match. No agent-read text calls this ID an address. A structured worker's mail is still keyed by its minted handle. * test(orchestration): check CLI help and status for "address" wording from a CLI test The node project cannot compile src/cli, so the guard over CLI help, specs and status text moves to src/cli; both halves share one pattern. Also brings two comments and the long-prompt repro's coordinator-line regex to the Orca session ID wording. * fix(native-chat): keep the Orca session ID tooltip within the tooltip primitive's typography Drops a restyle the design-system gate refuses on TooltipContent, keeps "Agent" untranslated in the Japanese tooltip as that catalog does, and types the test's tooltip mock without an assertion. * fix(native-chat): the Orca session ID tooltip names the agent CLI's own session ID in the singular |
||
|
|
0b79720c2e |
feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip
The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.
* feat(native-chat): the chat strip reads the host's child records with its parent's verdict
The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.
Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.
* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open
- The view decoder ignores unknown keys, degrades unknown kinds, states,
outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
roster of finished children and never the views themselves; a stop-only
reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.
* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered
* test(native-chat): type the switch tests' mocks instead of asserting them
* test: remote clients advertise reading child views
* docs(agent-status): the structured row folds the store's child records
* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary
The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.
* refactor(native-chat): the status summary's broadcast equality gets its own module
The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.
* fix(native-chat): command admission reads the strip's child records
A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.
Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.
* refactor(native-chat): command admission takes only what it reads of a turn
* fix(native-chat): the session list drops a session's children when the store does
A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.
The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.
* test(native-chat): write the Codex frame script's parent row out step by step
Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.
* fix(native-chat): the idle sweep and the restart snapshot read the host's child records
The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.
The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.
* test(native-chat): the child-record tests follow the merged command lifecycle
A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.
Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.
* refactor(native-chat): the status feed's journal projection cache gets its own module
The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.
* test(native-chat): the admission test's compaction resolves with a real outcome
Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.
* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished
The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.
This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.
* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source
`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.
A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.
* test(native-chat): the switch test passes the startup child key main's status bar takes
* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own
Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.
Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.
* fix(native-chat): a background Stop reaches the tasks the child records show
The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.
The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.
* fix(native-chat): one rule for a finished child that still owns live work, at any depth
The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.
* fix(native-chat): an older client sees a Codex child's shell as it did before views
Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.
* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives
The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.
* fix(native-chat): the strip channel forgets a closed conversation's roster
It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.
* docs(native-chat): rewrap the retention comment
* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays
Two lifecycle gaps from the round-1 fixes.
A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.
A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.
Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.
* fix(native-chat): the strip keeps one empty list for a roster that omits one
A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.
* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent
The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.
* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once
A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.
The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.
Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.
* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader
CI on
|
||
|
|
46d6b76ae8 |
fix(orchestration): retry worker_done while the Orca runtime is briefly unreachable (#23984)
* fix(orchestration): retry worker_done while the Orca runtime is briefly unreachable A worker reports worker_done once and ends its turn, so a few-minute app outage silently stranded finished work at dispatched. The CLI now retries worker_done on runtime_unavailable for about two minutes with backoff, reusing one request id so the host's mutation ledger replays rather than double-applies it, then prints the existing recovery command. The contract probe no longer caches a failed status.get, which otherwise made every retry fail without reaching the app. Part of STA-8833. * refactor(cli): pass the worker_done retry window in the mutation options bag * Revert "refactor(cli): pass the worker_done retry window in the mutation options bag" The options bag is forwarded to client.call as-is; the retry window is not a client.call option, and folding it in needed a value scan to keep the no-options call shape. * test(cli): fold the explicit retry-request case and drop a vacuous timing assert * fix(cli): keep worker_done recovery when the last retry fails before sending Also skip the Unix-socket retry test on Windows and remove its temp profile. |
||
|
|
2ca4ecbc61 |
feat(orchestration): let a structured chat run orchestration as itself (#22568)
* feat(orchestration): inject the Orca session id into structured children and let the CLI act as it Every structured session's child (native Claude, native Codex, and the terminal view) carries ORCA_AGENT_SESSION_ID and reaches the Orca CLI. The CLI sends the id in the orchestration envelope; when present it is the caller, and a caller flag naming anyone else is refused before any request. The id is stripped from inherited PTY env and from the SSH host-CLI passthrough, and crosses into WSL so the host can refuse the cross-host claim. * test(orchestration): pin session id injection for native Claude, native Codex, the terminal view, WSL, PTY inheritance and SSH * test(orchestration): pin one caller precedence rule across every CLI verb that names its caller Adds the per-verb table (flagless acts as the session; a conflicting --from or --terminal is refused before any request; the session's own spellings are accepted), the enumerated guess population with its positive control, the structured worker's own handle, the identity-less refusal for an older child, the unchanged terminal agent, and the envelope. dispatch-show's --from only fills preview text, so it passes through unfenced and a session's flagless preview names the address the real dispatch writes. * refactor(orchestration): keep the identity-less marker reader to the marker; the id is checked first * test(orchestration): pin that a host refusal of the session surfaces verbatim from the CLI * fix(orchestration): keep the identity-less marker beside the id for CLIs that predate it A CLI older than the id, reached through a global install when a shell rc resets PATH, would otherwise guess a sibling's terminal in a chat that no longer carries the marker. It refuses on the marker instead; a current CLI checks the id first, so the marker never makes a session with an id identity-less. * fix(orchestration): refuse a conflicting --from on gate-list and task-list scoped by --run A --run listing needs no caller, so both handlers skipped the resolver and a --from naming another actor was dropped silently under a session. The conflict check now runs on that branch too; terminal callers are unchanged. * fix(orchestration): name this app's CLI by absolute path for a structured session's login shells A provider can run each command in a login shell: Codex runs zsh -lc, and the profile rebuilds PATH, putting a global install (possibly an older Orca) ahead of the directory Orca prepended. ORCA_CLI_COMMAND, which an agent resolves the CLI from first, is now the absolute launcher in that directory (the native launcher on Windows), so no shell's startup files can swap it. The PATH prepend stays for shells that read no profile. Found by the live coordinator run of the next PR. * test(orchestration): pin a structured worker's CLI command as this app's absolute launcher * test(orchestration): run the zsh login-shell arm in the real-shell lane that installs zsh The ordinary Linux unit lane has no /bin/zsh, so the zsh arm failed there with ENOENT. It moves to a live-shell file registered in the shell-contracts lane; the bash arm keeps running in every lane. The lane guard's detector now also sees a zsh spawned through the ProcessSpec program field, which is how this test escaped it. * fix(orchestration): omit a structured child's CLI command when no launcher resolves, and pin its instance A bare `orca` fallback named GNOME's screen reader on packaged Linux, and an inherited value named another app's CLI. The builder now deletes any inherited value, sets the absolute launcher only when one resolved, and pins ORCA_USER_DATA_PATH so a current CLI dials the instance that minted the id. Renames the marker reader to hasStructuredSessionMarker and records why the terminal view carries the id without the marker. * fix(terminal): name this app's CLI launcher by absolute path in every local terminal ORCA_CLI_COMMAND meant three things by lane: an absolute launcher for a structured session, a bare name for WSL, and nothing for any other terminal, so a structured session's terminal view lost it. Local terminals now get the same absolute launcher the structured lane gets; WSL keeps its guest command name, and a terminal whose launcher does not resolve still gets none. * feat(cli): hand a command to the session's own CLI when another Orca CLI was invoked A login shell can reorder PATH behind a global install, and an agent or its helper script can run bare `orca`, so the binary that answered depended on the agent following instructions. Orca's packaged launchers and bare-orca shims now export ORCA_CLI_SELF (outermost wins). At the CLI entry, when it names a different launcher than ORCA_CLI_COMMAND, the command re-runs once through the named launcher with ORCA_CLI_REEXEC=1 and exits with its status; both variables are consumed so no child inherits them. Dev launchers export no self on purpose, WSL and SSH names never qualify, and a launcher that cannot start leaves the command to run here. The Windows launcher no longer rewrites ORCA_CLI_COMMAND; the legacy ask protocol normalizes its resume command itself. * refactor(orchestration): declare which flag names the caller on each spec and refuse at the CLI entry Each handler hand-classified its --from/--terminal as the caller or a target, and the refusal of a conflicting caller flag ran inside the caller resolver plus two standalone calls for --run listings, so a new verb that read its flag raw would pass a sibling's handle to a pre-session host. Specs now declare identityFlagRoles, the CLI entry refuses a conflicting caller flag once from the spec, the resolver only applies the id-wins rule, and a test fails any orchestration verb that accepts --from or --terminal without classifying it. * perf(cli): keep the session caller check off the actor codec's module graph The check runs at the CLI entry for every command, and the actor codec pulls zod through the session record. Compare the session's own spellings as plain strings instead. * refactor(cli): spell a session's address from the one prefix constant, off the codec's module graph The Orca session address prefix moves to a leaf module with no imports, re-exported by the address codec, so the CLI entry check derives `session:<id>` from that constant instead of re-typing it and still stays off the codec's zod graph. Prose and test names say caller or Orca session id, not actor. * refactor(orchestration): drop the session id's terminal-view spawn now that the handoff is gone The terminal handoff was removed, so no terminal is ever a structured session: - delete the terminal-view identity env and its WSL passthrough, and their tests; - strip the session caller keys from every terminal's env unconditionally; - the CLI's own-address spelling moves beside the injected id in src/shared, with a test pinning it to the address the host's party resolver gives that session. * fix(terminal): run the Codex launch preflight through the CLI the terminal names Packaged Linux names the userData shim in ORCA_CLI_COMMAND, while the preflight ran the bundled launcher behind it. The CLI saw a different launcher and handed the preflight off to the shim, booting Electron twice before every codex launch. * revert(terminal): keep terminals on main's ORCA_CLI_COMMAND and Codex preflight Only a structured session needs an absolute ORCA_CLI_COMMAND; local terminals go back to naming none (WSL keeps its guest command), and the Codex launch preflight goes back to the bundled launcher. The CLI handoff is scoped to sessions, so a terminal's preflight can no longer be handed off and start Electron twice. This reverts commit |
||
|
|
17690e6b9a |
style: settle oxfmt 0.70 drift and stop formatting vendored licences (#23377)
The oxfmt 0.65 -> 0.70 bump landed without a repo-wide reformat, so 36 files already in the tree no longer matched what the new version emits. Anyone running `pnpm format` picked all of them up alongside their own change. Also excludes `resources/licenses/**`: `oxfmt --write .` was rewriting the vendored PCRE2 licence, turning its `*` redistribution bullets into `-`. Third party licence text has to be reproduced verbatim, so formatting must not touch it. |
||
|
|
82412dab8b |
Persist profile state in SQLite with background writes (#22612)
Migrate profile state to SQLite and move writes and backups into a background worker. Acknowledge terminal, SSH and automation changes only after durable saves. Preserve JSON import, recovery, rollback and compatibility exports. Validate migration, worker failures, maintenance, cross-profile moves and terminal lifetime races with unit, integration and end-to-end coverage. |
||
|
|
9af6a3d798 |
fix(cli): report a denied runtime connection instead of a dead Orca (#22341)
* fix(cli): report a denied runtime connection instead of a dead Orca Inside Codex's macOS Seatbelt sandbox, connect() on the runtime socket fails with EPERM. The CLI dropped the errno and reported "Could not connect ... Restart Orca", appended "Orca is not running. Run 'orca open' first.", and `orca status` answered ok:true with `starting` (its pid probe also gets EPERM). An agent following that advice restarts a healthy app, which cannot help. EPERM/EACCES on the metadata read or the socket/pipe connect now fails with a CLI-local `runtime_access_denied` error: ok:false, non-zero exit, operation/systemCode/processState:"unverifiable"/retryable:false and nextSteps that say to re-run with escalated permissions and not restart. CODEX_SANDBOX only picks the wording. `orca open` stops before launching. The status pid probe is unchanged: a refused or missing socket proves the caller reached the endpoint, so a later EPERM probe is another uid and keeps #20098's `starting`. Missing, refused, stale-pid and timeout paths are unchanged. Adapted from the diagnosis and tests in #20487 (and #19605, #13583). Co-authored-by: lifeodyssey <zhenjiazhou0127@outlook.com> * docs(skills): tell agents runtime_access_denied means escalate, not restart The shared CLI-resolution block told every bundled skill to run `orca open` when a command says Orca is not running. Add the counterpart for the new access-denied code so sandboxed agents re-run with escalated permissions instead of launching or restarting Orca. Regenerated stubs and manifest. * refactor(cli): classify only a denied runtime connect, with a leaner error A denied metadata read was never observed under a sandbox, and it turned an unreadable user-data path (the Linux launch contract's root-owned HOME) into runtime_access_denied instead of "Orca is not running". Keep metadata reads as on main and classify only the socket/pipe connect. One helper now maps a socket errno to the error or null; the error data keeps only systemCode and nextSteps. Tests drop cases already pinned by status.test.ts. * fix(cli): give not-running advice when a denied socket belongs to a dead Orca A crashed Orca leaves its metadata and socket file behind, and a sandbox denies the connect with EPERM before the CLI can see ECONNREFUSED. The sandbox still reports ESRCH for a gone pid, so a denied connect now probes the metadata pid and falls through to the ordinary unavailable path when the pid is proven gone. isProcessRunning moves to its own module so transport and status share it. * refactor(cli): inline the runtime_access_denied code like other CLI error codes --------- Co-authored-by: lifeodyssey <zhenjiazhou0127@outlook.com> |
||
|
|
981a4821da |
fix(cli,relay): stop reading an unsignalable pid as a dead one (+ unverifiable-collapse sweep result) (#20098)
* fix(cli): stop reporting an unsignalable Orca pid as a stale bootstrap `orca status` falls back to a `kill(pid, 0)` probe when `status.get` cannot be reached, and a bare catch read every refusal as absence. EPERM means the pid exists under another uid -- an Orca reached via ORCA_USER_DATA_PATH, or one started with sudo -- so a live app was reported `running: false`, `pid: null`, `runtime.state: stale_bootstrap`, `graph.state: not_running`. Only ESRCH proves the pid is gone, which is the rule every other liveness probe in the repo already applies (`isProcessAlive` in relay/pty-shell-utils.ts, pack-refs-lock-ownership.ts, runtime-metadata-ownership-watch.ts, and agent-session-process-identity-probe.ts). See docs/reference/ssh-execution-boundary.md. * fix(relay): keep a revived pane whose pid only refuses the liveness probe `revive` gated each serialized pane on a hand-rolled `process.kill(pid, 0)` in a bare try/catch, so any refusal retired the pane. EPERM means the process exists under another uid; only ESRCH is evidence of absence. The file already imports `isProcessAlive`, whose ESRCH-only contract `reapPtyProvenExited` documents 450 lines earlier -- this call site just did not use it. Reuse it rather than keeping a second implementation of the same concept. Malformed pids still skip, as before. See docs/reference/ssh-execution-boundary.md. * fix(lint): clear the casting gate on the pid-probe changes main tightened typescript/consistent-type-assertions to assertionStyle: never, which the rebase brings onto these added lines. The CLI probe narrows instead of casting; the relay test keeps the file's serialize idiom behind a SAFETY-annotated suppression. |
||
|
|
2626e2eca4 |
Make the structured turn lifecycle row durable so completed durations survive (#19695)
* Make the structured turn lifecycle row durable so completed durations survive A structured-chat turn used to end by tombstoning its running lifecycle item, which threw away the only durable record of when the turn ended. Completed "Worked for" labels therefore depended on the renderer having observed the turn finish, and vanished on reopen. The lifecycle item is now revised in place, never tombstoned: - running, with startedAt, at the provider's turn start - completed or interrupted, with completedAt, at the provider's terminal frame, a user stop, or a child exit the host observed - unverifiable, with no end, when a cold acquire finds a running row from a generation whose exit nobody observed Both timestamps are the execution host's clock at receipt, captured before the deferred sink, so the completed value is identical on every client and needs no client clock. Codex history restore uses the provider's own second-granular endpoints for turns that predate this change. Desktop and mobile read settled durations off the journal through one shared selector, and anchor the live counter on the host start with the client's local receipt so a skewed client clock never leaks into the label. Locally observed durations remain the fallback for hosts that still tombstone. Timestamps live inside the existing turnLifecycle field, which old clients strip, and every working-state consumer keys on state === 'running', so no capability negotiation is needed. * native-chat: avoid stale working status on settled turns * test: align settled turn status expectations * Name settled lifecycle rows by their terminal state An interrupted or unverifiable turn must not read as completed for any consumer that renders status text raw. One shared helper builds the text for both providers from the lifecycle state. * test: deduplicate turn lifecycle suites Each behavior keeps one test; duplicated harnesses and restated cases go. * Key lifecycle rows to their user item and record the provider's measured duration A lifecycle row now names the user item that opened the turn by its provider key, so clients attribute timing explicitly and fall back to journal order only for rows from older hosts. A provider-initiated turn with no prompt can no longer claim the previous prompt's duration. When the provider measures the turn itself (Codex turn.durationMs, Claude result.duration_ms) the terminal row records it and clients prefer it over the host interval, so a turn shows the same number live and after a history restore. Host receipt times remain the live-counter anchor and the fallback. * Record a turn as a first-class journal item The turn record is now its own item kind rather than a status row carrying a lifecycle field: no text to misuse, and the fold matches the durable turn record other systems keep. Rows that carry it are stamped journal schema v3; every other row stays v2, so an older host keeps reading them and latches read-only at the first v3 row instead of truncating the epoch. Clients that predate the item would paint an unknown kind as a text bubble, so the host publishes the legacy status form to any client that does not advertise agent-session.turn-item.v1, through the same per-client seam background tasks use. The downgrade is transitional and goes once no supported release lacks the capability. The shared projection now renders unknown item kinds as nothing, so later kinds need no gate. One shared reader handles both forms for old journals and old hosts. * Preserve observed turn end across settlement retries * Retain turn attribution for loaded chat history * Preserve Codex exit receipt across close retries * Register completed turn duration reliability gate * Keep earlier turns through a Codex rewind and count a mid-turn attach from the real start Findings from an independent adversarial review of the typed turn record: - A Codex rewind adopted the provider's item list as the new epoch, and the provider never returns the host's own turn rows, so every duration before the rewind point vanished. The host's turn rows are now spliced back beside the item each followed, and recovery no longer expects the provider to prove rows it never owned. - The epoch row was stamped with the current schema version, so an older host latched read-only at row 1 of every new session, defeating the mixed version design. It carries no body and stays at v2; a stored-row test now reads SQLite directly, because the reader upcasts every row on read. - A send Codex folds into a running turn shares the opening prompt's provider key, and the alias map credited the duration to the later prompt. The earliest submission naming a key now wins. - The live counter anchored on first sight, so a client attaching mid-turn counted from zero. Published frames now carry the host's clock, the reducer keeps the last sample with its local receipt time, and both clients anchor on how long the host says the turn has run. * Correct turn duration gate assertion reference * Respect authoritative unknown native chat duration * Preserve unverifiable timing across older host upgrade * Record final completed turn duration reliability evidence * Fix the CI failures the merge left behind - A merged import list named the same module twice, which the native code quality plugin fails on. - A running turn is now reported by the host with no duration, so the settled map carries an explicit null for it; the hook test still expected the entry to be absent. - main gave the older-page action a cursor with a head-trim guard, so the retention test's epoch-only action no longer typechecks; it now passes an unbounded sequence, which is what the old shape meant. - The roster comparator moved into the extracted module, leaving its import unused in the reducer. * Split two files back under the line cap after the merge Merging main put both one effective line over 300, and the cap forbids a disable or a shave. The wire module's refusal vocabulary moves to its own file and is re-exported, so its consumers are untouched; the host's four thin mutation delegates move next to the functions they call. * Advertise the turn-item capability on every client transport Local IPC and mobile advertised it; the remote and web transports did not, so a desktop paired to a remote host, the CLI, and web silently ran on the legacy carrier forever and the canonical row was never exercised there. The renderer that paints it is the same build on every transport. * Update the web auth-frame expectation for the new capability --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
f2af92b2fa |
feat(native-chat): show live background work and name each row by kind (#19705)
* fix(codex): reserve the label's share of a qualified command row
A child's label is raw provider text and was spliced into the command
row unbounded, then the pair clipped to the description cap. A label at
or past that cap clipped the command away entirely, leaving a row of
kind 'command' that named an agent and showed no command - the failure
qualification exists to remove, inverted. The same clip could also cut a
surrogate pair, which boundSubagentField already guards against on the
agent row two lines away.
Give the label a reserved share and clip it the way the agent row does.
* feat(native-chat): show live background work and name each row by kind
The strip suppressed itself in three places: the Claude tracker blanked
its roster for the whole of any turn, the Codex tracker returned nothing
while a primary turn was open, and the renderer view gated on
`turnId === null`. Between them, work in flight was never shown — and a
task backgrounded in an earlier turn vanished from the strip as soon as
the next prompt was sent. Claude additionally dropped every foreground
subagent, so a fan-out reported nothing at all.
Report work while it is live, in all three layers. Foreground Claude
work is turn-scoped, so `result` retires it — that is the provider's own
outcome for a task it marked foreground, not a roster sweep. Nothing
settles a Codex child on turn end: those keep reporting well past their
parent, so turn frames only prompt a republish.
Name each ROW by kind — Subagent, Shell command, Workflow, Monitor —
instead of a generic "Background <kind>", each drawing the glyph the
shared tool-icon table already uses for that category. A row that
carries a provider description still shows it unchanged. The collapsed
header summary is deliberately untouched; it is owned elsewhere.
The conversation-command gate is unchanged in effect: an open turn
already refuses first, and Claude foreground work never reaches the
backgrounded set the gate reads.
* fix(native-chat): withhold the row stop Claude foreground work cannot honour
The strip now publishes foreground rows, but `stoppableTaskIds` still filters
on `backgrounded`, so `stopClaudeBackgroundTasks` resolved an empty target list
and returned `{ cancelled: false }` that no renderer reads: the user clicked
"Stop Subagent" and nothing ever happened.
Carry stoppability per row instead of widening the stop to a target the SDK has
no way to reach. `AgentSessionBackgroundTask.stoppable` is absent-means-yes, so
hosts that predate it keep their working control, Claude emits `false` only on
foreground rows, and the strip hides that row's button the same way it already
hides the stop-all a provider cannot honour.
* fix(claude): scope aggregate-roster authority to the work it enumerates
`background_tasks_changed` lists BACKGROUNDED tasks, so a foreground subagent
can never appear in it. Treating it as the whole world meant any such frame
cleared every live foreground row mid-flight and then dropped every later
foreground `task_started` for the rest of the session, killing the in-turn
fan-out the strip exists to show in any session that ever backgrounds anything.
Decide `backgrounded` before the staleness guard and apply the guard only to a
backgrounded start, and retain live foreground entries across a roster replace.
Retained rows count against MAX_TRACKED_TASKS, so the map stays bounded, and a
stale backgrounded start the roster no longer lists is still dropped.
* test(native-chat): pin the strip's monitor amber to the constant that defines it
`MONITOR_GLYPH_COLOR`'s comment claimed a test held it and AgentStateDot's amber
together, but no test imported it — the assertions hardcoded 'text-yellow-500',
so the two could drift with every test still green. Read the colour from the
module, which is what the comment always said was happening. Drop the unused
`BackgroundTaskGlyph` export too: nothing outside the module names it.
* fix(native-chat): keep the task list open across a gap in live work
The strip is now mounted on live work, so a sequential fan-out unmounts it
between one subagent finishing and the next starting: local `useState` meant
the expanded list collapsed itself on every such gap, on top of the strip
flickering above the composer.
Hand the disclosure to the session, keyed by session id so it does not leak
across a session switch. The strip is now controlled and holds no state of its
own, which is what makes it survive its own mount churn.
* fix(codex): route every command-row cut through one surrogate-safe clip
`boundLabel` avoided splitting a pair, then `qualifiedDescription` re-cut the
COMPOSED string with a raw slice: label (<=96) plus separator plus description
(<=512) is up to 611 chars, so that second cut landed at an arbitrary index
inside the description and could publish a lone high surrogate — lossy through
any non-JSON UTF-8 hop. `parse` had the identical hazard on an unqualified
primary-thread command.
One `boundText` helper now owns all three cuts, so no path in the file can emit
a lone surrogate from well-formed input.
* fix(claude): keep terminal evidence for ids an aggregate roster never lists
Narrowing the admission guard to backgrounded starts left a finished FOREGROUND
id with no defence: `replaceAggregateRoster` wiped `terminalTaskIds` wholesale,
so after any `background_tasks_changed` a replayed `task_started` revived a task
whose completion had already been seen — and only a later `result` could settle
it again.
Scope the wipe the same way the guard was scoped: delete only the ids the
incoming roster actually enumerates. A roster still overrules terminal evidence
for the work it lists, which is what that behaviour was added for.
* fix(claude): keep retained rows in place and evict the stalest, not the newest
Re-adding retained foreground entries after the roster made a live row the user
is reading jump below the backgrounded rows on every `background_tasks_changed`,
and the cap `break` kept the STALEST retained rows while dropping the newest.
Merge in the tracked map's own order so a surviving row holds its position, and
count the overflow up front so eviction takes the oldest retained rows. Roster
entries are never starved and the map stays bounded either way.
* fix(claude): retire leftover foreground rows when the next turn starts
A foreground `task_started` arriving with no turn open has no `result` coming
to retire it, so it sat in the strip indefinitely — with no per-row stop, since
foreground rows are not stoppable — and refused conversation commands behind an
instruction nobody could follow.
Settle on turn start as well as on `result`. This is cleanup only: visibility
never consults `startsTurn`, so a missed one degrades to today's behaviour and
can never switch the feature off. It shortens the row's life to the next turn;
the case where no further turn is ever sent is filed separately.
* fix(agent-session): withhold unstoppable rows from readers that predate them
Rule 3 of remote-wire-compatibility: changing what the host publishes reaches
old clients with no wire change. The Claude host published no foreground rows
before this feature; it does now, and a client that cannot read `stoppable`
draws a per-row Stop on every one of them — Claude always sets
`supportsTaskStop` — which filters to the backgrounded ids, stops nothing, and
returns a result no renderer inspects. That is the dead button `stoppable` was
added to remove, reappearing across a version skew.
Negotiate it. A client can advertise the existing background-task-stop
capability and still predate `stoppable`, so this needs its own constant.
Readers that do not advertise it get unstoppable rows dropped, and a state whose
every row is dropped becomes no strip — exactly their pre-feature view.
RUNTIME_PROTOCOL_VERSION is not bumped: this adds an optional field and a new
negotiated capability, and changes no existing field's meaning, which is the
explicit do-not-bump case in protocol-version.ts.
* test(agent-session): name the projected rows so the fixture typechecks
An indexed lookup into the fixture's task list is possibly-undefined under
`pnpm tc`; the rows are more readable named anyway.
* test(web): advertise the row-stop capability in the e2ee auth expectation
The web e2ee handshake started sending
AGENT_SESSION_BACKGROUND_TASK_ROW_STOP_CAPABILITY, and this test asserts the
advertised list by deep equality, so it went red on CI while every targeted
test run stayed green. Add the capability in the position the router sends it.
* test(claude): pin why the roster empties mid-turn in a sequential fan-out
The strip unmounting between two sequential subagents is truthful, not a swept
row: A leaves on the provider's own terminal frame, B does not exist yet, and
backgrounded work spanning the same gap holds the roster open — so an empty
roster is never work the strip is hiding.
Also pins the previous-turn rule against the one the subagent roster already
applies on the same frame: a still-working FOREGROUND child becomes
`unverifiable` there and a backgrounded one is left alone, so the strip drops
the first and keeps the second rather than asserting `live` for either.
---------
Co-authored-by: Merge Sim <merge-sim@users.noreply.github.com>
Co-authored-by: Merge Sim <sim@local>
|
||
|
|
5868fdc9e3 |
feat(native-chat): report Codex background tasks in the chat strip (#19346)
* feat(native-chat): report Codex background tasks in the chat strip The background-tasks strip works for Claude only; a structured Codex session shows nothing in it. Feed it from the Codex app-server stream. The strip stands for work that OUTLIVED a turn, which is what the monitoring header, Claude's foreground suppression, and the conversation command gate all already assume. Codex has no `is_backgrounded` flag, so that fact is derived from the turn boundary: a `subAgentActivity` child or a primary-thread `commandExecution` becomes visible once the turn it belongs to completes and it is still unsettled. `turn/completed` only reveals a task here, never settles one — measured on `codex app-server` 0.153.4, a spawn_agent child reported `completed` 95.8s after its parent turn ended. Only a child's own activity kind settles it. Codex exposes no honest stop: `turn/interrupt` on a child ends its turn without emitting a terminal activity item and leaves its shell running. So the state carries a new optional `supportsStopAll: false`, the strip hides a control that could not act, and the blocked-command message asks the user to wait rather than to press a button that does not exist. * refactor(codex): move session teardown out of the structured adapter Merging main crossed the 300-line cap on `codex-structured-session-adapter.ts`: the rewind backend (#19235) and this branch's close-time strip clear both landed in it. The four close paths move verbatim into `codex-structured-session-teardown.ts`, where they funnel through one `settled` helper instead of repeating the notification-retry and background-task cleanup at each call site. No ratchet bump. Also normalize a background task's description once at receipt rather than on every projection; the roster is re-projected on each observed frame. * fix(codex): drop the shell row the journal already settles A `commandExecution` still `inProgress` when its turn ends was reported as a `command` task. But `settleCodexJournalTurn` writes exactly those items to the journal as `state: 'failed'` on `turn/completed` and forgets them, so the strip row would have claimed a shell was still running at the same instant Orca recorded that it was not — two surfaces contradicting each other about the same process. A subagent is the opposite case and stays: the roster pointedly does not sweep at a turn boundary, because children measurably outlive it. That leaves the producer making exactly one claim — these spawn_agent children are still live after their turn — which the durable roster row corroborates. * fix(native-chat): track Codex background execution lifetimes * fix(native-chat): keep running tool groups from claiming completion * Fix runtime catalog and capability expectation * fix(codex): keep a child's name on the command row that outlives it A child agent's commands stay hidden behind its agent row while the child works. Once the child's turn settles with a command still running, that command surfaces as its own row labelled from the raw command string, so 'long_probe' became "/bin/zsh -lc 'ping -c 300 127.0.0.1 > /dev/null'" at the moment that row was the only remaining signal for the work. Qualify a child's command row with the child's label. Resolved on read, so a label registered after the command still lands, and bounded by the existing description cap so admission accounting stays valid. Primary- thread commands are left unqualified: they have no child to name. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
06a607a1d7 |
feat(orchestration): make multi-agent workflows durable (#16904)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 225 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$21666 | $\color{#cf222e}{\Huge{\mathbf{−}}}$2820 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$18846 |
| Prod | 348 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$17107 | $\color{#cf222e}{\Huge{\mathbf{−}}}$4706 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$12401 |
<!-- /orca-pr-loc -->
## ELI5
Orca now treats orchestration like a durable control plane instead of inferring success from terminal keystrokes. Agents can tell whether a prompt was accepted or a turn started, replay an ambiguous request without sending twice, and recover coordinator mail after a crash. Completed workers can be inspected, released, or retained, and their panes no longer auto-resume as if the work were still running.
## What changed
- **Run receipts** from `run-create/use/current/show/list` are the row without routing plumbing (`home_database`, `coordinator_pane_key`) and without the duplicate `binding` object.
- **`terminal send` receipts are honest and idempotent.** `input_accepted` and `turn_started` are the only stages; `--wait-submit` observes without resending; `--retry-request <uuid>` replays the exact request against the same process incarnation. A transport timeout keeps the retry ID; only a different runtime answering strips it. Value-less or non-UUID `--retry-request` is rejected on the CLI and the SSH shim.
- **Mailbox delivery is committed before wakeup.** Pointer writes are staged in the DB before any PTY byte, replayed once after restart, and never emit a naked Enter. The watermark that parks concurrent deliveries is released with the DB reservation. Restart rescans pointer-pending and `dispatch:` mailboxes.
- **Lifecycle is a guarded transition graph** (`lifecycle-transition.ts`) with a table-driven test over every caller edge. Task reopen/overturn stays in the public contract. A PTY exit during `worker-stop` is the stop succeeding, not a failure.
- **Worker lifecycle CLI:** `worker-start` (`--spec` creates Task + attempt in one call), `worker-show`, `worker-read` (provider transcript first, bounded terminal fallback with a typed reason, local/WSL/SSH), `worker-stop`, `worker-abandon`, `worker-release`, `worker-retain`, `worker-list` (rowid-fenced pagination, fleet liveness, `attention`, literal `nextAction`).
- **Release is an explicit ownership table** (`decideWorkerTerminalRelease`): only an `owned` resource can be settled, the archive is mandatory where reachable, and an owner whose process is proven exited can always get out of `retained` via `archive_status: unavailable`. User-taken-over, external, and transferred panes stay retained.
- **Settled-worker resume fence** (folds in #17651): a settled dispatch whose pane is still open is fenced at settlement, on stop/abandon/exit, and at startup; lifted on release, retain, takeover, and pane reuse.
- **Liveness is `live` / `unverifiable` / `exited` only**, from execution-host evidence. Fleet projection reads the evidence clock, not the relay delivery clock. A host-certified exit outranks the worker's settled state. `unverifiable` never authorizes stop, abandon, retry, or release, in code or in the guide.
- **Federation:** structured reads negotiate by `method_not_found` so every shipped host keeps transcript-first output; exited remote workers are closed before being reported closed; epoch fencing holds across peer restart, downgrade, and pairing rotation; no per-second forced capability probe.
- **Schema v35:** repairs databases stamped v34 by the pre-fix branch (mailbox_handle default, index predicates), drops the write-only `lifecycle_transition_receipts` ledger and five never-read v31 identity columns.
- **Schema v36:** `dispatch:<id>` mailboxes get a real consumer generation on `dispatch_contexts` and `remote_dispatch_attachments`, bumped and fenced in the same transaction on every re-attach (manual inject, worker-start, federated attach). A stale worker whose Dispatch moved to another process now gets `consumer_fenced` instead of silently acking the new worker's Delivery. Run mailboxes already worked this way.
- **Schema v37:** `dispatch_contexts` records its creator (`creator_handle`, `creator_pane_key`), so a coordinator's context-only self-dispatch is bookkeeping rather than a nesting parent; before this, one self-dispatch made every later `worker-start` from that coordinator fail the depth cap. Pre-v37 rows keep counting (fails closed).
- **Dispatch-mailbox ownership is checked, not inferred.** A `check` from a process whose pane no longer holds the Dispatch, or whose last Attempt was abandoned/failed and moved to another terminal, gets `consumer_fenced` instead of an empty inbox that reads as "no mail yet". `--peek`/`--all` stay readable. A paneless caller still gets `stable_pane_required` with the rebind recovery.
- **Liveness certification is stricter:** a `process_exited` stage whose termination reason is `unknown` (a stop that was issued but never observed) projects `unverifiable`, not `exited`. Federated `worker-show` carries the execution host's verdict and host kind instead of a local guess. A live, ready worker with nothing pending has `nextAction: none` rather than pointing at the `worker-show` that produced it.
- **Wire:** `workerShow` keeps `dispatch.task_id` next to `taskId` for shipped CLIs. `ask --json` uses the standard `{ok, result}` envelope like every sibling verb.
- **Migration start-version detection** treats the two v32 recovery columns as versioned. Before this, every shipped database stamped below 32 resolved to the v6 floor and replayed the whole chain (the v23 backfill synthesized 68 phantom retained workers on a real v30 profile). Verified on a copy of a real 62 MB v30 profile: starts at 30, no row delta, integrity ok, 11 ms.
- **Skill guide** rewritten as a ≤200-line kernel plus seven references, to the outcome-first standard (Result / Done / Safe failure first, conditions not case lists, one done bar, references loaded at the point of use). The canonical loop uses `worker-start --spec`, names `worker-list` for completion accounting, documents `--retry-request` / `request-show` / `--wait-submit`, and requires positive evidence before any stall action. The other seven guides get the same treatment in #18724, split out so this PR stays orchestration-only.
- **`rpc/methods/orchestration-*`** (126 flat files) regrouped into `orchestration/{worker,federation,messaging,runs,gates}/`.
## Why
User reports showed the same boundary failures: false `agent_prompt_stalled` causing duplicate sends (#15180), coordinators unable to trust screen scrapes, cold-parked terminals receiving a pointer without the submit, settled workers accumulating as live tabs and auto-resuming after restart, and no way to tell a stalled worker from a working one.
## Linked issues
Fixes #15180. Fixes #17935 (orchestration skill description is 866 characters; a guard now caps every bundled skill at 1,024). Supersedes #17651 (fence folded in). Advances #16660, #16522, #14907, #13047.
## Review record
This PR was reviewed adversarially after revival: eight independent lenses (lifecycle, mailbox, send, worker, federation, transcript, complexity, live ergonomics), each required to prove findings with a failing test. That produced 16 proven blockers, all fixed with red-then-green regression tests, followed by two re-review rounds and a third fix wave that caught 3 regressions introduced by the fixes and 7 fixes that missed their target; all closed. A final pass (five lenses incl. a live built-runtime smoke, then a re-review of the fix wave) found and fixed seven more, chiefly the stale-worker mailbox steal, the self-dispatch depth wedge, and the unproven-exit certification. Three independent Codex (gpt-6-astra) passes followed: the first found nothing new, the second found and fixed 3 defects (task-status reachability, WSL-local host classification, peer-capability epoch), the third found and fixed 6 (production PTY controller never installed settled writes, ambiguous in-flight pointer failures allowed duplicate replay, SSH/relay deadlines cut off a valid `--wait-submit`, stop-vs-exit race during inspection, and two release-recovery paths for vanished or exited terminals). The full record (findings, proof tests, triage, declines with reasons) is archived outside the repo.
**Rework after the live smoke.** A first live cross-host run on the shipped adhoc build (this Mac, a paired Windows host on the same build, a paired Mac on 1.4.195, and an SSH host) found a P1: a running local worker read `unverifiable`/`missing_status` because the fleet snapshot rows lacked the terminal handle the matcher keyed on. A 59-row failure table over every bug fixed during review showed the same two classes recurring: a fact dropped in transit through optional fields, and two authorities for one fact. Two blind designs (Opus, Codex) converged on the same mechanisms, and the scoped tranches landed here with red-then-green seam tests from the real producer to the real consumer, faults injected only at the transport or hook-ingest boundary:
- **Settlement (data-loss class):** one three-valued `WriteSettlement` (`accepted | refused{reason} | unverifiable{reason, bytesHandedToTransport}`) from the SSH multiplexer through daemon client, providers, controller, to pointer staging. No boolean, no rejection-as-third-state. The two silent degrades that fabricated a handoff are deleted; a provider that cannot settle refuses before any effect. Pointer text and Enter share the contract; a partial flush is `unverifiable`, never `refused`.
- **Evidence identity (false-liveness class):** fleet agent-status evidence is a tagged union (`binding: worker | pane | unresolved{reason}`, `clock: observed | delivery`) minted once at ingest, so a hook row captured on one process incarnation can never bind to a later dispatch on the same pane. The matcher's `!worker.paneKey ||` defaults are gone. One host-scope parser replaces two.
- **Small pre-merge items:** `capability_unsupported` from an old peer is no longer relabelled `host_unavailable`; a producer census test asserts every agent-status consumer path projects a pane-only hook row as `live`.
Two ergonomics defects the second live run surfaced on a real database are fixed here too: a pre-v3 dispatch already marked `completed` projected as `outcome_unknown` / `requiresAction: true` forever (three copies of the outcome ladder disagreed on legacy rows; now one resolver, legacy `completed` reads `succeeded` with nothing to act on, legacy `failed` stays actionable on the failure), and an unscoped `worker-list` enumerated the entire database (now defaults to the Run bound to the calling terminal, `--run` overrides, and the receipt's additive `scope` field says which).
A third live round on the shipped adhoc build of `b082443e1f` (same four hosts) plus an unscripted run in the user's own prompt style (a plain Claude Code shell, `/orchestration`, three workers, zero errors, bound-Run default confirmed) found two more branch defects, fixed with red-then-green tests: a worker freshly started on a paired server projected `unverifiable`/`host_indeterminate` with `requiresAction` for ~3 minutes, including after its own `worker_done`, because the host's federation observation returned `missing_liveness_verdict` for any PTY the liveness register had not yet swept (the host now reads a connected pane it owns locally as `live`; disconnected or SSH-scoped panes stay `unverifiable`); and six pre-v3 completed rows still carried an `input` category because settling through the task-status path or `failDispatch` never closed the Dispatch's pending question threads (both paths close them now, and schema v38 closes threads already pending on settled rows). The guide's `worker-start` examples now show `--model sonnet`, since an omitted model inherits the launcher's default.
A Codex adversarial pass on the tranche diff found one real design hole (identity minted at read time instead of ingest, now closed) and two daemon settlement paths that threw instead of settling (fixed). Two `@ts-nocheck` runtime mixins on these paths were extracted into checked modules; the repo-wide `@ts-nocheck` count is unchanged at 171.
Deletions during review: ~1,900 lines (write-only ledger, unread columns, dead v1 archive path, test harnesses shipped in prod, duplicated liveness and state-machine copies, self-capability checks that were compile-time true).
## Testing
- `pnpm typecheck:tsc:node|cli|web` clean
- `pnpm run check:code-quality:changed` 0 findings; `check:react-doctor:changed` 0
- `pnpm verify:bundled-skill-guides`, `verify:skill-bundle-manifest`
- full `pnpm test` on the integrated head: 72,332 pass / 292 skipped; the only failures were three non-PR files (two zsh live-shell suites hit a node-pty spawn-helper ENOENT while a concurrent native rebuild ran, 44/44 in isolation; `release-checkout.unit.test.ts` is a known 30 s load timeout that passes in isolation on `origin/main` too).
- CI on
|
||
|
|
7b44f3c0e3 |
perf: decode fragmented CLI replies without repeated scans (#18909)
* perf: decode fragmented CLI replies without rescanning accumulated text * bench: require an explicit CLI framing baseline |
||
|
|
f37d2fec97 |
fix(linux): land the reviewed Linux packaging stack on main (#18100)
* fix(linux): give the CLI one entrypoint by extracting the AppImage once
* refactor(linux): trim AppImage CLI registration seams
* test(cli): assert registration lock serialization
* fix(linux): fence AppImage terminal shim mounts
* fix(linux): accept extracted AppImage runtimes with APPDIR only
* docs(linux): make headless AppImage extraction runnable
* refactor(linux): import bundled launcher directly
* fix(linux): reclaim superseded AppImage payloads and packaged symlinks
Pruning removed 3215 of 3216 files from a superseded generation and always
stranded resources/app.asar, leaking ~105 MB per version update. Electron's
asar shim reports a *.asar file as a directory, so the recursive remove tried
to rmdir a real file and failed with ENOTEMPTY; the .catch(() => {}) hid it.
Reproduced end to end on Ubuntu 24.04: 519M -> 623M across one update, and
519M again once the payload is actually reclaimed.
removeExtractedAppImagePayload holds process.noAsar for the removal, counted
so overlapping removals cannot hand the shim back early, and the prune site
now warns with the path instead of swallowing the rejection. All three
removal sites use it -- staging cleanup and displaced roots leaked the same
way.
Also reclaim symlinks left by a packaged deb/rpm install, which the
extracted-cache-only rule turned into a hard conflict on a deb -> AppImage
migration, and name the remedy in the conflict error.
* fix(linux): bound the CLI registration lock wait
`retries: 1000` caps the attempt count, not elapsed time, so at up to 1s per
attempt an IPC-driven registration could hang ~16 minutes against a wedged
holder with no feedback.
A legitimate holder is bounded by the extraction timeout, so wait that plus
slack and then fail with a message naming the lock file, rather than hanging.
`maxRetryTime` is forwarded verbatim to the `retry` package by proper-lockfile.
* fix(linux): stop re-extracting the AppImage on inode metadata churn
The extracted-payload cache key hashed ctime alongside dev/ino/size/mtime.
ctime moves on any inode metadata write -- `chmod +x`, which every AppImage
user is told to run, plus `chown`, an ACL or SELinux relabel, and a backup
restore -- none of which alter a byte of the payload.
Measured on Ubuntu 24.04: `chmod +x` leaves dev, ino, size and mtime
identical and moves ctime alone, so the key changed and the next launch paid
a full ~519 MB re-extraction and a multi-second stall to rebuild a payload it
already had, then pruned the old generation.
Key on content identity instead. An in-place content change moves mtime and
almost always size; a replacement moves the inode. The existing
replace-in-place test still passes.
* fix(linux): stop CLI commands from falling through to Chromium startup
* refactor(cli): remove redundant command membership check
* test(cli): cover command-named project selectors
* fix(cli): redirect the open-url command before startup
* test(linux): cover AUR serve wrapper flags
* fix(linux): tighten CLI launch detection
* fix(linux): respect CLI flag value boundaries
* fix(linux): strip injected Chromium switches from CLI args
* fix(linux): report a missing display instead of dying in uv_close
* refactor(linux): read display locks without a preflight race
* fix(linux): preserve unverified external displays
* chore: format reliability gate manifest
* test(packaging): split runtime resource checks
* fix(linux): fail serve when no display is available
* fix(linux): do not treat a lockless X socket as a dead display
An X server writes its lock beside its socket and both survive a crash
(verified against Xvfb under SIGKILL), so a socket with no lock was never
left by a crashed server. It is an endpoint published from elsewhere: a
container bind-mounting only /tmp/.X11-unix, WSLg, or a foreign PID
namespace. Declaring those dead made the desktop gate exit(1) on displays
that work, with no workaround, and the serve gate refuse to start.
Liveness now splits by ownership. A foreign DISPLAY trusts a lockless
socket; Orca's own :99 does not, because removeStaleDisplayArtifacts
unlinks the lock before the socket and so manufactures that state itself --
adopting it would resurrect the orphan-socket bug and stop the cleanup from
self-healing. The stale-lock rejection is unchanged.
Also correct four doc statements this behaviour falsified.
* fix(linux): fail closed when a stale socket blocks the Xvfb rebind
Readiness only checked that /tmp/.X11-unix/X99 exists. A stale socket we
could not unlink still exists after our own Xvfb refused to bind, so Orca set
DISPLAY to a dead server and Chromium died in Ozone init.
Measured on Ubuntu 24.04 against the pre-fix build: with a leftover :99
socket and no lock, serve exits 139 (SIGSEGV), the socket inode is unchanged
before and after, and no lock is recreated -- it neither cleaned up nor
respawned. To a user that is a crash, not a misconfiguration.
This is reachable in the documented topology, where orca-xvfb.service has no
User= and runs as root while serve runs as User=orca: /tmp is sticky, so the
orca uid cannot unlink a root-owned socket, rmSync fails, and Xvfb exits with
the display already active.
Readiness now requires the display to actually be live -- our socket plus a
lock naming a running process -- so the same state reports an unusable
display and exits 1 with the existing diagnosis.
* fix(linux): recognise abstract X sockets and inherited Wayland fds
Two display setups this gate could not prove were refused outright, and on the
desktop path that is app.exit(1) with no workaround.
An X server may bind only the abstract namespace (`@/tmp/.X11-unix/X0`), which
leaves no filesystem socket to stat. Abstract addresses are kernel-owned and
vanish the moment the owner exits, so an entry in /proc/net/unix is proof of a
live server -- no lock file needed and no stale entry possible. Verified on
Ubuntu 24.04, where 139 such addresses were present.
WAYLAND_SOCKET is an already-connected fd handed over by the compositor, so
there is no path to stat and WAYLAND_DISPLAY may be unset entirely. Its
presence is the display.
Both are consulted only after the filesystem-socket check fails, so no
existing verdict changes.
* fix(linux): never treat Orca's own display number as a foreign endpoint
Recognising a lockless X socket as live is correct for an endpoint published
from elsewhere -- a container bind mount, WSLg -- because an X server writes
its lock beside its socket and both survive a crash. It is wrong for
VIRTUAL_DISPLAY_NUMBER, because Orca's own teardown unlinks the lock before
the socket and so manufactures that exact state.
The managed branch was already strict, but a caller that sets DISPLAY=:99
explicitly takes the foreign path and skipped it, accepting a dead display
left by Orca's own interrupted cleanup. Route the managed number through the
strict probe on both paths.
Found by an adversarial audit of the asymmetry introduced earlier in this
branch; the documented systemd topology is unaffected because its Xvfb writes
a real lock.
* test(linux): add a packaged-artifact contract for the CLI launch paths
* test(linux): avoid buffered serve readiness detection
* test(linux): signal AppImage serve owner directly
* test(linux): tolerate readiness timeout boundary
* test(linux): add startup margin to shutdown oracle
* ci(linux): give package contracts timeout headroom
* fix(ci): route all Linux packaging contract changes
* test(linux): poll shutdown readiness without tail leaks
* test(linux): bound shutdown cleanup grace
* test(linux): assert on CLI output, not the harness's own control lines
run-cli-case.sh echoes `RESULT status=N case=<name>`, and the two cases named
*-skills asserted `expectOutput: 'skills'`. That substring was satisfied by
the case name in the harness's own line, so 2 of 8 cases asserted nothing
about the command -- gutting `skills` entirely would still have gone green.
Control lines are now excluded before matching, and both cases assert the
rendered help header, which only real help output produces. Verified on an
Ubuntu 24.04 host: 8/8 still pass against a stack-tip AppImage.
Also register the gate in reliability-gates.jsonc, which #15085 added a CI
Docker gate without. Red/green is recorded from a stock release AppImage
failing 4 of 8, three of them at status 133 (SIGTRAP).
* fix(linux): require static AppImage runtimes (#17319)
* test(linux): reject a wrong-architecture native binary at packaging time
Cross-building the arm64 slice on an x64 host silently packed an x86-64
`pty.node` -- the rebuild logged "Forcing native rebuild for linux-arm64" and
shipped the host's binary anyway. Every gate here inspects symbol versions,
which are perfectly valid on the wrong architecture, so nothing noticed.
Observed on a Raspberry Pi 5: the packaged app loaded, then failed with
"Failed to load native module: pty.node", and the launch contract reported
3 of 8 cases crashed rather than naming the cause. Swapping in the aarch64
`pty.node` took the same build to 8/8.
Compare ELF `e_machine` against the slice being packaged and fail with the
offending path. Checked before the glibc pass, because a wrong-architecture
binary's symbol versions are valid but meaningless and would send the reader
down the wrong path.
Release CI builds arm64 on a native runner, so this guards local and future
cross-builds rather than a shipped artifact.
* test(linux): judge per-arch vendored binaries against their own path
The first CI run of the architecture gate failed the x64 package job on
`@parcel/watcher-linux-arm64-glibc/watcher.node`. That binary is arm64 on
purpose: the package ships every architecture and its loader picks the match,
so its presence in an x64 build is correct.
Judge a binary against the architecture its own path names, falling back to
the slice when the path names none. That keeps the case this gate exists for
-- `bin/linux-arm64-*/node-pty.node` holding an x86-64 binary, which is what
shipped to a Raspberry Pi 5 -- while letting multi-arch dependencies through.
Dry-run over the real dependency tree flags nothing for either target arch.
* fix(linux): move deb/rpm update installation outside Orca (#17318)
* fix(linux): complete deb/rpm package metadata
* fix(linux): preserve CLI link during package upgrades
* docs(linux): document local RPM build prerequisites
* fix(linux): move deb/rpm update installation outside Orca
* fix(updater): preserve Linux recovery across stale events
* fix(updater): fence stale downloaded events by active target
* fix(updater): preserve active Linux package recovery
* test(linux): keep workflow order assertion in scope
* test(updater): assert stale recovery stays silent
* fix(updater): preserve Linux package recovery after checks
* refactor(updater): keep Linux marker message with status
* fix(linux): describe the right manual update path for deb/rpm hosts
A remote host installed from .deb or .rpm now reports
manual-service-update-required, and the guidance told the operator to
"update through the service manager that starts this server" -- which is
correct for unsupported-headless-serve but wrong for a package install,
where nothing about the remedy involves the service manager.
Say both, keyed on how the host was installed.
* docs(linux): document orcad update restart safety
* docs(linux): scope restart census omissions
* docs(linux): use absolute service CLI launcher
* fix(serve): validate in-process serve options before startup (#17683)
* fix(linux): stop offering updates a distro-managed install cannot apply (#17918)
Closes #17702.
The resources/package-type marker is authoritative but never checked against
the host, so any repackager that unpacks Orca's .deb -- AUR, Nix, a container
rebuild -- inherits `deb` verbatim. Install feasibility was then computed
after a ~165 MB download, so those users got check -> download -> a card
promising an install command -> a dead end.
Validate the marker against the host: a deb/rpm marker with no matching
package manager in the trusted directories means a package manager owns this
install. This reuses the exact lists and resolver that
buildLinuxPackageInstallCommand already loops over, so a false positive is
impossible by construction -- any host flagged here would have failed with
no-package-manager after the download anyway. The gate only moves that
verdict earlier. Verified across Debian 12, Ubuntu 24.04, Arch, Fedora 40 and
openSUSE Leap: no false positive on a real deb host, correct on every
repackaging host.
The release is still reported, because the user does want to know 1.4.194
exists and to update through their distro; only the download path is closed.
`externallyManaged` is an additive optional field on the existing `available`
status, so older paired clients decode it unchanged. downloadUpdate() refuses
authoritatively, since main owns this verdict rather than the card, and
unwinds any pinned-build state first -- a Linux pinned jump resolves to
'release', and stranding isPinnedBuildActive would silently kill every
background check for the rest of the process.
Note the fix the issue suggests cannot work: electron-updater builds a
PacmanUpdater whose doDownloadUpdate looks for a .pacman asset Orca does not
publish, then dereferences undefined.
* style(cli): restore prettier wrapping on install error copy
* test(linux): re-pin the child-process ratchets and the batch-shim allowlist after the merge
|
||
|
|
aabcc57366 |
fix(runtime): publish remote control outages to host surfaces (#17531)
* fix(runtime): publish remote control diagnostics to renderer * test(runtime): account for diagnostics bridge listener * fix(i18n): add runtime connection state labels * test(runtime): clean up shared control connection * fix(runtime): fence diagnostics by shared-control capability * fix(runtime): preserve authoritative transport state * fix(runtime): preserve diagnostic overlay lifecycle * fix(runtime): avoid publishing unchanged diagnostics state --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
c3aceacc7b |
Fix PR unlink for auto-detected reviews (#16898)
* fix: make PR unlink hide auto-detected reviews * Type the empty-content test double against the real model The literal narrowed suppressedGitHubPR to number and typed the callback as Mock, so neither direction was comparable and tsconfig.tc.web.json failed on TS2352. Keeping the 'as' cast preserves checking of the fields the double does supply. * Add localization keys for the unlinked checks-panel state The unlinked title, relink action, and the remote-runtime upgrade notice introduced untranslated keys that static analysis requires in en.json. * Advertise PR suppression capability in the transport test The client capability list is pinned by websocket-transport.test.ts, and adding WORKTREE_GITHUB_PR_SUPPRESSION left the expected list stale. * Fix stale PR suppression in Checks * fix: harden PR unlink suppression state * refactor: extract PR unlink state handling * fix: show PR relink recovery in source control * fix: add unlinked PR localization * Clarify workspace-scoped PR unlinking --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
913509edeb |
fix(orchestration): prevent slow worker-start stalls (#16300)
* Extend orchestration agent submission timing budgets * fix(orchestration): preserve mutation recovery identity * fix(orchestration): preserve recovery executable identity * fix(orchestration): keep worker starts and recovery commands safe * test(orchestration): cover federated worker preflight * fix(orchestration): harden mutation recovery * fix(orchestration): redact dispatch recovery credentials * chore: preserve upstream skill dialog formatting * test(orchestration): stabilize agent prompt submit e2e * fix(orchestration): validate federated start receipts * perf(runtime): cache unchanged prompt verification tail * fix(orchestration): reject worker-start timer overflow * fix(orchestration): normalize worker-start timeout defaults * fix(orchestration): normalize worker-start readiness budgets * fix(orchestration): normalize federated readiness timeout * test(runtime): tolerate current-main degradation exports * chore: preserve current-main orcad formatting * chore: drop unrelated formatting carryover |
||
|
|
0f522c35e5 | fix(remote): gate empty session inventory on host authority (#16546) | ||
|
|
0e10fc5925 | fix(browser): retire helpers with page owners (#16564) | ||
|
|
cda2280d63 |
Show all automations (#16532)
* Add all-host automations with scoped ownership and multi-authority suppo
Enable automations to run on multiple hosts (SSH targets and local) with
owner-fenced mutations, scoped list queries per host, and conflict
resolution. Introduces desktop and runtime authorities as distinct
automation storage owners, with per-host caching, invalidation, and
retry scheduling on the renderer. Captures registration generations for
SSH hosts to survive re-adoption. Adds CLI support for destination
selection and conflict recovery.
* Filter automation create projects by destination host
Only offer projects available on the selected destination, preventing
the mismatches that would fail at submit time. Auto-adjust the project
selection if it becomes unavailable when the destination changes.
* Add runtime storage authority support for automations
- Support both runtime and desktop as automation storage authorities
- Make owner preconditions optional for legacy-client compatibility
- Cache automation list projections to improve performance
- Add per-row repo/worktree resolution for cross-authority collisions
- Extend automation.list RPC to always include owner metadata
* Replace child_process.execFile with runProcess for external automations
- Migrate external-manager to use cross-platform runProcess wrapper per child-process safety policy
- Abstract electron app/ipcMain APIs in orca-runtime via environment accessors
- Install fake app environment in automation tests for consistent setup
- Reorganize imports to use specific module paths (ssh-target-registry, agent-detection, browser-error)
- Remove external-manager from child-process import allowlists (no longer violates direct import)
* Unify desktop automation CRUD onto the local runtime RPC surface
The desktop authority now speaks the same automation.* RPC contract as
remote runtimes, via callRuntimeRpc({kind:'local'}) -> runtime:call ->
the shared RpcDispatcher. The automations:list/listRuns/create/update/
delete/runNow IPC arms, their preload members, and every renderer
desktop-vs-runtime transport fork are retired; the runtime methods are
the single implementation of scoped lists, owner fencing, and change
publication for both transports (mobile clients already exercised them).
The desktop probe scheduler's priority lease survives the move as an
AutomationService hook the IPC registration installs and the runtime
methods take, so Orca's own automation traffic still parks queued
external-manager probes.
External-manager scope arms and dispatch-loop plumbing stay on IPC by
design; automation change events keep their existing channels (renderer
ingestion already converges them by authority).
* Remove automation ghost SSH tombstone scanning
This functionality for synthesizing tombstones for automation-referenced SSH
targets is no longer needed as part of the automation system refactoring.
* Refuse orphan automations at dispatch time, not migration time
Remove migration-time disabling of orphan automations and the `enabledDecidedBy` field. Dispatch now refuses orphans at runtime instead, simplifying state management and UI. Orphans are left unstamped and enabled; dispatch refuses to run them via `resolveAutomationRunTarget`.
* Show all automations in flat table with unified filter menu
- Replace host picker component with comprehensive Filters menu supporting status, last run, agent, and host filters
- Flatten automation list layout to single table instead of host-grouped sections
- Add Host column to display execution host for each automation
- Display active filters as removable pills below toolbar
- Delete unused AutomationHostPicker* components
* Add automation owner fencing and destination validation
- New AUTOMATION_OWNER_FENCING_RUNTIME_CAPABILITY for owner preconditions; legacy clients get owner metadata snapshotted at RPC boundary for compatibility
- Editor captures and revalidates automation destination before save, preventing silent retargeting if SSH infrastructure changes mid-edit
- SSH target types now isolate renderer-authored fields; generation is server-owned and stripped by IPC handlers
* Route automation recovery actions to the origin host
When an automation action fails due to owner fencing, recovery verbs
("Update server", "Reconnect") must run on the host where the refusal
originated: the row's captured owner for row operations, or the
destination the create dialog captured, not the list's filtered host.
* Remove external manager scope limitation notices
Consolidate create destination eligibility checks with a unified predicate
and fix the bug where desktop repo IDs could be sent to runtime hosts where
they cannot resolve.
* Persist only store-derived automation contexts, not client-perspective o
Store contexts must never be based on client-provided runContext or sourceContext
values—clients speak a different perspective (e.g., 'runtime:<id>' for host IDs
they assign), and persisting those makes the store projection orphan automations
it actually owns. Derived contexts now take precedence in create and update paths,
with explicit null still honored to clear a value. Tests verify this by simulating
drift after storage and confirming that moves re-derive while toggles preserve.
|
||
|
|
09048c63d4 |
feat(orcad): add headless browser providers (#16193)
* feat(orcad): add headless browser providers * fix(orcad): merge the duplicate runtime-browser type import |
||
|
|
a61b39a9a6 |
fix(runtime): stamp a runtime's own project setups as local, and report remote status about the remote (STA-4792) (#15376)
* fix(runtime): stamp a runtime's own project setups as local, and report remote status about the remote (STA-4792) Two independent frame-of-reference bugs, both from code describing one machine while labelled as another. #15366 — projectHostSetup.* persisted the caller's host id verbatim. Those `runtime:<environment-id>` ids are minted by the calling client's own pairing store, so they name a machine only relative to that client. A client sending one is addressing this runtime, and runtimes do not proxy these calls onward, so the host it names is us. Storing the client's spelling made one machine look like a different host to every other client, hid its rows from them, and defeated the (projectId, hostId) duplicate check — two laptops paired to one server each created their own setup for the same checkout. Re-spell it as `local` at the RPC boundary. Rows written earlier keep their old stamp; readers already project `local` back to `runtime:<their-id>`, so the client-visible model is unchanged and no ids are rewritten. STA-4792 defect 4 — `status --environment <name>` hardcoded app.running:false to mean "no desktop on THIS machine" while every other field in the same object described the target, including a desktopWindowStatus echoed straight from it. The result contradicted itself and read as "that run was headless" when the remote GUI was up. `app` now describes the target, keyed off the one window status that requires a live renderer, and the result names its own subject so the frame can't be misread again. The remote pid is not knowable, so it stays null. STA-4792 defect 2 gets a regression test rather than a fix: routing already made the client remote, which is what stops a Windows destination being joined to the local cwd. The test pins the exact reported invocation. * fix(status): share the remote app projection with the SSH host passthrough, and name the version gap on project host setup Two review follow-ups. The SSH host passthrough answered `app.running: true` unconditionally for the Orca host a caller reached over SSH, claiming a desktop app even for a headless `serve`. That is the same defect as the paired-server path, one transport over, so the projection moved to shared and both now answer the question the same way. `--host runtime:<id>` routes project commands to a paired server, which means a client can reach a server that predates project host setup without meaning to. That answered a raw `method_not_found`, which reads as an Orca bug rather than a version gap; the CLI now names it the way the desktop already does. Reverted a third change: making the persistence duplicate check treat `local` and `runtime:*` as one machine. That assumption holds at the RPC boundary, where a `runtime:` host means the runtime being addressed, but not in the store, which also records independent provisioning metadata for machines that are not itself. An existing test covers exactly that, and it was right. The duplicate convergence therefore stays bounded to rows written after the normalization. |
||
|
|
fa9b20cb41 | feat(skills): reland private bundle sharing safely (#14934) | ||
|
|
763b1febeb |
Revert "feat(skills): add private bundle sharing (#14401)" (#14913)
This reverts commit
|
||
|
|
757fae28d7 |
feat(skills): add private bundle sharing (#14401)
Co-authored-by: E2E Test <e2e@test.local> |
||
|
|
78d5920446 |
fix(orchestration-cli): point dropped mutations at --retry-request (#14586)
* fix(orchestration-cli): guide dropped mutations to idempotent retry * test(orchestration-cli): preserve read-only drop message * fix(orchestration): harden mutation replay identity * fix(orchestration): preserve replay across remints * fix(orchestration): defer local mutation identity |
||
|
|
83e2123582 |
Add global worktree visibility source defaults (#14276)
* Add global external worktree visibility defaults * Expand global worktree visibility source defaults * Fix host-scoped visibility settings races * Fix global worktree visibility integration * Enable source visibility defaults on mobile * Polish external worktree settings navigation * Clarify inherited worktree visibility settings * feat(sidebar): replace the inherited-visibility switch with a Show/Hide picker Each source row now shows a two-segment Show / Hide control preselected to the global setting, and explains itself only where the project actually disagrees: an "Overriding global setting: <value>" card names the value being ignored. Picking the segment global already holds drops the override instead of pinning a duplicate, so the same control both overrides and reverts, retiring the separate "Use global" link. The dialog footer now lists every inheritable source with its global value. * fix(sidebar): preserve reset for matching visibility overrides |
||
|
|
4882eeb8ac |
rm git shim: neutralize stale wrappers without a host gate (#14255)
* Revert "fix terminal attribution shim removal edge cases (#14187)"
This reverts
|
||
|
|
585dd6d3a9 |
fix terminal attribution shim removal edge cases (#14187)
* fix(terminal): fully retire attribution shim * fix(terminal): harden shim tombstone path lookup |
||
|
|
cd8c66551a |
fix(agent-hooks): resumed Claude Code session gets its sidebar agent row at SessionStart (STA-3386) (#12859)
* fix(agent-hooks): give resumed Claude sessions a sidebar row at SessionStart (STA-3386) Claude's hook set never registered SessionStart and normalizeClaudeEvent dropped it at ingest, so a resumed session that idled produced zero hook traffic and earned no sidebar agent row until the first prompt. - Register SessionStart in CLAUDE_EVENTS (local + remote installs). - Map lead SessionStart (startup/resume/clear) to an idle 'done' row, resetting stale roster/task/cron/tool/prompt state like the Codex path; compact restarts and child-attributed SessionStart stay dropped. - Thread hookEventName through the agent-status IPC payload so the completion coordinator can tell a session connect from a turn result; a SessionStart 'done' no longer raises agent-task-complete. * fix(agent-hooks): mark SessionStart rows as session boundaries, not completions (STA-3386) Review follow-up: represent the idle connect as a first-class sessionBoundary flag on the status payload instead of gating one renderer consumer on hookEventName. - sessionBoundary rides AgentStatusPayload/AgentStatusEntry (done-only, clamped like interrupted); drops the hookEventName IPC threading. - Completion-reactive consumers ignore session boundaries: the completion coordinator (task-complete notifications), automation dispatch observers (a connecting agent no longer completes the run and closes its tab), activity unread counts, and the dashboard finished timestamp; the status slice keeps boundaries out of stateHistory and preserves the flag across done->done repaints. - SessionStart sources are allowlisted (startup/resume/clear) so compact restarts or unknown sources fail closed mid-turn. - A live SessionStart now un-retires a reusable pane like a fresh prompt, so resume-in-reused-pane earns its row too. * fix(agent-hooks): keep session-boundary dones out of teardown and completion history (STA-3386) Review round 2: - A boundary done no longer deletes the pane's launch-config registry entry, so a resumed idle TUI keeps its registered-launch-agent identity evidence. - A boundary landing on a REAL done pushes that completion into stateHistory so the finished timestamp and unread badge survive a resume//clear right after a finish. - The done->done flag carry yields to turn evidence (assistant message or changed prompt) so a genuine completion can never be suppressed. - Star-nag value-moment observer and the server's OSC-equivalence dedupe now discriminate the flag. * fix(agent-hooks): keep a displaced completion unread in the sidebar badge (STA-3386) Review round 3: sidebar-badge mode counts only the live entry, so a session boundary landing on an unacknowledged completion silently dropped the sidebar badge while the agent-events count kept it. Count the displaced completion from history for boundary rows, and pin the behavior with countActivityUnread tests. * fix(agent-hooks): prevent SessionStart completion side effects (STA-3386) * fix(agent-hooks): preserve SessionStart through renderer IPC (STA-3386) |
||
|
|
8c65dd5094 |
perf(runtime): keep PowerShell ACL work and a second auth off the remote command path (#12451)
* perf(runtime): keep PowerShell ACL work and a second auth off the remote command path Two costs sat on the remote authentication path on Windows: - The E2EE handshake persisted `lastSeenAt` inline, and every secure-file write spawns PowerShell synchronously twice to reapply the registry ACL, so the client's `e2ee_authenticated` waited on both spawns. - Every remote CLI command except `status.get` opened a second full WebSocket connection just to re-read status for the protocol-compat check, doubling the authentications per command. The first sighting of a device still persists inline (rotation drops entries disk says were never scanned); later refreshes update memory now and coalesce onto one deferred write. The compat verdict is saved against the runtime's per-launch `runtimeId`, so a restarted or upgraded runtime retires it. * fix(runtime): preserve compatibility on one remote auth * fix(runtime): flush registry after transport shutdown |
||
|
|
73c5009b82 |
chore(dead-code): drop ~2k lines of unreachable exports and orphan modules (#12077)
* chore(dead-code): drop 2k lines of unreachable exports and orphan modules Ran knip across every build entry (main, preload, renderer, popout, web, cli, relay, workers, forked sidecars, config scripts) and removed what no entry graph can reach. - 11 orphan modules nothing imported, plus one test that only covered them - 159 unused exports/types, with their now-dead helpers, imports and tests Each candidate was verified against dynamic references before deletion. 42 knip hits were false positives and are kept: shared modules consumed by the mobile/ workspace, the src/shared/plugins/** public API, vendored shadcn primitives, and relay wire-protocol constants held for compatibility. Adds knip.json + `pnpm audit:dead-code` so this stays measurable. Verified: pnpm typecheck, pnpm lint, and 2081 tests across the 73 affected test files all pass. * chore(dead-code): move knip config under config/ Root-level additions are blocked by the root directory guard. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
363e478909 |
fix(orchestration): preserve active workers across updates (#11271)
* fix(orchestration): preserve active workers across updates * test(ssh): model absent legacy adoption * test(orchestration): align compatibility contracts * fix(windows): escape updater PowerShell booleans * fix(windows): restore stock uninstall process check * fix(orchestration): keep recovery off renderer startup barrier * fix(orchestration): harden legacy recovery migration * fix(orchestration): close recovery review gaps * fix(orchestration): complete legacy worker cutover recovery * fix(orchestration): preserve legacy workers across updates --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
badf91101b |
fix(quality): enforce performance-safe lint baseline (#11074)
* fix(quality): clear safe existing lint findings * fix(quality): keep lint cleanup allocation-free * fix(quality): enforce performance-safe baseline * test(terminal): drain deferred confirmation cleanup |
||
|
|
6677b5f171 |
perf(cli): construct the runtime client only when a command needs it (#10919)
src/cli/index.ts was the only eager value-import of RuntimeClient, and five other eager modules imported just RuntimeClientError / RuntimeRpcFailureError from the runtime-client barrel -- dragging in client -> pairing -> zod -> ws -> e2ee on every invocation. Those error classes live in runtime/types.ts, which has zero children, so the five imports now point there and the client loads through the existing (already lazy by design) ctx.client getter. Eager modules 199 -> 46, with node_modules dropping 94 -> 0. `orca --help` 2.04x (59.6 -> 29.2 ms); the same for help, no-args, and both error paths, which return before constructing a client. Commands that DO construct one still gain 1.10-1.12x from not eagerly parsing the transport the local path never uses. Correction to an earlier note: websocket-transport alone is ~24 modules / ~8 ms, not the 107 / 28 ms once recorded -- that figure wrongly charged it for zod, which enters through shared/pairing on a different edge. Marginal cost, never isolated cost. Co-authored-by: Orca <help@stably.ai> |
||
|
|
24706ccff0 | fix(terminals): negotiate explicit close intent for paired runtimes (#10129) | ||
|
|
cd05f2ff93 | Implement robust orchestration primitives and connected-server workers (#9925) | ||
|
|
76b2a3b44d |
fix(cli): bound orchestration ask timeouts (#10689)
* fix(cli): bound orchestration ask timeouts * fix(cli): harden remote timeout boundaries |
||
|
|
9ae8f340ae |
fix(cli): explain SIGABRT serve exits instead of naming the signal (#10464)
* fix(cli): explain SIGABRT serve exits instead of naming the signal (#10461) `orca serve` reported only "Orca serve exited via SIGABRT", which sent a P0 investigation down a code-signature path while a diagnostic crash report sat unread on disk. On darwin + SIGABRT the signal-exit path now names the macOS application-startup abort, its usual sandbox/SSH/CI causes, and points at ~/Library/Logs/DiagnosticReports/Orca-*.ips via the existing nextSteps channel. Other platforms and signals get a clear message with no invented cause. * fix(cli): stop asserting the SIGABRT exit happened at startup * fix(cli): stop steering macOS SIGABRT users away from SSH serve |
||
|
|
aab112933e |
Revert "fix(memory): bound OOM-prone accumulators (#10179)" (#10255)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
8f40ddf328 | fix(memory): bound OOM-prone accumulators (#10179) | ||
|
|
0326594d52 | Update paired Orca servers from the active client (#9839) | ||
|
|
34c160442f | Fix headless Linux serve pairing readiness (#9785) | ||
|
|
1fef1e1ddd | Relaunch macOS orca serve safely after updates (#9634) | ||
|
|
1d2aaf1bf5 |
Fix recipe serve desktop promotion (#8646)
* fix(runtime): preserve terminals during headless desktop activation * rm design doc * Fix desktop activation launch ordering and blocked-window status resolut - Check desktopWindowStatus before spawning the Orca app so a blocked runtime no longer launches a doomed second instance. - Reuse resolveDesktopWindowStatus for remote runtime status so it honors the same authoritativeWindowId fallback as local status. - Re-check the authoritative window at spawn time instead of trusting a possibly-stale snapshot, since it can be destroyed mid-await. - Harden the e2e activation spec against silent spawn failures. --------- Co-authored-by: bbingz <zzb@gxsmjx.com> |
||
|
|
73d83a9fb4 |
fix(cli): stop ELECTRON_RUN_AS_NODE leaking into orca claude-teams child (#8513)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
a5faf19631 |
fix(cli): wait for valid serve recipe JSON (#8361)
* fix(cli): wait for valid serve recipe JSON * fix(cli): harden recipe output diagnostics Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Siddharth Ahire <siddharth@Siddharths-MacBook-Air.local> Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
60c26af8b8 |
fix(linear): preserve mixed-version RPC filtering compatibility (#8192)
* fix(linear): guard mixed-version RPC filtering * fix(linear): surface filter capability failures correctly Prevent capability checks from pinning to rejected compatibility cache entries, and rethrow typed attribute-filter unsupported errors from the Linear store so TaskPage can show an upgrade message instead of an empty filtered list. * fix(runtime): refresh cached capability verdicts * test(linear): mock isLinearIssueAttributeFilterUnsupportedError Prevents the invalidation slice test from failing after the runtime client gained this export, which was otherwise undefined in the mock. * Fix cold-cache capability probes firing duplicate status.get calls Coalesce concurrent status.get requests for the same environment by publishing the in-flight probe to the compatibility cache before awaiting it, so parallel capability checks share one RPC call. On failure, drop the cache entry immediately since this probe always re-fetches and must not leave a stale cached verdict. |
||
|
|
e2b4bc2c2c |
feat(cli): make the CLI self-correcting and self-describing for agents (#6303)
* feat(cli): make the CLI self-correcting and self-describing for agents Agents build a generalized model of how CLIs work and apply it to every tool. When orca diverged — `rm` where git uses `remove` — a reasonable first guess (`orca worktree remove`) dead-ended on a bare "Unknown command" with no path forward. This makes the CLI degrade gracefully when the orca-cli skill isn't loaded in context. - First-class CommandSpec.aliases, resolved to the canonical path before dispatch (no new handler registrations). `worktree remove`/`delete` now resolve to `rm`; the ad-hoc `terminal focus` duplicate spec/handler is migrated onto the mechanism. - Did-you-mean suggestions on unknown commands and unknown flags, ranked by edit distance over the live registry, surfaced in both stderr and --json error.data (reusing the existing nextSteps channel). - `orca agent-context [--json]`: a versioned, machine-readable dump of the command schema. Pure local read (no RPC), so it works over SSH and when the app isn't running. - CI guards: specs<->handlers parity, and a vocabulary policy that fails on new off-policy deletion/read verbs (existing ones grandfathered). * Address PR review feedback (#6303) - agent-context now emits each command's effective flag set (globals + conditional --page), not just allowedFlags, so the schema no longer under-reports --json/--help. Shared as effectiveAllowedFlags() between validation and the schema. - Collision check now covers alias paths too, so a duplicate alias that would silently shadow a real command fails the build. * fix(cli): harden agent recovery and introspection Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
39964149c8 |
Per-Workspace Environments (on-demand disposable runtimes) + Add Project remote host setup (#6320)
Co-authored-by: Orca <help@stably.ai> |