mirror of
https://github.com/stablyai/orca.git
synced 2026-10-03 16:02:11 +00:00
cd8d03bc06398b856f6db734aef11df4e2afe15a
61
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6e7e964705 |
feat(orchestration): tell each agent its own orchestration address (#22636)
* feat(orchestration): report the caller's host-resolved orchestration address in orca status orca status --json gains a caller block: the calling agent's address as the host resolved it from the identity its environment carries. A structured session is session:<id>; a terminal agent is its handle, with whether the host still knows it. A session the host refuses reports that refusal instead. The host answers through a new read-only orchestration.callerShow, so the session claim runs through the same dispatch-entry resolver every verb uses. An older host leaves caller unresolved. The help footer and the run/check specs stop describing identity only in terminal terms. * docs(orchestration): tell agents their address and give chat coordinators a non-waiting loop The orchestration guide now states that a chat session's address is session:<id> (never the provider's id), that orca status --json reports it, and that no caller flag should name another agent. A consuming check no longer tells every caller to name itself with --terminal. A chat coordinator starts its wave, ends the turn, and on each turn Orca starts for new mail runs a non-waiting check and ack; it never blocks in check --wait. The guide also names ORCA_CLI_COMMAND as the executable in chat sessions. * feat(native-chat): add Copy Orchestration Address to a structured chat's context menu Copies session:<id>, the Orca-minted address other agents message the chat by. The existing Copy Session ID still copies the provider's id and is left as is; the new action is labelled so the two cannot be confused. Strings are added to every locale catalog. * feat(orchestration): tell every dispatched worker its own orchestration address The worker preamble names the coordinator's address rather than a terminal handle, and states the worker's own address. A structured worker is told it is session:<id>, that its coordinator reaches it there or at its dispatch mailbox, and that mail arriving while it is idle starts a new turn. Its commands invoke the CLI through ORCA_CLI_COMMAND in its own shell's form, the same rendering the pointer turn uses, because a bare orca in a login shell can reach a different Orca. * docs(orchestration): give the ORCA_CLI_COMMAND form for POSIX shells and PowerShell A chat session's shell reads the variable as "$ORCA_CLI_COMMAND" in a POSIX shell (Git Bash included) and as & $env:ORCA_CLI_COMMAND in PowerShell, the same two forms the pointer turn and worker preamble render. The chat coordinator loop now runs the check its pointer turn names. * docs(orchestration): say that /clear gives a chat a new address and Orca moves its Runs * fix(orchestration): keep CLI resolution in the shared skill stub and the orchestration kernel in budget The guide-contract tests own two rules this PR broke: only the shared skill stub may describe how to resolve the CLI, and the always-loaded orchestration kernel stays within 202 lines. The ORCA_CLI_COMMAND text moves to the stub's resolver block, which now covers chat sessions and login shells beside WSL and gives the POSIX and PowerShell forms; every skill projection and the bundle manifest are regenerated. The kernel keeps one line each for the caller's address, the environment-resolved check caller and the chat coordinator's non-waiting loop; the loop steps and the address details move to the coordinator-loop and messaging references. The two kernel pins now assert the new check contract and refuse the old --terminal <your_handle> shape. * fix(orchestration): refuse a blocking check --wait from a native chat session A chat runs turn by turn through a shell tool with its own timeout, so a blocking wait is killed mid-wait and retried. The host now refuses it with wait_requires_terminal and the turn-loop recovery, keyed on the session's lease: a session a terminal view holds still runs in a PTY and may block. * fix(orchestration): resolve orca status's caller with the verbs' ladder, host-side callerShow now answers a terminal caller the way the coordinator verbs act: the carried handle while it is live, else the handle its pane was reminted as. The CLI always asks, so the host decides that a process has no identity from the same envelope every verb sends; a pane key alone now resolves. * fix(orchestration): show a structured worker as session:<id> wherever agents read mail A structured worker was session:<id> in orca status and its preamble, but structworker_<uuid> in check rows, banners, reply hints, its own check label and a sub-worker's coordinator line. The minted handle is now only the mailbox key: mailbox reads, the check label and preamble coordinator lines spell the worker session:<id>, which the host binds back to that mailbox. Send receipts still echo the stored row, whose sender key worker_done settlement matches. * fix(orchestration): teach a chat worker the turn loop and pin preamble parity at the contract The worker preamble was byte-identical across modes except its address, so a chat worker was taught a 600s blocking ask its shell tool kills before the message ID for --resume prints, heartbeat exemptions for check --wait, and to keep a shell open. Parity now pins the contract (sections, verbs, flags, lifecycle ids); interaction discipline follows the mode: a chat asks with a 5s wait and ends its turn, owns sub-workers through the turn loop, and names itself session:<id> in every command. The guide says a chat's address survives /clear and that Orca refuses a chat's check --wait. * test(orchestration): pass the db to preamble delivery and fence a terminal-view waiter The coordinator line maps a structured coordinator's handle through the orchestration db, so delivery takes it from its caller. The consumer-fencing waiter test now waits as a terminal-view session, the only session kind that may still block in check --wait. * test(orchestration): read the Run id with the fixture's checked accessor * feat(orchestration): copy a chat's conversation address, which /clear keeps Copy Orchestration Address copied session:<live id>. A chat's address is its conversation's, derived by the host from the session records, so the menu now asks the host for it at copy time through orchestration.sessionAddress, the same derivation a verb acting as that session binds to. A host that predates the method has no /clear lineage, so there the live id is the address. The guide's /clear text says the address survives and nothing moves. * test(orchestration): pin that a cleared chat's successor copies its conversation's root address * test(orchestration): give the mode-opacity fixture's record store the listing a lineage lookup reads A structured worker's agent-visible address now resolves through its conversation's lineage, which lists the session records; the fixture's partial store lacked that listing, so the sub-worker start failed at dispatch input. * refactor(orchestration): format a chat's copied and reported address from its root Orca session id Carries the Orca session id rename into the self-address surfaces. orchestration.sessionAddress, the copy action's fallback, callerShow and the address a structured worker is shown now format `session:<id>` from the conversation's bare root Orca session id with formatOrcaSessionAddress, and ids arriving as strings are checked with isOrcaSessionId first. The CLI status line, the check caller label and the dispatch preamble spell the prefix from the one exported constant. * refactor(orchestration): resolve a session's reported address through the party resolver, and refuse every session's check --wait - orchestration.sessionAddress, and the agent-visible spelling of a structured worker, resolve through the party resolver, so they format the lineage root the one id hook derives; sessionAddress.sessionId is classified as a target. - With the terminal handoff gone every structured session runs turn by turn, so check --wait is refused for any session caller, a worker included, and the session caller no longer carries its lease's runtime kind. - The coordinator loop no longer mentions a terminal view, and the messaging reference says a chat takes messages but is refused as a Dispatch assignee. * fix(orchestration): cap a session caller's blocking wait below its shell tool instead of refusing it A chat or structured worker runs each command under its provider's shell-tool timeout, so check --wait was refused for every session caller and chats were taught a separate loop. The host now caps check --wait and ask for a session caller below that timeout (Codex 10s one-shot exec default, Claude Code Bash 120s) and answers the normal timed-out result, so the terminal coordinator loop runs unchanged in a chat. A terminal caller's wait is untouched. * refactor(orchestration): teach a chat worker the terminal worker's preamble, byte for byte but the address One preamble for both modes: the chat variant (short ask, end your turn, this chat stays available, ORCA_CLI_COMMAND invocation) is deleted. A structured worker's only difference is its address, session:<id>; the byte-parity test between modes is restored with just that substituted. * docs(orchestration): drop every chat-specific instruction; name the address once, generically The guide, its references, the shared CLI-resolution stub and the help return to main's text, with one kernel line saying `orca status --json` shows your address (the kernel stays at main's length). The status caller block reports only the opaque address, the same shape for a chat and a terminal agent. Guides regenerated. * test(orchestration): pin that a chat and a terminal agent see the same preamble, pointer and guide * test(orchestration): key the wait-cap fixture's records by plain session id strings * chore(i18n): add the copy-address strings at the head of native-chat, clear of main's catalog edits * test(orchestration): fail the capped-wait test on the settle, not on the test timeout * test(orchestration): keep main's takeover assertions on a session coordinator's waiting check With the session wait capped rather than refused, the test main extended runs as it is: the restack re-added the shorter pre-main version over it. * fix(orchestration): show a /clear-ed chat its lineage root's address everywhere it reads its own check labelled a session caller with session:<live id>, while orca status and the preamble show the conversation's root. The CLI cannot read the lineage, so the label now comes from the same host answer orca status prints (orchestration.callerShow), asked only when there are messages to render, and falling back to the live id only when the host cannot say. The host also spells a session address it shows an agent with the lineage root: a dispatch preview filled in from the chat's own address, and the provider-id refusal that names a session's address. * test(orchestration): the parity test's gate facts resolve like the host's * test(orchestration): the parity test's gate facts carry the submissions main's pointer lane reads * fix(orchestration): wait a chat's check --wait and ask exactly as long as a terminal's The host capped a session caller's blocking wait (Codex 6s, Claude 100s) so the provider's shell tool would not kill it. Neither provider kills a long shell call: default Codex's exec tool yields and keeps the command running, and Claude Code moves a timed-out Bash call to the background. Terminal agents run the same tools uncapped, so the cap only made a chat coordinator re-poll every few seconds. A session caller's check --wait and ask now wait the budget asked for. * fix(orchestration): show every agent one address, the mailbox address its mail is keyed by A structured worker was told `session:<id>` in orca status and its preamble, but its own send receipts, inbox, worker-list, dispatch previews and task rows still showed the `structworker_` handle its mail is stored under; only some reads were re-spelled. Instead of re-spelling reads, callerShow, sessionAddress and the preamble now report the caller's stored mailbox address (mailboxAddressOf): a terminal's handle, a structured worker's handle, and a chat's `session:<lineage root>`. The read-side re-spelling layer (withAgentVisibleAddresses and its check/banner/preamble call sites) is gone. Dispatch previews spell the coordinator by its party's mailbox address, so a `/clear`ed chat's dispatch-show still names its root. * refactor(orchestration): label check output from what the CLI already knows check asked the host for orchestration.callerShow after every non-empty check by a session, only to fill a label used when a legacy row lacks to_handle, which host rows never do. The label is again the caller's handle or its injected mailbox address, with no second round trip after mail is consumed. * refactor(native-chat): offer Copy Orchestration Address on chat tabs only No mount passes both terminal-pane actions and an orchestration address: a chat shown inside a terminal pane is that terminal's agent, copied by its terminal ID. Drop the unreachable terminal-pane placement and its tests. * fix(orchestration): have orca status report the handle the agent's own check reads After a window reload a terminal agent keeps ORCA_TERMINAL_HANDLE=term_old while its pane is reminted as term_new. callerShow reminted and advertised term_new, but check, send and ask act as the carried handle and never remint, so mail sent to the advertised address was never read by that agent. callerShow now answers the carried handle with its liveness, and null for a pane key alone, from which the mailbox verbs have no identity. Resolving terminal callers once on the host for every verb is a separate follow-up. * fix(orchestration): read the renamed coordinator line in the long-prompt repro, and trim round-one leftovers The reliability repro's fake worker parsed "Your coordinator's terminal handle is:", which the preamble now spells "Your coordinator's address is:", so it silently skipped worker_done; it accepts both. Dispatch and its dry-run go back to main's coordinator line (their `from` is already bound at the entry); only dispatch-show, whose `from` is unbound, resolves it. Also drops a stale status-caller comment, trims the wait test to its one uncapped-wait case, and reverts comment-only churn in the worker opacity test. * docs(orchestration): keep worker obligation 1 as main words it The guide grows by the one caller.address line; the parity test bounds the kernel at main's length plus that line instead of forcing a reword. * test(native-chat): prove a structured chat tab offers Copy Orchestration Address Renders the pane-commands hook as a structured chat tab and selects the item: it asks orchestration.sessionAddress with the tab's target and session id. Also corrects the menu item's comment to what it copies. * test(orchestration): D5's tests expect the orca_session_id prefix and 'Orca session ID' wording * fix(orchestration): name a session by its Orca session ID, and leave terminal agents as main has them Terminal agents keep main's exact wording: a terminal worker's preamble is byte-identical to main's, and orca status prints nothing new for them. A session is named by its Orca session ID (orca_session_id:<id>, its /clear root's): a structured worker's preamble says "Your Orca session ID is: …" and its commands use that ID, and a session coordinator is "Your coordinator's Orca session ID is: …". orca status shows a session caller's `caller.orcaSessionId`; callerShow answers null for anyone else. The chat tab menu item becomes "Copy Orca Session ID" with a tooltip saying what the ID is, and its toasts match. No agent-read text calls this ID an address. A structured worker's mail is still keyed by its minted handle. * test(orchestration): check CLI help and status for "address" wording from a CLI test The node project cannot compile src/cli, so the guard over CLI help, specs and status text moves to src/cli; both halves share one pattern. Also brings two comments and the long-prompt repro's coordinator-line regex to the Orca session ID wording. * fix(native-chat): keep the Orca session ID tooltip within the tooltip primitive's typography Drops a restyle the design-system gate refuses on TooltipContent, keeps "Agent" untranslated in the Japanese tooltip as that catalog does, and types the test's tooltip mock without an assertion. * fix(native-chat): the Orca session ID tooltip names the agent CLI's own session ID in the singular |
||
|
|
3047353017 |
fix(orchestration): worker-abandon settles a stuck worker and records who did it (#23983)
* fix(orchestration): worker-abandon settles a stuck worker and records who abandoned it worker-abandon refused or no-oped in the states it exists to escape: a stop stranded by a dead runtime, an active attempt that was no longer the Task's latest, and settled workers whose terminal release was stuck at requested or unknown. It now settles every non-terminal worker except this runtime's own in-flight stop, records who abandoned it, and retains (never closes) an owned terminal whose release is not already in flight. Part of STA-8833. * refactor(orchestration): abandon retains through worker-retain's rule; record cancellations - One retain helper serves worker-retain and worker-abandon. A committed release (releasing, unknown) keeps its state and archive, since the tab may already be closed; a retained terminal drops its stale archive. - worker-abandon keeps the published stale field (this attempt was not the Task's current one) on the worker path. - task-list shows a failed Task's reason, whitespace-collapsed; task-update help and the recovery guide document cancel = failed + --result cancelled. * test(orchestration): drop duplicate abandon asserts; leave task-list output unchanged Cancellation stays documented in task-update help and the recovery guide. * fix(orchestration): abandoning an already-settled worker changes nothing Only the settling path retains an owned terminal. * fix(orchestration): attribute abandon only to a verified caller and keep the prior diagnostic * fix(orchestration): report an already-settled abandon as stale, as main did |
||
|
|
9afd1101ff |
fix(orchestration): stop minting and printing the dispatch capability (#23994)
* fix(orchestration): authorize worker reports without the dispatch capability Worker lifecycle reports and questions no longer depend on the per-dispatch capability token that lives only in the agent's conversation. The host now: - ignores capability_hash/capability_revoked_at for authorization on every row and checks the exact worker process instead (ask gains that check); - refuses a report whose calling terminal is provably another orchestration party (a Run coordinator or another Dispatch's worker), treating env that names no live pane here as absent; - applies one worker-state rule locally and remotely: a stop in flight refuses, while stop_unknown and start_unknown accept and settle. Minting and printing the flag are unchanged, so an older host and older preambles keep working. * fix(orchestration): stop minting the dispatch capability Dispatches no longer mint a per-Dispatch token, and preambles, the bundled skill guide and the ask resume hint stop printing --dispatch-capability. The consumer-generation bump and delivery fence that minting carried stay, now as setDispatchConsumer. Readers that inferred meaning from capability_hash read what they meant instead: worker-show's injected stage comes from the attached consumer, and a failed start copies custody identity only when no authority was ever attached. The CLI keeps accepting and forwarding the flag for older hosts. Cancelling a Task is recorded as failed with a reason; task-list now shows that reason and the guide and task-update notes document the recipe. * fix(orchestration): name the fenced party without implying which Dispatch it owns * refactor(orchestration): one worker report rule, fence only a different party - One module owns the unproven/settleable worker states and the refusal rule; local send records it, ask and remote throw it. A stale process is worker_identity_changed on every path. - The caller fence passes the worker's own terminal when its --from handle went stale. - Document the shared-tmux-server limit; drop the dead dispatch_capability_invalid rejection member; tests assert dispatch state, not the capability column. * refactor(orchestration): drop setDispatchConsumer and the dead capability retention - dispatch --inject no longer re-points the row createDispatchContext just wrote; worker-show reports every worker-less Dispatch as context_only, since Orca keeps no record of the paste. Tests re-point through a fixture. - failWorkerStart always records when the lifecycle closed; nothing authorizes on it. - Restore the ask resume hint's echo of a passed --dispatch-capability: an old host checks it before --resume. - Move the cancellation convention to #23983. * test(orchestration): drop capability-era assertions other tests already cover * refactor(orchestration): drop the host-side capability field and no-op test fixtures - RpcRequest and the SSH bridge stop carrying orchestrationCapability; the CLI's wire field stays for older hosts. - Fixtures pass identity to createRootDispatch instead of re-pointing to the same values; drop absence checks for a flag that can no longer be produced. * test(orchestration): cover a current process whose terminal moved to another pane * chore(orchestration): finish the capability cleanup in test stubs and skill wording * test(orchestration): drop needless response casts; mark the db stub cast safe |
||
|
|
e03870403e |
fix(orchestration): accept worker reports without the dispatch capability (#23982)
* fix(orchestration): authorize worker reports without the dispatch capability Worker lifecycle reports and questions no longer depend on the per-dispatch capability token that lives only in the agent's conversation. The host now: - ignores capability_hash/capability_revoked_at for authorization on every row and checks the exact worker process instead (ask gains that check); - refuses a report whose calling terminal is provably another orchestration party (a Run coordinator or another Dispatch's worker), treating env that names no live pane here as absent; - applies one worker-state rule locally and remotely: a stop in flight refuses, while stop_unknown and start_unknown accept and settle. Minting and printing the flag are unchanged, so an older host and older preambles keep working. * fix(orchestration): name the fenced party without implying which Dispatch it owns * refactor(orchestration): one worker report rule, fence only a different party - One module owns the unproven/settleable worker states and the refusal rule; local send records it, ask and remote throw it. A stale process is worker_identity_changed on every path. - The caller fence passes the worker's own terminal when its --from handle went stale. - Document the shared-tmux-server limit; drop the dead dispatch_capability_invalid rejection member; tests assert dispatch state, not the capability column. * test(orchestration): drop capability-era assertions other tests already cover * test(orchestration): cover a current process whose terminal moved to another pane |
||
|
|
52a1e2875b |
feat(orchestration): accept Muse model and effort for supervised workers (#22383)
* feat(orchestration): accept Muse model and effort for supervised workers `worker-start --agent muse` already launched, but `--model` was refused because Muse had no session-option catalog. Add one that maps worker preferences to `muse --model <id>` and `--reasoning-effort <level>`; it seeds no models, so native-chat surfaces show no picker. opencode stays without `--model`: the opencode 2 TUI (now shipped as `opencode`) rejects the flag, so the refusal now tells callers to rely on the agent's own config. Help, skill guide, and docs list valid `--agent` ids and the agents that accept `--model`. Refs #19823 * test(mobile): repin session route closure for the Muse option catalog |
||
|
|
eb92222e7f |
feat: support Antigravity as supervised worker (#21705)
* feat: add supervised Antigravity worker support * fix: address Antigravity worker review findings * fix: stabilize Antigravity readiness detection * fix: allow Antigravity resume footer after readiness * fix(antigravity): make agy reach worker_done as a supervised worker Three defects each blocked `orchestration worker-start --agent antigravity --worktree new-child` at the agent_readiness stage. 1. Readiness never fired. The composer check required the trimmed line to be exactly one character, but agy 1.2.7 launches in accept-edits mode and paints it into the caret row (`> Accept-edits mode: ...`). Widened narrowly to a bare `>` or `> <name> mode:`; matching any `> <text>` would make every menu dialog read as ready, since they all prefix their highlighted row the same way. 2. No trust artifact for agy. Added markAntigravityWorkspaceTrusted, writing ~/.gemini/antigravity-cli/settings.json under `trustedWorkspaces` — verified empirically against agy 1.2.7, and distinct from the Gemini CLI's trustedFolders.json, which agy does not consult. Trust is exact-path and not inherited by subdirectories, so each child worktree needs its own entry. 3. The orchestration path skipped the preset. Orca has two trust dispatch chains: the renderer's preflightAgentTrust and the main-process markLocalWorktreeTrusted. worker-start only takes the second, which matched cursor/copilot/codex and fell through for antigravity, so the trust write never happened while renderer-side tests passed. Verified live end to end: the dispatch settles `succeeded` with worker_done carrying the right task and dispatch ids, and the worktree is appended to agy's settings with sibling keys untouched. Known gap: remote-agent-trust-presets.ts has no antigravity branch. The SSH artifact path is unverified, so agy over SSH still stalls at agent_readiness. Recorded in a comment there rather than guessed at. * fix(antigravity): wire trust preset through preload safely * fix: preserve Antigravity readiness across transcript tails --------- Co-authored-by: Neil <neil@stably.ai> Co-authored-by: LielinaH <lielinah@gmail.com> |
||
|
|
3336933cc8 |
fix(orchestration): list worker Dispatches newest first and warn when the page truncates (#21523)
* fix(orchestration): list worker Dispatches newest first and warn when the page truncates `worker-list` paged `ORDER BY d.rowid ASC` with a 100-row cap, so a Run with more than 100 Dispatches answered with its OLDEST 100. The workers a coordinator had just started, and the rows carrying `projection.attention.requiresAction`, were on a page nobody fetched, while `counts` and `page.total` covered the whole Run so the receipt read as complete. One ordering, flipped: the detail query and the terminal-state scan it pages by both order `d.rowid DESC`, and the cursor fence walks down (`d.rowid < anchor`). The snapshot fence is unchanged — `d.rowid <= snapshot` still means "nothing created after the first call". When the page truncates the receipt now carries a `warnings` string, the same shape `worker-output` already uses, alongside `page.hasMore`. Text output keeps its `More: --cursor` line and prints the warning through the block it already had for partial-host errors. Refs STA-7861 * fix(orchestration): make the worker-list truncation warning true on every page The warning said "Showing the N newest of T Dispatches" unconditionally, but `hasMore` is true on every page except the last, so page 2 of a 300-Dispatch Run claimed to be the newest 100 while showing rows 200..101. This PR exists because a receipt read as complete when it was not; that warning shipped a receipt that read as the newest page when it was not. The page count and the ordering are separate facts, so state them separately: "Showing N of T Dispatches, newest first; more are on later pages." True on page one and page N alike, no extra state. The 105-row case only ever reached the last page, where `hasMore` is false, which is why it missed this; a new case walks 6 Dispatches at `--limit 2` so a page that is truncated AND not page one is covered. Also: the `worker-list` --help note and the recovery-and-cleanup reference still described the oldest-first contract; both now say newest first. The snapshot test is renamed to the property it actually proves — under DESC a later insert is unreachable by arithmetic, so what the `d.rowid <= snapshot` fence still earns is pinned `page.total` and `counts`, not row exclusion. The continuation comment says "below the anchor" next to `d.rowid < ?`, and the two SAFETY rationales now say what they are: an unchanged cast the gate flagged because the diff moved inside its span. Refs STA-7861 |
||
|
|
12d744f253 |
fix(skills): keep computer-use off filesystem and shell tasks (#21069)
* fix(skills): keep computer-use off filesystem and shell tasks STA-7615: "On my desktop create a folder" was matching computer-use because discovery copy said OS/window-level and neighboring skills advertised desktop UI. Scope the trigger to visible GUI with no CLI path, and exclude files/folders/git/shell. * fix(skills): prefer programmatic paths over computer-use State the last-resort rule in discovery copy instead of enumerating files/folders/git/shell. computer-use prefers shell, filesystem, git, HTTP, CLIs, and Playwright/CDP; neighboring skills route to Computer Use only when a visible window needs GUI control those cannot do. * fix(skills): stop advertising computer-use from orchestration Orchestration coordinates workers; it does not drive a GUI. Drop Computer Use and Playwright/embedded-browser routing from its discovery description so those tools are not pulled in from a coordination skill. * fix(skills): drop Playwright from orca-cli discovery orca-cli should not prescribe Playwright or CDP. Those tools may not be installed, and page automation is not this skill's job. * fix(skills): drop the page-only ban from computer-use discovery Page automation is a preference, not a prohibition. If Playwright or CDP is not available, a visible browser window is valid Computer Use. Keep the hard split for Orca's embedded browser (`orca-cli`) only. |
||
|
|
3631a1e77f |
feat(cli): show orca search now the settings toggle ships (#20677)
* feat(cli): orca search over the agent session index `orca search <query>` calls PR 5's `aiVault.searchSessions` over the CLI's existing runtime RPC, against the host `--environment` / `--pairing-code` selects and no other. `orca search --index-status` calls `aiVault.searchStatus`. It is the proof the contract works with no panel. Every flag maps onto a contract field and nothing else: `--scope`, `--fresh`, `--limit`, `--cursor`, repeatable `--agent` and `--path`, `--since`, `--sort`, `--debug`, `--json`. No fan-out, no merged output, no `--host`. One command rather than a `search status` subcommand: the query is a bare positional, so `orca search status` could not be told apart from searching for the word "status". `--status` is unavailable because `orchestration task-list --status <state>` already owns the name as a valued flag. No new runtime capability. PR 5 decided an explicit `method_not_found` refusal maps to `unavailable/no-service`, so reusing `createSessionSearchClient` gives an old host a plain "this host runs no session search service" answer at exit 0 instead of a raw JSON-RPC error. `CommandSpec.repeatableFlags` scopes repeatability per command, because `--agent` must repeat for search and stay single-valued for `worktree create`. `help.ts` sat exactly at max-lines, so `skills-command-flag-help.ts` becomes `command-scoped-flag-help.ts` carrying both tables at the same call-site size. * refactor(cli): drop the search type assertions main's casting gate now rejects Main gained a `consistent-type-assertions: never` scan in the changed-code gate after this branch was cut, and it reported twelve assertions in the new files. The four in the argument parser were avoidable. `readEnum` now keeps the value `find` returns, which already carries the narrow type, and the agent filter goes through an `isAiVaultAgent` predicate over a `Set<string>` instead of widening the agent tuple. The test now narrows the printed envelope by shape and re-reads the printed result through `AiVaultSearchResponseSchema`, so the JSON assertions are checked rather than claimed, and the flag table is typed so its callback needs no cast. One assertion is left, for the structural fake client, with the SAFETY rationale AGENTS.md requires. * fix(cli): sanitize host strings and scope pre-command repeatable flags Route every host-supplied string the search formatter prints through the escape stripper, and resolve the repeatable-flag set from the command tokens ahead when a flag sits before the command. * refactor(cli): resolve repeatable flag rules once per command * fix(cli): clarify session search availability and SSH scope * feat(cli): hide orca search until the settings toggle ships `orca search` stays dispatchable but leaves every discovery surface: root help, group help, unknown-command suggestions, and `agent-context --json`. `buildAgentContext` did not filter hidden specs, so it also stops leaking the hidden `terminal stop`. * feat(cli): show orca search now the settings toggle ships * docs(skills): teach the orca-cli guide the search command One section: what orca search covers, one host at a time, scope and narrowing flags, index status before searching, and that a human turns search on. * docs(skills): shape the search section like the other command sections |
||
|
|
4b1b7178ad |
fix(orchestration): scope @ group addresses to the sender's Run (#19783)
* fix(orchestration): scope @ group addresses to the sender's Run `@all`, `@idle`, and the agent-name groups (`@claude`, `@codex`, ...) resolved against every terminal on the host. A coordinator meaning "my three reviewers" reached 126 agents across every open project, twice in one day, and every unrelated agent burned a turn discarding mail that was never for it. Every group except `@worktree:<id>` now means the live Dispatches of the sender's own Run, each addressed as `dispatch:<id>` so delivery is durable even when the worker terminal is not attached yet. A sender bound to no Run is refused with `invalid_argument` naming `run:<id>` / `dispatch:<id>`; there is no host-wide fallback and the host's terminals are never enumerated for it. `@idle` and the agent-name groups filter within that set by the same terminal status and host-resolved identity as before. `ask --to @group` returns the same code and points at the owning Run mailbox. Federated Dispatches read relayed control mail rather than a local mailbox, so a Run-scoped fan-out skips them with a `recipient_unreachable` warning naming the direct `dispatch:<id>` address. Group addresses are resolved host-side, so no RPC or stream shape changes; an older CLI sending `@all` to a new host gets the Run-scoped meaning. Claude-Session: run-scoped-group-addresses * fix(orchestration): revalidate legacy takeover before the recipient verdict A legacy coordinator taken over while `listTerminals` was in flight reported `runtime_error` instead of `legacy_read_only`: Run scoping made "no live Dispatch in this Run" the first thing the group send could fail on, and that threw before the takeover check ran. Takeover is a precondition, not a commit-time detail — the sender must be told it is read-only whatever else is wrong with its recipient set. Revalidation moves to immediately after the only `await` in the path. Everything below it is synchronous, so the commit-time window it used to guard is unchanged; only the error paths now see it. The legacy partition test gave `term_current_worker` no Dispatch, so under Run scoping it is correctly not a recipient. It now holds a real current-contract Dispatch in the same adopted Run, which is what the test is named for: one `legacy_direct` and one `current_delivery` recipient in one fan-out. Claude-Session: run-scoped-group-addresses * fix(orchestration): address the Run a nested coordinator created, not its parent A nested coordinator is both a worker of its parent Run and the coordinator of the Run it created. `resolveMessageRun` answers with the parent, correctly, because that is where its own `worker_done` belongs — but audience is a different question. Scoping `@all` to that Run sent a nested coordinator's "shared context" to the siblings it was started beside instead of the workers it started, and reported success, so it never learned its sub-workers heard nothing. Before Run scoping the host-wide fan-out reached the sub-workers by accident; this turned an over-broad delivery into a wrong-audience one, the exact failure class the change exists to remove. Group audience now resolves off the Run the sender coordinates, falling back to its Dispatch's Run. A leaf worker coordinates nothing and is unaffected. This is a separate question from `routing.run`, not a second answer to the same one, so `resolveMessageRun` keeps its meaning for point-to-point mail. Also: when every live Dispatch in a Run is federated, the fan-out skipped them all and threw a bare `Error` that discarded the warnings naming those remote workers and how to address each one. The sender was told "no recipients" while three remote workers existed. That throw now carries a code and the skip explanations. Claude-Session: run-scoped-group-addresses * docs(orchestration): say that no group address reaches a coordinator A coordinator is not a Dispatch, so Run-scoped groups never include one. That follows from the rule, but nothing said it, and the old host-wide meaning did include the coordinator — a worker sending `@all` to raise a blocker would be heard by its siblings and by nobody who can act. The guide, the CLI note, and the docs page now say to use `run:<id>` for that, and that a worker which created its own Run addresses that Run's workers. Also restores the `@cursor` case dropped when the group tests moved: a Claude pane titled "Fix the text cursor blink" must not receive Cursor's mail. That hazard was recorded from real titles and `@droid` alone did not cover it. Claude-Session: run-scoped-group-addresses * fix(orchestration): preserve group audience and mailbox identity * fix(orchestration): validate group scope before dispatch routing * fix(orchestration): preserve pane identity and exclude coordinator dispatches |
||
|
|
98fdbc4ade | fix(orchestration): file federated worker mail under the coordinator Run (#19542) | ||
|
|
9fed61e5c2 |
Persist agents sidebar search visibility as pairing-local preference (#19313)
* Persist agents sidebar search field visibility as pairing-local preferen - Add `agentsShowSearch` to workspace UI state with default on - Include in pairing-local fields so preference syncs across clients - Convert search from menu action to checkbox menu item for explicit toggle - Update activity thread options menu to reflect checkbox state - Add localization strings across all supported languages - Update RPC schemas and preference persistence layer - Includes readiness validation reports confirming feature is clean * rm review * fix documentation |
||
|
|
fb322046e8 |
skills: rewrite and trim the seven non-orchestration guides (#19128)
* skills: rewrite the seven non-orchestration guides to one outcome-first standard
Every guide leads with Result / Done / Safe failure, states conditions instead of case lists, keeps one done bar and one autonomy envelope, and loads references at the point of use via `skills get <topic> --full`. orca-cli drops from 424 to 260 always-loaded lines with three references; orca-per-workspace-env from 794 to 397 with five.
Defects fixed in shipped guides: `emulator camera` (no such command), iOS `permissions` (backend refuses it), Android pane described as in development, `relayGracePeriodSeconds: 0` documented as immediate teardown (it is unbounded), doctor `ok: true` hiding `warn`, an SSH exemplar setting both `jumpHost` and `proxyCommand`, a provisioned-root fetch from `origin`, and the Linear unconfirmed-write rule keyed on four verbs when ten emit it.
The resolver ladder, placeholder rule, and older-binary fallback shared by every installable SKILL.md now come from one skill-stubs/_shared/cli-resolution.md fragment composed by the generator, which also bundles per-guide references into --full. New guards: every ORCA invocation and flag resolves against COMMAND_SPECS, descriptions carry no angle-bracket tokens, reference routing is checked both ways, and an always-loaded size ratchet (300 lines) that guides may leave but never join.
* skills: address review on the SSH recipe and the parity guard
- ssh-host create script: route the bootstrap ssh through the chosen jump host or proxy command, refuse both at once, use StrictHostKeyChecking=accept-new instead of a blind ssh-keyscan append, and pass gh_token/project_root/repo_url/repo_ref to the remote bash via printf %q so a quote in a value cannot break out of the command.
- per-workspace-env envelope: the step-10 workspace test the user asked for is no longer forbidden by the same paragraph.
- linear guides: name the full verb, ORCA linear list-issues.
- parity guard: a prefix reference such as ORCA linear --help or ORCA emulator --webcam now has its flags checked against every command under that prefix; only an exact path or an explicit ... was checked before.
* skills: tighten prose in the seven rewritten guides
Shorter outcome spines, one idea per sentence, no restated rationale after a rule. No rule, command, or pinned phrase changes; 47 net lines fewer across the guides and references.
* skills: route orca-cli and per-workspace-env gates through --reference
Both guides told agents to load --full at a gate because the per-reference
selector did not exist when they were written. Now that main serves
`skills get <topic> --reference references/<file>.md`, load only the
named file and keep --full as the fallback for an older CLI, matching the
orchestration kernel.
* skills: drop outcome-spine boilerplate from the CLI-wrapper guides
The Result/Done/Safe-failure preambles and Next Action closers restated
rules the body already carries. Agents stop fine without them, and for
a CLI wrapper the command surface is the guide. Keeps the one substantive
rule computer-use's Done block added (never report unverified as success)
inside Action Rules. orchestration and per-workspace-env keep theirs:
those are multi-step workflows where the done bar is load-bearing.
(cherry picked from commit
|
||
|
|
b8311d509a |
Revert "skills: rewrite the seven non-orchestration guides to one outcome-first standard (#18724)" (#19126)
This reverts commit
|
||
|
|
15d0f8aedf |
skills: rewrite the seven non-orchestration guides to one outcome-first standard (#18724)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 6 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$544 | $\color{#cf222e}{\Huge{\mathbf{−}}}$49 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$495 |
| Prod | 36 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$1719 | $\color{#cf222e}{\Huge{\mathbf{−}}}$1703 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$16 |
<!-- /orca-pr-loc -->
## ELI5
Orca ships eight skill guides that agents read before running the CLI. Seven of them (everything except `orchestration`, which #16904 rewrites) were command catalogs that had drifted from the binary. This PR rewrites them so an agent reads the outcome, the done bar, and the safe-failure rule first, loads reference material only at the step that needs it, and never sees a command or flag the installed CLI does not define.
## What changed
- **Seven guides rewritten** to one standard: outcome spine first (Result / Done / Safe failure), conditions instead of case lists, one done bar, one autonomy envelope, references loaded at the point of use via `skills get <topic> --full`, every runnable invocation spelled `ORCA`. `orca-cli` is 424→260 always-loaded lines with three references (browser, automations, publishing); `orca-per-workspace-env` is 794→397 with five (provider-vercel, ssh-host, docker-ssh, windows-scripts, failure-modes).
- **Defects fixed in shipped guides:** `emulator camera` (no such command), iOS `permissions` (backend refuses it), Android pane described as "in development" (shipped in June), `relayGracePeriodSeconds: 0` documented as immediate teardown (it is unbounded), doctor `ok: true` hiding `warn`, an SSH exemplar setting both `jumpHost` and `proxyCommand`, a provisioned-root fetch from `origin`, the Linear unconfirmed-write rule keyed on four verbs when ten emit it. Linear and emulator descriptions dropped embedded commands and angle-bracket placeholders (651→329, 732→404 chars).
- **Generator bundles references.** `skill-guides/<name>/references/*.md` is appended to `--full`; `skills get` help says compact by default, full with references.
- **Stubs single-authored.** The resolver ladder, placeholder rule, and older-binary fallback shared by all eight installable `SKILL.md` files come from one `skill-stubs/_shared/cli-resolution.md` fragment composed by the generator. Projections were byte-identical before the content fixes.
- **Guards:** every `ORCA <cmd>` and flag in every guide and reference resolves against `COMMAND_SPECS` (this found the camera defect); descriptions ≤1024 chars with no angle-bracket tokens; reference routing checked both directions; an always-loaded size ratchet (300 lines) that guides may leave but never join. `orchestration` (440 lines on main) is recorded as an exception until #16904 lands its kernel.
## Relationship to #16904
Split out of #16904 so that PR carries only the orchestration guide. On main, `terminal send` has no `--wait-submit` / `--retry-request` and the orchestration kernel still carries the resolver ladder and worktree-selector rule, so this branch pins `accepted: true` for handoff receipts and leaves the orchestration pins where main has them. The merge in either direction is mechanical: #16904 rebased on this becomes a one-file `orchestration.md` change plus dropping the two exceptions.
## Standard
Compound Engineering's portable skill-authoring guidance (outcome spine, conditions not cases, pinned fragile commands with an ordered hatch, references at point of use). NVIDIA SkillEvaluator Tier 1 (`schema,pii,license,quality,unicode,lint`) was run on every guide; its deterministic checks pass, its template nudges (Instructions/Examples sections, 50–150 char descriptions) do not apply to Orca's stub architecture and were not applied.
## Testing
- `pnpm typecheck:tsc:cli` clean; `check:code-quality:changed` and `check:react-doctor:changed` 0 findings
- `pnpm verify:bundled-skill-guides` and skill-bundle manifest verify clean
- vitest over `config/scripts`, `src/cli/skill-guide-cli-parity.test.ts`, `src/cli/skills.test.ts`, `src/cli/specs/skills.test.ts`, `src/cli/help.test.ts`, `src/main/skills`: 240 files / 2,019 pass
- Live smoke on the built CLI of every `skills get <topic>` and `--full`, every emulator, linear, and vm verb named in the guides, and every projection's resolver, GNOME warning, and bounded fallback (done on the #16904 branch before the split; the guide bodies are identical here except the send-receipt vocabulary noted above)
## Deferred product decisions
Merging `orca-emulator` and `orca-emulator-android` into one skill with a platform branch; collapsing `linear-tickets` to a guide alias; a `skills get --reference <name>` selector so a gate table can load one file; a fresh-agent routing eval before trimming the `orca-cli` (1,015 chars) and `orchestration` descriptions, whose quoted triggers each fixed a routing misroute.
|
||
|
|
2283f8ba4e |
docs(orchestration): never pick a worker model the user did not name (#19109)
The sonnet examples were added for a test cohort. Orchestration must not choose a model on the user's behalf: pass --model only when the user named one, otherwise inherit the configured agent default. |
||
|
|
06a607a1d7 |
feat(orchestration): make multi-agent workflows durable (#16904)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 225 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$21666 | $\color{#cf222e}{\Huge{\mathbf{−}}}$2820 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$18846 |
| Prod | 348 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$17107 | $\color{#cf222e}{\Huge{\mathbf{−}}}$4706 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$12401 |
<!-- /orca-pr-loc -->
## ELI5
Orca now treats orchestration like a durable control plane instead of inferring success from terminal keystrokes. Agents can tell whether a prompt was accepted or a turn started, replay an ambiguous request without sending twice, and recover coordinator mail after a crash. Completed workers can be inspected, released, or retained, and their panes no longer auto-resume as if the work were still running.
## What changed
- **Run receipts** from `run-create/use/current/show/list` are the row without routing plumbing (`home_database`, `coordinator_pane_key`) and without the duplicate `binding` object.
- **`terminal send` receipts are honest and idempotent.** `input_accepted` and `turn_started` are the only stages; `--wait-submit` observes without resending; `--retry-request <uuid>` replays the exact request against the same process incarnation. A transport timeout keeps the retry ID; only a different runtime answering strips it. Value-less or non-UUID `--retry-request` is rejected on the CLI and the SSH shim.
- **Mailbox delivery is committed before wakeup.** Pointer writes are staged in the DB before any PTY byte, replayed once after restart, and never emit a naked Enter. The watermark that parks concurrent deliveries is released with the DB reservation. Restart rescans pointer-pending and `dispatch:` mailboxes.
- **Lifecycle is a guarded transition graph** (`lifecycle-transition.ts`) with a table-driven test over every caller edge. Task reopen/overturn stays in the public contract. A PTY exit during `worker-stop` is the stop succeeding, not a failure.
- **Worker lifecycle CLI:** `worker-start` (`--spec` creates Task + attempt in one call), `worker-show`, `worker-read` (provider transcript first, bounded terminal fallback with a typed reason, local/WSL/SSH), `worker-stop`, `worker-abandon`, `worker-release`, `worker-retain`, `worker-list` (rowid-fenced pagination, fleet liveness, `attention`, literal `nextAction`).
- **Release is an explicit ownership table** (`decideWorkerTerminalRelease`): only an `owned` resource can be settled, the archive is mandatory where reachable, and an owner whose process is proven exited can always get out of `retained` via `archive_status: unavailable`. User-taken-over, external, and transferred panes stay retained.
- **Settled-worker resume fence** (folds in #17651): a settled dispatch whose pane is still open is fenced at settlement, on stop/abandon/exit, and at startup; lifted on release, retain, takeover, and pane reuse.
- **Liveness is `live` / `unverifiable` / `exited` only**, from execution-host evidence. Fleet projection reads the evidence clock, not the relay delivery clock. A host-certified exit outranks the worker's settled state. `unverifiable` never authorizes stop, abandon, retry, or release, in code or in the guide.
- **Federation:** structured reads negotiate by `method_not_found` so every shipped host keeps transcript-first output; exited remote workers are closed before being reported closed; epoch fencing holds across peer restart, downgrade, and pairing rotation; no per-second forced capability probe.
- **Schema v35:** repairs databases stamped v34 by the pre-fix branch (mailbox_handle default, index predicates), drops the write-only `lifecycle_transition_receipts` ledger and five never-read v31 identity columns.
- **Schema v36:** `dispatch:<id>` mailboxes get a real consumer generation on `dispatch_contexts` and `remote_dispatch_attachments`, bumped and fenced in the same transaction on every re-attach (manual inject, worker-start, federated attach). A stale worker whose Dispatch moved to another process now gets `consumer_fenced` instead of silently acking the new worker's Delivery. Run mailboxes already worked this way.
- **Schema v37:** `dispatch_contexts` records its creator (`creator_handle`, `creator_pane_key`), so a coordinator's context-only self-dispatch is bookkeeping rather than a nesting parent; before this, one self-dispatch made every later `worker-start` from that coordinator fail the depth cap. Pre-v37 rows keep counting (fails closed).
- **Dispatch-mailbox ownership is checked, not inferred.** A `check` from a process whose pane no longer holds the Dispatch, or whose last Attempt was abandoned/failed and moved to another terminal, gets `consumer_fenced` instead of an empty inbox that reads as "no mail yet". `--peek`/`--all` stay readable. A paneless caller still gets `stable_pane_required` with the rebind recovery.
- **Liveness certification is stricter:** a `process_exited` stage whose termination reason is `unknown` (a stop that was issued but never observed) projects `unverifiable`, not `exited`. Federated `worker-show` carries the execution host's verdict and host kind instead of a local guess. A live, ready worker with nothing pending has `nextAction: none` rather than pointing at the `worker-show` that produced it.
- **Wire:** `workerShow` keeps `dispatch.task_id` next to `taskId` for shipped CLIs. `ask --json` uses the standard `{ok, result}` envelope like every sibling verb.
- **Migration start-version detection** treats the two v32 recovery columns as versioned. Before this, every shipped database stamped below 32 resolved to the v6 floor and replayed the whole chain (the v23 backfill synthesized 68 phantom retained workers on a real v30 profile). Verified on a copy of a real 62 MB v30 profile: starts at 30, no row delta, integrity ok, 11 ms.
- **Skill guide** rewritten as a ≤200-line kernel plus seven references, to the outcome-first standard (Result / Done / Safe failure first, conditions not case lists, one done bar, references loaded at the point of use). The canonical loop uses `worker-start --spec`, names `worker-list` for completion accounting, documents `--retry-request` / `request-show` / `--wait-submit`, and requires positive evidence before any stall action. The other seven guides get the same treatment in #18724, split out so this PR stays orchestration-only.
- **`rpc/methods/orchestration-*`** (126 flat files) regrouped into `orchestration/{worker,federation,messaging,runs,gates}/`.
## Why
User reports showed the same boundary failures: false `agent_prompt_stalled` causing duplicate sends (#15180), coordinators unable to trust screen scrapes, cold-parked terminals receiving a pointer without the submit, settled workers accumulating as live tabs and auto-resuming after restart, and no way to tell a stalled worker from a working one.
## Linked issues
Fixes #15180. Fixes #17935 (orchestration skill description is 866 characters; a guard now caps every bundled skill at 1,024). Supersedes #17651 (fence folded in). Advances #16660, #16522, #14907, #13047.
## Review record
This PR was reviewed adversarially after revival: eight independent lenses (lifecycle, mailbox, send, worker, federation, transcript, complexity, live ergonomics), each required to prove findings with a failing test. That produced 16 proven blockers, all fixed with red-then-green regression tests, followed by two re-review rounds and a third fix wave that caught 3 regressions introduced by the fixes and 7 fixes that missed their target; all closed. A final pass (five lenses incl. a live built-runtime smoke, then a re-review of the fix wave) found and fixed seven more, chiefly the stale-worker mailbox steal, the self-dispatch depth wedge, and the unproven-exit certification. Three independent Codex (gpt-6-astra) passes followed: the first found nothing new, the second found and fixed 3 defects (task-status reachability, WSL-local host classification, peer-capability epoch), the third found and fixed 6 (production PTY controller never installed settled writes, ambiguous in-flight pointer failures allowed duplicate replay, SSH/relay deadlines cut off a valid `--wait-submit`, stop-vs-exit race during inspection, and two release-recovery paths for vanished or exited terminals). The full record (findings, proof tests, triage, declines with reasons) is archived outside the repo.
**Rework after the live smoke.** A first live cross-host run on the shipped adhoc build (this Mac, a paired Windows host on the same build, a paired Mac on 1.4.195, and an SSH host) found a P1: a running local worker read `unverifiable`/`missing_status` because the fleet snapshot rows lacked the terminal handle the matcher keyed on. A 59-row failure table over every bug fixed during review showed the same two classes recurring: a fact dropped in transit through optional fields, and two authorities for one fact. Two blind designs (Opus, Codex) converged on the same mechanisms, and the scoped tranches landed here with red-then-green seam tests from the real producer to the real consumer, faults injected only at the transport or hook-ingest boundary:
- **Settlement (data-loss class):** one three-valued `WriteSettlement` (`accepted | refused{reason} | unverifiable{reason, bytesHandedToTransport}`) from the SSH multiplexer through daemon client, providers, controller, to pointer staging. No boolean, no rejection-as-third-state. The two silent degrades that fabricated a handoff are deleted; a provider that cannot settle refuses before any effect. Pointer text and Enter share the contract; a partial flush is `unverifiable`, never `refused`.
- **Evidence identity (false-liveness class):** fleet agent-status evidence is a tagged union (`binding: worker | pane | unresolved{reason}`, `clock: observed | delivery`) minted once at ingest, so a hook row captured on one process incarnation can never bind to a later dispatch on the same pane. The matcher's `!worker.paneKey ||` defaults are gone. One host-scope parser replaces two.
- **Small pre-merge items:** `capability_unsupported` from an old peer is no longer relabelled `host_unavailable`; a producer census test asserts every agent-status consumer path projects a pane-only hook row as `live`.
Two ergonomics defects the second live run surfaced on a real database are fixed here too: a pre-v3 dispatch already marked `completed` projected as `outcome_unknown` / `requiresAction: true` forever (three copies of the outcome ladder disagreed on legacy rows; now one resolver, legacy `completed` reads `succeeded` with nothing to act on, legacy `failed` stays actionable on the failure), and an unscoped `worker-list` enumerated the entire database (now defaults to the Run bound to the calling terminal, `--run` overrides, and the receipt's additive `scope` field says which).
A third live round on the shipped adhoc build of `b082443e1f` (same four hosts) plus an unscripted run in the user's own prompt style (a plain Claude Code shell, `/orchestration`, three workers, zero errors, bound-Run default confirmed) found two more branch defects, fixed with red-then-green tests: a worker freshly started on a paired server projected `unverifiable`/`host_indeterminate` with `requiresAction` for ~3 minutes, including after its own `worker_done`, because the host's federation observation returned `missing_liveness_verdict` for any PTY the liveness register had not yet swept (the host now reads a connected pane it owns locally as `live`; disconnected or SSH-scoped panes stay `unverifiable`); and six pre-v3 completed rows still carried an `input` category because settling through the task-status path or `failDispatch` never closed the Dispatch's pending question threads (both paths close them now, and schema v38 closes threads already pending on settled rows). The guide's `worker-start` examples now show `--model sonnet`, since an omitted model inherits the launcher's default.
A Codex adversarial pass on the tranche diff found one real design hole (identity minted at read time instead of ingest, now closed) and two daemon settlement paths that threw instead of settling (fixed). Two `@ts-nocheck` runtime mixins on these paths were extracted into checked modules; the repo-wide `@ts-nocheck` count is unchanged at 171.
Deletions during review: ~1,900 lines (write-only ledger, unread columns, dead v1 archive path, test harnesses shipped in prod, duplicated liveness and state-machine copies, self-capability checks that were compile-time true).
## Testing
- `pnpm typecheck:tsc:node|cli|web` clean
- `pnpm run check:code-quality:changed` 0 findings; `check:react-doctor:changed` 0
- `pnpm verify:bundled-skill-guides`, `verify:skill-bundle-manifest`
- full `pnpm test` on the integrated head: 72,332 pass / 292 skipped; the only failures were three non-PR files (two zsh live-shell suites hit a node-pty spawn-helper ENOENT while a concurrent native rebuild ran, 44/44 in isolation; `release-checkout.unit.test.ts` is a known 30 s load timeout that passes in isolation on `origin/main` too).
- CI on
|
||
|
|
54a8afc91d |
fix(orchestration): typed error codes for dispatch and worker-start refusals (#18902)
* fix(orchestration): typed error codes for dispatch and worker-start refusals orchestration dispatch (and worker-start, which composes it) surfaced task not found, task not ready, and inject rejected as the same bare runtime_error, so an agent reading the receipt could not choose between creating the task, waiting on dependencies, or picking another terminal. Add task_not_found (data.taskId), task_not_ready (data.status, data.unmetDependencies), and inject_rejected (data.terminal, data.reason), each carrying data.nextSteps so every shipped CLI already prints the recovery. worker-start's not-ready refusal moves from task_not_startable to task_not_ready with the same detail. runtime_error stays for genuinely unexpected failures. Proven red-first from RpcDispatcher through the CLI's own failure formatting, plus an SSH bridge test that the host CLI's typed refusal relays unchanged. * test(orchestration): load CLI formatter at runtime in the dispatch-code test The composite node typecheck (config/tsconfig.node.json without --composite false, as CI runs it) rejects a static import of src/cli from a main test with TS6307. Load the formatter and error class dynamically behind narrow structural types, as the CLI/runtime boundary test does. * fix(orchestration): keep task_not_startable and split the CLI-format proof Review on #18902: - Drop task_not_ready. worker-start already published task_not_startable for a not-ready Task, so renaming it would change an existing receipt value under old clients. dispatch now emits task_not_startable too (it was a bare runtime_error before, so this is purely additive), with the new data.status / data.unmetDependencies / data.nextSteps. - Move the refusal receipts (code, message, data) into src/shared/orchestration-dispatch-refusal-contract.ts so the runtime emits them and the CLI test formats the identical envelope. The RPC test under src/main asserts toEqual against the contract; the new src/cli/orchestration-dispatch-refusal-format.test.ts feeds those same receipts to formatCliError / reportCliError. Neither tsconfig widens and the composite typecheck CI runs is clean. * fix(orchestration): keep published refusal messages and type the DB claim guards Codex review of #18902: - Every call site keeps the exact message it published on main ("Task not found: <id>", "only a ready Task can start.", "cannot retry from Dispatch"); the shared contract now takes the message per site and only owns the code and data. Baseline strings are pinned as literals. - createDispatchContext's own missing/non-ready guards, including the atomic-claim loser, now emit the same typed receipt instead of a bare Error, so a dispatch that races a status change no longer flattens to runtime_error. Covered by a dispatcher-level race test. - Invalid --retry-of keeps task_not_startable but now carries status, unmetDependencies, retryOf, and a retry-specific next step. - Dependency recovery text distinguishes waiting on running deps from retrying/unblocking failed ones. - CLI test adds an unknown-code case so the old-client claim rests on an assertion, not a comment; SSH test asserts exact stdout. - Guide table narrowed to the covered preflight cases; occupancy stays runtime_error and is named as such. |
||
|
|
3e4fd4a7af |
Shorten orchestration skill description under the Agent Skills 1024-char limit (#18683)
* Shorten orchestration skill description under the Agent Skills 1024-char limit The folded description was 1038 chars, so spec-conforming installers such as SkillStar rejected the bundled orchestration skill. Drop the two clauses already covered elsewhere in the same description: "decomposing work across agents" (implied by "structured multi-agent coordination") and "automation of the browser embedded inside Orca" (restated by the locked `orca-cli` embedded-pages sentence). Every routing trigger asserted by orchestration-skill-guidance.test.mjs, the orca-cli handoff boundary, and the Computer Use boundary are unchanged. Result: 958 chars. Add config/scripts/skill-description-length.test.mjs, which parses every skills/*/SKILL.md frontmatter with `yaml` and fails on an empty or >1024 char description, so the regression cannot return. orca-cli sits at 1015 and is left as is. Fixes #17935 * Keep the embedded browser in the orchestration description's orca-cli routing Restores the word "browser" in the orca-cli sentence ("and the Orca embedded browser") so agents scanning for it still route embedded-browser control to orca-cli. Description is 985 chars, 39 under the spec limit. |
||
|
|
573537ecd4 |
feat(cli): make terminal close the canonical workspace teardown (#18073)
* fix(runtime): recover stale session owners and await retirement * fix(runtime): preserve session hydration and smoke compatibility * test(runtime): cover empty and unindexed session owners * feat(cli): make terminal close the canonical workspace teardown * fix(preload): align ssh termination result type * test(runtime): assert folder hydration owner * fix(runtime): fence legacy terminal stop by worktree host * fix(preload): reconcile ssh result import with main * fix(runtime): keep same-id sibling hosts out of workspace close The stale-owner fallback in the session controller re-routed any worktree whose catalog partition had no tabs to whichever other partition held tabs. Only `runtime:` environment ids rotate across relay restarts; `repoId::path` legitimately repeats across hosts, so an SSH workspace close could retire the local copy's tabs and resume records, or flip owners mid-close and strand the SSH PTY. Restrict the fallback to runtime hosts, and pin the session partition once per workspace close so record clearing targets the partition that owned the tabs. * test(runtime): give the cross-host close fixture a real resume record * fix(preload): take main's ssh-bridge import order so the merge stays duplicate-free |
||
|
|
b44ef1e59d |
fix(skills): narrow computer-use discovery boundary (#17736)
* fix(skills): narrow computer-use discovery boundary * chore: remove merge-formatting noise * fix(skills): name browser page automation surfaces |
||
|
|
fbe94ceff6 |
fix: close readiness gaps found by merged-change audit (#17159)
* fix(ssh): fence stale kills and retired pane replay * fix(ssh): support cancellable interactive authentication * fix(ssh): await remote catalog before snapshot adoption * fix(pty): contain Windows ConPTY input failures * fix(power): avoid redundant macOS display blocking * perf(editor): narrow markdown override subscriptions * fix(quick-open): close directory handles after reads * refactor(linux): remove unused proc socket scanner * fix(usage): apply flat Sonnet 4.6 pricing * ci: prime Node next native test cache * docs(skills): resolve snapshot cleanup data path * fix(ssh): recover install locks after host reboot * test(ssh): recognize boot-aware install locks * test(ssh): prove previous-boot lock recovery live * test(wire): pin pre-metadata release coverage * fix(terminal): preserve remote tab ownership through recovery races * test(runtime): fence replaced terminal handles in agent guard * fix(ssh): preserve remote snapshot authority across polls * fix(pty): contain late ConPTY output EPIPE * test(pty): register Windows exit watcher before kill * fix: close SSH and tab readiness race gaps * fix(tabs): retain headless order and placeholder titles * fix(build): avoid parallel electron-vite config race * test(windows): avoid MSYS temp path rewriting * test(windows): avoid killing exited PTY * fix(pty): avoid late ConPTY input teardown race * fix(terminal): sync reconnect error ownership after commit * fix(runtime): use canonical worktree identity comparison * test(ssh): assert complete cold-hydration baseline * test(windows): invoke quoted retention fixture via PowerShell * test(windows): read ConPTY grid through mode con * fix(terminal): publish PTY replacements atomically * fix(terminal): infer stale identity on reattach * fix(terminal): fence stale pane PTY callbacks * fix(terminal): fence stale pane binds after rebind * fix(terminal): reject stale pane transport callbacks * fix(terminal): fence mirrored reattach spawn callbacks * fix(terminal): replace stale pane PTYs on remount * fix(ci): size the Windows launcher-compile test budget from measurement `native-smoke (windows-latest)` fails ~4.5% of runs on `preserves a multiline argument through the compiled remote launcher` with "Test timed out in 15000ms" — on unrelated PRs, for reasons that have nothing to do with them. Across 176 sampled attempts it is the only red that job produced, and it hit seven different PRs in two days: #16900, #16904, #16915, #16955 (twice), #16979, #17014, #17085. The test is six process creations: powershell.exe forks csc.exe, then the freshly compiled orca.exe forks node.exe, twice. Hosted Windows runners periodically slow process creation down, and this test amplifies that far harder than anything else in the job. Comparing the 80 attempts where it ran under 3s against the 12 where it ran over 12s, its own median goes 2198ms -> 15917ms (7.2x) while the same file's powershell-only test moves 556 -> 686ms (1.2x), the cmd.exe and Git Bash process tests in the neighbouring file move 1.4x, and the other 35 files put together move 1.5x. Measured across those 176 attempts: 1881ms to 35438ms, p50 4264ms, correlation +0.881 with the job's total Vitest duration. 8 of 176 (4.5%) exceeded the 15s cap; 2 of 176 (1.1%) also exceeded the shared 30s testTimeout, so deleting the override and inheriting the config is not enough on its own. 60s clears all 176 with 1.7x headroom on the worst. This is slow, not hung. Every body here is synchronous spawnSync, so Vitest cannot interrupt one — the timer fires only after the body returns and the reported duration is real elapsed time. That is why a failure reads `× ... 22464ms` under `Test timed out in 15000ms`. The work finished; the stopwatch was short. Seven reruns at one identical head measured 2053 / 4680 / 5551 / 8732 / 13506 / 14868 / 21937ms — the last of those would have been red on code that had not changed. The 15s came from #8897, which raised this test off Vitest's built-in 5s default because the job then ran bare `pnpm vitest run`. #8909 landed 3h27m later and pointed the job at config/vitest.config.ts, which is the real fix for that. The constant stayed behind and has been the binding budget ever since. * fix(terminal): fence stale remount reattach ownership * fix(terminal): reconcile mounted pane identity after replacement * fix(terminal): fence stale reattach fallback ownership * fix(terminal): fence deferred SSH reattach ownership * fix(terminal): fence stale split pane ownership callbacks * fix(terminal): keep stale spawns from consuming startup --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
cc6b600e21 |
Fix orchestration CLI recovery, settled-Dispatch mail, and guide defects (#16919)
* Fix orchestration CLI recovery, settled-Dispatch mail, and guide defects Five reported orchestration CLI defects, verified individually before fixing. Two were real code defects, one was a docs error, one was correct as-is, and one was correct on both ends except for its recovery wording. - Mail addressed to a settled `dispatch:<id>` was accepted and silently dropped. Local sends bypassed the settlement check the federated branch already had, so the caller was told success for a delivery no worker would ever read. Reject with `dispatch_inactive` and name the Run mailbox to use instead. - A lost mutation response offered no read-only way to ask whether it took effect. `--retry-request` does dedupe correctly, but the recovery guidance emitted a query command only when the payload carried a dispatch id, which is exactly what a lost response lacks. Add read-only `orca orchestration request-show --request <id>` over the durable receipt ledger, and always emit a read-only step before the keyed retry. - The bundled `orca-cli` guide documented `check --unread --inject`, a flag the parser rejects. Correct it to `--format` and add a ratchet that runs every orchestration invocation in the bundled guides through the real CLI parser. - `check --json` is one stdout document and its keepalives are stderr-only; the reported `Extra data: line 2` came from merging the streams. Document the contract rather than changing the wire. - A rejected lifecycle message is loud on both ends already, but the rejection never named the flag that supplies the missing capability. Name it. * Harden orchestration mutation recovery guidance |
||
|
|
2b0ee06205 |
docs(env-recipes): warn that snapshotting a started runtime bakes its identity (#17001)
Snapshotting a VM on which `orca serve` has already run captures the runtime's user-data dir into the image. Every VM booted from that image then shares one pairing identity and one agent-session-authority key, which defeats the per-device token design. Confirmed by booting two VMs from one such snapshot: both emitted identical deviceToken and pairedDeviceId. Adds the rule to the base-snapshot section and repeats it for the agent-auth layer, which is the likelier place to start the runtime by hand while smoke-testing. Says to delete the whole user-data dir rather than a named file list, since that list drifts as Orca adds state. |
||
|
|
c4b39295c1 |
style: format codebase (#16935)
* style: format codebase * style: format codebase * refactor: extract skill install dialog footer and content Extract footer and content sections from SkillInstallDialog and SkillInstallManagementDialog into separate components for improved maintainability and clarity of component responsibilities. |
||
|
|
9135b6f004 |
feat(orchestration): surface nested worker depth and propagate it across hosts (#16669)
* feat(orchestration): surface nested worker depth and propagate it across hosts Builds on the depth enforcement in the previous commit, which shipped with the setting reachable only by editing settings.json and with workers never told they could nest. Adds the Settings -> Agents control (a 1/2/3 select rather than a free-form number, which bounds the value without inventing a numeric input primitive). The key stays absent from the SettingsUpdate RPC schema, matching agentSkillSharingEnabled: settings.update is reachable from the CLI, so an RPC-writable depth would let a worker raise its own cap. Adds a SUB-DISPATCH block to the dispatch preamble, emitted only when the worker actually has budget left. A worker told it "usually cannot" delegate still tries and then reports the refusal as a blocker, so the section is omitted entirely rather than softened. Propagates depth to federated worker hosts. Previously the home side computed and stored a depth the remote host never received, so a remote attachment always read as depth 1. That is correct at the default cap and wrong as soon as the cap is raised — precisely when someone starts relying on nesting. The field is optional, so an older Run home simply omits it and the attachment's NOT NULL DEFAULT 1 keeps the fail-closed behaviour. Enforcement still runs on the executing host against that host's own cap, consistent with the SSH execution boundary. * fix(orchestration): close nested depth readiness gaps * fix(settings): defer nested depth translations * fix(orchestration): drop federated depth keys that main already landed The enforcement PR's review pass added the same federated depth propagation before it merged, so replaying this branch onto main produced duplicate object keys. Keep main's versions -- its schema entry validates an integer >= 1 rather than any finite number. * fix(settings): label nested worker depth select * fix(settings): move nested depth to orchestration * fix(settings): refine nested depth placement |
||
|
|
8a07bbd8cf |
fix(orchestration): enforce nested worker depth instead of an accidental fence (#16668)
* fix(orchestration): enforce nested worker depth instead of an accidental fence Orca documented that "dispatched workers cannot spawn their own sub-workers (worker-start is coordinator-fenced)". No such check existed. What existed was a single Run-binding check in the workerStart RPC: a worker's terminal is not bound to a Run, so worker-start happened to fail. The rule was emergent, asserted by no test, and written in no doc — and it leaked. A worker could run-create its own Run, task-create, and worker-start: now bound, the check passed. Replace it with a real, configurable depth cap. Depth is derived from the caller's own active Dispatch rather than from Run binding, which is what dissolves the run-create bypass: creating a Run does not stop you being a worker. Enforcement lives in a single dispatch-row writer that owns all three INSERTs that mint a live worker — the generic claim, the supervised worker-start path (including every retry), and the remote attachment. Two of those were missed by earlier drafts of this change, so `creator` and `maxDepth` are required parameters: a new spawn path cannot compile without deciding, and a boundary test refuses the SQL anywhere else. Schema v30 adds depth to dispatch_contexts and remote_dispatch_attachments, NOT NULL DEFAULT 1 and backfilled to 1 so an unstamped or pre-upgrade row fails closed rather than reading as a root coordinator. The attachment pane indexes widen to the five states in which a remote worker may still be running: loss of contact is not evidence of process death, so an unverifiable worker still counts as a nesting parent. Also adds the caller-evidence assertion that workerStart was the only Run-scoped verb to skip, so a declared --from cannot name another terminal's pane and inherit its depth. Default is 1, so behaviour is unchanged unless the new setting is raised. Two limitations are deliberate and documented rather than papered over: this is a guardrail and not a security boundary, since a caller whose launch evidence is unverifiable (any ordinary restored terminal) can declare another handle; and it is enforced at supervised dispatch creation, so a settled worker whose process is still alive counts as a root again. * fix(orchestration): share caller resolution and pin worker gaps * refactor(orchestration): make the caller resolver's pane contract explicit Overloads so requireStablePane callers get a non-null string instead of casting, and rename the attestation opt-out to say what it means: the caller asserts it itself. A flag called assertEvidence:false reads as "attestation optional", which is the hole this helper exists to close. * fix(orchestration): propagate dispatch depth to federated workers * chore(cli): refresh bundled orchestration guide |
||
|
|
a9781a4118 |
STA-4150: client-hosted remote browser (consolidated) (#15448)
Co-authored-by: Jinwoo-H <jinwoo@stably.ai> |
||
|
|
3fca1d1648 |
fix(linear): unbound list-issues by default, surface truncation, bind cursor workspace (#15824)
Fixes STA-5076. list-issues capped at 50 by default and hard-clamped at 250, with hasMore buried under result.meta and no stderr warning for --json, so a page that stopped early read as a complete answer. Omitting --limit now walks Linear's pages until they run out (meta.limit is null), and --limit <n> is the only cap, paging past Linear's 250-per-request maximum to reach it. result.truncated sits next to result.issues and is set only when a cap actually held results back; human output prints "truncated: showing N". The read still has to fit the CLI's 60s RPC budget, so a 20s wall-clock deadline and a 200-page ceiling stop the walk early and report truncated with a continuation cursor rather than failing the command. Also: - issued --cursor values bind the resolved workspace, so call -> nextCursor -> call works without --workspace; raw Linear cursors still need one and now carry nextSteps - issued cursors whose payload smuggles back `all` or an empty workspace are rejected at decode, since either would widen the read past the bound workspace - JSON issue rows carry priorityLabel (none/urgent/high/medium/low), matching orca linear priority set - truncated and priorityLabel are optional on the wire, so a host that predates either is not read as "complete"; readers fall back to meta.hasMore - the truncation line prints the rows actually rendered, so a remote result with no meta.returned cannot print "showing undefined" |
||
|
|
2a760e310b |
fix(computer): report unasserted accessibility actions (#15028)
* fix(computer): report unasserted accessibility actions * fix(computer): fail closed on missing action metadata * Fix merged tab search test fixture |
||
|
|
fc8b92e507 |
docs(computer): explain screenshot file requirements (#15054)
* docs(computer): clarify screenshot output requirements * fix(cli): do not advertise an unshipped --probe flag The capabilities help line referenced --probe, which does not exist yet; it ships in a later change. Advertising it here would be false until then. * fix(cli): align computer-use screenshot guidance * docs(computer): document inline screenshot fallback * docs(computer): keep screenshot summary accurate * docs(computer): keep screenshot guidance general |
||
|
|
0bedeea642 | fix(orchestration): expose unsupervised dispatch lanes (#15105) | ||
|
|
fa9b20cb41 | feat(skills): reland private bundle sharing safely (#14934) | ||
|
|
763b1febeb |
Revert "feat(skills): add private bundle sharing (#14401)" (#14913)
This reverts commit
|
||
|
|
757fae28d7 |
feat(skills): add private bundle sharing (#14401)
Co-authored-by: E2E Test <e2e@test.local> |
||
|
|
66dfdc456f |
feat(computer-use): support macOS middle click and stop the silent left-click fallback (#14721)
* feat(computer-use): support macOS middle click and gate the AX click path `--mouse-button middle` already validated end-to-end through the CLI, the zod schema, and the provider validator, and both the Windows and Linux providers honored it. Only the macOS provider rejected it outright with "middle-click is not yet supported", so the flag was a dead end on the one platform that has no fallback. Two changes: - Add `.middle` to the macOS button mapping. macOS has no dedicated middle event family, so it rides `otherMouseDown`/`otherMouseUp` with the button number carried by `mouseButton: .center`; that constructor argument is honored for exactly the `otherMouse*` types, so no extra field write is needed. - Validate the requested button before the accessibility fast path, and skip that path for buttons it cannot express. Previously the raw string was read unvalidated, and `performClickAction` only special-cased `right`, so `click --mouse-button middle --element-index N` (no modifiers, count 1) fell through to `AXPress` — a left click — and reported success with `path: "accessibility"`. Any unrecognized button string did the same. This matches guards the Windows and Linux providers already had. The button enum moves into `OrcaComputerUseMacOSCore` so it is unit-testable; `main.swift` keeps only the CoreGraphics mapping. Also documents `--mouse-button` in the computer-use skill guide, which never mentioned the flag, so agents on Windows and Linux had no way to discover it. * test(computer-use): cover macOS middle click in the real-desktop e2e suite * test(computer-use): prove macOS middle-click delivery |
||
|
|
500b72d8ef |
fix(vm): harden provisioned root ownership and cleanup (#14477)
* fix(vm): verify provisioned root ownership * test(vm): retry transient removal menu * test(vm): stabilize provisioned root teardown * fix(vm): clarify recipe-owned cleanup * fix(vm): pin provisioned root source commit * fix(vm): make runtime cleanup user-cancellable |
||
|
|
77b37d85e2 |
feat(vm): create workspaces from provisioned SSH roots (#14359)
* feat(vm): use recipe-provisioned SSH roots * fix(vm): preserve ordinary create failure timing * test(vm): prepare provisioned root SSH fixture * ci(vm): enable SSH setup for provisioned root E2E |
||
|
|
1f4b731f7c |
fix(skills): use exported recipe id in environment guide (#14280)
* fix(skills): use exported recipe id in environment guide * fix(skills): keep recipe-derived Vercel names valid --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> |
||
|
|
cbca291aa7 |
fix(orchestration): preserve direct user authority after worker_done (#14192)
* fix(orchestration): preserve direct user authority * test(orchestration): assert settled dispatch boundaries |
||
|
|
d349f9a972 | Use provider-neutral Opus alias in orchestration guide (#14119) | ||
|
|
3ec48a74d5 |
Gate artifact publishing behind off-by-default capability (#13368)
* fix(artifacts): gate agent artifact publishing behind an off-by-default capability Public artifact sharing was reachable by any agent through `orca artifacts share`: the Artifacts settings toggle only controlled sidebar visibility, and nothing in the main process checked a capability before minting a public URL. Add `artifactSharingEnabled` (default off) and enforce it in ArtifactCloudService.share/update — before auth, network, or the share-record write — so the CLI, relay-forwarded remote CLI, and IPC paths are all denied. The denial carries a stable `artifact_sharing_disabled` code plus next steps through the RPC error allowlist, so the CLI prints actionable guidance. list, unshare, and delete stay ungated: turning publishing off must not strand already-published links. The capability is absent from the `settings.update` RPC schema, so an agent cannot grant it to itself — only the desktop UI can. Co-authored-by: Orca <help@stably.ai> * fix(artifacts): gate agent artifact publishing behind an off-by-default Publishing is blocked until enabled in Settings → Artifacts. CLI preflights the capability before reading files to avoid unnecessary uploads. RPC surface rejects capability grants so callers cannot self-grant. UI shows opt-in workflow and recovery path when publishing is off. Web clients mirror the host's setting read-only. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
c991bb27d3 | Add account-backed artifact sharing (#13012) | ||
|
|
b0ba51831c | Add per-worker model and effort overrides (#12851) | ||
|
|
39c3c58d55 |
perf(runtime): gate terminal.list visual layouts (#12450)
* perf(runtime): gate terminal.list visual layouts and stop the false writable claim visualLayouts is ~31% of a large terminal.list payload (44,208 B of 137,412 B on a live 134-terminal remote runtime) and has exactly one consumer: the human-readable CLI formatter. Gate it behind an includeVisualLayouts request param that defaults to included, so pre-flag clients are unaffected, and have every --json/internal caller opt out. Also drop the record-backed builder's writable, which was a verbatim copy of connected. terminal.show now states writability explicitly as exactly what terminal.send's PTY gate enforces. * test(runtime): type the payload-size fixture arrays for tsc * fix(runtime): preserve terminal list compatibility * test(runtime): guard terminal list optimization * fix(cli): preserve agent access to terminal layouts |
||
|
|
f4b2b782b5 |
feat(orchestration): coordinator-driven release of settled worker terminals (STA-905) (#12355)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
9a2676023c | fix(orchestration): prefer current authority over legacy fallback (#11737) | ||
|
|
d0f341ad69 |
fix(computer-use): make modifier clicks interruption-safe (#11451)
* fix(computer-use): make modifier clicks interruption-safe * fix(computer-use): pace modified Windows multiclicks * fix(computer-use): address modifier safety review |
||
|
|
78b8a37aed | fix(cli): keep automated worktree creation in background (#11445) | ||
|
|
363e478909 |
fix(orchestration): preserve active workers across updates (#11271)
* fix(orchestration): preserve active workers across updates * test(ssh): model absent legacy adoption * test(orchestration): align compatibility contracts * fix(windows): escape updater PowerShell booleans * fix(windows): restore stock uninstall process check * fix(orchestration): keep recovery off renderer startup barrier * fix(orchestration): harden legacy recovery migration * fix(orchestration): close recovery review gaps * fix(orchestration): complete legacy worker cutover recovery * fix(orchestration): preserve legacy workers across updates --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |