An action gate names one reference, but --full was the only way to reach it and
returned the kernel plus every reference. The selector serves one document, and
the orchestration guide now teaches it with --full as the older-CLI fallback.
The seven non-orchestration guide rewrites, the shared stub fragment, and their guards move to a separate PR. The orca-cli cross-guide pin follows the worktree-selector rule into the orchestration placement reference.
Apply the Change findings from the skill review.
Correctness:
- relayGracePeriodSeconds now states that 0 is unbounded (the relay stays up
until explicitly terminated) and names the accepted domain, 0 or 60..604800
from EphemeralVmRecipeSshTargetSchema. Removed from the SSH exemplar so the
omitted default and a written value cannot disagree.
- The SSH exemplar carries only the required fields. jumpHost/proxyCommand
exclusivity is stated once, where the choice is made.
- The provisioned-root snippet fetches "$ORCA_REPO_URL", the remote Orca
resolved the base ref against, not origin.
- portForwards entries name localPort/remoteHost/remotePort and optional label
against the strict SavedPortForwardSchema.
- The free doctor gate is clear only with no fail and no warn; buildDoctorResult
leaves ok true with warnings.
- The worked auth check uses the status command's exit code by default via a
sentinel, since a provider CLI need not propagate a remote exit code. The grep
recipe stays as the named fallback and matches a shell variable, so no
pipefail/SIGPIPE hazard remains.
Structure:
- One outcome spine with a joined done bar, and one autonomy envelope replacing
the four money restatements, the five checkpoint rationales, and Boundaries.
- Quick-start deleted; the numbered sequence is the single copy.
- Kernel keeps what must fire without a read; the four worked examples and the
failure modes move to skill-guides/orca-per-workspace-env/references/.
- Description replaced with the review's proposal; the ORCA placeholder rule is
stated once and the Claude Code bang prefix is fenced as a harness adapter.
The generator test's reference assertions are generalized per guide, and the
Vercel name-building pins are repointed to the file that now carries them.
Emulator: drop the camera-injection verb (it does not exist in
src/cli/specs/emulator.ts), drop iOS permissions (the iOS backend
declares permissions: false so the bridge throws emulator_unsupported),
fix the Android permissions positional order and remove the nonexistent
list op, and delete the stale 'visual pane in development' status and
the 'once the attach/active flow lands' qualifier. State the wrapped-verb
and backend-capability conditions once, add an outcome spine, and drop
the ASCII diagrams and identity prose.
Linear: delete the 27-line Common Commands mirror of --help, replace the
verb-keyed unconfirmed-write rules with the payload-keyed condition that
covers every write verb, add a skill-level done bar for all five
branches, and move examples onto the ORCA placeholder.
All four descriptions rewritten off body-owned detail.
Applies the Change findings from the orca-cli / computer-use skill review.
orca-cli guide:
- Delete the duplicated mobile-emulator tail and the second `Next Action`; the
emulator now routes through one conditional-references row.
- State the handoff done bar once, and state the `terminal wait` gate with its
failure direction beside the recipe: `terminal wait` prints an ordinary result
envelope on timeout and signals the unsatisfied wait only through the exit
code, so an agent could send a brief into a half-started TUI.
- State the one-agent-handle invariant once in Worktrees; the Terminals copy and
the two restatements are gone.
- Drop the executable-resolution ladder the discovery stub owns and keep one
placeholder rule, in the shape `skill-guides/orchestration.md` uses.
- Move the reconstructible command catalogs (browser, automations, artifact and
skill publishing) behind `skills get orca-cli --full`; the routing paragraph,
the untrusted-page-content rule and the artifact publish gate stay inline.
- 424 always-loaded lines to 260.
computer-use guide: promote the verification vocabulary to a done bar, and drop
its copy of the resolver ladder.
Repointed the generator's resolver-phrase assertions for these two guides to the
stub projections that now own the ladder, and generalized the bundled-reference
assertions beyond orchestration.
Applies the Change findings from the orchestration skill review.
- Outcome names the next consumer (the requesting user) and the turn-end report
contract, and states the positive-proof condition once beside Safe failure.
- The three wait/release gates cite that condition instead of carrying divergent
case lists.
- Drops the kernel's third copy of the worker_done command; the runtime preamble
owns it at the point of use and worker-contract.md owns it behind the gate.
- Drops messaging-and-gates.md's duplicate coordinator delivery loop block; the
paragraph below it already states the --terminal delta as a condition.
- worker-start gets an ordered failure hatch naming failedStage and
residualResources.
- Conditional references describes what --full returns instead of promising
selective loading; the worker-contract gate row states the condition the kernel
does not decide.
- Moves the worktree-selector form to placement-and-remote.md, where an exact
selector is consumed, and aligns worker-contract.md on the kernel's
capability-first question-TUI wording.
Pins move with the facts: the worker_done flag spellings to worker-contract.md's
recipe case, the worktree-id form to the placement reference.
The superseded-Attempt fence ran ahead of the stable_pane_required guard, so a
terminal that had only lost its pane binding was told it had been re-attached and
to stop, and lost the run-use recovery data. The fence message and the worker
contract now also cover the no-successor case they already fired on.
The worker contract calls an empty check a checkpoint rather than a failure, so
a fenced worker had no sentence telling it that consumer_fenced, not silence, is
how it learns the Task moved. Pinned next to the existing consumer_fenced pin.
A coordinator that never loads a reference read unverifiable for every
--on <environment> worker. Fits the existing wrap, so the 202-line budget
is unchanged.
For a healthy in-progress worker the projection returns nextAction
{kind:'inspect', argv:['orchestration','worker-show','--dispatch',<same id>]}
with attention.requiresAction false, and the kernel told the coordinator to
follow the literal argv. Raises the kernel line budget 200 -> 202: the
paragraph had zero slack and no existing guidance was worth cutting.
check --types without --wait returned an unmatched type. wakeTypes is an
existence probe in getOrCreateMailboxDelivery; the batch query that
follows has no type predicate, so a Delivery is never filtered. Behavior
is correct; only the guide and --help were silent.
The stall path told coordinators to enumerate with plain worker-list, but
a worker started --on <environment> reads unverifiable without
--include-remote. Also names page.nextCursor for fleets past 100 rows.
The branch's rewrite dropped main's verbatim phrases ("hand off", "handoff",
"handover", "give this to another agent", "another worktree", threaded
messages, worker_done/escalation waits, decision gates, reading or waiting on
terminals) — the only text a model sees when choosing this skill. Restored in
both the kernel frontmatter and the identical stub, still shorter than main's,
and pinned by a routing test.
The guide still documented queued_pending_turn and submission_observed, and the
send-submit repro gated on the removed submission_observed literal, so its
unsubmitted assertion passed vacuously. Both now use input_accepted then
turn_started.
The four prompt-delivery warnings were built inside the text formatter, so
--json callers (every agent) saw none of them. The receipt now carries a
warnings array built from the same function, and the swallowed-Enter warning is
ordered ahead of the unsupported-observation arm so an agent provider always
gets the recovery command; a plain shell keeps its cannot-report-delivery text.
The kernel's stall exit fired on "not live", which includes every unverifiable
arm — all of which are absence — contradicting the safety floor and the recovery
table. It now names the positive signals (exited liveness, the worker's own
observation of exit, a final agent turn with no worker_done) and states that
unverifiable never authorizes stop, abandon, retry, or release. Also corrects the
projection.* field paths the worker-list row actually nests.
D1: name the two liveness layers (worker-list projection.liveness is the fleet
verdict, worker-show observation.status is PTY-only) and give the supervised
loop a bounded stall procedure instead of an unbounded wait.
D2: put worker-list, attention, requiresAction and nextAction in the loop and in
completion accounting.
D4: document request-show / --retry-request / terminal send --wait-submit.
D5: check names its caller with --terminal, never --from.
D7: give a dispatched worker a concrete follow-up read cadence.
D10: document the real folder-workspace route (project setup-existing-folder).
worker-start --spec is now the canonical loop's default.
Guidance pins are contracts via squash() instead of reflow-fragile prose.
* fix(runtime): recover stale session owners and await retirement
* fix(runtime): preserve session hydration and smoke compatibility
* test(runtime): cover empty and unindexed session owners
* feat(cli): make terminal close the canonical workspace teardown
* fix(preload): align ssh termination result type
* test(runtime): assert folder hydration owner
* fix(runtime): fence legacy terminal stop by worktree host
* fix(preload): reconcile ssh result import with main
* fix(runtime): keep same-id sibling hosts out of workspace close
The stale-owner fallback in the session controller re-routed any worktree whose
catalog partition had no tabs to whichever other partition held tabs. Only
`runtime:` environment ids rotate across relay restarts; `repoId::path` legitimately
repeats across hosts, so an SSH workspace close could retire the local copy's
tabs and resume records, or flip owners mid-close and strand the SSH PTY.
Restrict the fallback to runtime hosts, and pin the session partition once per
workspace close so record clearing targets the partition that owned the tabs.
* test(runtime): give the cross-host close fixture a real resume record
* fix(preload): take main's ssh-bridge import order so the merge stays duplicate-free
* fix(ssh): fence stale kills and retired pane replay
* fix(ssh): support cancellable interactive authentication
* fix(ssh): await remote catalog before snapshot adoption
* fix(pty): contain Windows ConPTY input failures
* fix(power): avoid redundant macOS display blocking
* perf(editor): narrow markdown override subscriptions
* fix(quick-open): close directory handles after reads
* refactor(linux): remove unused proc socket scanner
* fix(usage): apply flat Sonnet 4.6 pricing
* ci: prime Node next native test cache
* docs(skills): resolve snapshot cleanup data path
* fix(ssh): recover install locks after host reboot
* test(ssh): recognize boot-aware install locks
* test(ssh): prove previous-boot lock recovery live
* test(wire): pin pre-metadata release coverage
* fix(terminal): preserve remote tab ownership through recovery races
* test(runtime): fence replaced terminal handles in agent guard
* fix(ssh): preserve remote snapshot authority across polls
* fix(pty): contain late ConPTY output EPIPE
* test(pty): register Windows exit watcher before kill
* fix: close SSH and tab readiness race gaps
* fix(tabs): retain headless order and placeholder titles
* fix(build): avoid parallel electron-vite config race
* test(windows): avoid MSYS temp path rewriting
* test(windows): avoid killing exited PTY
* fix(pty): avoid late ConPTY input teardown race
* fix(terminal): sync reconnect error ownership after commit
* fix(runtime): use canonical worktree identity comparison
* test(ssh): assert complete cold-hydration baseline
* test(windows): invoke quoted retention fixture via PowerShell
* test(windows): read ConPTY grid through mode con
* fix(terminal): publish PTY replacements atomically
* fix(terminal): infer stale identity on reattach
* fix(terminal): fence stale pane PTY callbacks
* fix(terminal): fence stale pane binds after rebind
* fix(terminal): reject stale pane transport callbacks
* fix(terminal): fence mirrored reattach spawn callbacks
* fix(terminal): replace stale pane PTYs on remount
* fix(ci): size the Windows launcher-compile test budget from measurement
`native-smoke (windows-latest)` fails ~4.5% of runs on
`preserves a multiline argument through the compiled remote launcher`
with "Test timed out in 15000ms" — on unrelated PRs, for reasons that
have nothing to do with them. Across 176 sampled attempts it is the only
red that job produced, and it hit seven different PRs in two days:
#16900, #16904, #16915, #16955 (twice), #16979, #17014, #17085.
The test is six process creations: powershell.exe forks csc.exe, then
the freshly compiled orca.exe forks node.exe, twice. Hosted Windows
runners periodically slow process creation down, and this test amplifies
that far harder than anything else in the job. Comparing the 80 attempts
where it ran under 3s against the 12 where it ran over 12s, its own
median goes 2198ms -> 15917ms (7.2x) while the same file's
powershell-only test moves 556 -> 686ms (1.2x), the cmd.exe and Git Bash
process tests in the neighbouring file move 1.4x, and the other 35 files
put together move 1.5x.
Measured across those 176 attempts: 1881ms to 35438ms, p50 4264ms,
correlation +0.881 with the job's total Vitest duration. 8 of 176 (4.5%)
exceeded the 15s cap; 2 of 176 (1.1%) also exceeded the shared 30s
testTimeout, so deleting the override and inheriting the config is not
enough on its own. 60s clears all 176 with 1.7x headroom on the worst.
This is slow, not hung. Every body here is synchronous spawnSync, so
Vitest cannot interrupt one — the timer fires only after the body
returns and the reported duration is real elapsed time. That is why a
failure reads `× ... 22464ms` under `Test timed out in 15000ms`. The
work finished; the stopwatch was short. Seven reruns at one identical
head measured 2053 / 4680 / 5551 / 8732 / 13506 / 14868 / 21937ms — the
last of those would have been red on code that had not changed.
The 15s came from #8897, which raised this test off Vitest's built-in 5s
default because the job then ran bare `pnpm vitest run`. #8909 landed
3h27m later and pointed the job at config/vitest.config.ts, which is the
real fix for that. The constant stayed behind and has been the binding
budget ever since.
* fix(terminal): fence stale remount reattach ownership
* fix(terminal): reconcile mounted pane identity after replacement
* fix(terminal): fence stale reattach fallback ownership
* fix(terminal): fence deferred SSH reattach ownership
* fix(terminal): fence stale split pane ownership callbacks
* fix(terminal): keep stale spawns from consuming startup
---------
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
* Fix orchestration CLI recovery, settled-Dispatch mail, and guide defects
Five reported orchestration CLI defects, verified individually before fixing.
Two were real code defects, one was a docs error, one was correct as-is, and
one was correct on both ends except for its recovery wording.
- Mail addressed to a settled `dispatch:<id>` was accepted and silently dropped.
Local sends bypassed the settlement check the federated branch already had, so
the caller was told success for a delivery no worker would ever read. Reject
with `dispatch_inactive` and name the Run mailbox to use instead.
- A lost mutation response offered no read-only way to ask whether it took
effect. `--retry-request` does dedupe correctly, but the recovery guidance
emitted a query command only when the payload carried a dispatch id, which is
exactly what a lost response lacks. Add read-only
`orca orchestration request-show --request <id>` over the durable receipt
ledger, and always emit a read-only step before the keyed retry.
- The bundled `orca-cli` guide documented `check --unread --inject`, a flag the
parser rejects. Correct it to `--format` and add a ratchet that runs every
orchestration invocation in the bundled guides through the real CLI parser.
- `check --json` is one stdout document and its keepalives are stderr-only; the
reported `Extra data: line 2` came from merging the streams. Document the
contract rather than changing the wire.
- A rejected lifecycle message is loud on both ends already, but the rejection
never named the flag that supplies the missing capability. Name it.
* Harden orchestration mutation recovery guidance
Snapshotting a VM on which `orca serve` has already run captures the
runtime's user-data dir into the image. Every VM booted from that image
then shares one pairing identity and one agent-session-authority key,
which defeats the per-device token design.
Confirmed by booting two VMs from one such snapshot: both emitted
identical deviceToken and pairedDeviceId.
Adds the rule to the base-snapshot section and repeats it for the
agent-auth layer, which is the likelier place to start the runtime by
hand while smoke-testing. Says to delete the whole user-data dir rather
than a named file list, since that list drifts as Orca adds state.
* style: format codebase
* style: format codebase
* refactor: extract skill install dialog footer and content
Extract footer and content sections from SkillInstallDialog and
SkillInstallManagementDialog into separate components for improved
maintainability and clarity of component responsibilities.
* feat(orchestration): surface nested worker depth and propagate it across hosts
Builds on the depth enforcement in the previous commit, which shipped with the
setting reachable only by editing settings.json and with workers never told they
could nest.
Adds the Settings -> Agents control (a 1/2/3 select rather than a free-form
number, which bounds the value without inventing a numeric input primitive). The
key stays absent from the SettingsUpdate RPC schema, matching agentSkillSharingEnabled:
settings.update is reachable from the CLI, so an RPC-writable depth would let a
worker raise its own cap.
Adds a SUB-DISPATCH block to the dispatch preamble, emitted only when the worker
actually has budget left. A worker told it "usually cannot" delegate still tries and
then reports the refusal as a blocker, so the section is omitted entirely rather
than softened.
Propagates depth to federated worker hosts. Previously the home side computed and
stored a depth the remote host never received, so a remote attachment always read
as depth 1. That is correct at the default cap and wrong as soon as the cap is
raised — precisely when someone starts relying on nesting. The field is optional,
so an older Run home simply omits it and the attachment's NOT NULL DEFAULT 1 keeps
the fail-closed behaviour. Enforcement still runs on the executing host against
that host's own cap, consistent with the SSH execution boundary.
* fix(orchestration): close nested depth readiness gaps
* fix(settings): defer nested depth translations
* fix(orchestration): drop federated depth keys that main already landed
The enforcement PR's review pass added the same federated depth propagation
before it merged, so replaying this branch onto main produced duplicate object
keys. Keep main's versions -- its schema entry validates an integer >= 1 rather
than any finite number.
* fix(settings): label nested worker depth select
* fix(settings): move nested depth to orchestration
* fix(settings): refine nested depth placement
* fix(orchestration): enforce nested worker depth instead of an accidental fence
Orca documented that "dispatched workers cannot spawn their own sub-workers
(worker-start is coordinator-fenced)". No such check existed. What existed was a
single Run-binding check in the workerStart RPC: a worker's terminal is not bound
to a Run, so worker-start happened to fail. The rule was emergent, asserted by no
test, and written in no doc — and it leaked. A worker could run-create its own
Run, task-create, and worker-start: now bound, the check passed.
Replace it with a real, configurable depth cap.
Depth is derived from the caller's own active Dispatch rather than from Run
binding, which is what dissolves the run-create bypass: creating a Run does not
stop you being a worker. Enforcement lives in a single dispatch-row writer that
owns all three INSERTs that mint a live worker — the generic claim, the supervised
worker-start path (including every retry), and the remote attachment. Two of those
were missed by earlier drafts of this change, so `creator` and `maxDepth` are
required parameters: a new spawn path cannot compile without deciding, and a
boundary test refuses the SQL anywhere else.
Schema v30 adds depth to dispatch_contexts and remote_dispatch_attachments,
NOT NULL DEFAULT 1 and backfilled to 1 so an unstamped or pre-upgrade row fails
closed rather than reading as a root coordinator. The attachment pane indexes
widen to the five states in which a remote worker may still be running:
loss of contact is not evidence of process death, so an unverifiable worker still
counts as a nesting parent.
Also adds the caller-evidence assertion that workerStart was the only Run-scoped
verb to skip, so a declared --from cannot name another terminal's pane and inherit
its depth.
Default is 1, so behaviour is unchanged unless the new setting is raised. Two
limitations are deliberate and documented rather than papered over: this is a
guardrail and not a security boundary, since a caller whose launch evidence is
unverifiable (any ordinary restored terminal) can declare another handle; and it
is enforced at supervised dispatch creation, so a settled worker whose process is
still alive counts as a root again.
* fix(orchestration): share caller resolution and pin worker gaps
* refactor(orchestration): make the caller resolver's pane contract explicit
Overloads so requireStablePane callers get a non-null string instead of casting,
and rename the attestation opt-out to say what it means: the caller asserts it
itself. A flag called assertEvidence:false reads as "attestation optional",
which is the hole this helper exists to close.
* fix(orchestration): propagate dispatch depth to federated workers
* chore(cli): refresh bundled orchestration guide
Fixes STA-5076.
list-issues capped at 50 by default and hard-clamped at 250, with hasMore buried
under result.meta and no stderr warning for --json, so a page that stopped early
read as a complete answer. Omitting --limit now walks Linear's pages until they
run out (meta.limit is null), and --limit <n> is the only cap, paging past
Linear's 250-per-request maximum to reach it. result.truncated sits next to
result.issues and is set only when a cap actually held results back; human output
prints "truncated: showing N".
The read still has to fit the CLI's 60s RPC budget, so a 20s wall-clock deadline
and a 200-page ceiling stop the walk early and report truncated with a
continuation cursor rather than failing the command.
Also:
- issued --cursor values bind the resolved workspace, so call -> nextCursor ->
call works without --workspace; raw Linear cursors still need one and now carry
nextSteps
- issued cursors whose payload smuggles back `all` or an empty workspace are
rejected at decode, since either would widen the read past the bound workspace
- JSON issue rows carry priorityLabel (none/urgent/high/medium/low), matching
orca linear priority set
- truncated and priorityLabel are optional on the wire, so a host that predates
either is not read as "complete"; readers fall back to meta.hasMore
- the truncation line prints the rows actually rendered, so a remote result with
no meta.returned cannot print "showing undefined"
* docs(computer): clarify screenshot output requirements
* fix(cli): do not advertise an unshipped --probe flag
The capabilities help line referenced --probe, which does not exist yet;
it ships in a later change. Advertising it here would be false until then.
* fix(cli): align computer-use screenshot guidance
* docs(computer): document inline screenshot fallback
* docs(computer): keep screenshot summary accurate
* docs(computer): keep screenshot guidance general
* feat(computer-use): support macOS middle click and gate the AX click path
`--mouse-button middle` already validated end-to-end through the CLI, the
zod schema, and the provider validator, and both the Windows and Linux
providers honored it. Only the macOS provider rejected it outright with
"middle-click is not yet supported", so the flag was a dead end on the one
platform that has no fallback.
Two changes:
- Add `.middle` to the macOS button mapping. macOS has no dedicated middle
event family, so it rides `otherMouseDown`/`otherMouseUp` with the button
number carried by `mouseButton: .center`; that constructor argument is
honored for exactly the `otherMouse*` types, so no extra field write is
needed.
- Validate the requested button before the accessibility fast path, and skip
that path for buttons it cannot express. Previously the raw string was read
unvalidated, and `performClickAction` only special-cased `right`, so
`click --mouse-button middle --element-index N` (no modifiers, count 1) fell
through to `AXPress` — a left click — and reported success with
`path: "accessibility"`. Any unrecognized button string did the same. This
matches guards the Windows and Linux providers already had.
The button enum moves into `OrcaComputerUseMacOSCore` so it is unit-testable;
`main.swift` keeps only the CoreGraphics mapping.
Also documents `--mouse-button` in the computer-use skill guide, which never
mentioned the flag, so agents on Windows and Linux had no way to discover it.
* test(computer-use): cover macOS middle click in the real-desktop e2e suite
* test(computer-use): prove macOS middle-click delivery